Reflection AI Is About to Answer DeepSeek — the Open-Weights Race Just Became a Two-Sided War


The announcement hiding inside a scoop

Yesterday, Axios reported that Reflection AI — the Nvidia-backed lab founded by two former Google DeepMind researchers — is about to release its first open-weight foundation model, expected this month. No name, no weights, no license, no benchmarks yet. On paper, it’s a nothing-burger: a rumored model with no specs.

In practice, it might be the most consequential AI release of the quarter. Here’s why.

The “DeepSeek of the West”

Reflection AI was founded in 2024 by Misha Laskin and Ioannis Antonoglou, DeepMind veterans from the lineage that taught machines to master games no human had taught them. The Wall Street Journal has called Reflection the “DeepSeek of the West”: a lab explicitly trying to do for the US what DeepSeek did for China — frontier-scale models with open weights.

The money says this is real. Reflection has raised close to $2.6 billion, with Nvidia itself investing $800 million; Sequoia, Lightspeed, Eric Schmidt, and Citi are in alongside existing backers. Its last known valuation: $25 billion pre-money. And compute — the actual currency of this war — is secured on a scale few startups can touch: $150 million a month to SpaceX for Nvidia GB300 chips at the Colossus 2 data center, running from July 2026 through 2029, plus more than $1 billion in capacity from Nebius. Add it up and you’re past $7 billion in committed compute.

This isn’t a research-demo budget. It’s a national-champion budget.

Why Washington is in the room

Ahead of the launch, Reflection has been meeting with parties in Washington to explain the release and something called the “AI factory” — Jensen Huang’s long-pushed vision where companies and governments combine their own data, open models, and their own compute into customized AI systems. In March, Reflection signed a memorandum of understanding with Korea’s Shinsegae Group to build a 250-megawatt AI factory in South Korea.

The politics matter. In July, Nvidia joined a letter urging Washington not to restrict open-weight AI. And the administration has been pressuring Anthropic and OpenAI to restrict their most powerful models — which exposed, in Reuters’ words, “the risks of relying on providers that can be cut off overnight.” When your most capable model can be throttled by policy, open weights stop being an ideology and start being business continuity.

The demand side: why enterprises suddenly care about open weights

Three forces are pushing enterprises toward open weights at the same time:

Cost. The AI bill is the story of the year. Open models, which are typically easier to customize and cheaper to run than closed-weight rivals, have drawn growing interest as enterprises hunt for ways to cut inference spend. My earlier piece on Fireworks’ Ember-1 showed just how hard vendors are working to trim the token bill — and owning the weights is the ultimate cost-control lever.

Control. Data retention concerns are surging. An open-weight model running on your own infrastructure means your data never leaves the building — a requirement in finance, healthcare, and government that no closed API can fully satisfy.

Continuity. The cutoff risk is no longer theoretical. When frontier labs can be pressured to restrict their models, enterprises need weights they own, on servers they control, with no kill switch in someone else’s data center.

The hard part: what GLM-5.3 just taught us

Here’s what Reflection has to get right — and why this release deserves scrutiny, not just celebration. Open weights don’t just remove the vendor’s kill switch. They remove the vendor’s guardrails too.

Late last month, Anthropic published a detailed analysis of GLM-5.3, Z.ai’s open-weight model and the current leader on open-weight capability leaderboards. The findings were stark: using techniques within reach of virtually anyone, Anthropic’s team bypassed the model’s safeguards in simulated tests at rates of 64% (a deceptive prompt posing as a red-team exercise), 92% (prefilling the model’s thinking tokens), and 100% (after “abliteration” — editing the weights to strip out refusal behavior). The whole abliteration job took about 2,200 GPU hours and cost roughly $4,400. After that edit, refusal rates fell from above 90% to about 3%, while general capability scores were unchanged.

A separate, independent assessment by NIST reached a similar conclusion and labeled GLM-5.3 the most cyber-capable open-weight model it had evaluated. Z.ai itself had reported faster-than-expected growth in cyber capability between versions — GLM-5.3 shares GLM-5.2’s base model, yet post-training alone turned it into a far more capable exploit developer.

None of this is an argument against open weights. But it is the reality Reflection is launching into. The company has framed openness as the path to broader safety research and oversight, rather than concentrating decisions inside closed labs. That’s a legitimate position — and a hard one to defend if the first release ships with the same safeguard profile as its Chinese rival. The license, the evals, and the safeguards report will matter as much as the benchmark scores.

What builders should do now

The model isn’t out yet. But the preparation is the same whether it lands this week or next month:

1. Build your eval harness before the weights drop. The gap between “good benchmark” and “good for your workload” is where every open-weight disappointment lives. Have your task-specific tests ready so you can compare Reflection’s release against DeepSeek-V4.1-Flash and Qwen on your own data within hours of the drop.

2. Watch the license, not just the leaderboard. “Open weights” is not “open source.” MIT, Apache, or a custom commercial license changes everything about redistribution, indemnity, and what you can legally build on top. DeepSeek’s permissive MIT licensing of V4.1-Flash is the bar — see what Reflection matches.

3. Plan your serving path. A frontier-scale open model is a GPU commitment. Know whether you’re serving locally, renting inference, or using the AI-factory model Reflection is pitching — the infrastructure decision is bigger than the model decision.

4. Treat version-level risk seriously. Z.ai’s own data carries the lesson: GLM-5.3 shared GLM-5.2’s base but had a completely different security profile after post-training. An approval tied to one version should never auto-carry to the next. If you adopt Reflection’s model, version-pin and re-evaluate every release.

The two-sided war

Step back and the shape of 2026 is clear. The open-weight leaderboards have been a Chinese stronghold — DeepSeek and Qwen leading on coding, math, and cost efficiency. Now the US has a funded champion, with $7 billion in committed compute, Nvidia’s explicit backing, and Washington’s attention, about to fire back.

This is no longer a debate about whether open weights matter. It’s an arms race over who writes the rules for them — the capability leaderboards, the safety norms, the licenses, and the geopolitics. The closed labs will keep building the biggest models. But the models most of the world actually runs — the downloadable, customizable, affordable ones — are where the next chapter of the AI race is being decided. Reflection’s release is the moment that chapter goes two-sided.

The weights aren’t out. The war already is.