A US Lab Finally Shipped Frontier Open Weights. It Measured Them Against China, Not Against OpenAI.
In July we said the only move that would contest the Chinese open-weight default was a strong American open release. On October 5, Reflection AI shipped a 501 billion parameter one and benchmarked it against Chinese open models, the same weekend Aleph Alpha released a 78 billion parameter model it calls sovereign and Korea proposed 4.7 trillion won for a frontier track. Here is why Western open weights are now priced against the Chinese floor rather than the American ceiling, and the one licence line to check before you build on any of them.
On July 27 we wrote that lobbying would not contest the Chinese open-weight default. Only a strong American open release would, and if the next quarter passed without one, the argument was already lost. Ten weeks later, the release arrived. Its own launch post measured it against China.
## What shipped between Friday and Monday
On October 5 Reflection AI unveiled Beam, a 501 billion parameter sparse mixture-of-experts model with open weights. The lab says Beam matches or approaches the top Chinese open models on coding and reasoning while using three to four times less inference compute. Read the comparison set again: not GPT, not Claude, not Gemini. The Chinese open models are the yardstick. Axios, the day before, described the business around it: Nvidia-backed Reflection wants to sell an "AI factory" in which institutions combine its models with their own data on rented Nvidia clusters from partners such as Nebius and SpaceX. The capacity was bought months ago, through a $1 billion Nebius compute deal in July and a $6.3 billion Blackwell pact with SpaceX.
Two days earlier Aleph Alpha released Kolibri-1, a 78 billion parameter mixture-of-experts model tuned for German and English, with reasoning modes and tool calling, full weights on Hugging Face under Apache 2.0. The Frontier's teardown puts the active parameters at 3.46 billion per token and the validated context at up to one million tokens, and reads the positioning as sovereign European deployment. Aleph Alpha's own announcement calls it a sovereign open-weight model.
And on Monday South Korea's science ministry said its sovereign foundation-model contest will run a third round and gain a frontier track, with about 4.7 trillion won of mostly equity-based public investment proposed for 10,000 Nvidia GPUs and training data, pending the National Assembly. Upstage, which took 560 billion won from the Korea fund in May, is the precedent for what that equity looks like.
## The floor moved, so the comparison moved
For two years a Western model launch compared itself to the closed frontier. Beam compares itself to the Chinese open models because that is who it competes with for the deployment slot: the developer who needs weights they can run themselves. Closed labs still set the ceiling. Chinese open models set the floor. A Western open model is now priced against the floor.
The Chinese labs know which way the comparison runs. Alibaba's Qwen-Image-2.1 shipped in September with a benchmark that puts it narrowly ahead of several closed models, under a research-only licence. The week before, Anthropic's threat dossier accused Alibaba, Moonshot AI and DeepSeek of distilling Claude, one operation allegedly routing about 151 million API interactions through proxy "transfer stations". Each side measures against the other: Chinese open models against the Western ceiling, Western open models against the Chinese floor.
## Sovereign and open are becoming one product
In August we argued the sovereign buying question had changed from "can you match the frontier" to "can you run on premise, in our language, under our law." Kolibri-1 is that answer written as a model card: German, Apache 2.0, weights you can hold. Korea's frontier track is the same answer written as a budget: a GPU count and a data line, equity rather than grants.
The closed labs are answering the same procurement question without releasing anything. On October 5 Anthropic turned on in-country Claude inference on Amazon Bedrock's India endpoints, so inference runs in Mumbai and Hyderabad for customers with data residency rules, delivering an August pledge. Residency is the closed lab's substitute for weights. Whether a regulator accepts the substitute is the open question for every sovereign tender in 2027.
## Hedges
Beam's benchmark claims are Reflection's own, "matches or approaches" is doing a lot of work, and independent evaluations are not in our records yet. Kolibri's one-million-token figure comes from one published analysis. Korea's 4.7 trillion won is a proposal to parliament. And licences diverge: Kolibri-1 is Apache 2.0, Qwen-Image-2.1 is research-only, Kimi K3 shipped under its own licence rather than a standard open one, and our record of Beam does not state its licence terms.
## What to do with this
If you evaluate open models, run Beam and Kolibri-1 against your current Chinese default on your own workload this month, and report the result as cost per completed task, the unit Reflection chose when it claimed three to four times less inference compute. That claim is the one number in this week's launches worth checking yourself.
Before any of them goes into production, read the licence line, because that is where open models now differ most. The July test is settled: an American lab did release frontier-scale open weights. The next test is whether a sovereign tender accepts a residency endpoint in place of them. Reflection AI, Aleph Alpha, Anthropic and Upstage are on our tracker.