Skip to content

Reflection AI unveils Beam, a 501B-parameter open-weight model tuned for coding and agents

Reflection AI introduced Beam, a 501-billion-parameter sparse mixture-of-experts model it says matches Chinese open models on reasoning benchmarks at far lower inference cost, with open weights arriving later this month.

By VibecodedThis 2 min read
Benchmark chart from Reflection AI's Beam announcement comparing the model's coding and agentic scores against rival open-weight models
Reflection AI

Reflection AI pulled back the curtain on Monday on Beam, its first open-weight model and the centerpiece of the Brooklyn startup's bet that Western labs can compete with China's open-model builders on efficiency rather than raw scale.

Beam is a sparse mixture-of-experts model with 501 billion total parameters, of which only 23 billion are active for any given token. It is text-only, carries a 1 million token context window, and was pretrained on 23.8 trillion tokens drawn from curated web data and licensed datasets, according to the company's launch post.

The efficiency pitch

The efficiency claim is the whole story. Reflection says Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using three to four times less inference compute, and that it beats leading Western open models on coding tasks. In the company's own results table, Beam scores 80.9 on SWEBench Verified against 77.6 for Thinking Machines Lab's Inkling and 70.7 for Nvidia's Nemotron 3 Ultra, and 65.5 on SWE-Bench Pro v1 against 62.1 for GLM-5.2. The same table shows Beam trailing newer Chinese models on several tests: it logs 80.1 on Terminal Bench v2.1 while GLM-5.3 reaches 88.2 and Moonshot's Kimi K3 reaches 88.3. None of these numbers have been independently verified, and the weights are not yet public, so outside researchers cannot reproduce them.

Reflection is upfront that the 3-4x efficiency figure is an estimate. Its methodology multiplies active parameters by generated tokens and leaves out prompt prefill, context-dependent attention operations, and serving overhead. That makes it an approximate comparison of forward-pass compute rather than a measured cost figure, a caveat worth keeping in mind before citing it as a price.

Training at unusual scale

The training scale is striking either way. Pretraining ran on 6,144 Nvidia GB300 NVL72 GPUs and finished in under four weeks at what the company calls 92.3% goodput. The reinforcement learning phase used 10,500 GB300 GPUs over four weeks, generated more than 100 million rollouts, and ran roughly 1.3 billion sandbox executions across nearly a million synthetic coding, agentic, and STEM environments. Reflection says RL is the real source of Beam's efficiency: the model learned to reach correct answers in fewer reasoning steps, and users get a reasoning-effort parameter to trade response length against quality on demanding tasks.

The company is not shy about where this is aimed. It calls Beam a "workhorse model" for enterprises, the public sector, and developers, and frames it as the foundation for "AI factories" that let institutions train its models on proprietary data. Nvidia, which backs Reflection and whose chips trained Beam, has been pushing that same AI-factory narrative. Reflection has already started testing a sovereign AI factory partnership with Shinsegae Group in South Korea, TechCrunch reported, and signed compute deals worth more than $7 billion with SpaceX and Nebius this summer.

What you cannot do yet is run it. Beam is still going through final red-teaming and evaluations, with early access limited to a waitlist. The company says it will release the weights under an Apache 2.0 license, along with a technical report and model card, later this month. Until then, the benchmark story is Reflection's word alone. If the weights land on schedule and the numbers hold up under independent testing, Beam becomes the strongest US-built open-weight coding model to date. That "if" is doing real work.