TL;DR
Reflection AI has launched Beam, a 501B-parameter open-weight model claiming 3 to 4x better inference efficiency than China's best open models, but the weights are not yet public and every benchmark is self-reported.
What happened
- Beam announced October 5 by Brooklyn startup Reflection AI, ending months of missed timelines and vague guidance.
- 501 billion total parameters, 23 billion active via sparse mixture-of-experts architecture, trained on 23.8 trillion tokens.
- 6,144 Nvidia GB300 NVL72 GPUs ran the pretraining in under four weeks at 92.3% goodput; reinforcement learning used another 10,500 GB300 chips across 100 million rollouts.
- Benchmark claims: 80.9 on SWE-Bench Verified, 80.1 on Terminal-Bench v2.1, 97.8 on AIME 2026, matching Z.ai's GLM-5.2 at 3 to 4x lower inference compute.
- Full weights promised under Apache 2.0 license later in October; current access is waitlisted and red-teaming is still in progress.
Why it matters
- Efficiency is the actual battleground: DeepSeek and Qwen have proven Chinese labs can hit frontier benchmark scores; the fight is now over cost-per-token at scale, and Beam's 3 to 4x compute advantage is the entire strategic pitch.
- Nvidia's $800M stake in Reflection is not passive cheerleading: Beam is a live proof-of-concept for the "AI factory" model Nvidia sells to hyperscalers and sovereign cloud customers.
- Governments and enterprises are the real target, not developers: Beam is already validated on Dell's AI Factory platform, certified in the Department of Energy's Genesis Mission, and anchors a 250-megawatt sovereign AI cloud in South Korea with Shinsegae.
- $4.7 billion raised at a $25 billion valuation (backers include Nvidia, Sequoia, Lightspeed) with over $7 billion in SpaceX and Nebius compute committed, making Beam the most expensive open-model bet in American AI history.
- If the efficiency claim holds, Beam is the first credible American open-weight answer to the Chinese models governments have been defaulting to for on-premise deployments.
What to watch next
- Independent evals the moment full weights drop: self-reported numbers from a company that just raised $2 billion to out-build DeepSeek demand outside verification before the efficiency claim is treated as fact.
- Reflection's October deadline: the company has already missed one self-imposed timeline this year; a second slip would deepen the "well-funded but unshipped" narrative.
- Sovereign cloud traction in South Korea and DOE: whether enterprise and government customers actually standardize on Beam is the commercial signal that separates a benchmark win from a business.
Originally published on Present of AI, a daily source-linked AI news timeline. Read the full timeline or browse the open dataset.