A Western challenger enters the open-model race.
Nvidia-backed Reflection AI is challenging Chinese open-weight models with a pitch focused on reducing the compute required for enterprise coding and AI agents.
Reflection AI, a Brooklyn-based lab founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, announced Beam on Oct. 5. The 501-billion-parameter system is the company's first open-weight model, engineered specifically to automate coding, handle multi-step reasoning, and power autonomous software agents.
Competing with Chinese open-weight models
Beam enters an ecosystem heavily dominated by Chinese developers. Andreessen Horowitz partner Martin Casado estimates that 80% of developers worldwide use open-source AI tools build with Chinese models. Cloud platform Vercel reported that open-weight architectures processed 56% of tokens routed through its AI Gateway in August 2026, up sharply from just 7% in December 2025, with systems from DeepSeek, Moonshot AI, Alibaba, and Zhipu AI (Z.ai) driving the surge.
That dominance has raised national security alarms in Washington. US Treasury Secretary Scott Bessent warned Congress last month against letting a small group of closed labs dominate the domestic market while foreign open models fill enterprise infrastructure.
Reflection positions Beam as a sovereign alternative for enterprises that require total data ownership. "The only way to own intelligence is, by definition, if it’s open," Laskin said on the Sources Podcast.
Sparse architecture and training scale
Beam operates as a sparse mixture-of-experts model. While containing 501 billion total parameters, it routes tasks through specialized internal networks to activate only 23 billion parameters per token. By comparison, Z.ai's GLM-5.2 activates 40 billion out of 744 billion parameters.
The pretraining phase ingested 23.8 trillion tokens across 6,144 Nvidia GB300 NVL72 GPUs in under four weeks. Reflection then executed an aggressive reinforcement learning campaign using 10,500 Nvidia GB300 GPUs, generating over 100 million rollouts across 1.3 billion evaluation sandboxes.
To sustain development, Reflection has raised $4.7 billion, reaching a $25 billion valuation, backed by an $800 million check from Nvidia. Over the summer, the startup locked in compute agreements worth more than $7 billion with SpaceX and Nebius to access Nvidia GB300 clusters through 2029.
The model is currently undergoing red-teaming and is slated for public release under an Apache 2.0 license later this month.
What eWeek found: Lower compute does not establish lower deployment costs
Rather than establishing outright dominance on standard benchmarks, Reflection's disclosed evaluation figures reveal a strategic tradeoff: trading peak capability ceilings for dramatic inference efficiency.
On heavy agentic benchmarks, Beam trails leading Chinese models. On DeepSWE v1.1, Beam scored 44.4, well behind Alibaba's Qwen3.8-Max at 51.0, Moonshot's Kimi K3 at 68.0, and DeepSeek V4.1 Flash at 74.2. Similarly, on Terminal-Bench 2.1, Beam scored 80.1, compared to 86.6 for Qwen3.8-Max, 88.3 for Kimi K3, and 90.6 for DeepSeek.
However, Beam's architectural moat lies in compute consumption. Reflection reports that Beam achieves reasoning parity with GLM-5.2 while burning three to four times less inference compute. The startup does acknowledge that these metrics approximate generation FLOPs without factoring in prompt prefill or operational serving overhead.
For enterprise teams, Beam could offer another model to evaluate for coding and agent workloads. Buyers should compare cost per successfully completed task, latency, reliability, and hosting requirements before assuming compute efficiency will reduce their bills. Open weights and US development alone do not establish security or affordability.
Read more: For a practical example of balancing data control with deployment costs, explore how Latham & Watkins uses private Nvidia infrastructure alongside cloud AI services.


