Reflection AI Unveils Beam: 501B-Parameter MoE Built for Agentic Coding
Reflection AI's Beam is a 501B MoE with 23B active parameters that the company says matches GLM 5.2 on Terminal Bench using 3–4x less inference compute.

Updated
Why it matters
- Beam is a 501B-parameter MoE with 23B active parameters per token, built for coding and agentic workloads.
- Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1, versus GLM 5.2 at 81.0.
- Pretraining used 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs in under 4 weeks.
- The RL run generated over 100 million rollouts on 10.5K GB300 GPUs over 4 weeks.
- Apache 2.0 weights are planned for release later in October 2026; early access runs via waitlist.
Reflection AI has introduced Beam, a sparse Mixture-of-Experts model with 501B total parameters and only 23B active per token, designed for coding, reasoning and agentic workloads. The company claims Beam competes with larger open models like GLM 5.2 while using 3 to 4x less inference compute on reasoning benchmarks — a pitch that matters as enterprises weigh the serving costs of frontier-sized open models against the capability of closed ones.
Beam is not available for self-hosting yet. The model is in final red-teaming, and early access runs through a waitlist on the Reflection platform. Apache 2.0 weights are planned for later in October 2026.
What is Reflection Beam?
Beam is a general agent model trained from scratch by Reflection AI, targeting enterprise coding and agentic workloads. Reflection positions it as advancing the Western open-weight frontier. The team is candid about where it stands: Kimi K3 remains ahead on raw capability, so Beam's argument is efficiency at inference time rather than absolute capability.
Users get a reasoning effort parameter. Lower settings favor short answers; higher settings allow longer reasoning on hard tasks. Teams can tune effort to task difficulty and compute budget.
How was Beam pretrained?
The pretraining corpus totaled 23.8 trillion tokens drawn from the web, public sources and proprietary licensed datasets. Reflection states its curation pipeline removed about 95% of raw internet tokens while retaining roughly 1.8 trillion high-quality tokens that conventional filters would have discarded.
The architecture interleaves local and global attention with fine-grained routed experts. Load balancing builds on the auxiliary-loss-free balancing method from DeepSeek-V3, adding cosine decay of expert-bias updates. The busiest expert reached just 1.04x average load by the end of pretraining — a sign that routing stayed evenly distributed. Across all 52 layers, residual norms stayed bounded through depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation.
The compute budget is concrete: pretraining finished in under 4 weeks on 6,144 NVIDIA GB300 NVL72 GPUs, with goodput reaching 92.3% near the end and 9 semi-automatic rewinds. Midtraining extended effective context to 1M tokens.
What did the reinforcement learning run look like?
RL is Beam's central scaling axis. The run used 10.5K NVIDIA GB300 GPUs for 4 weeks and generated over 100 million rollouts, with a maximum rollout context of 256K tokens. Training and grading consumed about 1.3 billion sandboxes across nearly 1 million coding, agentic and STEM environments.
Reflection trained with fully asynchronous policy gradients: every token is tagged with the policy version that produced it. New algorithms kept learning stable even at one-day staleness — 107 weight versions behind the current policy — and the team reports no plateau as RL compute increased.
The infrastructure numbers stand out on their own:
- 110K concurrent rollouts sustained on average.
- New weights reached the inference fleet in a median of about 12 seconds.
- 71 inference incidents were handled without stopping training.
A controllable length penalty taught Beam to solve tasks with fewer tokens. Browsing skills also improved without browsing tasks in the RL mix, which the team reads as evidence of transfer across agentic domains.
How did Reflection handle safety?
Reflection trained a separate safety and alignment teacher from the pretrained checkpoint, then merged it with the RL teacher using multi-teacher on-policy distillation. Safety training used deliberative alignment. Full safety evaluation results will appear in the upcoming technical report.
How does Beam score?
The scores are Reflection-reported, with rival numbers sourced from Artificial Analysis and DataCurve.
On SWE-bench Verified, Beam scores 80.9 versus 70.7 for Nemotron 3 Ultra. On Terminal Bench v2.1, Beam scores 80.1, close to GLM 5.2 at 81.0 — while DeepSeek V4.1 Flash (90.6) and Kimi K3 (88.3) lead that benchmark.
How does Beam compare with its closest open-weight rivals?
| Feature | Reflection Beam | GLM-5.2 | Nemotron 3 Ultra | DeepSeek V4.1 Flash | Kimi K3 |
|---|---|---|---|---|---|
| Developer | Reflection AI (US) | Z.ai (China) | NVIDIA (US) | DeepSeek (China) | Moonshot AI (China) |
| Total params | 501B | ~753B | 550B | 552B backbone + 196B Engram | 2.8T |
| Active params | 23B | ~40B | 55B | 8B prefill / 16B decode | 104B |
| Context | 1M (effective) | 1M | Up to 1M | 1M | 1M |
| Input | Text | Text | Text | Text + image | Text + image |
| License | Apache 2.0 (planned) | MIT | OpenMDW-1.1 | MIT | Kimi K3 License |
| Weights | Later in Oct 2026 | Available | Available | Available | Available |
| Terminal Bench v2.1 | 80.1 | 81.0 | 56.4 | 90.6 | 88.3 |
Beam is the smallest model in this set by total parameters, and its 23B active count sits below GLM-5.2, Nemotron 3 Ultra and Kimi K3. Apache 2.0 and MIT are standard permissive licenses; Kimi K3's custom license adds attribution requirements for very large products.
Key takeaways
- Beam is a 501B MoE with 23B active parameters, text-only, with 1M effective context.
- Pretraining used 23.8T tokens on 6,144 NVIDIA GB300 GPUs in under 4 weeks.
- RL ran 100M+ rollouts on 10.5K GB300 GPUs over 4 weeks.
- Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1.
- Apache 2.0 weights are scheduled for release later this month.
The stakes are straightforward. With Kimi K3 and DeepSeek V4.1 Flash holding the capability lead among open weights, Beam's bet is that asynchronous RL plus a small active-parameter footprint can deliver near-frontier agentic performance at a fraction of the serving cost — a claim the market can test directly once the Apache 2.0 weights land later in October.
Original: platform.reflection.ai
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
207 articles
Related articles
- Reflection launches Beam, a 501B-parameter open-weight model it says beats Chinese rivals on cost
- OpenAI's Jalapeño Chip Beats Rivals on Power and Latency
- Anthropic and OpenAI Ship New Models With the Same Pitch: More for Less
- OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
- OpenAI Ships AgentKit, Expanded Evals, and Reinforcement Fine-Tuning for Agents