Models

Reflection AI Unveils Beam: 501B-Parameter MoE Built for Agentic Coding

Reflection AI's Beam is a 501B MoE with 23B active parameters that the company says matches GLM 5.2 on Terminal Bench using 3–4x less inference compute.

Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads
Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic WorkloadsAI-generated
By Elena Vasquez5 min read

Updated

Why it matters

  • Beam is a 501B-parameter MoE with 23B active parameters per token, built for coding and agentic workloads.
  • Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1, versus GLM 5.2 at 81.0.
  • Pretraining used 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs in under 4 weeks.
  • The RL run generated over 100 million rollouts on 10.5K GB300 GPUs over 4 weeks.
  • Apache 2.0 weights are planned for release later in October 2026; early access runs via waitlist.

Reflection AI has introduced Beam, a sparse Mixture-of-Experts model with 501B total parameters and only 23B active per token, designed for coding, reasoning and agentic workloads. The company claims Beam competes with larger open models like GLM 5.2 while using 3 to 4x less inference compute on reasoning benchmarks — a pitch that matters as enterprises weigh the serving costs of frontier-sized open models against the capability of closed ones.

Beam is not available for self-hosting yet. The model is in final red-teaming, and early access runs through a waitlist on the Reflection platform. Apache 2.0 weights are planned for later in October 2026.

What is Reflection Beam?

Beam is a general agent model trained from scratch by Reflection AI, targeting enterprise coding and agentic workloads. Reflection positions it as advancing the Western open-weight frontier. The team is candid about where it stands: Kimi K3 remains ahead on raw capability, so Beam's argument is efficiency at inference time rather than absolute capability.

Users get a reasoning effort parameter. Lower settings favor short answers; higher settings allow longer reasoning on hard tasks. Teams can tune effort to task difficulty and compute budget.

How was Beam pretrained?

The pretraining corpus totaled 23.8 trillion tokens drawn from the web, public sources and proprietary licensed datasets. Reflection states its curation pipeline removed about 95% of raw internet tokens while retaining roughly 1.8 trillion high-quality tokens that conventional filters would have discarded.

The architecture interleaves local and global attention with fine-grained routed experts. Load balancing builds on the auxiliary-loss-free balancing method from DeepSeek-V3, adding cosine decay of expert-bias updates. The busiest expert reached just 1.04x average load by the end of pretraining — a sign that routing stayed evenly distributed. Across all 52 layers, residual norms stayed bounded through depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation.

The compute budget is concrete: pretraining finished in under 4 weeks on 6,144 NVIDIA GB300 NVL72 GPUs, with goodput reaching 92.3% near the end and 9 semi-automatic rewinds. Midtraining extended effective context to 1M tokens.

What did the reinforcement learning run look like?

RL is Beam's central scaling axis. The run used 10.5K NVIDIA GB300 GPUs for 4 weeks and generated over 100 million rollouts, with a maximum rollout context of 256K tokens. Training and grading consumed about 1.3 billion sandboxes across nearly 1 million coding, agentic and STEM environments.

Reflection trained with fully asynchronous policy gradients: every token is tagged with the policy version that produced it. New algorithms kept learning stable even at one-day staleness — 107 weight versions behind the current policy — and the team reports no plateau as RL compute increased.

The infrastructure numbers stand out on their own:

  • 110K concurrent rollouts sustained on average.
  • New weights reached the inference fleet in a median of about 12 seconds.
  • 71 inference incidents were handled without stopping training.

A controllable length penalty taught Beam to solve tasks with fewer tokens. Browsing skills also improved without browsing tasks in the RL mix, which the team reads as evidence of transfer across agentic domains.

How did Reflection handle safety?

Reflection trained a separate safety and alignment teacher from the pretrained checkpoint, then merged it with the RL teacher using multi-teacher on-policy distillation. Safety training used deliberative alignment. Full safety evaluation results will appear in the upcoming technical report.

How does Beam score?

The scores are Reflection-reported, with rival numbers sourced from Artificial Analysis and DataCurve.

On SWE-bench Verified, Beam scores 80.9 versus 70.7 for Nemotron 3 Ultra. On Terminal Bench v2.1, Beam scores 80.1, close to GLM 5.2 at 81.0 — while DeepSeek V4.1 Flash (90.6) and Kimi K3 (88.3) lead that benchmark.

How does Beam compare with its closest open-weight rivals?

Feature Reflection Beam GLM-5.2 Nemotron 3 Ultra DeepSeek V4.1 Flash Kimi K3
Developer Reflection AI (US) Z.ai (China) NVIDIA (US) DeepSeek (China) Moonshot AI (China)
Total params 501B ~753B 550B 552B backbone + 196B Engram 2.8T
Active params 23B ~40B 55B 8B prefill / 16B decode 104B
Context 1M (effective) 1M Up to 1M 1M 1M
Input Text Text Text Text + image Text + image
License Apache 2.0 (planned) MIT OpenMDW-1.1 MIT Kimi K3 License
Weights Later in Oct 2026 Available Available Available Available
Terminal Bench v2.1 80.1 81.0 56.4 90.6 88.3

Beam is the smallest model in this set by total parameters, and its 23B active count sits below GLM-5.2, Nemotron 3 Ultra and Kimi K3. Apache 2.0 and MIT are standard permissive licenses; Kimi K3's custom license adds attribution requirements for very large products.

Key takeaways

  • Beam is a 501B MoE with 23B active parameters, text-only, with 1M effective context.
  • Pretraining used 23.8T tokens on 6,144 NVIDIA GB300 GPUs in under 4 weeks.
  • RL ran 100M+ rollouts on 10.5K GB300 GPUs over 4 weeks.
  • Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1.
  • Apache 2.0 weights are scheduled for release later this month.

The stakes are straightforward. With Kimi K3 and DeepSeek V4.1 Flash holding the capability lead among open weights, Beam's bet is that asynchronous RL plus a small active-parameter footprint can deliver near-frontier agentic performance at a fraction of the serving cost — a claim the market can test directly once the Apache 2.0 weights land later in October.

Original: platform.reflection.ai

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

207 articles

Related articles

  1. Reflection launches Beam, a 501B-parameter open-weight model it says beats Chinese rivals on cost
  2. OpenAI's Jalapeño Chip Beats Rivals on Power and Latency
  3. Anthropic and OpenAI Ship New Models With the Same Pitch: More for Less
  4. OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
  5. OpenAI Ships AgentKit, Expanded Evals, and Reinforcement Fine-Tuning for Agents

« Previous articleNext article »