AI News Feed
Market watch
Large Language Models

Reflection AI Introduces Beam, a 501B Open-Weight MoE Model With 23B Active Parameters

Reflection AI introduced Beam, a 501B open-weight MoE with 23B active parameters for coding and agentic workloads.

Reflection said Beam competes with larger open models such as GLM 5.2 while using three to four times less inference compute on reasoning benchmarks. The company positions Beam as an advance for the Western open-weight frontier, but its research team acknowledged a gap: Kimi K3 remains ahead on raw capability, so Beam's pitch is efficiency at inference time. Beam includes a reasoning effort parameter. Lower settings favor short answers, while higher settings allow longer reasoning on hard tasks, letting teams match effort to task difficulty and compute budget.

Beam was pretrained on 23.8 trillion tokens from the web, public sources and proprietary licensed datasets. Reflection said its curation removed about 95% of raw internet tokens and kept roughly 1.8 trillion high-quality tokens that conventional filters would have dropped. The architecture interleaves local and global attention with fine-grained routed experts. Load balancing builds on auxiliary-loss-free balancing from DeepSeek-V3 and adds cosine decay of expert-bias updates. The busiest expert reached just 1.04 times average load at the end of pretraining, and residual norms stayed bounded across all 52 layers through depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation. Pretraining finished in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, with goodput reaching 92.3% near the end and nine semi-automatic rewinds. Midtraining extended effective context to 1 million tokens.

Reinforcement learning was Beam's central scaling axis. The run used 10,500 NVIDIA GB300 GPUs for four weeks and generated more than 100 million rollouts, with a maximum rollout context of 256,000 tokens. Training and grading consumed about 1.3 billion sandboxes across nearly 1 million coding, agentic and STEM environments. Reflection trained with fully asynchronous policy gradients, tagging every token with the policy version that produced it. The company said new algorithms kept learning stable even at one-day staleness, 107 weight versions behind the current policy, and reported no plateau as RL compute increased.

The system sustained 110,000 concurrent rollouts on average, according to Reflection. New weights reached the inference fleet in a median of about 12 seconds, and 71 inference incidents were handled without stopping training. A controllable length penalty taught Beam to solve tasks with fewer tokens. Browsing skills also improved without browsing tasks in the RL mix, which the company said suggests transfer across agentic domains.

Reflection trained a separate safety and alignment teacher from the pretrained checkpoint. It merged that teacher with the RL teacher using multi-teacher on-policy distillation. Safety training used deliberative alignment. Reflection said safety evaluation results will appear in the technical report.

On SWE-bench Verified, Reflection reported Beam at 80.9, versus 70.7 for Nemotron 3 Ultra. On Terminal Bench v2.1, Beam scored 80.1, close to GLM 5.2 at 81.0. DeepSeek V4.1 Flash scored 90.6 and Kimi K3 scored 88.3 on that benchmark, according to the table in Reflection's announcement, which sources rival scores from Artificial Analysis and DataCurve.

The table lists Beam's total parameters at 501 billion and active parameters at 23 billion. It compares the model with GLM 5.2 at about 753 billion total parameters and about 40 billion active; Nemotron 3 Ultra at 550 billion total and 55 billion active; DeepSeek V4.1 Flash at 552 billion backbone plus 196 billion Engram, with 8 billion active for prefill and 16 billion for decode; and Kimi K3 at 2.8 trillion total and 104 billion active. Beam's effective context is 1 million tokens. It is text-only, while DeepSeek V4.1 Flash and Kimi K3 accept text and images. Reflection plans an Apache 2.0 license, compared with MIT for GLM 5.2 and DeepSeek V4.1 Flash, OpenMDW-1.1 for Nemotron 3 Ultra and a custom Kimi K3 license that adds attribution requirements for very large products. Reflection said weights are planned for later in October 2026.