Evidence at a glance

parameters501B; token 23BEvidence
23.8 token, Approx. 95% tokenEvidence
use 6,144 GB300, 4Evidence
RL use Approx. 10,500 GB300, 1 rolloutEvidence
Approx. 13 sandbox; RL 256KEvidence
1M; midtraining OutcomeEvidence

The mechanism in one line

InputReduce the input to a workable scale

Compress the visual or contextual input before the main reasoning path.

MechanismSpend compute where it matters

Route or verify the expensive step instead of repeating the full path.

OutcomeEnd with a measurable workflow result

Translate the mechanism into a bounded deployment or evaluation check.

An Open-Weight Announcement Without Downloadable Weights

On October 5, 2026, Reflection AI introduced Beam, its first open-weight model, aimed at enterprise coding, reasoning, and agentic workloads. It is a text model built with a sparse mixture-of-experts architecture: 501 billion parameters in total, with about 23 billion activated for each token. Reflection positions it as a contender in the open-model frontier, while acknowledging that Kimi K3 remains ahead in raw capability.

The gap in this announcement is that “open-weight” describes the intended release, not what users could access that day. Beam was still undergoing final red-team testing, early access required a waitlist, and the Apache 2.0 weights, technical report, model card, and developer tools were all planned for later release. For now, enterprises have a set of performance and efficiency claims to assess, not a model they can already deploy on their own infrastructure.

Model Size and Inference Compute Are Different Measures

Beam’s headline architecture numbers are easy to conflate. The 501B figure is the total parameter count; 23B is the number of parameters activated per token. Sparse MoE routing lets only some experts contribute to each step, making a large model with lower per-token computation plausible. But it does not mean deployment costs are equivalent to those of a 23B model: the expert weights still have to be hosted, and actual memory, throughput, and serving costs depend on implementation and hardware. The available material does not provide end-to-end figures for those costs.

Reflection says Beam comes close to GLM-5.2 on reasoning benchmarks while using three to four times less inference compute. That ratio is not a promise of lower online latency or a smaller bill. The company’s estimate draws on measures such as active parameters and generated tokens, while excluding prompt prefill, attention costs that vary with context length, and serving overhead. For technical leaders, it is best treated as a compute claim to reproduce, not a discount that can be applied directly to a deployment budget.

Training Scale Helps Explain the Agentic Focus

Beam’s efficiency story does not rest on MoE architecture alone. Reflection says pretraining used 23.8 trillion tokens from the web, public sources, and proprietary licensed datasets. Its curation removed about 95% of raw internet tokens while retaining roughly 1.8 trillion high-quality tokens that conventional filters might have discarded. The company also reports that pretraining took less than four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, followed by midtraining that extended the model’s effective context to one million tokens.

The more distinctive investment came during reinforcement learning. Reflection reports using about 10,500 GB300 GPUs for four weeks, generating more than 100 million rollouts, and training and grading across nearly one million coding, agentic, and STEM environments, with roughly 1.3 billion sandboxes. The maximum rollout context was 256K tokens, which is a different training figure from the one-million-token effective context reported after midtraining. Reflection also says a length penalty encouraged Beam to solve tasks with fewer tokens, and browsing improved even though browsing tasks were absent from the RL mix. That suggests possible transfer across agentic tasks, but the available evidence does not establish how broad or reliable that transfer is.

The Benchmarks Support “Close,” Not “Matched Across the Board”

The reported results paint a competitive picture, but not one in which Beam leads or ties everywhere. On AIME 2026, Reflection reports 97.8 for Beam versus 99.2 for GLM-5.2; on HLE without tools, the scores are 36.2 and 40.5; on GPQA Diamond, 90.5 and 91.2. Beam scores 80.1 on Terminal Bench v2.1, close to GLM-5.2 at 81.0, while DeepSeek V4.1 Flash and Kimi K3 score 90.6 and 88.3. On SWE-bench Verified, Beam’s reported 80.9 exceeds the 70.7 listed for Nemotron 3 Ultra.

These figures support the narrower claim that Beam is close to strong competitors on some evaluations. They do not establish parity across the board, or prove that its compute advantage holds across tasks. The scores come from Reflection’s published table, which draws some competitor results from Artificial Analysis and DataCurve; TechCrunch reports that the performance claims have not been independently verified. Comparisons also depend on matching task versions, tool access, and evaluation settings. Without that context, a single score can overstate capability or obscure weaknesses in a particular workflow.

What Enterprises Should Validate Before Betting on Beam

For teams evaluating models for coding or agentic workloads, Beam’s value can only be established against real tasks once the weights are released. The priority should not be a single benchmark score, but success rate, generated tokens, latency, context length, and hardware use on the same workload, with prompt prefill and serving costs included. The reasoning-effort control is also worth testing across task difficulty: it allows a choice between shorter answers and longer reasoning, but the available material does not show a cost-quality curve for its settings.

A deployment decision also depends on information that has not yet been published: whether the weights arrive as planned under Apache 2.0, what the model card and technical report say about data and safety evaluation, and what large-scale inference actually requires in an enterprise environment. Reflection says Beam is still in final red-teaming and that safety results will appear in the technical report, so safety and compliance should not yet be treated as established. The practical course is to keep Beam on the shortlist and define replication criteria, then decide whether it can handle production workloads at lower total cost once the weights, full report, and independent tests are available.