Evidence at a glance

parameters501B, token 23BEvidence
use 23.8T tokensEvidence
OCR PDFEvidence
4 , use 6144 GB300Evidence
RL 4 , use Approx. 10500 GB300Evidence
RL 1 rolloutsEvidence

A New Model Changes the Supply Picture First

On October 5, 2026, Reflection AI announced Beam, which it describes as its first open-weight model, built for coding, reasoning, and agent tasks. Trained from scratch, Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token. Reflection promised to release the weights under Apache 2.0, along with a technical report and model card, later in the month. None of those materials had been released at the time of the announcement.

The launch is worth examining not because Beam has already overtaken leading models, but because a US team that had stayed largely out of public view has now put forward a model and training program that others can assess. Reflection says Beam is competitive with GLM 5.2, while the supplied coverage says GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead. Beam first expands the set of US-trained open-model options. Its position on capability remains an open question.

Sparse Activation Does Not Make It a Small Model

Beam activates 23 billion parameters per token, meaning each computation uses only part of the model’s experts rather than all 501 billion parameters. Sparse activation can help limit computation, but the total parameter count still matters for storing and deploying the weights. The available material does not specify the hardware needed to run Beam, so its 23 billion active parameters should not be read as proof that it can be deployed like a 23-billion-parameter model.

Reflection also says Beam interleaves global and sliding-window attention. In broad terms, local windows focus computation on nearby content, while global attention provides a path for relating information across longer spans. This is a tradeoff between computation and context handling, not simply a bid to maximize the context-length headline. The announcement says midtraining extended the “effective context length” to one million tokens, while the maximum context length in RL training was 256,000. The public materials do not explain how those figures relate, so they do not establish a one-million-token context window for ordinary use.

Behind the Training Scale Is a Task Factory

Reflection’s published figures put Beam’s pretraining at 23.8 trillion tokens, completed in under four weeks using 6,144 GB300 GPUs. The data came from web sources, public materials, and licensed sources. The team also says it used OCR to process hundreds of millions of PDFs and applied filtering to STEM literature. The point of that pipeline is not just to add more text, but to bring scientific material into the corpus used for later training.

The more distinctive effort came in RL. Reflection says the phase ran for four weeks on about 10,500 GB300 GPUs, producing more than 100 million rollouts across roughly one million training environments. Code, terminal, science, search, and tool-use environments let the model learn from task outcomes and feedback rather than static text alone. The company describes an asynchronous system that generates trajectories while updating the model, tagging each trajectory with the model version that produced it to manage the lag between data generation and policy updates. Reflection says the system remains stable when using interaction data more than a day old, but that claim has not yet received public independent verification.

Benchmark Scores and Efficiency Claims Have Limits

The announcement’s scores form a useful evidence card for what Reflection wants to demonstrate: 80.9 on SWE-bench Verified, 80.1 on Terminal Bench 2.1, and 44.4 on DeepSWE v1.1. These are company-reported results. The technical report and full evaluation settings were not yet public, so the scores show the team’s targets and initial claims, but cannot replace independent testing with a consistent harness. Results for coding agents can depend on inference settings, toolchains, and test procedures. A single score is not enough to predict performance in an engineering workflow.

The efficiency claim also depends on its measurement method. Reflection says that on some reasoning benchmarks, Beam used three to four times less inference computation than GLM 5.2 at comparable performance. That figure estimates forward-pass computation from generated token counts and excludes prompt prefill, context-dependent attention costs, and serving overhead. It supports a comparison under a particular estimate, not an equivalent reduction in latency, hardware bills, or service prices. External assessments also differ: some expect strong token efficiency, while others place Beam behind newer leading models.

For Engineering Teams, Treat It as an Option to Test

The part of Beam most worth examining may not be its 501-billion-parameter headline, but the engineering system that combines large numbers of environments and rollouts with asynchronous RL. For teams building coding or tool-using agents, it raises a concrete question: when feedback can be programmatically scored and environments can run at scale, can training throughput and fresher data improve task performance more directly than simply increasing pretraining? Reflection’s disclosures make that question tangible, but do not yet establish the return on investment of its approach.

A practical evaluation should wait for the weights and license text, then test Beam against the team’s own codebases, tools, and latency constraints. Track more than benchmark scores: measure the generated tokens, retries, hardware resources, and context settings needed to meet a target level of performance. If a team needs a US-trained model and intends to inspect or deploy open weights, Beam may merit a place on its shortlist. If the decision depends on Beam being a capability leader or three to four times cheaper, the available evidence is not enough. Open-weight delivery, independently reproducible results, and real deployment costs are three separate acceptance criteria.