Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Reflection Unveils Beam, a 501B Open-Weight Model Built for Coding and Reasoning

Reflection's Beam packs 501B parameters with 23B active, offering coding, reasoning, and agentic prowess at a fraction of rivals' compute.

By mitch·3 min read
A glowing digital beam illuminates a row of server racks in a dim data center.

A reflection has brought Beam’s first open-weight model into view, a creation bearing 501 billion total parameters, of which 23 billion remain active. This design serves coding, reasoning, and agentic work, and it is built to operate far more efficiently than the rival systems.

The Design Behind Beam

The design behind Beam is a sparse Mixture-of-Experts approach. Its training involved 23.8 trillion tokens drawn from the web and proprietary licensed datasets, a scale that matches or exceeds comparable open base models of similar size. At the same time, Reflection built the algorithms, training settings, and infrastructure required to support high-compute reinforcement learning (RL). Over four weeks of training, a run based on reinforcement learning generated more than 100 million rollouts using 10.5K NVIDIA GB300 GPUs. That work produced competitive open-weight performance alongside frontier-level efficiency for running inference.

Model Parameters Focus
Beam 501B total, 23B active Coding, agentic workloads
GLM 5.2 Not stated Reasoning
Qwen 3.8-Max Not stated Agentic tasks

Final testing and evaluations are still underway for Beam, though early access has already begun. The full weights, technical report, model card, and developer artifacts are all set to be released later this month.

Advertisement

The Efficiency Advantage

On coding and agentic tasks, Beam matches Qwen 3.8-Max, though models like Kimi K3 still hold an advantage in sheer capability. When tested on complex reasoning tasks, Beam performs at the same level as GLM-5.2 while requiring 3 to 4 times less processing power for inference. The savings grow larger when compared to much bigger models, such as those in the 2T+ parameter range, including Qwen 3.8-Max.

The findings mean each token carries more intelligence, which gives the model strong capabilities while keeping costs down.

The Reinforcement Learning Push

RL became the central scaling axis for Beam through reflection. The firm put money into RL science, data, and infrastructure, turning more compute into better capabilities. Longer rollouts help with multi-step reasoning, tool use, and adapting to environment feedback.

Over a span of four weeks, the firm put into service 10.5K NVIDIA GB300 GPUs, which produced well over 100 million rollouts while keeping the context length at no more than 256K tokens. The training and grading process drew upon roughly 1.3 billion sandboxes. One million coding, agentic, and STEM environments of high quality were drawn from across the industry to support the operation.

There was no sign of a ceiling as RL compute rose across the evaluation suite; instead, capabilities kept getting better the whole way through.

Generalization Beyond Training

The RL training for Beam aimed to build reasoning and agentic abilities that could be applied past the specific tasks used to train it. Through a stage focused on reasoning, software engineering, and terminal tasks, browsing skills rose steadily even though browsing itself never showed up in the RL mix. Beam was given permission to browse the web, which allowed it to look up and ask other large language models for information, as well as to call upon OCR APIs so it could read through documents on its own.

A demo was constructed to produce a real-time NYC subway dashboard from public information. A second effort produced interactive applications and prepared model fine-tuning notebooks. Beam itself developed a fine-tuning notebook for the newest and smallest Gemma-4 model on a Text2SQL task.

What This Means for Users

The system combines beam pairs coding with highly efficient reasoning and agentic capabilities. On coding and agentic tasks, its performance matches larger open models like GLM 5.2, and it is closing the gap with Qwen 3.8-Max.

The model’s strength lies in its efficient performance during inference, which is what gives Beam its value as a workhorse for enterprise coding and agentic workloads. Later this month, the company has promised to release its full technical details, with weights, technical report, model card, and developer artifacts all included. Early access is already available.

Beam is a significant step forward for open models.

See the video the story is built around at reflection.ai.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.