A reflection has brought Beam’s first open-weight model into view, a creation bearing 501 billion total parameters, of which 23 billion remain active. This design serves coding, reasoning, and agentic work, and it is built to operate far more efficiently than the rival systems.
The Design Behind Beam
The design behind Beam is a sparse Mixture-of-Experts approach. Its training involved 23.8 trillion tokens drawn from the web and proprietary licensed datasets, a scale that matches or exceeds comparable open base models of similar size. At the same time, Reflection built the algorithms, training settings, and infrastructure required to support high-compute reinforcement learning (RL). Over four weeks of training, a run based on reinforcement learning generated more than 100 million rollouts using 10.5K NVIDIA GB300 GPUs. That work produced competitive open-weight performance alongside frontier-level efficiency for running inference.
| Model | Parameters | Focus |
|---|---|---|
| Beam | 501B total, 23B active | Coding, agentic workloads |
| GLM 5.2 | Not stated | Reasoning |
| Qwen 3.8-Max | Not stated | Agentic tasks |
Final testing and evaluations are still underway for Beam, though early access has already begun. The full weights, technical report, model card, and developer artifacts are all set to be released later this month.
The Efficiency Advantage
On coding and agentic tasks, Beam matches Qwen 3.8-Max, though models like Kimi K3 still hold an advantage in sheer capability. When tested on complex reasoning tasks, Beam performs at the same level as GLM-5.2 while requiring 3 to 4 times less processing power for inference. The savings grow larger when compared to much bigger models, such as those in the 2T+ parameter range, including Qwen 3.8-Max.
The findings mean each token carries more intelligence, which gives the model strong capabilities while keeping costs down.
The Reinforcement Learning Push
RL became the central scaling axis for Beam through reflection. The firm put money into RL science, data, and infrastructure, turning more compute into better capabilities. Longer rollouts help with multi-step reasoning, tool use, and adapting to environment feedback.
Over a span of four weeks, the firm put into service 10.5K NVIDIA GB300 GPUs, which produced well over 100 million rollouts while keeping the context length at no more than 256K tokens. The training and grading process drew upon roughly 1.3 billion sandboxes. One million coding, agentic, and STEM environments of high quality were drawn from across the industry to support the operation.
There was no sign of a ceiling as RL compute rose across the evaluation suite; instead, capabilities kept getting better the whole way through.
Generalization Beyond Training
The RL training for Beam aimed to build reasoning and agentic abilities that could be applied past the specific tasks used to train it. Through a stage focused on reasoning, software engineering, and terminal tasks, browsing skills rose steadily even though browsing itself never showed up in the RL mix. Beam was given permission to browse the web, which allowed it to look up and ask other large language models for information, as well as to call upon OCR APIs so it could read through documents on its own.
A demo was constructed to produce a real-time NYC subway dashboard from public information. A second effort produced interactive applications and prepared model fine-tuning notebooks. Beam itself developed a fine-tuning notebook for the newest and smallest Gemma-4 model on a Text2SQL task.
What This Means for Users
The system combines beam pairs coding with highly efficient reasoning and agentic capabilities. On coding and agentic tasks, its performance matches larger open models like GLM 5.2, and it is closing the gap with Qwen 3.8-Max.
The model’s strength lies in its efficient performance during inference, which is what gives Beam its value as a workhorse for enterprise coding and agentic workloads. Later this month, the company has promised to release its full technical details, with weights, technical report, model card, and developer artifacts all included. Early access is already available.
Beam is a significant step forward for open models.
See the video the story is built around at reflection.ai.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

