Beam: Reflection's 501B open-weight model

Hacker News by 17 min read 51x views
Beam: Reflection's 501B open-weight model

Share Post

We are introducing Beam, Reflection’s archetypal open-weight model. Beam is a sparse Mixture-of-Experts example alongside 501 milliard total parameters, 23 milliard active, built for coding, reasoning, and agentic workloads.

Beam’s capabilities arrive from important investments in the two pretraining and reinforcement learning (RL). We pretrained the example on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming accessible similar-sized open basis models. In parallel, we developed the algorithms, training environments, and infrastructure needed to prolong high-compute RL at exceptional scale. Our high-compute RL run generated complete 100 myriad rollouts on 10.5K NVIDIA GB300 GPUs complete 4 weeks of training.

Together, these efforts produced rivalrous open-weight achievement alongside frontier conclusion compute efficiency.

Beam is undergoing final red-teaming and evaluations. You can sign up here for first admission to the model. We volition publish the weights, specialized report, example card, and developer artifacts afterward this month.

Model Capability

We trained Beam alongside a particular concentration on coding and agentic performance. Beam advances the Western open-weight frontier and is rivalrous alongside larger open models akin GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Where frontier open models akin Kimi K3 remain onward on raw capability, Beam's advantage is effectiveness at conclusion time.

The below fig shows Beam's achievement throughout a range of coding, agentic, reasoning, and STEM benchmarks. NR denotes scores that have not been reported.

Beam pairs coding and agentic capabilities alongside extremely productive reasoning. On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 during using 3–4× small conclusion compute. Efficiency gains are equal additional pronounced whenever comparing to models in the 2T+ indicator family akin Qwen 3.8-Max, which necessitate considerably additional conclusion compute per token.

These results translate into additional intellect per token, delivering powerful example capabilities at lesser cost, making Beam a mighty workhorse example for endeavor coding and agentic workloads.

Figure 2: Beam demonstrates frontier-level conclusion efficiency, the two whenever measured in conditions of FLOPS and token figure throughout DeepSWE, Humanity’s Last Exam (HLE), and Terminal Bench 2.1. We used data from Artificial Analysis and DataCurve, estimating generation forward-pass compute as FLOPs ≈ 2 × energetic indicator figure × average generated tokens per attempt, counting all multiply-add as two operations. Generated tokens contain the two reasoning and the final answer. For mixture-of-experts models, we used the parameters activated per token fairly than the total example size. These estimates exclude immediate prefill, context-dependent notice operations, and serving overhead, so they portray an approximate compute difference fairly than measured conclusion cost. We use Artificial Analysis and DataCurve as sources for another model’s evals.

High-Compute Reinforcement Learning

We made high-compute reinforcement learning a chief scaling axis for Beam, investing in RL science, data, and infrastructure to rotate additional compute into stronger capabilities. Scaling RL enables additional extended exploration of problem-solving strategies, during longer rollouts assistance multi-step reasoning, tool use, and adaptation to surroundings feedback.

To measure reinforcement learning, we deployed 10.5K NVIDIA GB300 GPUs for four weeks generating additional than 100 myriad rollouts alongside a maximum environment dimension of 256K tokens. Training and grading used about 1.3 milliard sandboxes. To prolong a run of this magnitude, we originated one myriad high-quality coding, agentic, and STEM environments. We accept this is among the largest measure RL runs conducted by any open lab to date. Across our evaluation suite, capabilities continued to enhance as we risen RL compute, alongside no sign of a plateau.

Figure 3: Terminal-Bench 2.1, HLE, and DeepSWE scores as a function of cumulative RL rollouts during training of Beam’s reasoning expert, accounting for 80M of the complete 100M rollouts generated throughout the complete RL campaign. For comparison, Inkling was trained on 30M rollouts and MiMo on 753K.

We trained Beam alongside asynchronous guideline gradients. At scale, guideline staleness becomes a important origin of instability for these methods. Long operating rollouts have tokens that are generated by multiple example checkpoints, alongside before tokens becoming increasingly old related to the current policy. Numerical mismatch between training and conclusion engines additional compounds this challenge.

We developed new algorithms to keep stable learning under these conditions during systematically reducing training–inference mismatch throughout our pipeline. These advances allow completely asynchronous RL at measure that remains stable, equal whenever learning from interactions generated additional than a day earlier.

Figure 4: Stable learning continues as staleness builds up complete time. The top storyline shows the oldest example in the batch, alongside the base storyline demonstrating stable numerics. Even whenever training Beam alongside one-day staleness—107 importance versions rearward the current policy—the numerics remain stable.

Learning to logic efficiently

We trained Beam alongside a controllable dimension penalty that rewards prosperous solutions during discouraging unnecessary tokens. Early in RL, achievement improved equal as completion lengths fell: the example learned to resolve tasks additional efficiently alongside small reasoning. Later, as Beam developed stronger agentic capabilities, completion lengths grew again, but those additional tokens supported additional gains in performance. Throughout training, RL improved the tradeoff between capability and token usage.

Figure 5: DeepSWE scores during part of the RL run. Each item corresponds to a distinct reasoning effort. The Pareto frontier moves in two phases. First, it contracts as the guideline learns to be additional token efficient. Then, the higher reasoning efforts develop outwards to accomplish elevated performance.

Users can authority this tradeoff through Beam’s reasoning attempt parameter: lesser settings favor shorter responses, during higher settings authorize longer reasoning to enhance achievement on demanding tasks. This gives users the elasticity to equivalent reasoning attempt to their project and compute budget.

How behavior generalizes alongside RL

We designed Beam’s RL training to create reasoning and agentic capabilities that generalize beyond its training tasks. During a phase of training on reasoning, application engineering, and terminal tasks, we saw accordant gains in browsing notwithstanding the deficiency of browsing tasks from the RL mixture. This transfer suggests that Beam was learning broader agentic capabilities that generalize throughout domains. When stated web access, it organically learned to hunt for and query another ample tongue models, and to use OCR APIs to peruse documents.

The demos below showcase Beam applying these capabilities throughout research, use development, gameplay, and device learning workflows. The examples range from construction a live NYC subway dashboard using community data to creating interactive applications and preparing example fine-tuning notebooks. Although Beam is text-only, it can activity alongside data from another modalities whenever represented as text. In another out of allocation domain, Beam additionally created a fine-tuning notebook for the latest and smallest Gemma-4 example on a Text2SQL task.

Together, these demos exemplify the breadth of tasks Beam can tackle by combining reasoning, coding, and tool use. Each example includes the first petition and resulting output, alongside alongside applicable setup and person iterations.

Scaling reinforcement learning environments

Frontier-scale reinforcement learning requires a ample quantity of difficult, high-quality tasks. We built a pond of nearly one myriad environments, chiefly through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.

We relied heavily on an iterative curation process. First, we synthesized or originated environments throughout a broad set of domains including application engineering, terminal use, rivalrous coding, STEM, web search, tool use, and broad cognition work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor unattainable for the model) and norm (e.g., not underspecified, misleading, guessable, hackable, or alternatively damaged or noisy). Third, we tested the tasks through RL, which allowed us to acknowledge additional norm or difficulty issues and communicate the next repeat of sourcing and filtering.

Throughout Beam’s development, we established that compromises in data norm led to capability plateaus and another training issues. Systematic improvements to project norm were essential to sustaining capability gains throughout the run, which ended alongside no sign of saturation.

Frontier RL infrastructure

High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the example at scale. We built an asynchronous phase that lets these processes run independently during coordinating the stream of cognition and example updates.

During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:

Fully asynchronous execution: Agents create rollouts during the trainer learns and publishes new example versions. Each token is tagged alongside the type that produced it, allowing the training algorithm to document for guideline staleness as completed rollouts stream into training.

Flexible compute allocation: We adjusted the balance between conclusion and training as the workload evolved, functioning at inference-to-training GPU ratios from 3.9:1 to 5.4:1. We additionally resized the trainer throughout five GPU mesh configurations inside the identical training lineage without losing training state.

Fast example updates: New weights reached the conclusion fleet in a median of about 12 seconds. Hierarchical allocation transfers weights throughout racks complete RoCE, afterward shares them locally complete NVLink. Compared alongside all replica pulling weights directly, this reduced cross-rack traffic by 75% and made fleet-wide acceptance of new weights 2.2× faster.

Resilience to conclusion failures: During the run, 71 conclusion incidents were handled without terminating the training job. Inference capability restored in a median of eight minutes, alongside misplaced capability accounting for fair 0.02% of elapsed serving GPU-minutes.

Environments at scale: We supported up to 170K concurrent sandboxes during the run. Across the platform, we processed additional than one milliard sandbox innovation requests, spanning complete 20 clusters, two clouds, and four regions. 90% of new sandboxes were prepared in under 10 seconds.

Efficient Trainer Packing: Dynamic packing kept training batches 99.99% complete on average, holding per-GPU trainer throughput inside 1.5% as average rollout dimension grew nearly 70%.

Observability and reward integrity: Per-token records enabled numerical consistency checks between training and conclusion at all step. Independent judges re-screened passing solutions for verifier exploits, during replayable records made rewards and their use in training inspectable.

Together, these capabilities enabled us to train on longer interactions and additional demanding environments during maintaining throughput, recovering from failures, and checking the integrity of the learning process.

Pretraining a Foundation for Reasoning

Reinforcement learning builds on top of a sturdy basis model. To facilitate reasoning, we ensured Beam’s basis had affluent cognition in coding domains, innate agentic capabilities that could be amplified, and stable MoE optimization dynamics.

While evolving Beam, we pretrained a sequence of iteratively bigger models to established and verify our scaling recipe. Making their achievement predictable required carefully designing example tiers, curating varied in-house code and web validation sets, and aggressively decontaminating all training data against them. The scaling held; the final Beam Base matches its predicted performance, and additionally matches or outperforms accessible similar-sized open-source basis models.

Figure 6: Beam’s pretraining formula scales predictably throughout four orders of dimension in compute. It achieves Pareto-optimal defeat – compute achievement on decontaminated code and web validation data among all comparable open basis models.

Stable and stable MoE optimization

The Beam architecture and optimization formula emphasizes a numerically fit basis for sustained downstream RL. It combines interleaved local and earth attention, fine-grained routed experts, a controlled residual stream, and multiple forms of burden balancing to justify stable, stable expert use and fit indication propagation through residuals.

For expert utilization, we built on auxiliary-loss-free burden balancing (DeepSeek-AI et al., 2024), introducing cosine decay of expert-bias updates to decrease routing perturbations afterward in training. Sequence-level balancing additional encourages stable expert use on data exterior the pretraining distribution, preparing the example for the changing allocation of downstream RL. As a result, the final pretrained basis has almost-perfect uniform utilization, ensuring all experts can be used for learned reasoning.

Figure 7: Beam’s expert use is near-uniform. The busiest expert’s load, averaged throughout MoE layers, reaches fair 1.04× at pretraining completion.

For the residual stream, we developed a depth-based scaling method that counteracts activation growth as sublayer outputs accumulate, assisting keep residual norms stable as example degree increases. Combined alongside SandwichNorm, elementwise notice gating, and FP32 residual gathering – which reduces rounding error whenever adding small updates to the stream – this formula controls activation growth and outliers, thereby ensuring fit indication propagation throughout all layers of Beam. This stability persists throughout pretraining, reinforcement learning, and alignment.

Figure 8: Residual stream RMS remains bounded throughout pretraining throughout all 52 layers of Beam, alongside smooth, depth-dependent trajectories and no sustained activation growth.

Quality-centric data curation

Beam was pretrained on 23.8 trillion varied high-quality tokens from the web, community sources, and proprietary licensed datasets. Our data pipeline was designed to provision Beam a basis for downstream agentic coding: origin code, specialized explanations, and mathematical and specialized knowledge, preserved through all phase of curation. We train on nearly all publically accessible and unrestrictively-licensed code and code records on the web.

We trained our own norm classifiers for web, code, and STEM content, divided data into fine-grained norm tiers, and weighed training toward stronger material. After extended specialized iteration, we optimized the two precision and recall of data curation considerably beyond accepted web filters used in state-of-the-art OSS data frameworks. On one hand, concerning 95% of raw Internet tokens are eliminated through parsing, deduplication, and curation. On the another hand, we established that accepted techniques would have missed approximately 1.8 trillion high-quality tokens we retain, including 87% of our curated web-code tokens.

Code modeling requires its own curation for the highest performance. For all language, we applied individually tuned filters, removed low-quality autogenerated and unlearnable content, and trained classifiers to acknowledge corrupted satisfied or code that may hurt training stability. Overrepresented languages and document types are rebalanced to additional broaden exposure.

We additionally developed a high-throughput pipeline for handling PDF artifacts, to justify the basis Beam has cognition throughout a broad range of STEM topics. It integrates a vision-language OCR example alongside in-house norm classifiers and relic detectors that capture malformed reconstructions, distributed throughout thousands of GPUs to procedure petabytes of specialized data.

We repeated code and specialized satisfied multiple times to addition the model’s visibility complete the way of its training horizon. This required careful notice to fuzzy deduplication, packing algorithms, and the discipline of overtraining, to justify all repeated data origin helps fairly than hurts generalization.

Frontier-grade pretraining infrastructure

Beam was pretrained end-to-end in under four weeks on a collection of 6,144 NVIDIA GB300 NVL72 GPUs. To accomplish the reliability, performance, and betterment speed required to train Beam at scale, we built nearly the complete infrastructure stack in-house. This included a novel topology-aware, Kubernetes-based scheduler throughout our clusters; an inner node lifecycle scheme alongside uninterrupted health monitoring and alerts; and a silent data misconduct (SDC) finding scheme capable of semi-autonomous rewinds and restarts. Together, these systems gave us the achievement and operational authority required to train our own frontier example efficiently and alongside greater stability.

As a outcome of extended investments in the training formula stability, infrastructure, and data quality, the general pretraining run completed alongside an extremely gentle trajectory. We executed nine semi-automatic rewinds throughout, attributed either to non-deterministic gradient norm spikes or to suspected SDCs. In addition, the run’s goodput (the portion of wall-clock period spent on training steps retained in the final model) reached 92.3% towards the end gratitude to improvements in checkpointing, error detection, and node health management.

Figure 9: Beam training defeat complete the way of its pretraining run. We observed no instabilities or ample irrecoverable spikes.

Building a powerful previous for RL in Midtraining

Our midtraining phase was designed specifically as a basis for high-compute RL, evolving the knowledge, reasoning, and tool-use capabilities that authorize Beam to study from additional demanding tasks.

We built multi-stage data curation pipelines that grasp the construction and complexity of real-world tasks during expanding safety of capabilities that are difficult to study from raw data alone. Starting from carefully selected real-world examples, these pipelines transform, combine, and broaden matter into training data designed to instruct particular capabilities. This includes long, reasoning-rich documents that disclose the example to extended chains of logic.

Midtraining additionally extends Beam’s productive environment dimension to 1M tokens. We merge organized code repositories, long-horizon tasks, and high-quality long-form documents to instruct the example to identify, retain, and nexus applicable data throughout lengthy sequences.

Together, these capabilities provision RL a stronger starting point, enabling Beam to examine additional complex solutions, activity through longer interactions, and study from tasks that would alternatively be out of reach.

Safety and Alignment

Our safety and alignment activity engaged training a second example from our pretrained checkpoint using a distinct SFT and RL pipeline on data specifically targeting the principles by which the example should abide. We merged the capabilities of the two teachers – a large-scale RL instructor and a dedicated safety and alignment instructor – via multi-teacher on-policy distillation (MOPD).

We organized Beam's safety and alignment principles into three tiers:

(1) Rules that Beam should not break: following our safety policies and maintaining its character as an AI agent.

(2) Qualities that Beam should consistently satisfy, specified as: using the environment it is given, making exact claims, acknowledging uncertainty, and transparently following the user’s request.

(3) The default manner for how Beam should interact: direct, thorough, efficient, and proactive in anticipating what the person may need next.

Figure 10: Alignment & safety RL predictably shapes Beam’s behavior throughout non-verifiable domains. From the pre-RL checkpoint, we are capable to accurately forecast (dashed line) betterment in training reward (dots).

We designed the RL environments in the alignment phase to incentivize Beam’s adherence to these principles. Many of these environments engaged non-verifiable rewards judged by a generative reward model, yet we were capable to use them to predictably form the model’s behavior. Forecasting how rewards would power distinct behaviors before operating RL, we were capable to foretell RL gains improved than the Best-of-N ceiling solitary (r = 0.79 compared to r=0.46 alongside BoN ceiling). This allowed us to iterate on rubrics and reward design, squash behaviors akin hallucinations and excessive formatting, and enhance Beam’s general communication quality.

Our safety training used deliberative alignment (Guan et al., 2024) techniques to merge our safety guideline immediately into the model's reasoning. The dataset was built adversarially and iteratively: in all round, we trained a model, generated prompts that elicited harmful or over-refusing behavior from it, and folded the prosperous attacks rear into SFT blend for the next round. For safety RL, we likewise originated single-turn, multi-turn, jailbreak, and agentic scenarios in which a simulated adversary pressures a tool-using example to obtain unsafe actions, to simultaneously decrease over-refusals and harmful compliance.

We volition publish the results of our safety evaluations in our example specialized study and volition open-source safety evaluations we developed and used internally to create a shared, inspectable norm that the open ecosystem can test against and contribute to.

The Path Ahead

This preview shows what Beam can do today. We are making this first type of Beam accessible to a choose collection of users; you can sign up for the waitlist here.

We desire Beam to be extensively accessible and uncomplicated to build on. This month, we volition publish the weights under an Apache 2.0 license, alongside alongside records and the complete stack for running, evaluating, and fine-tuning the model. We volition be launching Beam alongside an ecosystem of allocation partners, as fine as integration alongside a broad range of open origin libraries and harnesses, so developers can use Beam throughout existing open-source workflows.

Beam is the archetypal example in a series, and the archetypal display of the open intellect our squad is committed to building. We are already training what comes next, alongside the goal of bringing the open frontier nearer to the frontier of intellect alongside all release.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads