TL;DR
- Beam Preview: Reflection AI unveiled Beam for coding and tool-using work, offering selected users access before its public weight release.
- Coding Results: Company benchmarks put Beam ahead of Western rival Inkling, with mixed results against Chinese model GLM-5.2.
- Reasoning Control: Beam’s lower effort settings favor shorter responses; higher settings allow longer reasoning for demanding tasks.
- Open Weights: Reflection promises downloadable model parameters and an Apache 2.0 license later in October, enabling developers to run and adapt Beam themselves.
Reflection AI unveiled Beam, its first model for writing code and carrying out tasks with software tools, on October 5, 2026. The startup is targeting businesses and governments that want a Western-built AI system they can run and customize themselves, but that control still awaits its promised release of downloadable model parameters. Selected users receive an early version, and others can request access through a waitlist.
Reflection’s chief executive Misha Laskin told Semafor that Reflection sees demand from organizations unwilling or unable to use Chinese models. Beam’s immediate pitch is a combination of coding capability and lower inference compute, the processing work needed to generate an answer.
Coding Gains and Current Alternatives
Reflection’s scorecard gives Beam 80.1 on Terminal Bench v2.1, a test of tasks performed through a computer’s command-line interface, against 63.8 for Inkling from Thinking Machines Lab. On SWE Bench Pro v1, which tests software-engineering work, Beam scores 65.5 against Inkling’s 54.3.
The comparison with Z.ai’s GLM-5.2 is closer: Beam’s 65.5 exceeds GLM-5.2’s 62.1 on SWE Bench Pro v1, while its 80.1 trails 81.0 on Terminal Bench v2.1. GLM-5.2 already offered public weights and local deployment options in June, giving developers an existing model they could run themselves.
Inkling’s developers used an internal coding harness, the software that supplies tools and runs the model’s task attempts, for Terminal Bench. Reflection’s technical report remains promised for later in October.
Beam accepts text, while Inkling accepts text, images and audio and has public weights for customization. Inkling from Thinking Machine Labs is a direct US rival.
Z.ai’s GLM-5.3-Flash supports visual input and local deployment, and reports 84.3 on Terminal Bench 2.1. It uses 18 billion active parameters: those are the learned numerical values used to process each token, or chunk of text. Beam uses 23 billion, so Flash uses fewer active parameters and has a higher published score. Z.ai ran the test through Claude Code with a six-hour timeout; without equivalent Beam settings, the figures remain separate vendor results rather than a controlled comparison.
Reflection also acknowledges that Moonshot AI’s Kimi K3 remains ahead on raw capability. Its own Terminal Bench v2.1 table lists Kimi K3 at 88.3 against Beam’s 80.1. The model’s public weight files are available, giving developers another current alternative to Beam’s limited preview.
How Beam Tries to Spend Less Compute
Beam contains 501 billion parameters in total, with 23 billion active for each token. Its mixture-of-experts architecture routes work through part of the model, while the complete set of weights retains all 501 billion learned values.
Reflection says Beam reaches comparable advanced-reasoning scores to GLM-5.2 using three to four times less inference compute.
Reinforcement learning rewards the model for successful task attempts. Reflection also applied a length penalty to discourage unnecessary generated tokens: early in training, it says performance improved while responses shortened. Later, longer reasoning accompanied further gains on more demanding tool-using tasks.
Beam’s reasoning-effort control lets users make that tradeoff themselves. Lower settings favor shorter responses; higher settings permit more reasoning on difficult work.
Tool use extends what a text-only model can attempt. Reflection reports that, when given web access, Beam learned to query other AI models and invoke text-recognition services to read documents. One company demonstration uses public data to build a live New York City subway dashboard.
Reflection reports pretraining on 23.8 trillion tokens, followed by a separate four-week reinforcement-learning run using 10,500 Nvidia GB300 GPUs and generating more than 100 million rollouts, the model’s task attempts. In an interview with Alex Heath, chief technology officer Ioannis Antonoglou explained why the team trained Beam from scratch: control of the architecture, data and initial training supplied a reasoning foundation and helped maintain stability as it increased reinforcement-learning compute.
Beam follows Reflection’s reported June compute agreement with SpaceX, made while the startup was still training its models.
Open Weights, Paid Deployment
Reflection promises Beam’s weights under Apache 2.0 later in October, alongside documentation and tools for running, evaluating and fine-tuning the model. Downloadable weights would let organizations operate Beam on their own infrastructure and adapt it to their data, moving those choices beyond access to Reflection’s preview service.
The founders’ idea is to sell the software and infrastructure around that control. They told Heath that customers need help deploying agents, customizing models and operating them at scale.
Beam remains in final adversarial safety testing and evaluation. Reflection plans to publish its safety results in the technical report and release the safety evaluations it developed internally, together with the broader developer materials promised for October.