Reflection AI announces Beam, open-weight model that promises to rival Chinese AI while spending less compute
Reflection AI, a Brooklyn startup backed by Nvidia, released Beam, an open-weight model that it says matches the performance of Chinese models while using up to 4 times less computing power at inference.
The announcement
Reflection AI, a Brooklyn startup founded in 2024, announced on Monday (5th) Beam, its first frontier model with open weights. The announcement confirms a weekend report from Axios that had already pointed to an imminent launch, and came accompanied by a long post on the company's blog detailing the model's architecture and benchmarks.
According to Reflection, Beam is a text-only model (no image or audio support) built on a mixture-of-experts architecture, trained with compute-intensive reinforcement learning to be efficient at reasoning, code generation, and agentic tasks. The company describes the model as a "workhorse" aimed at enterprises, the public sector, and developers.
What Beam is, in numbers
Beam has 501 billion total parameters, but only 23 billion active parameters per inference, thanks to the mixture-of-experts architecture. It was pretrained on 23.8 trillion tokens and supports a 1 million token context window.
Reflection did not announce immediate public access: the full weights and technical documentation are expected to be released later this month, with distribution via hyperscalers and neoclouds, plus integration with open source libraries at launch.
Beam versus GLM-5.2 and Inkling
The company states that Beam ties with GLM-5.2, from China's Z.ai, on advanced reasoning benchmarks, surpassing the leading current Western open models. The detail Reflection highlights is cost: according to the company, the result comes at a fraction of rivals' computational spend.
a fraction of the token cost and inference time compute of rivals.
Reflection AI, in an official post on the company's blog
For comparison, GLM-5.2 has about 744 billion total parameters, with 40 billion active:
| Model | Company | Total parameters | Active parameters | Context |
|---|---|---|---|---|
| Beam | Reflection AI | 501B | 23B | 1M tokens |
| GLM-5.2 | Z.ai | ~744B | 40B | - |
The most direct Western rival is Inkling, an open model from Thinking Machines Lab (Mira Murati's company), launched in July 2026. On the four coding tests for which both companies publish comparable results, Reflection claims Beam outperforms Inkling, but the comparison comes with a caveat: Inkling is multimodal while Beam is text-only.
None of these performance claims have been independently verified as of this writing. These are figures released by Reflection AI itself, without external audit.
Money, Nvidia, and the "AI factories" bet
Reflection was founded in 2024 by two former Google DeepMind researchers and has already raised approximately $4.7 billion in investment, with Nvidia, Sequoia Capital, and Lightspeed Venture Partners among its backers, according to PitchBook data. The most recent round valued the company at $25 billion (pre-money valuation).
This past summer (Northern Hemisphere summer), the startup closed deals totaling more than $7 billion with SpaceX and Nebius to secure access to Nvidia GB300 chips through 2029. That's an unusual amount of computing capacity for a two-year-old company, and shows the scale of the bet behind Beam.
Reflection's long-term product goes beyond releasing a model: the company wants to sell "AI factories," a service in which companies and governments train customized versions of Reflection's models on their own proprietary data. It's the same vision that Jensen Huang, Nvidia's CEO and an investor in Reflection itself, has publicly championed for some time, not by coincidence, since it also means selling more GPUs.
Axios reported that hedge funds and trading firms have already shown interest in this type of system, and Reflection is already testing a "sovereign AI factory" partnership with Shinsegae Group in South Korea.
What changes for those running AI locally, in Brazil
Here's the part that matters for anyone without a hyperscaler budget: Beam's mixture-of-experts architecture means that, although the model carries 501 billion parameters, each processed token only activates 23 billion of them. That brings Beam's inference compute cost close to that of a much smaller dense model, which is the basis for the company's claim of "3 to 4 times less compute."
But there's a technical distinction worth highlighting: less active compute per token is not the same as less memory required. To run Beam locally, you still need to have all 501 billion parameters loaded (or partitioned between GPU and disk, using the offloading techniques already used with other MoE models like Mixtral). In other words, the efficiency gain shows up in inference time and in cost per generated token, not necessarily in the amount of VRAM you need to buy.
In practice, that puts Beam in the same category as other large open MoE models: unfeasible to run on an ordinary developer machine, but potentially cheaper to operate on a dedicated server or through a cloud provider than a dense model of similar size to the Chinese models. For teams in Brazil that currently rely on paid APIs from closed models because they lack access to top-tier Nvidia clusters, the real appeal will only be confirmed once the weights are released and someone publishes the inference cost per token on accessible hardware, not hyperscale infrastructure.
What's still unresolved
Reflection did not respond to requests for additional information from TechCrunch as of the publication of the original report. Beam's weights are not yet available: the promise is release "this month," with no exact date and no details on the usage license (open weights do not automatically guarantee a permissive license for commercial use or redistribution).
There is also, so far, no independent benchmark confirming the performance and efficiency numbers released by the company. For anyone considering adopting Beam in production, the prudent path is to wait for the official release of the weights and run their own tests before swapping out any model currently in use.
Translated from the Brazilian Portuguese original · Read the original
ChatGPT gets a feature to build, host, and publish sites and apps right inside the chat
OpenAI has documented ChatGPT Sites, a feature that takes you from conversation to published site without leaving the chat, with access control, login, and a custom domain. What's missing is what runs under the hood.