Open-weight models Beam and Mistral Large 4 challenge OpenAI and Anthropic on cost
Reflection AI launched Beam and Mistral updated Large 4, promising top-tier performance at a fraction of the price of closed models. For those deciding on AI stacks in production, the math has already changed.
Two launches, the same pitch
On October 7, 2026, Reflection AI, an American startup founded in 2024 by two former Google DeepMind researchers, announced Beam, according to a report by Brazil Journal, a Brazilian business news outlet. The model was designed for the corporate market and for code developers, with a specific focus on AI agents that perform autonomous tasks in engineering pipelines.
Around the same time, French company Mistral launched Large 4 (ML4), internally nicknamed "le Chonk," optimized for cybersecurity, programming, and finance. The pitch from both companies is nearly identical: deliver performance equivalent to the best open Chinese models, such as DeepSeek and Moonshot, while charging a fraction of what Anthropic and OpenAI charge for their closed models.
It's no coincidence. The LLM market is splitting between those who sell closed access via API, like OpenAI and Anthropic, and those who distribute open weights for anyone who wants to run, fine-tune, or embed the model wherever they want. Beam and Large 4 are the latest bets from the second group, arriving just as that group is winning the fight for usage volume.
What "open-weight" really means (and what Beam still doesn't deliver)
Open-weight is not synonymous with open source. In LLMs, the equivalent of open source code is releasing the model's architecture and trained parameters, allowing any developer to download the weights and run inference on their own infrastructure, without depending on the provider's cloud. This doesn't necessarily mean opening up the training dataset or the full pipeline used to get there.
Beam doesn't yet meet even the minimum promise: Reflection AI stated that it will release the weights and the full technical details of the model "soon." In other words, the launch announced this week is, for now, an announcement of capability, not the actual release of the artifact that an engineering team could download and put into production today.
For those evaluating a stack, this distinction matters in practice. A closed model like Claude or GPT can be consumed via API from day one; an open-weight model only becomes a real self-hosting option once the weights actually exist. Until then, Beam competes on narrative, not on the shelf.
The numbers behind the hype
The strongest argument in favor of open models isn't a promise, it's a consumption trend. According to Brazil Journal, open-weight models already account for nearly 80% of the tokens (the units of information processed by LLMs) used in the market. In early 2026, the proportion was reversed: closed models concentrated around 80% of the volume.
That's a shift that took less than a year. Much of it is driven by those who need to run high-volume, cost-sensitive inference, a typical case for products with coding agents, RAG pipelines, or automations that trigger thousands of calls per day. When the cost per token drops significantly, the open model stops being an ideological choice and becomes a spreadsheet decision.
Valuations are tracking the dispute: Reflection AI closed a new round at a pre-money valuation of US$25 billion (the amount raised wasn't disclosed), with Nvidia and Sequoia on the cap table. Mistral, considered the only cutting-edge AI startup in Europe, raised €3 billion in a Series D round in September 2026, according to CNBC, at a valuation of €21 billion.
How the models stack up
| Model | Company | Open weights today | Declared focus |
|---|---|---|---|
| Beam | Reflection AI | No (promised "soon") | Code agents, corporate use |
| Mistral Large 4 | Mistral | To be confirmed in official documentation | Security, programming, finance |
| DeepSeek / Moonshot | Chinese companies | Yes | Reference benchmark for open cost-performance |
| Claude / GPT | Anthropic / OpenAI | No | Closed API, highest revenue in the sector |
What changes for those building on top of this
For an engineering team that currently runs everything via the OpenAI or Anthropic API, the arrival of open models that are competitive in performance changes three concrete decisions:
- Marginal cost per token: self-hosting eliminates the API provider's margin, but requires owned or contracted GPUs, with a fixed cost that only pays off at high volume.
- Fine-tuning and sensitive data: open weights allow the model to be fine-tuned with proprietary data without that data traveling to a third-party cloud, relevant for regulated sectors such as finance and healthcare.
- Agent lock-in: AI products built on agents that make many chained calls (coding assistants, multi-step automations) feel the effect of cost per token in an amplified way, and are the first to migrate when an open model delivers equivalent performance for less.
None of these advantages come for free. Running a cutting-edge LLM on your own infrastructure requires an MLOps team, high-cost GPUs, and an SLA the team has to build itself, something a closed provider's API solves with a curl and an authentication key.
The counterpoint: closed models still dominate revenue
Brazil Journal's token data measures usage volume, not revenue. Anthropic and OpenAI continue to concentrate most of the AI market's revenue, because they charge a premium for reliability, enterprise support, and ready-made integration, not just for model capability. A company running a critical product in production pays for that, even knowing a cheaper alternative exists on paper.
This is the real bet behind Beam and Large 4: it's not about beating the incumbents on revenue in the short term, it's about eroding the margin that sustains their valuation, pushing the market price down token by token. If the argument holds, the effect doesn't show up first in OpenAI's revenue, it shows up in the pressure it feels to cut the price of its own API.
When it's worth migrating (and when it isn't)
A team that already operates GPU infrastructure, has a high enough volume of calls to dilute the fixed cost, and doesn't depend on a formal enterprise SLA has real reason to test Beam as soon as the weights are released, or to evaluate Large 4 right now via Mistral's API. An early-stage startup, with low volume and a need for integration speed, is probably still better off staying with the closed API, at least until the tooling ecosystem around open models matures as much as OpenAI's and Anthropic's.
Translated from the Brazilian Portuguese original · Read the original
Claude API pricing: why cache and tokenizer matter more than the chosen model for margin
The official Claude API pricing documentation, covering Haiku 4.5, Sonnet 5.5, Opus 5.5, and Fable 5.1, shows that cache, tokenizer, and fast mode weigh on margin as much as the choice of model.