Mistral Large 4 arrives in preview with 1-trillion-parameter MoE architecture and open weights expected in October
ML4, nicknamed Le Chonk, is a mixture-of-experts model with 1 trillion total parameters and 49 billion active per token. The preview API is already live; open weights are coming by the end of October 2026.
Mistral announced on Tuesday, October 6, 2026, the public preview availability of Mistral Large 4, internally nicknamed "Le Chonk." The model can already be tested via API on Mistral Studio, but the full weights will only be released by the end of this month, when the model actually becomes open-weight for those who want to run it on their own infrastructure.
What ML4 really is
According to Mistral's announcement, ML4 is a natively multimodal model built on a mixture-of-experts (MoE) architecture: it has 1 trillion parameters in total, but activates only 49 billion per processed token. This is an important technical distinction for anyone evaluating inference cost: the model doesn't process the full trillion on every call, so latency and computational cost per token tend to approach those of a much smaller dense model, even while holding a much larger volume of knowledge.
Training ran from scratch on 3,800 NVIDIA Grace Blackwell GPUs, in Mistral's own datacenters in Europe, and that's the same infrastructure serving the public preview today. The company describes the launch as the first milestone of the roadmap funded by its €3 billion Series D round, presented as the largest capital raise ever made by a European technology company.
For anyone pricing the API now: Mistral charges $1.36 per million input tokens and $4.18 per million output tokens in preview.
Open, but not yet today
It's worth separating two things the announcement blends together: the preview API is live now, but the weights (what actually enables self-hosting, local fine-tuning, and on-premise deployment) only arrive by the end of October. Until then, according to Mistral, the model is being red-teamed in real-world environments with cybersecurity partners and state authorities, who receive a version with reduced moderation and expanded cyber capabilities.
This matters for anyone planning adoption: if the promise of running outside third-party cloud is the reason for interest, the real timeline is "weights in weeks," not "today." The company hasn't yet detailed the full architecture, additional benchmarks, or post-training methodology, promising to publish all of that alongside the weights.
The numbers that matter for people who write code
In tests released by Mistral itself, ML4 shows up in competitive positions against the current set of open and closed models on the market in 2026, including GPT-6 Astra, Claude Opus 5.5, GLM-5.3, DeepSeek V4 Pro, and Kimi K3.
| Benchmark | ML4 Result | Comparison |
|---|---|---|
| DeepSWE v1.1 | 61.7% | coding agent reference |
| SWE-Atlas-QnA | 59.4% | repository comprehension |
| Terminal-Bench 4.0 | 28.3% | terminal workflows |
| Coding Agent Index (combined) | 49.8% | ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
| AutomationBench | 59.9% | ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro |
| AA-Briefcase (Elo) | 1,393 | ahead of DeepSeek V4 Pro |
Mistral also commissioned a blind evaluation with professional human annotators via Surge AI, focused solely on code quality. In that evaluation, on a scale of 1 to 5, ML4 Preview scored 3.74, placing second among five models tested: behind only Claude Opus 5 (4.22) and ahead of GLM-5.3 (3.60), Kimi K3 (3.59), and GLM-5.2 (3.40).

The cybersecurity bet that changes practical use
The most specific angle of the announcement is cybersecurity, and here the material brings a concrete data point: on the Artificial Analysis Cyber Index, described by Mistral itself as an independent evaluation, ML4 ranks among the top five models in the world, and on an index test that asks the model to reproduce a real open-source software vulnerability and then fix it, it scored 82%, the highest score recorded on the test so far. On Cybench, a set of 40 exercises from security competitions, the model solved 93% of the challenges.
The counterpoint cited by Mistral itself is revealing: closed models like Claude Opus 5.5 and GPT-6 Astra score close to zero on the same test because they refuse to perform the task, even though it's legitimate vulnerability research. It's this refusal bottleneck in closed models that Mistral uses to justify selling ML4 to security teams that need an auditable, self-deployable model, running on private cloud or on-premise.
At the same time, the company reports that ML4 has the highest refusal rate among open models when the request is actually malicious, as measured on the JailbreakBench, StrongREJECT, and AgentHarm benchmarks. On the B3 AI Security Benchmark, a public benchmark from Lakera (a third party), the model resists 93.3% of indirect prompt injection attempts, the highest score reported among competitors such as GLM-5.2, GLM-5.3, Kimi-K2.6, Kimi-K3, and DeepSeek V4 Pro 0813.
Legal, finance, and multimodality
Beyond the purely technical axis, Mistral reports third-party evaluations via vals.ai in which ML4 outperforms GPT-6-Astra on both legal and financial tasks, and on Harvey AI's legal agent benchmark the model ranks ahead of every open-source model tested. In computer vision, Mistral highlights that ML4 outperforms GPT-6-Astra on Dense 200, a visual grounding test (42% versus 41%), and describes use cases ranging from inspecting gigapixel satellite imagery to verifying parts in engineering technical drawings.
What's still left open
Most of the numbers above come from Mistral's own internal benchmarks. There are, however, evaluations conducted or documented by third parties: the legal and financial analyses via vals.ai, the blind code evaluation via Surge AI, the Artificial Analysis Cyber Index (which Mistral itself describes as an independent evaluation), and the B3 AI Security Benchmark, which is a public benchmark from Lakera.
Outside of those cases, as of this preview's publication, there is no additional independent benchmark run by third parties on ML4. The exact license under which the weights will be published also wasn't detailed in the source, nor were the hardware requirements to run the full model locally, something that should only become clear once Mistral publishes the architecture documentation promised for the end of the month.
For anyone deciding on an AI production stack in Brazil, ML4 is still, for now, a preview API hosted in Europe, useful for testing code quality and agents before the weights exist to evaluate real self-hosting costs. The promise of running on-premise is concrete in the company's messaging, but it only becomes an actual possibility once the weights file comes out of the oven.
Translated from the Brazilian Portuguese original · Read the original
Mistral launches Le Chonk, open 1-trillion-parameter model to rival OpenAI and Anthropic
France's Mistral unveiled Mistral Large 4, nicknamed Le Chonk: an open-weight, 1-trillion-parameter model that the company says is the best outside China, amid US restrictions on access to frontier models.