Mistral Large 4 arrives in preview with 1 trillion parameters, focus on security and code
Mistral this week opened API access to Large 4, an open-weight 1-trillion-parameter model trained in European data centers, with strong benchmarks in code and cybersecurity.
What Mistral announced
On October 6, 2026, Mistral AI published the public preview of Large 4, officially referred to as "le Chonk" and internally nicknamed ML4. The model is already available via preview API in Mistral Studio, but the open weights are only expected to be released by the end of this month, according to the company's announcement.
ML4 is a natively multimodal model with 1 trillion parameters in total and 49 billion active parameters per inference, characteristics of a mixture-of-experts architecture. Mistral describes Large 4 as its most capable model to date and claims it outperforms any open-weight model developed in the US or Europe, with performance competitive against the strongest open-source models on the global market.
Until the weights are released, the company says it is red-teaming the model with cybersecurity leaders, vetted partners, and state authorities, who receive access to the same base model with reduced moderation and expanded cyber capabilities.
The numbers that matter to developers
For those building software, the most relevant point of the announcement is the code and agentic task benchmarks. Mistral reports 61.7% on DeepSWE 1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. The combined Coding Agent Index comes in at 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.
The company also ran a blind human evaluation with Surge AI, in which professional annotators rated code output quality on a 1-to-5 scale, without knowing which model generated each response.
| Model | Score (1-5) |
|---|---|
| Claude Opus 5 | 4.22 |
| Mistral Large 4 | 3.74 |
| GLM-5.3 | 3.60 |
| Kimi K3 | 3.59 |
| GLM-5.2 | 3.40 |
Large 4 placed second among five models evaluated, behind only Claude Opus 5. On broader agentic tasks, the model scores 59.9% on AutomationBench, a set of 657 corporate workflows involving Gmail, Google Sheets, Slack, and Salesforce, surpassing Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro. On AA-Briefcase, which evaluates long-horizon knowledge work (spreadsheets, slides, PDFs), the model reaches 1,393 Elo, also ahead of DeepSeek V4 Pro.
Cybersecurity without the guardrails of closed models
The most unusual angle of the announcement is Large 4's explicit positioning as an offensive and defensive cybersecurity tool. On the Artificial Analysis Cyber Index, an independent evaluation of how well a model finds and fixes security flaws in real-world software, ML4 ranks among the top five in the world and leads by a wide margin among open-weight models developed outside China.
In one of the index's tests, which asks the model to reproduce a real vulnerability in open-source software and then fix it, Large 4 scores 82%, the highest score recorded by any model so far. On Cybench, a set of 40 security competition exercises, the model solves 93% of the challenges.
Mistral notes that competing closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to carry out the task. The company argues that defending software often requires proving that a flaw is real, exactly the kind of work that safety filters in closed models block, and that this matters even more as attackers already jailbreak these same models for offensive use.
In the company's internal tests, the model also proved useful for malware analysis, vulnerability prioritization, and writing detection rules, tasks it was not explicitly trained for. For organizations that need sovereign, auditable AI in security operations, the promise is to run the model in a private cloud or on-premise.
European infrastructure and the money behind it
Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs, in Mistral's own data centers in Europe, and the public preview runs on the same infrastructure. The company states that this marks its long-term investment in infrastructure, research, and product, delivering state-of-the-art performance via open weights to give organizations control over their AI.
The model will be available across multiple regions, including a European deployment operated by Mistral itself, independent of other digital service providers and subject to European law. A significant share of the training data is multilingual, covering more than 160 languages, including all official languages of the European Union.
Large 4 is the first milestone on the roadmap funded by Mistral's €3 billion Series D round, the largest funding round ever raised by a European technology company. The company says the reinforcement learning run behind the preview is still ongoing, with no signs of saturation, and expects rapid improvements in the coming weeks and months as compute capacity expands.
Pricing, access, and what's still missing
The preview API charges $1.36 per million input tokens and $4.18 per million output tokens, already available in Mistral Studio. For those deciding today which model to use in an agentic or automation project, it's a competitive price against other frontier models, but it still doesn't come with weights for self-hosting.
This matters for teams in regulated sectors, such as finance and healthcare, that tend to prefer running models locally for data residency reasons and compliance with laws like Brazil's LGPD (Brazil's data protection law): the promise of on-premise deployment only becomes verifiable once the weights are released, expected by the end of October 2026. Mistral also did not detail, in this announcement, the license under which the weights will be distributed, nor has it yet published the full architecture paper, the additional promised benchmarks, or the post-training methodology, items the company itself says will come together with the weight release.
It's also worth noting that Large 4 shows higher refusal rates than other open-weight models on malicious cybersecurity prompts (as measured via JailbreakBench, StrongREJECT, and AgentHarm), and resists 93.3% of attacks on Lakera's public B3 benchmark, the highest rate among the competitors cited. For teams evaluating whether to adopt the model in legitimate offensive security pipelines (red teaming, bug bounty), this combination of high technical capability with selective refusal of malicious use is the differentiator Mistral is selling.
Translated from the Brazilian Portuguese original · Read the original
Anthropic expands Cyber Verification Program and opens Claude Opus 5.5 and Mythos to more security teams
The company merged two restricted-access programs into one, the Cyber Verification Program, giving vetted security teams access to Claude Opus 5.5, Sonnet 5.5, and Mythos 5.1 with fewer safety guardrails.