Aleph Alpha launches Kolibri, an open 78B-parameter model built to run on-premise
German company releases full weights under Apache 2.0, with a focus on data sovereignty, but still without Portuguese support
The German company Aleph Alpha released the full weights of Kolibri under the Apache 2.0 license, an open-source alternative designed for those who need to run language models locally, without sending data to third-party APIs. The model is not yet focused on Portuguese, but its technical design is of interest to those thinking about data sovereignty in Brazil.
The German company Aleph Alpha released the full weights of Kolibri under the Apache 2.0 license, an open-source alternative designed for those who need to run language models locally, without sending data to third-party APIs. The model is not yet focused on Portuguese, but its technical design is of interest to those thinking about data sovereignty in Brazil.
Aleph Alpha, a German AI company, released the full weights of Kolibri on October 3, 2026, a bilingual English-German Mixture-of-Experts language model, under the Apache 2.0 license. The download is available on Hugging Face, according to the company's official blog. The date was not chosen at random: October 3 is German Reunification Day, and the company made a point of marking the launch on the national holiday.
What Kolibri Is
Kolibri has 78.1 billion total parameters, but activates only 3.46 billion per token thanks to its MoE (Mixture of Experts) architecture, with 384 experts of which 6 are activated on each pass. According to Aleph Alpha, the model supports contexts of up to 1 million tokens, although the technical table published in the same post shows that effective training reached 256 thousand tokens (262,144), with additional extension via long-context adaptation.
The model was designed for regulated sectors such as public administration, industry, and aerospace, with an emphasis on reasoning, mathematics, and agentic behavior. The central proposal, according to the company, is to allow clients to run the model on-premise, without sending internal data to third-party inference services.
From Origin to Kolibri in Three Months
Kolibri is the second generation of a training pipeline built by Aleph Alpha throughout 2026. The first, named Kolibri Origin, finished pre-training on June 11 and never had a public release; it served to validate the training pipeline. Kolibri finished pre-training on September 11, just three months later.
| Characteristic | Kolibri Origin | Kolibri |
|---|---|---|
| Total parameters | 30.6B | 78.1B |
| Active parameters/token | 3.27B | 3.46B |
| Pre-training tokens | 7.51T | 20T |
| Maximum trained context | 65,536 | 262,144 |
| Reasoning modes | 1 | 4 (none, low, medium, high) |
In that span, the team tripled the number of experts, swapped the routing algorithm, changed the attention design, and more than doubled the environment tasks used in reinforcement post-training. For those working with ML infrastructure, the most relevant figure is operational: over 21 days of pre-training, Aleph Alpha recorded 38 unplanned interruptions (roughly one every 10 thousand GPU hours), all automatically recovered by the pipeline, which restarts the job on a different set of nodes from a checkpoint no more than 250 steps back.
Training ran on 768 B200 GPUs, totaling nearly 24 trillion tokens across pre-training (20T), mid-training at 64k context (3.44T), and long-context adaptation at 256k (200B). The company also introduced its own exact-quantile expert balancing technique, building on an idea that Kimi K3 had implemented in approximate form.
Efficiency Before Size
Aleph Alpha tested scaling the model up to 123B total parameters and saw quality gains, but the cost of serving it made that choice unviable: according to the company, a 123B model processes only 3 simultaneous 256k-token requests on two H100 GPUs, compared to 18 simultaneous requests for the 78B Kolibri, which also decodes 28% faster. This trade-off (fewer active parameters, more throughput) is the launch's central efficiency argument.
In the public benchmarks released, Kolibri appears on par with models that have up to four times more active parameters, such as Nemotron 3 Super (120B total, 12B active):
| Benchmark | Kolibri | Kolibri Origin | Nemotron 3 Super |
|---|---|---|---|
| AIME 2025 | 96.9 | 81.9 | 91.7 |
| GPQA (diamond) | 84.3 | 68.1 | 78.0 |
| HumanEval+ | 92.7 | 76.8 | 94.7 |
| LiveCodeBench v6 | 85.9 | 59.2 | 82.0 |
The numbers show a large jump over Kolibri Origin, but a mixed result against larger competitors: Kolibri loses to Nemotron on HumanEval+ (92.7 vs 94.7), but wins on LiveCodeBench v6 (85.9 vs 82.0), even while using a fraction of the active parameters.
Sovereignty by Design
Aleph Alpha describes Kolibri as built from the ground up to comply with the European AI Act, the Code of Practice for general-purpose AI, and the GDPR, with specific attention to copyright.
Our teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.
Aleph Alpha, official blog
The model was also trained to acknowledge uncertainty: through what the company calls the Merlin-Arthur protocol, Kolibri is encouraged to answer "I don't know" when the provided context does not support an answer, rather than hallucinating. It's a feature designed for corporate document search applications, where a made-up answer is costly.
What This Means for Developers in Brazil
Kolibri is not a model for Portuguese: the tokenizer and training data prioritize English (62%) and German (about 21%, or 4.3 trillion tokens), with only 6% of translation used in a deliberately limited way. Anyone who wants to use the model in products for the Brazilian public will need their own fine-tuning, and the Apache 2.0 license allows exactly that: downloading the full weights and adapting them.
The point that matters to engineering teams here is the cost design. With just 3.46B active parameters, Kolibri was optimized to fit into smaller infrastructure than equivalent dense models, which lowers the barrier for companies and public agencies that want to stop depending on APIs like OpenAI's for reasons of cost, latency, or data policy. In Brazil, this discussion already comes up in debates about cloud sovereignty and the hosting of sensitive data within government agencies, even though no national project has announced anything equivalent to Kolibri.
The practical difference lies in two distinct paths: paying per token on a closed API versus running an open model on one's own infrastructure, taking on the costs of GPU and maintenance. Kolibri delivers that second option with competitive benchmarks, but without native support for the language of the world's largest Portuguese-speaking market.
What Remains Open
Aleph Alpha says it is considering scaling the pipeline used in Kolibri even further, but acknowledges that the jump from 30B to 78B parameters was "the easy direction": more data, more parameters, already-known architecture. For now, there is no announcement of a version covering Portuguese or other languages beyond English and German, nor any public breakdown of hosting costs for those who run the model outside the company's own testing environment.
Translated from the Brazilian Portuguese original · Read the original
Critical GitLab vulnerability allows unauthenticated data exfiltration and is already under attack
CVE-2026-85706 has the maximum severity score (CVSS 10.0) and allows anyone, without logging in, to read arbitrary files from self-managed GitLab instances. The fix has existed since September, but CISA confirms active exploitation.