📑 Table of contents

Aleph Alpha open-sources Kolibri-1: a bilingual DE/EN 78B MoE (3.46B active) under Apache 2.0 for sovereign on-premise AI

LLM & Modèles 🟢 Beginner ⏱️ 14 min read 📅 2026-10-04

Aleph Alpha open-sources Kolibri-1: a bilingual DE/EN 78B MoE (3.46B active) under Apache 2.0 for on-premise sovereign AI

🔎 On October 3, 2026, Germany dropped an open-weight 78B. The timing is no accident.

On October 3, 2026 — German Unity Day — the Heidelberg-based lab Aleph Alpha released Kolibri-1: a mixture-of-experts model with 78.1 billion parameters, of which 3.46B are active per token, under the Apache 2.0 license. Choosing that specific date for a model branded as "sovereign" is a political message as much as a software release.

On paper, the spec sheet holds up: a native 262,144-token context (1M extrapolated), native DE/EN bilingualism, a controllable reasoning mode, tool calling, and serving via vLLM. The weights are online on Hugging Face, in FP8 and BF16, with no gating and no commercial restrictions.

Why now? Because the 2026 open-weight wave has changed in nature: DeepSeek under the MIT license, Qwen on code, Alphabet on robotics — and now a European lab publishing weights, pipeline, and data provenance. The release is notably covered by AICoder News, and Orcarouter's analysis sums up the angle well: here, disclosure is part of the product.


The essentials

  • MoE 78.1B total / 3.46B active (4.4%): ratio of roughly 22.6:1, inference cost close to a 3B dense model.
  • Apache 2.0 license on the weights, non-gated repos, commercial use allowed — but the training pipeline is not published.
  • Natively bilingual DE/EN: 21.3% of pre-training in German. No other language, French included.
  • Context: 262,144 tokens native, 1,048,576 extrapolated (262k recommended for serving).
  • Benchmarks: AIME 2026 (EN) 96.0, GPQA Diamond 84.3, LiveCodeBench v6 85.9, SWE-Bench Verified 66.4, TerminalBench 2.1 only 27.7.
  • Minimum hardware: 2x A100 80GB, 2x H100 SXM5, 1x H200 or 1x B200/B300; footprint of about 78 GB in FP8.
  • Serving: vLLM with the aleph-alpha-inference plugin.
  • Knowledge cutoff: June 18, 2026.

Tool Main use Price (October 2026) Ideal for
Kolibri-1 (FP8) Quantized weights, ~78 GB Free (Apache 2.0) On-prem production deployment
Kolibri-1-BF16 Full-precision bfloat16 weights Free (Apache 2.0) Fine-tuning, research
vLLM + aleph-alpha-inference plugin High-performance serving Open source, free Enterprise GPU inference
Kolibri product sheet Specs, license, hardware requirements — Check before buying GPUs

The weights are free. The real cost is the hardware — and it's non-negotiable, as we're about to see.


A 78B MoE that activates only 3.46B parameters per token

The whole idea behind Kolibri-1 comes down to this ratio: 78,103,074,560 parameters in total, but only 3.46B activated per token, i.e., 4.4% of the model. You get the capacity of a 78B with an inference cost close to that of a dense 3B — it's the trade-off that makes on-premise economically debatable, then viable.

The architecture: 50 MoE blocks, 384 experts per layer, of which 6 are routed plus 1 shared expert per token. This ~22.6:1 ratio places Kolibri-1 in the same family as Qwen3-Coder-Next, the 80B MoE with 3B active parameters that rivals Claude Sonnet — the formula that proved itself in 2026 on the capacity-to-cost ratio.

A detail that matters for long context: the attention is hybrid, with a 512-token sliding window on four out of five layers, and a full attention layer one out of five. It's this design that keeps the KV cache footprint contained over long windows, without blowing up the memory bill.

One last, tasty detail: the predecessor, Kolibri Origin (30.6B total / 3.27B active), was never released. Aleph Alpha iterated internally before shipping a "production-ready" version — the opposite strategy to the ritual, media-heavy releases of some American labs.

The claimed Pareto frontier

The tech report claims the quality/serving-cost Pareto frontier among the evaluated open models. In other words: at equal quality, Kolibri-1 would be cheaper to run than the open competition — or, at equal cost, would perform better.

A claim to take with a grain of salt, since it depends heavily on the models being compared and the serving conditions. But it is consistent with the MoE design and the hybrid attention: this model was designed with operating cost in mind, not just for leaderboards.


Bilingual by design: 21.3% German in pre-training

Kolibri-1 is not an English-language model translated after the fact. German accounts for 21.3% of its pre-training mix, drawn from a single pool of 2.4T tokens, 80% of which were curated or generated in-house. That is a rare industrial choice: virtually all labs optimize for English and treat other languages as an add-on.

That "80% in-house" figure deserves a closer look. It means total control over data quality and provenance — the sovereignty argument taken all the way down to the dataset — but also probably a significant share of synthetic data, whose real effect on generalization remains to be documented over the long term.

Bilingualism is, moreover, openly embraced as a limitation: DE and EN, nothing else. No French, no Spanish, no Chinese. Aleph Alpha prefers two excellent languages to thirty average ones — a defensible trade-off for its target market, a penalizing one for the rest of the world.

The scores confirm it: Overall EN 75.5 vs DE 70.8. The gap exists, but it is contained. To put things in honest perspective: the dense Qwen3.8 27B remains ahead on both languages (80.2 / 79.9), according to the benchmark tables relayed by SignalStack. Kolibri-1 is therefore not the best open model, full stop — it's the best open sovereign and European model in its niche.

If your needs are in French, move along: our selection of the best LLMs in French lists far more relevant options for France.


Benchmarks: brilliant at reasoning, lagging on agentic

Kolibri-1 excels at pure reasoning — 96.0 on AIME 2026, 84.3 on GPQA Diamond — but drops off sharply on agentic execution, with a TerminalBench 2.1 score of 27.7. In short: this model thinks better than it acts.

Benchmark Kolibri-1 Score Takeaway
AIME 2026 (EN) 96.0 Olympiad-level math — excellent
GPQA Diamond (EN) 84.3 Expert-level science — very good
LiveCodeBench v6 85.9 Code generation — solid
SWE-Bench Verified 66.4 Real-world bugs — decent, not dominant
TerminalBench 2.1 27.7 Terminal agents — the weak spot
Overall EN / DE 75.5 / 70.8 Behind Qwen3.8 27B dense (80.2 / 79.9)

The post-training explains this profile: SFT followed by RL on over 1.2M internal tasks covering reasoning, tool use, code, and retrieval. The reasoning mode is controllable — you can switch it on or off depending on the task, a real plus for balancing latency and cost in production.

But let's be honest: 66.4 on SWE-Bench Verified and 27.7 on TerminalBench is a long way from the frontier models. On our agentic leaderboard, OpenAI's GPT-5.5 sits at 98.2 and Claude Opus 4.7 (Adaptive) at 94.3. If your use case is autonomous coding agents, take a look at our comparison of the best LLMs for coding or our guide to LLMs for AI agents instead.

Kolibri-1 isn't playing that game. It's playing the regulated deployment game — and on that field, a point on SWE-Bench matters less than compliance, auditability, and cost of ownership.


Apache 2.0: what the license really covers (and what it doesn't)

The Apache 2.0 license covers Kolibri-1's weights — commercial use included, repo not gated — but not the training pipeline. It's "open weight", not full open source. The distinction matters, especially when marketing sells you "sovereign".

In practice, you can download, modify, deploy, and monetize the model with no royalties and no negotiation. You cannot audit the training data line by line, nor reproduce the full pipeline. Aleph Alpha does, however, publish a disclosure that is rare in the industry: data provenance, 392k GPU-hours, energy figures.

This release is part of a powerful 2026 trend. DeepSeek opened V3.1 under the MIT license and is now pushing V4 Pro, while Alphabet open-sourced its entire Intrinsic Core robotics stack under Apache 2.0. Open weight has become a strategic weapon, not an activist gesture.

My take: Aleph Alpha's license is more honest than the vocabulary that sometimes surrounds it. "Open weight under Apache 2.0" is an exact, verifiable fact; plain "open source" would be an exaggeration. The two are not equivalent, and a sovereign buyer would do well to know that before signing.


Context: 262k native, 1M extrapolated — what you should actually use

The card advertises 1,048,576 tokens of context, but the recommended usage in serving remains ≤262,144. The training trajectory explains it: 16k at pre-training, extended to 65k in mid-training, then 262,144 natively. The 1M is validated by extrapolation — it works, but it wasn't learned.

For latency-sensitive workloads, 262k is therefore the right target. For offline batch — corpus analysis, large-scale retrieval, contract review — the 1M is worth a try, with the KV cache in FP8 to keep the memory footprint in check. The hybrid attention (512 sliding window, one full layer in five) is precisely designed to make these windows affordable.

For comparison, most frontier models advertise contexts of 1M to 2M, with serving costs that climb quickly beyond 100k tokens. Kolibri-1 makes the opposite bet: an honest, quantified context, paired with an explicit usage recommendation. I'll admit I prefer this candor to the race for round numbers.


Running Kolibri-1 at home: vLLM, 78 GB and datacenter GPUs

Kolibri-1 is served via vLLM, but plan for datacenter hardware: minimum 2x A100 80GB, 2x H100 SXM5, 1x H200 or 1x B200/B300. With ~78 GB of FP8 weights, a 24 GB consumer GPU is ruled out from the start — and even a single H100 80GB won't do: the weights would leave no room for the KV cache.

Configuration Verdict
1x RTX 4090 / 5090 (24-32 GB) No — weights too heavy
1x H100 80GB No — no headroom for the KV cache
2x A100 80GB Official minimum
2x H100 SXM5 Yes
1x H200 (141 GB) Yes
1x B200 / B300 Yes

Installation goes through Aleph Alpha's official plugin for vLLM:

pip install "aleph-alpha-inference>=1.0"
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8

The official blog details the recommended serving flags, including the FP8 KV cache. Plan for memory headroom: the 78 GB of weights don't include the KV cache at 262k context. And if you're going for fine-tuning, the BF16 version weighs about twice as much, i.e. ~156 GB.

No datacenter at hand? Two honest options: rent H100/H200 GPUs by the hour from a cloud provider, or settle for a lighter model. For the second route, our guide to installing an LLM locally covers Ollama and LM Studio, and our comparison of the best local LLMs helps you pick a model that fits on normal hardware. To experiment with these lighter models on a VPS, Hostinger offers an affordable entry point (pricing to be checked on hostinger.com, October 2026). And to build agents on less demanding local models, our feature on open source AI agents with Ollama remains the best starting point.


Sovereignty: provenance as a product

Kolibri-1's real differentiator isn't a benchmark, it's the provenance chain. 24T tokens documented, 392k GPU-hours on 768 NVIDIA B200s (21 days of pre-training), energy figures published, pipeline described. No American frontier lab offers this transparency: OpenAI, Google, and Anthropic publish neither data, nor costs, nor training budgets.

The positioning is explicit: "mission critical", EU AI Act compliant, on-premise. For a German hospital, a government agency, or a bank, the question isn't "which model has the best score on AIME?" but "which model can I audit, host behind my firewall, and get certified?". On this question, Kolibri-1 answers with verifiable arguments, not slides.

The knowledge cutoff of June 18, 2026 is recent — less than four months old at release — which limits the reliance on RAG for recent facts. Tool calling and retrieval are part of post-training, not band-aids added after the fact. And the training budget, 392k GPU-hours, remains modest by frontier lab standards: proof that a European player can play in the big leagues without a colossal budget.

My take as a writer: Kolibri-1 will never beat GPT-5.5 or Gemini 3 Pro Deep Think on general-purpose leaderboards, and that's not its goal. It's a political and industrial product — the demonstration that a European lab can deliver a top-tier model, open, documented, and deployable on private infrastructure. For the regulated European market, that may be worth more than a point on SWE-Bench.


Is Kolibri-1 right for you?

Yes if you check at least two of these boxes: regulated organization (healthcare, finance, public sector, defense), need for on-premise or sovereign cloud, working languages English and/or German, inference volume justifying dedicated GPUs.

No if you work in French, if you're looking for the best coding agent on the market, or if you don't have access to datacenter GPUs. In those cases, the monthly comparison of the best LLMs or the selection of free LLMs will point you in the right direction faster than a 78 GB download.

Use cases where Kolibri-1 makes sense today: internal DE/EN assistant on sensitive data, document extraction and analysis over long corpora, batch reasoning (math, science) on private infrastructure, Apache 2.0 fine-tuning base for a vertical domain.

Cases to avoid: high-intensity terminal agentic work (the 27.7 on TerminalBench 2.1 says it all), French-language production, experimentation on consumer GPUs.


❌ Common Mistakes

Mistake 1: confusing "open weight" and "open source"

The Apache 2.0 license covers the weights, not the training pipeline or the data. Solution: read the license and the tech report before communicating internally about "open source", and assess precisely what you can actually audit.

Mistake 2: underestimating the required hardware

"3.46B active" doesn't mean "lightweight": the 78B weights have to fit in memory, i.e. ~78 GB in FP8. Solution: start from the official minimum configuration (2x A100 80GB or 1x H200) and add the headroom needed for the KV cache.

Mistake 3: pushing the context to 1M in production

The 1M is extrapolated, not native; the card recommends ≤262,144 for latency-sensitive serving. Solution: 262k in production, 1M reserved for offline batch with the KV cache in FP8.

Mistake 4: expecting French

Kolibri-1 is bilingual DE/EN, period. Solution: for French-speaking clients or content, choose a model that is genuinely suited to it rather than forcing a half-effective system prompt.


❓ Frequently Asked Questions

Is Kolibri-1 really open source?

Technically, it's open weight: the weights are under Apache 2.0, usable commercially and without gating, but the training pipeline and data are not published. Aleph Alpha's transparency (provenance, GPU costs, energy) is above the industry average, without reaching full open source.

What is the minimum hardware to run it?

The official minimum is 2x A100 80GB, 2x H100 SXM5, 1x H200 or 1x B200/B300. The weight footprint is about 78 GB in FP8, to which the KV cache must be added. A 24 GB consumer GPU is not enough, even with additional quantization.

Does Kolibri-1 speak French?

No. The model is natively bilingual German/English, a deliberate choice by Aleph Alpha. Overall scores are 75.5 in English and 70.8 in German. For French, other open or commercial models are significantly better suited.

How does it stack up against GPT-5.5 or Claude Opus 4.7?

On agentic benchmarks, frontier models remain ahead: GPT-5.5 dominates the rankings (98.2), and TerminalBench 2.1 at 27.7 shows Kolibri's limits in execution. Kolibri-1 plays on a different field: on-premise, permissive licensing, EU AI Act compliance, and serving cost.

How do you serve Kolibri-1 in production?

Via vLLM with the official plugin: install aleph-alpha-inference, then run vllm serve with the KV cache in FP8. Recommended configurations are documented on Aleph Alpha's blog. Plan for datacenter-grade hardware and memory headroom for long context.

What is Kolibri-1's knowledge cutoff?

June 18, 2026, less than four months before the model's release. That's recent for an open weight model, which reduces the need for RAG for recent facts — without eliminating it for day-to-day news.


✅ Conclusion

With Kolibri-1, Aleph Alpha proves that a European lab can deliver an open, documented, and on-premise deployable 78B MoE, without claiming to outperform American frontier models on agentic capabilities. Before downloading the 78 GB, evaluate your hardware with our guide to installing a local LLM — or explore lighter alternatives if your use case allows it.