Germany has a new national champion in AI, and it is not a government agency. It is a private company called Aleph Alpha, and its new model is called Kolibri — German for hummingbird. The bird is small, fast and light, and so is the model. It is a large language model, or LLM, built entirely on German soil, trained in Germany and Finland, and released under the Apache 2.0 license so anyone can use it freely. The model card and its 189-page technical report explain exactly how it works.
The Hummingbird’s Weight Problem
A normal large language model reads every token through every parameter. That takes vast amounts of memory and computation. Kolibri solves this by splitting its mind into smaller minds, each with a special job. It has 50 layers, and each layer holds 384 experts plus one shared expert that every token passes through. The router picks six of those 384 experts for each token to process. That means 78.1 billion parameters do the work of a much smaller model — about 3.46 billion per token.
The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.” So Kolibri computes like a small model but needs the memory of a huge one.
The Tokenizer That Understands Compound Words
A model reads tokens, not words. Tokens are chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer trained mostly on English tends to chop them apart. Kolibri’s tokenizer, called UniBPE, fixes that. It keeps the bottom-up merging of byte-pair encoding, or BPE, but picks each merge with a different scoring rule — the Unigram objective — which respects how German builds words.
The result is a tokenizer that reads the German Federal Constitutional Court’s name as one compound word, while a standard tokenizer splits it into six separate tokens. On the Basic Law for the Federal Republic of Germany, the German constitution, Kolibri’s tokenizer needed 15% fewer tokens than GPT-5’s tokenizer, even more than Aleph Alpha’s own claimed 11.2% savings. In English, it tied with GPT-5.
Layers That Look Nearby
Forty of Kolibri’s 50 layers use sliding-window attention, meaning each token only looks at nearby tokens. That cuts down on the memory needed to hold the whole model in RAM. The remaining ten layers likely handle the longer-range connections that sliding windows miss.
Training Data and Filters
Kolibri was trained from scratch on infrastructure in Germany and Finland, with no foreign control over the build. Its training data drew on English web text rephrased with Google’s Gemma 4, German text filtered with Mistral-NeMo, and Qwen3-32B labeled data for quality filters. Aleph Alpha filtered the data for political bias, which it has measured in Chinese open models itself.
The model card says clearly that nothing from outside Europe went into its design, even though some training data came from outside sources. The sovereign claim applies to how it was built and deployed, not to what data it processed.
What the Model Does Well
Aleph Alpha’s own evaluation says Kolibri scores above every compared model of its size in both German and English. The company signed the European Union’s General-Purpose AI Code of Practice, which sets standards for how AI systems are used across Europe.
Running Kolibri
The model card and technical report are the primary source. The model runs on standard hardware, and its weights sit on Hugging Face for anyone to download. The tokenizer experiment was run using Hugging Face’s tokenizers library and OpenAI’s tiktoken.
Why It Matters
Kolibri is a direct answer to a worry voiced three years ago at SmashingConf New York. One speaker said the room heard that “you regulate, you don’t innovate,” and wished for a middle path between innovation and regulation around data privacy, data stewardship, environmental constraints and energy requirements. Kolibri is that answer: a German team built it with the EU AI Act in mind from the ground up, and it gives ministries and car suppliers full freedom to run it on their own servers without sending their data elsewhere.
The model is also a practical tool. Fewer tokens mean fewer steps to read or write the same German text, and more German fits in the same context window. That matters for legal documents, constitutional texts and everyday conversation.
The Verdict on Kolibri
Kolibri is a clever piece of engineering. It takes a fraction of its own weight to do real work, and it handles German better than almost anything else out there. The sovereign claim is honest about its limits — it was built with care, not with a wall around its data. It is a model worth watching, and it is a model worth running.
Here is what stands out:
- 78.1 billion parameters, but each token touches only about 3.46 billion of them
- 384 experts per layer, with a router picking six for each token
- 128,000-token vocabulary, trained with UniBPE
- Trained in Germany and Finland, with no foreign control over the build
- Scores above every compared model of its size in both German and English
- Weights released under Apache 2.0 on Hugging Face
Kolibri is a small model with a big ambition: to show that German engineering can hold its own on German soil. It does.
Source material: “Show HN: Germany's new sovereign AI model Kolibri,” tej.as.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

