Evidence at a glance
The mechanism in one line
Compress the visual or contextual input before the main reasoning path.
Route or verify the expensive step instead of repeating the full path.
Translate the mechanism into a bounded deployment or evaluation check.
The Problem Is Not Parameter Count but Deployment Cost
Aleph Alpha has released Kolibri, or Kolibri-1, an open-weight mixture-of-experts language model for English and German. It is aimed at regulated sectors such as public administration, industry, and aerospace. The model has 78.1 billion total parameters but activates 3.46 billion per token, or about 4.4 percent of the total. It supports contexts of up to 1,048,576 tokens and lets callers set reasoning effort per request. Its weights are available under Apache 2.0. The FP8 checkpoint is about 78 GB and is intended to run on one B200, B300, or H200, or on two H100 SXM5 GPUs, with vLLM and dedicated reasoning and tool-call parsers.
These specifications change how the model should be compared. Parameter count traditionally serves as a proxy for memory, compute cost, and capability, but an MoE system separates the amount of stored capacity from the amount of computation performed for each token. Kolibri keeps a large pool of experts while trying to make each generation resemble the compute burden of a 3.46B active model. For a sovereign deployment team, the question is therefore not only whether a larger model can be trained or accessed. It is whether that model fits the organization’s GPU inventory, data boundary, and operational model.
Sparse Experts Separate Capacity from Computation
Kolibri’s sparsity is not merely a matter of skipping random layers. The model has 50 Transformer blocks. In each MoE layer, a sigmoid router scores 384 routed experts, selects the top six, and always activates one shared expert. Expert load is managed with Exact Quantile Balancing and Load-Error Injection. The intention is to give different tokens more targeted parameter subsets while avoiding a situation in which a few experts become overloaded and the rest remain idle.
The gains and costs are inseparable. The gain is lower effective computation per token while retaining a large overall pool of expert capacity. The cost is that the serving system must handle routing, expert scheduling, and load balance. Actual throughput cannot be inferred from the 3.46B active-parameter figure alone. Different requests can select different expert combinations, and batching and parallelization policies will affect performance. Kolibri is therefore better understood as a model-plus-runtime system, not as a compressed 78.1B dense checkpoint.
That is why the release emphasizes vLLM support and dedicated parsers. For a team operating its own service, loading the weights is only the beginning. Expert routing, communication, KV-cache behavior, reasoning modes, and tool-call protocols all belong in the performance test plan. Open weights lower the access barrier, but they do not make the surrounding systems engineering disappear.
The Million-Token Context Comes from Layered Attention
Kolibri does not achieve its long context by applying full attention over one million tokens in every layer. It uses grouped-query attention with 48 query heads and four KV heads. Every fifth block applies full attention without positional encoding, while the other 40 blocks use RoPE-based sliding-window attention over the preceding 512 tokens. The sliding-window layers keep a fixed-size KV cache, so only the ten full-attention layers grow with context length.
This architecture confines the cost of long context to part of the network. Aleph Alpha reports that, at matched compute, the hybrid design supports sequences four times longer than a full-attention model. That helps explain how Kolibri can expose a one-million-token context window for long documents, extended conversations, or large retrieved collections.
But accepting one million tokens is not the same as using one million tokens reliably. The release explains the cache and sequence-length mechanics, but it does not provide independent evidence for long-context quality. A technical owner should treat the context window as a resource limit, not as automatic retrieval, memory, or reasoning capability. Tests still need to measure recall and citation stability when relevant information appears at different positions in the context.
German Tokenization Turns Language Coverage into a Product Decision
Kolibri’s language positioning is implemented in its tokenizer, not only in its training-data claims. It uses a 128,000-token vocabulary and UniBPE, whose merge process resembles BPE but scores candidate merges with Unigram loss. On German web text, Kolibri reaches 4.90 bytes per token versus 4.35 for the GPT-5 tokenizer, which the release translates into 11.2 percent fewer tokens. For English, Kolibri reaches 4.58 bytes per token versus 4.67 for GPT-5. The interactive material also says that 65 percent of Kolibri’s German token boundaries fall on morpheme boundaries, compared with 47 percent for the GPT-5 tokenizer.
These figures reflect a concrete engineering objective: German morphology should not be segmented inefficiently, causing the same document to consume more tokens and raising long-context and reasoning costs. Aleph Alpha added more than two trillion curated or synthetically generated German tokens during training. Pre-training covered 20 trillion tokens, followed by 3.44 trillion mid-training tokens and a 201-billion-token long-context stage. The main pre-training run used 768 NVIDIA B200 GPUs, while the long-context stage used sequences of up to 262,144 tokens.
Token efficiency is not, however, a complete proof of language capability. Fewer tokens can improve cost and context capacity, but they do not by themselves establish quality on specialist terminology, dialects, cross-language retrieval, or mixed German-English inputs. A European localization program should build evaluations around its business languages and document forms rather than relying on one av
Sovereign Deployment Still Requires Its Own Evidence
Kolibri’s training and governance story is aligned with its deployment target. Aleph Alpha says that teams in Germany controlled the full pipeline, including data, architecture, training infrastructure, post-training, and evaluation, with infrastructure located in Germany and Finland. The data pipeline removes personal data before training. The design is intended to align with the EU General-Purpose AI Code of Practice, the EU AI Act, and GDPR, and the company is a signatory to that code. Post-training combines supervised fine-tuning, MergeMix, and reinforcement learning on more than 1.2 million internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.
This makes Kolibri a candidate foundation model for regulated environments, not a guarantee of turnkey compliance. Apache 2.0 licensing and FP8 deployment at roughly single-GPU scale can help an organization retain control over its data and service, but compliant operation still requires data governance, logging, access control, tool-call boundaries, and domain-specific evaluation. Per-request reasoning effort can create useful tiers between latency, cost, and answer quality, but teams must decide which tasks can use low effort and which require a larger budget or human review.
The adoption decision should therefore be concrete. Kolibri belongs on the shortlist when an organization mainly handles English and German, requires local execution, wants open weights, and can validate MoE serving and long-context quality. If the business depends on broad multilingual