AI News Feed
Market watch
Companies

Aleph Alpha Releases Kolibri, an Open-Weight English-German MoE Model With 3.46B Active Parameters

Aleph Alpha has released Kolibri, a 78.1B-parameter open-weight English-German MoE model that activates 3.46B parameters per token, supports up to 1,048,576 tokens of context, and targets sovereign deployment in regulated sectors.

The release is deployable, the report says. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on two H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.

Kolibri, also called Kolibri-1, is a bilingual English-German MoE transformer developed end to end by teams in Germany, according to the technical research report cited by MarkTechPost. Aleph Alpha controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Germany and Finland. The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of the Code, and its data pipeline redacts personal data before training.

Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert. Expert load is balanced with Exact Quantile Balancing and Load-Error Injection. Attention uses grouped-query attention with 48 query heads and 4 KV heads. Every fifth block uses full attention without positional encoding. The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, Aleph Alpha reports the hybrid supports sequences four times longer than a full-attention model.

The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss. On German text it reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer. That means 11.2 percent fewer tokens on German web text. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.

Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length. A 201B-token long-context stage then trained on 262,144-token sequences. Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically. Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.

Aleph Alpha evaluated every model with the same eval-framework setup, according to the report. In English, Kolibri leads GPQA Diamond at 84.3, AIME 2025 at 96.9 and AIME 2026 at 96.0. It ties Qwen3.5 35B-A3B on the English agentic average at 63.4. It trails on BFCL v4, scoring 61.4 against 70.5 for Qwen3.5. The dense Qwen3.8 27B scores higher overall, at 80.2 English and 79.9 German, but activates about eight times more parameters per token. Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English. The report refers to Qwen3.5 35B-A3B in the benchmark text and Qwen3.6 35B-A3B in a comparison table.

In that comparison table, Kolibri-1 lists 78.1B total and 3.46B active parameters, 1,048,576-token maximum context, reasoning control of none, low, medium or high, and the Apache 2.0 license. Qwen3.6 35B-A3B lists about 35B total and 3B active parameters, 262,144 native context and about 1M with YaRN, thinking on or off, and Apache 2.0. Nemotron 3 Super lists 120B total and 12B active parameters, 1M tokens and the NVIDIA Nemotron Open Model License. Mistral Small 4 lists 119B total and 6.5B active parameters, 256k tokens, no reasoning control or high only, and Apache 2.0. The table gives overall scores of 75.5 English and 70.8 German for Kolibri, 71.4 and 67.3 for Qwen3.6 35B-A3B, 73.0 and 67.9 for Nemotron 3 Super, and 63.1 and 61.4 for Mistral Small 4. It also lists an agentic average of 63.4 for Kolibri, a German industry RAG average of 67.5, and an English code average of 89.3.

To deploy Kolibri, users install the aleph-alpha-inference package and serve with vLLM, per the report. The default context is 262,144 tokens. Users can pass --max-model-len 1048576 with a max_position_embeddings override for the full 1M window. Reasoning effort is set through chat_template_kwargs. Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128. The model uses 6 of 384 routed experts plus 1 shared expert per token, with 40 sliding-window and 10 full-attention blocks.