Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.
Is it deployable? Yes. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.
What is Kolibri?
Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany. According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Germany and Finland. The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.
Architecture: Sparse Experts and Hybrid Attention
Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert. Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.
Attention uses grouped-query attention with 48 query heads and 4 KV heads. Every fifth block uses full attention without positional encoding. The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.
A Tokenizer Built for German
The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss. On German text it reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer. That means 11.2% fewer tokens on German web text. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.
Training: 24T Tokens, Then SFT and RL
Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length. A 201B-token long-context stage then trained on 262,144-token sequences. Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically. Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.
Interactive Explainer: How Kolibri Works
/* 03 tokenizer */
var tl=”de”,bpt={de:{k:4.90,g:4.35},en:{k:4.58,g:4.67}};
function fmt(n){return n>=1e6?(n/1e6).toFixed(2)+’M’:Math.round(n/1e3)+’k’}
function tok(){var b=bpt[tl],mb=+$(‘#doc’).value,bytes=mb*1e6;$(‘#bK’).textContent=b.k.toFixed(2);$(‘#bG’).textContent=b.g.toFixed(2);$(‘#tK’).style.width=(b.k/5.2*100)+’%’;$(‘#tG’).style.width=(b.g/5.2*100)+’%’;$(‘#docL’).textContent=mb+’ MB of ‘+(tl===’de’?’German’:’English’)+’ text’;var k=bytes/b.k,g=bytes/b.g;$(‘#nK’).textContent=fmt(k);$(‘#nG’).textContent=fmt(g);var d=(k-g)/g*100;$(‘#dP’).textContent=(d>0?’+’:”)+d.toFixed(1)+’%’}
$$(‘[data-lang]’).forEach(function(b){b.onclick=function(){$$(‘[data-lang]’).forEach(function(x){x.classList.remove(‘on’)});b.classList.add(‘on’);tl=b.dataset.lang;tok()}});$(‘#doc’).oninput=tok;tok();
/* 04 pipeline */
var st=[[‘Pre-training’,’20T’,’tokens’,’A filtered bilingual corpus: about 43.4% English web, 23.4% German web, 13.6% code, 11.9% English instruction and reasoning data, plus STEM and specialised domains. More than 2T German tokens were curated or generated by Aleph Alpha itself.’,83],
[‘Mid-training’,’3.44T’,’tokens at 65,536 seq’,’The mix shifts toward instruction and reasoning (31.4%), STEM QA (25.4%), code (16.3%) and agentic code and tool use (15.4%). The pre-training mix fades out linearly over a transition window.’,14],
[‘Long context’,’201B’,’tokens at 262,144 seq’,’Sequences reach 256k tokens. OCR\’d PDFs make up 33.3% of this stage. The hybrid attention lets the released model accept up to 1,048,576 tokens.’,3],
[‘SFT’,’4,000′,’steps’,’Supervised fine-tuning on curated German and English data, mixed with MergeMix and labelled with reasoning-effort levels derived from trace length.’,0],
[‘RL’,’1.2M+’,’tasks’,’Reinforcement learning across reasoning, tool use, instruction following, code and retrieval. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.’,0]];
var pipe=$(‘#pipe’);st.forEach(function(s,i){var d=document.createElement(‘div’);d.className=”st”;d.innerHTML=’‘+s[0]+’
‘+s[1]+’
‘+s[2]+”;d.onclick=function(){sel(i)};pipe.appendChild(d)});
function sel(i){$$(‘.st’).forEach(function(x,j){x.classList.toggle(‘on’,j===i);x.querySelector(‘i’).style.width=j<=i?’100%’:’0′});$(‘#pdet’).innerHTML=’‘+st[i][0]+’: ‘+st[i][3];size()}
var pt=null;$(‘#play’).onclick=function(){var i=0;clearInterval(pt);sel(0);pt=setInterval(function(){i++;if(i>=st.length){clearInterval(pt);return}sel(i)},1600)};sel(0);
/* 05 effort */
var ef=[[‘none’,’none, enable_thinking=false’,’Reasoning is disabled. Proceed straight to answering according to the user\’s instructions.’,0],
[‘low’,’minimal, low’,’Reasoning effort is set to low. Think briefly through only the essential steps in the user\’s language, then proceed directly to the answer.’,1],
[‘medium’,’medium’,’Reasoning effort is set to medium. Think through the task methodically in the user\’s language, check key assumptions, and provide a well-supported answer.’,2],
[‘high’,’high, xhigh, max, default’,’Reasoning effort is set to high. Think carefully through the task in the user\’s language, validate key assumptions, consider plausible alternatives, and prioritize correctness and clarity.’,3]];
var effBox=$(‘#eff’),tt=null;ef.forEach(function(e,i){var b=document.createElement(‘button’);b.className=”eb”;var dots=””;for(var k=0;k<3;k++)dots+=’‘;b.innerHTML=e[0]+’
‘+dots+’
‘;b.onclick=function(){eff(i)};effBox.appendChild(b)});
function eff(i){$$(‘.eb’).forEach(function(x,j){x.classList.toggle(‘on’,j===i)});var e=ef[i],head=’request values: ‘+e[1]+’\nsystem prompt: ‘,txt=e[2],n=0,pre=$(‘#sys’);clearInterval(tt);
tt=setInterval(function(){n+=3;pre.innerHTML=head+txt.slice(0,n)+’ ‘;if(n>=txt.length){clearInterval(tt);pre.innerHTML=head+txt}},18);setTimeout(size,50)}
eff(3);
/* 06 bench */
var models=[‘Kolibri 78.1B-A3.46B’,’Qwen3.6 35B-A3B’,’Nemotron 3 Super 120B-A12B’,’Mistral Small 4 119B’,’GPT-OSS 120B’];
var M={‘Overall (EN)’:[75.5,71.4,73.0,63.1,72.3],’Overall (DE)’:[70.8,67.3,67.9,61.4,70.2],’Agentic avg (EN)’:[63.4,62.1,54.9,40.7,54.0],’Industry RAG avg (DE)’:[67.5,65.8,57.6,53.4,64.2],’GPQA Diamond (DE)’:[81.3,80.6,76.6,72.9,76.0],’Math avg (DE)’:[88.8,83.7,86.5,75.4,90.8],’Code avg (EN)’:[89.3,87.7,88.3,82.0,90.8]};
var cur=”Overall (EN)”,chips=$(‘#chips’);Object.keys(M).forEach(function(k){var c=document.createElement(‘button’);c.className=”chip”+(k===cur?’ on’:”);c.textContent=k;c.onclick=function(){cur=k;$$(‘#chips .chip’).forEach(function(x){x.classList.toggle(‘on’,x.textContent===k)});renderBench()};chips.appendChild(c)});
var bars=$(‘#bars’);models.forEach(function(m,i){var r=document.createElement(‘div’);r.className=”bl”;r.innerHTML=’‘+m+’0‘;bars.appendChild(r)});
function renderBench(){var v=M[cur],mx=Math.max.apply(null,v);$$(‘.bl’).forEach(function(r,i){var f=r.querySelector(‘.fill’);f.style.width=”0″;(function(f,i){setTimeout(function(){f.style.width=v[i]+’%’},30+i*70)})(f,i);var val=r.querySelector(‘.val’);val.textContent=v[i].toFixed(1);val.style.color=v[i]===mx?’var(–mag)’:’var(–ink)’;f.style.background=i===0?’var(–teal)’:(v[i]===mx?’var(–mag)’:’var(–magL)’)})}
renderBench();
window.addEventListener(‘load’,size);window.addEventListener(‘resize’,size);setTimeout(size,300);
})();
