Back

Aleph Alpha releases Kolibri, a German and English open-weight model under Apache 2.0

Aleph Alpha's Kolibri logo, a white hummingbird and the word Kolibri on a green background
Image: Aleph Alpha (cropped)

On Saturday, Oct. 3, Germany's Aleph Alpha released Kolibri, a model with 78 billion parameters that activates only about 3.5 billion at a time, with full weights on Hugging Face under the Apache 2.0 license. Aleph Alpha says it built the model for German and English work in public administration, industry and aerospace, and that it is small enough to run on a customer's own servers. Its benchmark results come from Aleph Alpha's own tests, mostly against open models from the spring, and no independent ranking exists yet.

Aleph Alpha, the Heidelberg AI company, released Kolibri on Saturday, Oct. 3, the Day of German Reunification. It is a mixture-of-experts model with 78 billion parameters in total, of which about 3.5 billion are active for each token, and it was trained on about 262,000 tokens of context, which the company says can be stretched to 1 million with extra settings. The full weights are on Hugging Face under the Apache 2.0 license, which the model card says covers only the weights and configuration files, not the training code or methods.

Aleph Alpha says Kolibri is built for German and English work in regulated areas such as public administration, industry and aerospace, and is small enough to run on a customer's own servers without sending data to outside services. The model card puts its memory footprint at about 78 GB, with a minimum of two Nvidia A100 (80 GB) or H100 chips, or one H200, B200 or B300. It needs the company's own plugin for the vLLM serving software, supports tool calls, and lets users choose how hard it reasons, from none to high. Aleph Alpha recommends staying under about 262,000 tokens when speed matters and for complex tasks.

In Aleph Alpha's own tests, Kolibri scored 96.9 on the AIME 2025 math test and 84.3 on the GPQA Diamond science test, ahead of Alibaba's Qwen3.6-35B-A3B, Nvidia's Nemotron 3 Super and Mistral Small 4. It trails Qwen3.6 on the BFCL v4 tool use test (61.4 against 67.2) and the LongBench Pro long-document test (64.5 against 70.8). The model card also shows Alibaba's Qwen3.8 27B, a model that uses all 27 billion of its parameters for every token, ahead on overall scores in both languages, 80.2 against 75.5 in English and 79.9 against 70.8 in German.

The comparisons have limits. The launch post's main charts compare Kolibri with three rivals from the spring. Its full table adds Alibaba's Qwen3.8 27B, which scores higher, but other newer open models, such as Z.ai's GLM-5.3 and Moonshot AI's Kimi K3, are not included. As of Saturday, Kolibri was not listed on the independent Artificial Analysis leaderboard. Aleph Alpha says the model was trained to answer "I don't know" when documents lack the answer, and on one public hallucination test it declined to answer, rather than answering wrong, 44 percent of the time, up from 15 percent for its unreleased predecessor, though Qwen3.6 scored 56.7 percent on the same measure.

Still unknown: how Kolibri performs in independent tests, and how its scores on Aleph Alpha's internal tests for sectors such as automotive suppliers and the German public sector, which are not public, carry over to real deployments. Aleph Alpha has also signed a binding agreement to merge with Canada's Cohere, with the combined company to operate under the Cohere name, so the model's longer-term roadmap is not yet clear.