Kolibri
Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.
Model overview
| Model | Kolibri 1 |
|---|---|
| Model Provider | Aleph Alpha GmbH |
| Model Developer | Aleph Alpha Research GmbH |
| Architecture | Mixture-of-Experts |
| Total parameters | 78B (78,103,074,560) |
| Active parameters / token | 3.46B (3,457,573,120) |
| Languages | German, English |
| Context length | 1,048,576 tokens; we recommend ≤262,144 tokens for serving efficiency and complex tasks |
| Precision | float8_e4m3fn weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16 |
| Reasoning mode | Yes |
| Tool calling | Yes |
| License | Apache 2.0 |
| Knowledge cutoff | EN: June 18, 2026, DE: June 18, 2026 This only affects implicit knowledge, the model may use more recent information through tool use. |
| Hardware requirements | Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300. |
| Best for | Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant |
| Release Date | 3rd of October 2026 |
| Code of Practice | Aleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai. |
| Training Data | Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension. Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases. |
| Training method | We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed. |
| Computing Resources | Pre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh) Mid-training: 5 days, 90k GPUh (same setup as above) Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6) FLOPS: 6.4e23 |
| Tech Report | https://aleph-alpha.com/downloads/tech-report.pdf |
Intended use
Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.
Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.
For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.
Design goals
Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.
AI system types
Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.
Sustainability
Energy consumption
9.5×10² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.
Energy measurement
For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.
Getting started
Kolibri requires the aleph-alpha-inference package that provides the
Kolibri vLLM plugin. You can either use the provided container image
ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from
Aleph-Alpha/aleph-alpha-inference,
which also installs the vLLM version it supports:
pip install 'aleph-alpha-inference>=1'
Serve the model with reasoning and tool-calling enabled:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.
The recommended sampling parameters for the model are temperature=1.0,
top_p=0.97 and top_k=128.
Querying the server (OpenAI-compatible API)
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[
{
"role": "user",
"content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist.",
},
],
extra_body={
"chat_template_kwargs": {
"reasoning_effort": "high",
"enable_thinking": True,
}
},
)
print(response.choices[0].message.content)
Reasoning mode
Kolibri supports explicit thinking effort levels that need to be configured
through the chat template. You can pass reasoning_effort values low,
medium and high to configure the amount of effort our model puts into
finding the answer. You can disable thinking altogether by setting
reasoning_effort to none or passing enable_thinking=false. The model
will then immediately respond.
Tool calling
The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the
standard tools field of the chat completions request; the model will emit
tool calls that the parser converts into structured output. Tool calling can
be combined with reasoning mode.
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Heidelberg",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
},
},
}
]
def get_weather(city, unit="celsius"):
return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}
messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]
first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)
for call in msg.tool_calls or []:
args = json.loads(call.function.arguments)
result = get_weather(**args)
messages.append(
{
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
}
)
final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)
Evaluation
The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.
Post-training
| Type | MoE | Dense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3B | 4-6B | 12B | 27B | 70B | |||||||||
| Ours | Baseline models | |||||||||||||
| Eval | Kolibri | Kolibri Origin | GLM-4.7 Flash 30B-A3B | Nemotron 3 Nano 30B-A3B | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Qwen3-Next 80B-A3B Thinking | Gemma 4 26B-A4B IT | GPT-OSS 120B | Mistral Small 4 119B-A6B | GLM-4.5 Air 106B-A12B | Nemotron 3 Super 120B-A12B | Qwen3.8 27B | Apertus 70B Instruct |
| Overall (EN) | 75.5 | 54.1 | 64.7 | 65.6 | 74.7 | 71.4 | 62.4 | 71.9 | 72.3 | 63.1 | 64.4 | 73.0 | 80.2 | – |
| Overall (DE) | 70.8 | 46.4 | 50.4 | 59.3 | 69.8 | 67.3 | 58.0 | 66.3 | 70.2 | 61.4 | 64.8 | 67.9 | 79.9 | – |
| Knowledge | ||||||||||||||
| Average (EN) | 50.1 | 39.7 | 45.5 | 46.0 | 52.7 | 52.1 | 48.4 | 51.4 | 50.0 | 47.5 | 45.7 | 52.0 | 56.8 | 22.8 |
| Average (DE) | 57.6 | 44.9 | 46.5 | 41.2 | 61.3 | 61.0 | 55.7 | 61.5 | 58.0 | 51.4 | 52.2 | 59.5 | 69.2 | 24.8 |
| GPQA Diamond (EN) | 84.3 | 68.1 | 73.1 | 73.9 | 83.8 | 83.4 | 76.1 | 81.1 | 76.4 | 74.7 | 73.2 | 78.0 | 89.2 | 29.5 |
| GPQA Diamond (DE) | 81.3 | 58.5 | 59.8 | 49.6 | 84.2 | 80.6 | 72.2 | 80.1 | 76.0 | 72.9 | 71.1 | 76.6 | 88.1 | 31.4 |
| Humanity's Last Exam (EN) | 21.5 | 9.4 | 15.4 | 12.1 | 20.4 | 21.1 | 11.6 | 19.2 | 19.4 | 9.7 | 8.7 | 20.6 | 35.6 | 5.2 |
| Humanity's Last Exam (DE) | 15.9 | 10.4 | 9.1 | 13.1 | 18.1 | 20.5 | 15.6 | 23.4 | 20.7 | 10.5 | 10.5 | 22.3 | 37.2 | 5.7 |
| AA-Omniscience Accuracy (public set) | 14.8 | 11.3 | 17.0 | 19.5 | 22.2 | 21.0 | 24.2 | 20.7 | 23.3 | 25.0 | 20.0 | 26.7 | 19.0 | 13.5 |
| AA-Omniscience Index (public set) | -32.8 | -64.2 | -62.8 | -45.7 | -46.2 | -12.5 | -42.3 | -47.3 | -35.2 | -24.0 | -28.8 | -36.5 | -5.8 | – |
| MMLU-Pro CoT (EN) | 80.0 | 70.1 | 76.5 | 78.3 | 84.6 | 84.3 | 81.7 | 84.5 | 80.8 | 80.4 | 80.9 | 82.7 | 85.0 | 43.0 |
| MMLU-ProX CoT (DE) | 75.5 | 65.7 | 70.7 | 61.0 | 81.7 | 81.9 | 79.4 | 81.1 | 77.2 | 70.7 | 74.9 | 79.7 | 82.4 | 37.3 |
| Math | ||||||||||||||
| Average (EN) | 96.5 | 81.7 | 88.8 | 88.8 | 90.1 | 87.8 | 86.3 | 87.4 | 90.7 | 81.4 | 82.8 | 91.1 | 97.8 | 0.6 |
| Average (DE) | 88.8 | 74.3 | 45.2 | 84.3 | 79.6 | 83.7 | 83.7 | 88.1 | 90.8 | 75.4 | 80.6 | 86.5 | 96.7 | 0.1 |
| AIME 2025 (EN) | 96.9 | 81.9 | 89.4 | 89.6 | 88.1 | 84.6 | 84.2 | 87.3 | 90.6 | 79.8 | 81.9 | 91.7 | 97.9 | 0.6 |
| AIME 2025 (DE) | 87.5 | 73.5 | 43.8 | 84.4 | 76.7 | 82.9 | 80.6 | 88.1 | 90.6 | 72.3 | 80.6 | 85.6 | 96.5 | 0.2 |
| AIME 2026 (EN) | 96.0 | 81.5 | 88.3 | 87.9 | 92.1 | 91.0 | 88.5 | 87.5 | 90.8 | 83.1 | 83.8 | 90.4 | 97.7 | 0.6 |
| AIME 2026 (DE) | 90.0 | 75.2 | 46.7 | 84.2 | 82.5 | 84.4 | 86.7 | 88.1 | 91.0 | 78.5 | 80.6 | 87.5 | 96.9 | 0.0 |
| Agentic | ||||||||||||||
| Average (EN) | 63.4 | 41.6 | 58.9 | 46.4 | 63.4 | 62.1 | 46.3 | 54.6 | 54.0 | 40.7 | 53.5 | 54.9 | 66.7 | – |
| TerminalBench 2.1 | 27.7 | – | 20.2 | 9.7 | 39.7 | – | 8.6 | – | 29.2 | 21.0 | – | 39.7 | 76.8 | – |
| Tau2-Bench (Telecom) | 94.7 | 67.5 | 95.9 | 45.9 | 97.7 | 99.1 | 43.9 | 45.3 | 73.1 | 41.5 | 53.8 | 68.1 | 82.5 | 10.8 |
| Tau2-Bench (Retail) | 69.9 | 58.5 | 57.9 | 64.9 | 70.8 | 71.6 | 60.8 | 71.3 | 60.5 | 62.9 | 61.4 | 67.5 | 68.7 | 9.6 |
| Tau2-Bench (Airline) | 76.7 | 58.7 | 68.7 | 52.7 | 76.0 | 70.7 | 65.3 | 73.3 | 72.7 | 40.0 | 70.7 | 72.7 | 83.3 | 40.0 |
| Tau3-Bench (Banking) | 38.1 | 5.7 | 7.2 | 5.7 | 11.3 | 10.6 | 5.4 | 16.0 | 14.7 | 5.7 | 6.4 | 15.5 | 50.0 | 2.1 |
| BFCL v3 (multi-turn) | 39.8 | 22.8 | 58.2 | 47.9 | 54.0 | 53.5 | 51.4 | 53.4 | 45.6 | 36.2 | 61.6 | 44.6 | 42.5 | 0.6 |
| BFCL v4 (overall) | 61.4 | 36.4 | 65.4 | 61.5 | 70.5 | 67.2 | 51.0 | 68.2 | 57.3 | 58.0 | 67.2 | 61.0 | 73.2 | – |
| BFCL v4 (non-live AST) | 79.1 | 78.1 | 83.3 | 85.0 | 85.8 | 88.2 | 83.6 | 83.7 | 35.8 | 83.6 | 85.5 | 45.0 | 85.3 | – |
| BFCL v4 (live) | 78.9 | 73.7 | 78.3 | 78.8 | 80.2 | 81.4 | 82.5 | 80.2 | 70.4 | 78.4 | 78.2 | 77.6 | 79.9 | – |
| BFCL v4 (multi-turn) | 47.5 | 27.5 | 62.7 | 53.5 | 59.9 | 58.1 | 56.0 | 61.4 | 55.4 | 40.4 | 65.2 | 51.7 | 55.5 | – |
| BFCL v4 (memory) | 62.8 | 19.4 | 41.5 | 39.1 | 62.6 | 53.8 | 35.3 | 52.9 | 50.7 | 39.1 | 43.4 | 59.6 | 79.6 | – |
| BFCL v4 (web search) | 62.5 | 10.5 | 69.0 | 66.0 | 75.0 | 68.5 | 12.5 | 75.0 | 57.0 | 69.0 | 71.0 | 71.5 | 82.0 | – |
| BrowseComp | 29.4 | 4.4 | – | 14.5 | 36.5 | 26.9 | 2.8 | 25.5 | 31.2 | – | – | 29.1 | 46.4 | – |
| Code | ||||||||||||||
| Average (EN) | 89.3 | 68.0 | 67.8 | 81.8 | 85.0 | 87.7 | 83.6 | 89.0 | 90.8 | 82.0 | 79.8 | 88.3 | 94.2 | 25.1 |
| LiveCodeBench v6 | 85.9 | 59.2 | 46.5 | 71.3 | 77.8 | 82.5 | 73.9 | 82.3 | 87.5 | 71.2 | 67.8 | 82.0 | 93.8 | 8.7 |
| HumanEval+ | 92.7 | 76.8 | 89.0 | 92.4 | 92.2 | 92.8 | 93.3 | 95.7 | 94.1 | 92.8 | 91.8 | 94.7 | 94.7 | 41.6 |
| SWE-Bench Verified | 66.4 | – | 51.0 | 38.6 | 71.6 | 73.8 | – | 57.8 | – | 60.8 | 11.6 | 60.2 | 72.6 | – |
| Instruction Following | ||||||||||||||
| Average (EN) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| IFBench (loose-prompt) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| Grounding / Hallucinations | ||||||||||||||
| Average (EN) | 59.4 | 42.7 | 57.6 | 55.9 | 67.5 | 68.3 | 61.0 | 62.5 | 62.6 | 61.5 | 63.8 | 64.4 | 67.5 | 31.5 |
| SQuAD (M/A Grounding Score) | 23.4 | 0.0 | 1.5 | 0.0 | 23.0 | 12.1 | 0.0 | 0.0 | 0.0 | 6.3 | 0.0 | 0.0 | 8.3 | 0.0 |
| SQuAD (Utility Accuracy) | 83.4 | 77.4 | 82.7 | 75.0 | 89.8 | 88.9 | 73.5 | 88.4 | 74.2 | 76.8 | 81.9 | 80.8 | 85.6 | 1.6 |
| RGB Closed-Book | 51.0 | 52.0 | 78.0 | 80.0 | 81.0 | 79.0 | 89.0 | 79.0 | 85.0 | 86.0 | 92.0 | 93.0 | 73.0 | 85.0 |
| RGB Negative (Abstention) | 85.6 | 73.9 | 79.6 | 76.9 | 86.6 | 79.6 | 81.3 | 86.0 | 79.6 | 82.3 | 89.0 | 74.6 | 70.6 | 55.9 |
| AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set) | 44.0 | 15.0 | 3.8 | 19.0 | 11.1 | 56.7 | 12.3 | 14.3 | 23.7 | 34.7 | 39.0 | 13.9 | 67.3 | 16.4 |
| RGB Fact-Check (Error Correction) | 34.0 | 14.0 | 68.0 | 58.0 | 74.0 | 74.0 | 85.0 | 61.0 | 77.0 | 60.0 | 74.0 | 90.0 | 53.0 | 14.0 |
| FRAMES (<24k) | 71.2 | 65.7 | 67.8 | 70.9 | 75.9 | 74.7 | 68.9 | 69.5 | 78.3 | 71.9 | 70.4 | 74.9 | 78.6 | 58.9 |
| FRAMES (>24k) | 73.0 | – | 73.9 | 74.3 | 81.5 | 78.8 | 68.5 | 76.6 | 78.4 | 76.6 | – | 78.8 | 83.3 | – |
| SealQA (no distractors, <24k) | 80.7 | 56.6 | 71.7 | 69.0 | 88.3 | 82.8 | 82.1 | 82.1 | 80.7 | 73.1 | 67.6 | 82.8 | 86.9 | 35.9 |
| SealQA (12 distractors, <24k) | 61.0 | 30.0 | 65.0 | 54.0 | 78.0 | 67.0 | 57.0 | 82.0 | 65.0 | 62.0 | 60.0 | 70.0 | 84.0 | 16.0 |
| SealQA (no distractors, >24k) | 100.0 | – | 61.1 | 66.7 | 100.0 | 88.9 | 83.3 | 77.8 | 83.3 | 77.8 | – | 94.4 | 94.4 | – |
| SealQA (12 distractors, >24k) | 65.1 | – | 57.1 | 39.7 | 73.0 | 71.4 | 50.8 | 68.3 | 63.5 | 49.2 | – | 68.3 | 82.5 | – |
| Agentic Retrieval | ||||||||||||||
| Average (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Average (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| MuSiQue (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Honeypot | 80.8 | 25.3 | 69.4 | 60.2 | 77.5 | 74.3 | 13.5 | 75.4 | 66.6 | 68.1 | – | 68.8 | 85.0 | – |
| Agentic Wiki QA (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| Industry RAG | ||||||||||||||
| Average (EN) | 89.7 | 53.8 | 75.9 | 61.3 | 87.0 | 86.0 | 62.7 | 79.4 | 82.8 | 74.9 | 82.3 | 80.3 | 93.3 | 23.9 |
| Average (DE) | 67.5 | 42.7 | 60.9 | 46.2 | 70.0 | 65.8 | 31.1 | 49.4 | 64.2 | 53.4 | 63.4 | 57.6 | 80.2 | 14.4 |
| Semiconductors | 80.4 | 35.3 | 64.7 | 39.2 | 79.4 | 79.4 | 41.2 | 63.7 | 70.6 | 62.7 | 72.5 | 69.6 | 89.2 | 11.8 |
| German Public Sector | 75.0 | 54.0 | 69.5 | 65.5 | 80.0 | 72.0 | 29.5 | 66.5 | 77.0 | 50.0 | 69.5 | 78.0 | 89.0 | 8.0 |
| Aerospace | 58.9 | 14.1 | 50.4 | 44.1 | 62.3 | 59.0 | 48.1 | 58.1 | 48.1 | 47.0 | – | 54.9 | 73.8 | – |
| Automotive Supplier | 99.0 | 72.4 | 87.2 | 83.3 | 94.6 | 92.6 | 84.2 | 95.1 | 95.0 | 87.1 | 92.0 | 91.0 | 97.3 | 35.9 |
| Industrial Drive Technology | 60.0 | 31.4 | 52.3 | 26.8 | 60.0 | 59.5 | 32.7 | 32.3 | 51.4 | 56.8 | 57.3 | 37.3 | 71.4 | 20.9 |
| Long Context | ||||||||||||||
| LongBench Pro | 64.5 | – | – | 53.2 | 70.3 | 70.8 | 63.8 | 64.3 | – | 56.4 | – | 62.9 | 76.9 | – |
| AA-LCR | 68.3 | – | 49.7 | 42.3 | 66.3 | 69.7 | 46.3 | 68.3 | – | 52.3 | – | 67.0 | 81.3 | – |
All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort
high. A model reports no score (–) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.
Pre-training
All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.
| Type | MoE | Dense | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 12B | 7B | 32B | 70B | |||||
| Ours | Baseline models | |||||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | GLM-4.5 Air 106B-A12B Base | Nemotron 3 Super 120B-A12B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| Overall (EN) | 81.1 | 58.4 | 58.1 | 77.5 | 73.8 | 77.1 | 83.1 | 57.6 | 67.9 | 48.6 |
| Overall (DE) | 81.5 | 61.1 | 61.1 | 76.0 | 76.6 | 79.4 | 85.0 | 48.0 | 65.2 | 50.5 |
| General Knowledge | ||||||||||
| Average (EN) | 81.3 | 70.8 | 77.2 | 80.5 | 82.7 | 82.5 | 87.2 | 67.8 | 77.8 | 73.8 |
| Average (DE) | 82.7 | 70.7 | 80.5 | 80.0 | 84.9 | 81.7 | 87.5 | 51.8 | 68.7 | 74.7 |
| MMLU (EN) | 81.0 | 68.9 | 78.2 | 79.0 | 84.7 | 82.8 | 86.8 | 67.1 | 76.3 | 69.3 |
| Global MMLU (DE) | 77.5 | 65.1 | 75.1 | 74.3 | 81.5 | 78.2 | 84.7 | 51.6 | 65.1 | 64.9 |
| ARC (EN) | 95.7 | 88.9 | 95.2 | 94.3 | 97.0 | 96.3 | 97.5 | 88.6 | 94.5 | 90.7 |
| ARC (DE) | 96.4 | 88.5 | 95.5 | 94.2 | 97.6 | 95.9 | 97.8 | 67.9 | 88.1 | 89.9 |
| PIQA (EN) | 89.5 | 76.9 | 88.0 | 89.4 | 91.9 | 89.8 | 95.1 | 77.7 | 86.8 | 80.1 |
| PIQA (DE) | 96.8 | 87.9 | 96.8 | 96.5 | 98.5 | 97.2 | 99.3 | 75.0 | 90.0 | 94.6 |
| HellaSwag (EN) | 83.1 | 80.7 | 84.8 | 85.4 | 85.3 | 87.2 | 88.8 | 76.1 | 83.5 | 84.4 |
| HellaSwag (DE) | 86.9 | 75.3 | 88.6 | 85.8 | 88.7 | 86.5 | 93.0 | 36.8 | 61.0 | 88.7 |
| MMLU-Pro (EN) | 61.1 | 39.3 | 51.2 | 55.6 | 62.6 | 55.3 | 68.9 | 38.0 | 50.9 | 40.6 |
| MMLU-ProX (DE) | 55.7 | 36.5 | 46.7 | 49.0 | 58.4 | 50.8 | 62.8 | 27.9 | 39.5 | 35.6 |
| TriviaQA (EN) | 77.2 | 69.9 | 65.6 | 79.5 | 74.6 | 83.8 | 86.4 | 59.0 | 74.7 | 77.4 |
| Wahl-O-Mat (DE) | 16.7 | 52.3 | 57.1 | 55.9 | 43.2 | 56.7 | 61.1 | 49.6 | 56.3 | 57.0 |
| Math | ||||||||||
| Average (EN) | 84.9 | 54.3 | 45.9 | 83.0 | 73.5 | 64.7 | 82.7 | 57.7 | 63.3 | 39.6 |
| Average (DE) | 76.8 | 55.8 | 47.0 | 72.5 | 77.5 | 69.8 | 81.2 | 43.8 | 62.9 | 39.0 |
| GSM8K (EN) | 89.8 | 71.6 | 65.7 | 86.3 | 89.5 | 82.8 | 87.5 | 75.3 | 81.1 | 62.3 |
| GSM8K Platinum (DE) | 90.5 | 71.7 | 64.1 | 86.3 | 88.0 | 86.5 | 93.1 | 58.7 | 80.8 | 60.2 |
| MATH Minerva (EN) | 80.1 | 37.0 | 26.1 | 79.8 | 57.6 | 46.5 | 77.8 | 40.1 | 45.6 | 16.9 |
| MATH Minerva (DE) | 63.2 | 39.9 | 29.9 | 58.8 | 67.0 | 53.2 | 69.3 | 28.9 | 45.0 | 17.7 |
| Code | ||||||||||
| Average (EN) | 77.0 | 50.1 | 51.1 | 69.0 | 65.2 | 84.1 | 79.3 | 47.4 | 62.5 | 32.5 |
| Average (DE) | 85.1 | 57.0 | 55.7 | 75.5 | 67.4 | 86.7 | 86.2 | 48.3 | 64.0 | 37.9 |
| HumanEval (EN) | 86.7 | 50.8 | 52.2 | 73.5 | 67.2 | 96.3 | 83.2 | 46.8 | 63.9 | 28.3 |
| HumanEval (DE) | 88.4 | 49.0 | 51.7 | 73.5 | 59.2 | 94.0 | 85.0 | 39.5 | 56.2 | 28.3 |
| MBPP (EN) | 67.3 | 49.5 | 50.0 | 64.5 | 63.1 | 71.9 | 75.4 | 48.0 | 61.0 | 36.8 |
| MBPP (DE) | 81.7 | 64.9 | 59.7 | 77.4 | 75.6 | 79.4 | 87.5 | 57.1 | 71.7 | 47.5 |
All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters
temperature = 0.6,top_p = 0.6,max_tokens = 1024and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommendedtemperature = 1.0,top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.
Long context
| Type | MoE | Dense | ||||||
|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 7B | 32B | 70B | ||||
| Ours | Baseline models | |||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| RULER | ||||||||
| 4k | 86.9 | 91.8 | 84.7 | 94.6 | 96.3 | 92.0 | 94.8 | 89.5 |
| 8k | 83.7 | 88.8 | 85.0 | 93.4 | 95.3 | 80.8 | 92.6 | 77.7 |
| 16k | 80.9 | 83.7 | 87.1 | 92.0 | 95.2 | 70.8 | 88.8 | 71.3 |
| 32k | 76.3 | 72.7 | 88.7 | 86.4 | 93.7 | 64.4 | 80.9 | 70.8 |
| 64k | 72.2 | – | 86.8 | 82.6 | 91.0 | – | – | – |
| 128k | 67.9 | – | 87.8 | 81.0 | 89.9 | – | – | – |
| 256k | 69.8 | – | – | 72.1 | 80.1 | – | – | – |
| 512k | 65.5 | – | – | 73.1 | 72.2 | – | – | – |
| 1M | 63.2 | – | – | 58.5 | 57.5 | – | – | – |
| HELMET | ||||||||
| 8k | 79.1 | 74.8 | 81.9 | 79.1 | 87.7 | 71.1 | 74.8 | 64.2 |
| 16k | 77.5 | 76.4 | 81.9 | 80.8 | 85.5 | 67.1 | 73.8 | 57.2 |
| 32k | 80.0 | 73.5 | 82.9 | 78.4 | 84.1 | 60.1 | 74.6 | 48.1 |
| 64k | 81.2 | – | 83.8 | 72.7 | 78.8 | 51.9 | 73.1 | – |
| 128k | 82.4 | – | 85.4 | 70.8 | 77.3 | – | – | – |
Each RULER row is the mean of its NIAH, VT, WE and QA tasks at that context length, sampled with
temperature = 0. A model reports no score (–) at lengths beyond its context window. The models with a 65,536-token window report none at 64k either: its prompts slightly exceed that under some tokenizers. Gemma 4 26B-A4B Base reports none at 256k: with the answer budget, its prompts exceed its 262,144-token window. At 512k and 1M, Kolibri Base and Qwen3.5 35B-A3B Base are served beyond their 262,144-token trained window, Qwen3.5 with static YaRN (factor 4) at every length; Nemotron 3 Nano 30B-A3B Base supports up to 1M.The HELMET rows are the mean of four HELMET tasks at that length: JSON KV and RAG QA scored by substring exact match, RULER MK3 by recall, and InfBench multiple choice by exact match. Apertus 70B Base and Kolibri Origin Base report none at 64k: some RAG QA prompts exceed their 65,536-token window.
Training Details
Model Dependencies
None. The model was trained from scratch.
Model Architecture
| Model detail | Value |
|---|---|
| Vocabulary size | 128,000 |
| Number of layers | 50 |
| Attention type | 4:1 causal sliding-window GQA to full causal GQA |
| Sliding window | 512 preceding tokens plus the current token |
| Hidden size | 2,560 |
| Attention heads | 48 |
| Key-value heads | 4 |
| Head size | 128 |
| QK normalization | Per-head RMSNorm |
| Expert hidden size | 512 |
| MLP type | SwiGLU |
| Routed experts per MoE layer | 384 |
| Experts selected per token | 6 |
| Shared experts per MoE layer | 1 |
| Routing | Token-choice, top-6 sigmoid routing |
| Pretraining sequence length | 16,384 |
| Maximum context length | 262,144 native; 1,048,576 by extrapolation |
| Position embeddings | RoPE, base 10,000, applied only in sliding-window layers |
Pre-training
Our tokenizer was trained on our pre-training mix with a vocabulary size of 128,000. With 23.9% of German share in our data, it achieves higher German compression (4.7 bytes/token) than leading models with larger vocabulary, without sacrificing English efficiency (4.2 bytes/token). We trained the tokenizer with a new algorithm, UniBPE, that respects the morphology of languages better than existing approaches, especially the compound structure of German. UniBPE keeps the greedy bottom-up approach of BPE and using the Unigram training objective for selecting which merge to add to the vocabulary.
We randomly initialized all model parameters and pretrained the model with a causal next-token-prediction objective on a large and diverse document corpus described above. Training examples consisted of 16,384-token sequences, with multiple documents potentially packed into a single sequence. The pretraining phase covered 20T tokens over 264,750 optimization steps. We used a global batch size of 4,608 sequences, corresponding to 75.497 million tokens per step, with two gradient-accumulation steps across 768 GPUs.
We linearly warmed up the learning rate over approximately 100 billion tokens, corresponding to 1,325 optimization steps, and then held it constant for the remaining 263,425 steps. AdamW was used for the embedding matrix, one-dimensional backbone parameters, router weights, and language-model head. All other two-dimensional backbone parameters were optimized with Nesterov Muon using a learning rate of 0.001.
The expert-balancing bias was updated separately using a recentered global-quantile target update. We additionally injected load-error gradients into the router scores with a weight of 1e-5 and a soft-clamp value of 1.0.
We applied independent weight decay of 2⁻¹², approximately 0.0002441, to all decay-eligible parameters. Embedding, normalization, and expert-balancing-bias parameters were excluded from weight decay. We used per-head RMS normalization of the query and key representations.
The mid-training phase increased the sequence length to 65,536 tokens and used a global batch size of 1,536 sequences, corresponding to 100.663 million tokens per step. Its 3.44 trillion token budget corresponds to 34,200 optimization steps. Because this phase resumed the pretraining optimizer state, it used no additional learning-rate warmup.
The subsequent long-context phase increased the sequence length to 262,144 tokens and uses a global batch size of 768 sequences, corresponding to 201.327 million tokens per step. Its 201-billion-token budget corresponds to 1000 optimization steps, again without additional warmup.
Training was conducted on 768 GPUs using our PyTorch-based TorchTitan distributed training infrastructure, with bfloat16 parameters, float32 reductions, and a maximum gradient norm of 1.0.
Pre-training data sources
Kolibri was trained on a filtered, bilingual (German/English) corpus combining curated web data, synthetic rephrasings and translations, and high-quality curated sources. We trained on 20T tokens, followed by 3.44T tokens in a mid-training stage and 201B tokens in a long-context adaptation stage.
Pre-training and mid-training data mix
| Domain | Pre-training | Mid-training | Long-context extension |
|---|---|---|---|
| English web and documents | 43.4% | 1.3% | 15.8% |
| German web and documents | 23.4% | 2.1% | 2.6% |
| English instruction and reasoning | 11.9% | 31.4% | 11.3% |
| German instruction and reasoning | 0.5% | 1.9% | 0.7% |
| Code | 13.6% | 16.3% | 13.2% |
| Agentic code and tool use | ~0% | 15.4% | 5.6% |
| STEM (documents) | 4.5% | 5.9% | 6.7% |
| STEM (QA-style) | 1.8% | 25.4% | 8.9% |
| Specialised Domains (Legal, Public Sector) | 1.1% | 0.2% | 1.9% |
| OCR'd PDFs | 0% | 0% | 33.3% |
Note: Rows do not sum to 100% due to rounding.
Synthetic data
We generated synthetic data using permissively-licensed LLMs.
English pre-training rephrases. Deduplicated English Common Crawl was rephrased with Gemma-4-26B-A4B following a Nemotron-CC-style recipe, using five prompt templates applied to each seed document: diverse QA pairs, distillation, knowledge list, knowledge extraction, and encyclopedia-style rewriting. Generation ran as a background job on 912 GPUs over the substring-deduplicated English corpus. The outputs were post-processed, shuffled, indexed, and included in the pre-training mix.
German pre-training rephrases. German synthetic rephrasings of web data were produced with Mistral-NeMo-12B.
Synthetic annotations. LLM-as-a-judge annotations using Qwen3-32B over randomly sampled English Common Crawl were generated as training data for our text-quality classifiers; see Data curation.
Data curation
The pre-training corpus was built from extracted Common Crawl snapshots. We applied a range of curation techniques, including but not limited to:
- URL filtering. To prevent training on illegal, harmful, infringing or pirated content, we filtered our data based on a URL filter list.
- Heuristic filtering and cleanup. Web documents passed a heuristic filter pipeline inspired by the approaches described in Dolma 3, DCLM and RefinedWeb: Gopher document and quality filters (length, symbol ratio, bullet and ellipsis lines, stop words, boilerplate, repeated lines and n-grams), fastText language identification routing documents into an English and a German branch with language-specific thresholds, RefinedWeb line-wise removal of navigation and counter lines, and, for English, the MADLAD-400 questionable-content rules. Unlike these filters, which apply to web data only, a final cleanup pass ran over every dataset in the pool: it collapsed runs of blank lines (at three newlines for code and math sources, two for all others) and capped horizontal whitespace runs, which removed extreme padding such as long runs of spaces. Filter statistics were recorded per filter, language, and source, for development purposes, e.g., understanding how much each step discards and tuning heuristics appropriately.
- Deduplication and shuffling. The Common Crawl corpus passed through exact/global (cross-dump) deduplication, MinHash fuzzy deduplication, and substring deduplication. When joining the data sources in our pre- and mid-training data pools, we globally shuffled the document indices. For mid-training, we additionally performed exact deduplication per dataset and also globally across all datasets.
- Quality classification. fastText and Luxical classifiers were trained both on LLM-as-a-judge annotations of sampled Common Crawl and on human-curated datasets. For English, the selected annotators consisted of 5 classifiers which had non-linear interaction terms. For German, we took the mean of the GermanWeb educational quality and grammar fastText model scores. Combined classifier scores were rescaled to [0, 1] in the annotation pipeline, and each language corpus was sorted into 20 equal-token-mass quantile buckets, grouped into five categories (low, medium-low, medium, medium-high, high) to enable quality-aware mixing. Bucket breakpoints were computed from per-shard equal-token-mass quantile documents.
- PII redaction. We removed personally identifiable information from
pre-training data using regular expressions, replacing
high-syntax-constraint entity types with special tokens, such as email
addresses (
<|pii-email-address|>), IP-addresses (<|pii-ip-address|>) and others. Documents in which more than 15% of bytes were replaced were dropped entirely. An ablation comparing redacted and unredacted versions of the same 750M-row sample showed no clear performance difference in either direction. - Harmful and unsuitable content. Illegal, harmful, unethical and infringing sources are excluded through the URL-level filtering described above, which also removed known adult and NSFW sites. For English web data, keyword-based filtering additionally removed sexually explicit content. These measures partially address child sexual abuse material (CSAM); no dedicated CSAM detection was applied. Personal data was redacted as described under PII redaction. The training corpus was text-only, so image-based CSAM or non-consensual intimate imagery (NCII) was not ingested.
- Decontamination. Documents in our mid-training and long-context extension datasets were scored for contamination by n-gram overlap against our evaluation suite, and contaminated documents were dropped in full. Measured contamination rates across the mid-training pool were low, between 0 and roughly 1e-4 of documents per dataset. RULER and MRCR were excluded from the reference set, as their synthetic needle-in-a-haystack construction made overlap checks uninformative.
- Bias mitigation. German is under-represented in web content, which lowers German performance. We counteracted this by upsampling German data where downstream evaluations showed gains, and by adding translated and synthetically rephrased German data to diversify style and register. Openly available fine-tuning data is often generated by models that carry political bias; we built a dedicated post-training dataset to counteract risk of contamination of our training data with politically biased material, grounded in curated reference material on politically sensitive topics.
- Mix selection. Pre- and mid-training mixes were chosen empirically using mix-search proxy models trained with different mixes to fit a function which then predicts the best mix based on our eval suite. The predicted mixes were subsequently confirmed using a separate validation proxy setup. In pre-training, this involved training thousands of small dense proxy models (30M params) with different mixes (3B tokens). For mid-training, a dedicated base proxy model was trained to match the tokens-per-parameter of the target scale. We then used the method described in MergeMix (weight-averaging of domain expert continued pre-trainings over nine domains) to find promising mixture candidates.
- Long-context. Long-context candidates were built by length-bucketing the mid-training pool and blending it with OCR'd PDFs. We included documents up to a sequence length of 256k tokens. We fixed the OCR'd PDFs at 34% of tokens, x% the optimal mid-training mix, and (66-x)% the mid-training pool weighted towards long-context documents. We ablated over multiple candidate mixes for different values of x and found the best mix to be at x=33.
Post-training
Post-training was split into two phases: A supervised fine-tuning stage that instilled reasoning and instruction following capabilities and a reinforcement learning stage that fine-tuned the long-context behavior across a broad set of environments. Across both stages we put particular focus on the key capabilities that shape the strengths of Kolibri.
Capabilities
German. We trained our model to reason in German to make it easier for our users to follow along. For the SFT side we generated translated prompts and generated language consistent reasoning answers and filtered the answers for correctness and language consistency to obtain 10.6B tokens. We continued this training in RL, where we used German environments and included a language consistency reward for both the reasoning trace and the answer.
Reasoning Effort. Kolibri can answer directly (reasoning effort none)
or reason at low, medium or high reasoning effort, which the chat
template requests through a fixed sentence in the system prompt. In SFT, we
assigned reasoning effort labels based on observed dataset-specific
reasoning length statistics and judged difficulty, while on RL side we
trained each task across all reasoning levels with different length
penalties.
Safety. Kolibri was trained in SFT to decline harmful and illegal requests while still providing partial answers where appropriate. The safeguards trained into the model do not replace safeguards at the solution level. When integrating Kolibri into an application, these should be supplemented with additional appropriate measures, such as content filtering and output validation.
Personally Identifiable Information. In addition to the mitigations during pre-training, we incorporated safety data in our SFT mix that provide refusals to PII extraction attacks. Refusals on such behavior were measured as part of our safety evaluations.
Long Context. We trained Kolibri on context windows up to 256k tokens, which enables reasoning over large documents. Since most datasets usually focus on short-context, we curated datasets that contained long-context documents based on publicly available data (e.g., from German legislative documents) with corresponding question/answering prompts. We similarly trained on long-context question-answer during RL.
Agentic Use and Tool Calling. We trained Kolibri to perform well in agentic settings involving tool calls and terminal interactions. For permissively available agentic SFT datasets, we regenerated the reasoning traces. During training, we masked out erroneous tool calls and terminal actions, so the model learned to recover from errors without learning the erroneous actions themselves. In addition to the open data, we generated two software-engineering SFT datasets: one for producing code patches and one for bug fixing as a Terminus 2 agent. During the RL stage, we trained on multiple tool-calling, software-engineering and terminal environments and randomized the harnesses to support generalization.
Retrieval. To make our model suited for customer information retrieval use-cases, we built environments that cover retrieval-augmented question answering over German and English corpora across general knowledge and specialised domains such as legal, electronics, hardware and aviation. The environments were heavily randomised across different harnesses, retrieval parameters, and languages. To warm-start the RL training phase we generate a curated set of 11k high-reward completions for SFT training.
Hallucinations. To address hallucinations, we trained with abstention data. We also constructed RL environments which implement our Merlin-Arthur protocol. The protocol shows the model each question with parts of the context hidden. One player, Merlin, hides parts such that the probability of the ground-truth answer increases; we trained the model to answer Merlin examples correctly. The other, Morgana, hides parts of the context which decreases the probability of the correct answer, and we trained the model to abstain on Morgana samples. We also trained with the original unchanged sample. To warm-start the RL training, we generated Merlin-Arthur rollouts for SFT and filtered out the data obtaining low rewards.
Supervised Fine-Tuning (SFT)
We fine-tuned for 4000 steps with a global batch size of 256 sequences of 262,144 tokens. We kept the optimizer split of pre-training, Muon for the two-dimensional backbone parameters and AdamW for everything else, and configured it at half the pre-training learning rates. The schedule warmed up over 50 steps and decays linearly over the last fifth of the run to a tenth of the peak.
Data Selection and Curation. We used different permissively available datasets as well as in-house generated data. We unified and aligned all datasets and decontaminated w.r.t. evaluation datasets and removed 0.006% of the rows. We filtered all of our data for political bias potentially inherited from generator models. In addition, we generated data centered on human dignity, universal human rights, liberal democracy, the rule of law and pluralism.
Mix Weighting. We adapted MergeMix to determine the relative weighting of our SFT data mix. We first validated the approach on math and code and then scaled the approach to 20 clusters. We trained one specialist per cluster on 4.5B packed positions, merged 69 weightings drawn from Dirichlet distributions around our hand-tuned mix, and scored every merge on 16 benchmarks in seven capabilities. Eight of these weightings we trained as data mixtures at proxy scale. Six beat the hand-tuned baseline, and the four best went into a test at target scale. Its weights were set to reach 1 epoch per cluster at 168B tokens, so the run saw the weighted mix about once.
Model Souping. To produce the final checkpoint we merged two checkpoints trained for the same token horizon on two different data mixtures. Their respective strengths carried over into the final model.
Reinforcement Learning (RL)
We trained in an asynchronous Reinforcement Learning setup, where vLLM inference ran in parallel to training. Updated model weights were synchronized to the inference engine without waiting for ongoing generations to finish. We kept the optimizer consistent with the SFT stage, but adapted the learning rate and additionally used quantization aware training to enable efficient low-precision inference with vLLM.
We trained for a total of 1000 steps with a total batch size of 2048 and a maximum sequence length of 256k tokens across all environment simultaneously. The different environments all contain verifiable rewards for the answer and additional format rewards to encourage general style and consistency. All tasks are heavily randomised over different harnesses, system prompts and languages to encourage generalisation. For hard tasks that the model can't solve, we additionally used on-policy self distillation to hint the model towards the correct solutions.
Responsible Use
Acknowledging the permissive nature of the license under which the weights of Kolibri are released, we nevertheless raise awareness and encourage users to refrain from engaging in unlawful activity. Kolibri should not be used for illegal or unlawful actions of any kind and with any illegal or unlawful content. This includes, in particular, prohibited practices according to Article 5 of Regulation (EU) 2024/1689 (EU AI Act) and other illegal activities such as engaging in terrorism, violence, human trafficking, illegal distribution of materials harmful to minors, sexual solicitation, harassment, discrimination, creating or promoting malicious code, any other criminal activities or activities risking death or harm, including those related to military or nuclear applications, and activities not in compliance with sanction regimes, technology export regulations, and other restrictions that may apply by law. Additionally, we ask users not to engage in any use of Kolibri which could constitute an infringement of any intellectual property rights, especially copyright. Kolibri should be used following ethical standards.
Risks and limitations
Note: The use of language models in high-stake environments, for critical decisions or to support a user's wellbeing should be performed with additional guardrails in place. All use of our model as part of a downstream AI system must meet all applicable regulatory and legal requirements, especially Regulation (EU) 2024/1689 (EU AI Act).
In the following sections, we describe risk categories and provide examples of completions we would consider inappropriate or harmful. We then describe steps to minimize these risks.
Harmful Language
Large language models can sometimes generate undesired outputs that are unsuitable for certain applications. This includes producing content with harmful language, discriminative content, inappropriate tone and style, or systemic biases. Our model has also not been optimized to represent a political opinion or take a specific point of view, and may generate outputs that contradict a user's opinion or expectation, including hateful or violent content. Such outputs can also include incorrect, outdated information, or material that is not suitable for all ages. While we constantly take efforts to reduce the likelihood of such undesired outputs, this possibility can never be fully ruled out. To minimize these issues, the following strategies can be employed:
- Abide by the guidance on Responsible Use provided for in this Model Card.
- Crafting prompts carefully to guide the model's output more effectively.
- Conducting additional validations at the application level to ensure output quality and appropriateness, e.g. via Red-Teaming or classifying the output.
Systemic Biases
Language models obtain world-knowledge from their pre-training data and may therefore exhibit the same systematic biases that are present in the data. Differing deployment scenarios (including differing cultural contexts) can expose systematic biases in different ways. We acknowledge the cultural diversity of communities and users worldwide.
Outdated World Knowledge
Our model's implicit knowledge reflects its training data cutoff (EN/DE June 18, 2026). Pretraining uses a fixed dataset compiled at a fixed point in the past, so the model's world knowledge is limited to what that data contained. This means, more recent events or facts may be unknown to it, or misunderstood if presented as input during live usage. This mainly matters for deployments without internet or tool access, since the model will otherwise use tool calling to work around potentially outdated information.
Risks include:
- Generation of unintended, irrelevant, or repetitive outputs. This includes the production of incorrect or outdated information.
Risks may be mitigated by:
- Injecting context, where relevant.
- Crafting prompts carefully to guide the model's output more effectively.
- Performing validations on the application layer, e.g., classifying the output.
- Using a repetition penalty or other parameters available in the API (see vLLM documentation).
Political Bias
We acknowledge the diversity of political contexts our model can be used in. Our model is not optimized to have a consistent political position across a large variety of political topics and may reproduce political biases present in its training data in some contexts.
Our model's training data also contains material generated with Chinese language models which are known to carry bias toward certain political positions (see Blog post). We actively reduced this type of political bias in our model, including through data filtering and dedicated alignment training. (See the Data Curation section for more details.)
Mistaken for a Human
Users may attribute human traits to AI models. This also includes the fact that content generated by the model is not explicitly detectable at this point. It is therefore required to design the system in a way that mitigates the impact of unintended interpretation of the output.
Other Errors
Any AI model can produce errors, even after implementing all legally required and additionally recommended measures. When integrating foundation language models into an application, users should:
- Be aware of the risk of (harmful) failure cases and implement the use case in a way that mitigates such risks.
- Be aware that foundation models do not contain application logic, e.g., content filters. Enforcement policies relevant to the use case need to be implemented in the application layer.
- Avoid unsupervised use in high-stakes environments.
- Validate output with adequate measures.
Reproducibility
Some inference parameters, e.g., temperature, lead to the random sampling of outputs, which precludes the reproducibility of outputs. Even when such parameters are not in use, outputs may diverge slightly on a numeric level for technical reasons. One may implement the following measures if needed:
- Logging of past model outputs on the application layer (Aleph Alpha is not storing any data and/or using any data provided in prompts for the training of its LLMs).
This list of risks, biases, and limitations may not be complete, as improving the understanding and behavior of language models is an ongoing research topic in the AI science community.
Public summary
| Summary about the content used for training (according to the European Commission's template) | aleph-alpha.com/downloads/data-summary.pdf |
|---|
License and terms
The model weights are published by Aleph Alpha GmbH under Apache 2.0 license. The rights granted thereunder only apply to the weights and configuration files published in this repository. For the avoidance of doubt, any other artifacts not included in the repository are excluded from the license. The license especially does not extend to underlying code, model architecture, parameter settings or any training method. Aleph Alpha retains all rights to its artifacts, code, model architecture, training methods, parameter settings and intellectual property rights.
Point of Contact for Rightsholders
Point of contact for rightsholders and their authorised representatives (Measure 1.5. General-Purpose AI Code of Practice): copyright-compliance@aleph-alpha.com
This model card was auto-generated by Savanna, our Model Factory, at commit 6f2108b924ac96a8e77e2536c4ef79dde0946b12.
- Downloads last month
- -
Model tree for Ishowbackup/Kolibri-1
Base model
Aleph-Alpha/Kolibri-1-BF16