Kolibri

Tech report | Tech blog

Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.

Model overview

Model Kolibri 1
Model Provider Aleph Alpha GmbH
Model Developer Aleph Alpha Research GmbH
Architecture Mixture-of-Experts
Total parameters 78B (78,103,074,560)
Active parameters / token 3.46B (3,457,573,120)
Languages German, English
Context length 1,048,576 tokens; we recommend ≤262,144 tokens for serving efficiency and complex tasks
Precision float8_e4m3fn weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16
Reasoning mode Yes
Tool calling Yes
License Apache 2.0
Knowledge cutoff EN: June 18, 2026, DE: June 18, 2026
This only affects implicit knowledge, the model may use more recent information through tool use.
Hardware requirements Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300.
Best for Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant
Release Date 3rd of October 2026
Code of Practice Aleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai.
Training Data Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.
Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases.
Training method We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed.
Computing Resources Pre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh)
Mid-training: 5 days, 90k GPUh (same setup as above)
Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6)
FLOPS: 6.4e23
Tech Report https://aleph-alpha.com/downloads/tech-report.pdf

Intended use

Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.

Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.

For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.

Design goals

Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.

AI system types

Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.

Sustainability

Energy consumption

9.5×10² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.

Energy measurement

For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.

Getting started

Kolibri requires the aleph-alpha-inference package that provides the Kolibri vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from Aleph-Alpha/aleph-alpha-inference, which also installs the vLLM version it supports:

pip install 'aleph-alpha-inference>=1'

Serve the model with reasoning and tool-calling enabled:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.

The recommended sampling parameters for the model are temperature=1.0, top_p=0.97 and top_k=128.

Querying the server (OpenAI-compatible API)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[
        {
            "role": "user",
            "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist.",
        },
    ],
    extra_body={
        "chat_template_kwargs": {
            "reasoning_effort": "high",
            "enable_thinking": True,
        }
    },
)
print(response.choices[0].message.content)

Reasoning mode

Kolibri supports explicit thinking effort levels that need to be configured through the chat template. You can pass reasoning_effort values low, medium and high to configure the amount of effort our model puts into finding the answer. You can disable thinking altogether by setting reasoning_effort to none or passing enable_thinking=false. The model will then immediately respond.

Tool calling

The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the standard tools field of the chat completions request; the model will emit tool calls that the parser converts into structured output. Tool calling can be combined with reasoning mode.

import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {
                        "type": "string",
                        "description": "City name, e.g. Heidelberg",
                    },
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
                "required": ["city"],
            },
        },
    }
]


def get_weather(city, unit="celsius"):
    return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}


messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]

first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)

for call in msg.tool_calls or []:
    args = json.loads(call.function.arguments)
    result = get_weather(**args)
    messages.append(
        {
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        }
    )

final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)

Evaluation

The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.

Post-training

TypeMoEDense
Active parameters3B4-6B12B27B70B
OursBaseline models
Eval
Kolibri
Kolibri Origin
GLM-4.7 Flash 30B-A3B
Nemotron 3 Nano 30B-A3B
Qwen3.5 35B-A3B
Qwen3.6 35B-A3B
Qwen3-Next 80B-A3B Thinking
Gemma 4 26B-A4B IT
GPT-OSS 120B
Mistral Small 4 119B-A6B
GLM-4.5 Air 106B-A12B
Nemotron 3 Super 120B-A12B
Qwen3.8 27B
Apertus 70B Instruct
Overall (EN)75.554.164.765.674.771.462.471.972.363.164.473.080.2–
Overall (DE)70.846.450.459.369.867.358.066.370.261.464.867.979.9–
Knowledge
Average (EN)50.139.745.546.052.752.148.451.450.047.545.752.056.822.8
Average (DE)57.644.946.541.261.361.055.761.558.051.452.259.569.224.8
GPQA Diamond (EN)84.368.173.173.983.883.476.181.176.474.773.278.089.229.5
GPQA Diamond (DE)81.358.559.849.684.280.672.280.176.072.971.176.688.131.4
Humanity's Last Exam (EN)21.59.415.412.120.421.111.619.219.49.78.720.635.65.2
Humanity's Last Exam (DE)15.910.49.113.118.120.515.623.420.710.510.522.337.25.7
AA-Omniscience Accuracy (public set)14.811.317.019.522.221.024.220.723.325.020.026.719.013.5
AA-Omniscience Index (public set)-32.8-64.2-62.8-45.7-46.2-12.5-42.3-47.3-35.2-24.0-28.8-36.5-5.8–
MMLU-Pro CoT (EN)80.070.176.578.384.684.381.784.580.880.480.982.785.043.0
MMLU-ProX CoT (DE)75.565.770.761.081.781.979.481.177.270.774.979.782.437.3
Math
Average (EN)96.581.788.888.890.187.886.387.490.781.482.891.197.80.6
Average (DE)88.874.345.284.379.683.783.788.190.875.480.686.596.70.1
AIME 2025 (EN)96.981.989.489.688.184.684.287.390.679.881.991.797.90.6
AIME 2025 (DE)87.573.543.884.476.782.980.688.190.672.380.685.696.50.2
AIME 2026 (EN)96.081.588.387.992.191.088.587.590.883.183.890.497.70.6
AIME 2026 (DE)90.075.246.784.282.584.486.788.191.078.580.687.596.90.0
Agentic
Average (EN)63.441.658.946.463.462.146.354.654.040.753.554.966.7–
TerminalBench 2.127.7–20.29.739.7–8.6–29.221.0–39.776.8–
Tau2-Bench (Telecom)94.767.595.945.997.799.143.945.373.141.553.868.182.510.8
Tau2-Bench (Retail)69.958.557.964.970.871.660.871.360.562.961.467.568.79.6
Tau2-Bench (Airline)76.758.768.752.776.070.765.373.372.740.070.772.783.340.0
Tau3-Bench (Banking)38.15.77.25.711.310.65.416.014.75.76.415.550.02.1
BFCL v3 (multi-turn)39.822.858.247.954.053.551.453.445.636.261.644.642.50.6
BFCL v4 (overall)61.436.465.461.570.567.251.068.257.358.067.261.073.2–
BFCL v4 (non-live AST)79.178.183.385.085.888.283.683.735.883.685.545.085.3–
BFCL v4 (live)78.973.778.378.880.281.482.580.270.478.478.277.679.9–
BFCL v4 (multi-turn)47.527.562.753.559.958.156.061.455.440.465.251.755.5–
BFCL v4 (memory)62.819.441.539.162.653.835.352.950.739.143.459.679.6–
BFCL v4 (web search)62.510.569.066.075.068.512.575.057.069.071.071.582.0–
BrowseComp29.44.4–14.536.526.92.825.531.2––29.146.4–
Code
Average (EN)89.368.067.881.885.087.783.689.090.882.079.888.394.225.1
LiveCodeBench v685.959.246.571.377.882.573.982.387.571.267.882.093.88.7
HumanEval+92.776.889.092.492.292.893.395.794.192.891.894.794.741.6
SWE-Bench Verified66.4–51.038.671.673.8–57.8–60.811.660.272.6–
Instruction Following
Average (EN)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
IFBench (loose-prompt)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
Grounding / Hallucinations
Average (EN)59.442.757.655.967.568.361.062.562.661.563.864.467.531.5
SQuAD (M/A Grounding Score)23.40.01.50.023.012.10.00.00.06.30.00.08.30.0
SQuAD (Utility Accuracy)83.477.482.775.089.888.973.588.474.276.881.980.885.61.6
RGB Closed-Book51.052.078.080.081.079.089.079.085.086.092.093.073.085.0
RGB Negative (Abstention)85.673.979.676.986.679.681.386.079.682.389.074.670.655.9
AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set)44.015.03.819.011.156.712.314.323.734.739.013.967.316.4
RGB Fact-Check (Error Correction)34.014.068.058.074.074.085.061.077.060.074.090.053.014.0
FRAMES (<24k)71.265.767.870.975.974.768.969.578.371.970.474.978.658.9
FRAMES (>24k)73.0–73.974.381.578.868.576.678.476.6–78.883.3–
SealQA (no distractors, <24k)80.756.671.769.088.382.882.182.180.773.167.682.886.935.9
SealQA (12 distractors, <24k)61.030.065.054.078.067.057.082.065.062.060.070.084.016.0
SealQA (no distractors, >24k)100.0–61.166.7100.088.983.377.883.377.8–94.494.4–
SealQA (12 distractors, >24k)65.1–57.139.773.071.450.868.363.549.2–68.382.5–
Agentic Retrieval
Average (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Average (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
MuSiQue (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Honeypot80.825.369.460.277.574.313.575.466.668.1–68.885.0–
Agentic Wiki QA (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
Industry RAG
Average (EN)89.753.875.961.387.086.062.779.482.874.982.380.393.323.9
Average (DE)67.542.760.946.270.065.831.149.464.253.463.457.680.214.4
Semiconductors80.435.364.739.279.479.441.263.770.662.772.569.689.211.8
German Public Sector75.054.069.565.580.072.029.566.577.050.069.578.089.08.0
Aerospace58.914.150.444.162.359.048.158.148.147.0–54.973.8–
Automotive Supplier99.072.487.283.394.692.684.295.195.087.192.091.097.335.9
Industrial Drive Technology60.031.452.326.860.059.532.732.351.456.857.337.371.420.9
Long Context
LongBench Pro64.5––53.270.370.863.864.3–56.4–62.976.9–
AA-LCR68.3–49.742.366.369.746.368.3–52.3–67.081.3–

All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort high. A model reports no score (–) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.

Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.

Pre-training

All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.

TypeMoEDense
Active parameters3-4B12B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
GLM-4.5 Air 106B-A12B Base
Nemotron 3 Super 120B-A12B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
Overall (EN)81.158.458.177.573.877.183.157.667.948.6
Overall (DE)81.561.161.176.076.679.485.048.065.250.5
General Knowledge
Average (EN)81.370.877.280.582.782.587.267.877.873.8
Average (DE)82.770.780.580.084.981.787.551.868.774.7
MMLU (EN)81.068.978.279.084.782.886.867.176.369.3
Global MMLU (DE)77.565.175.174.381.578.284.751.665.164.9
ARC (EN)95.788.995.294.397.096.397.588.694.590.7
ARC (DE)96.488.595.594.297.695.997.867.988.189.9
PIQA (EN)89.576.988.089.491.989.895.177.786.880.1
PIQA (DE)96.887.996.896.598.597.299.375.090.094.6
HellaSwag (EN)83.180.784.885.485.387.288.876.183.584.4
HellaSwag (DE)86.975.388.685.888.786.593.036.861.088.7
MMLU-Pro (EN)61.139.351.255.662.655.368.938.050.940.6
MMLU-ProX (DE)55.736.546.749.058.450.862.827.939.535.6
TriviaQA (EN)77.269.965.679.574.683.886.459.074.777.4
Wahl-O-Mat (DE)16.752.357.155.943.256.761.149.656.357.0
Math
Average (EN)84.954.345.983.073.564.782.757.763.339.6
Average (DE)76.855.847.072.577.569.881.243.862.939.0
GSM8K (EN)89.871.665.786.389.582.887.575.381.162.3
GSM8K Platinum (DE)90.571.764.186.388.086.593.158.780.860.2
MATH Minerva (EN)80.137.026.179.857.646.577.840.145.616.9
MATH Minerva (DE)63.239.929.958.867.053.269.328.945.017.7
Code
Average (EN)77.050.151.169.065.284.179.347.462.532.5
Average (DE)85.157.055.775.567.486.786.248.364.037.9
HumanEval (EN)86.750.852.273.567.296.383.246.863.928.3
HumanEval (DE)88.449.051.773.559.294.085.039.556.228.3
MBPP (EN)67.349.550.064.563.171.975.448.061.036.8
MBPP (DE)81.764.959.777.475.679.487.557.171.747.5

All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters temperature = 0.6, top_p = 0.6, max_tokens = 1024 and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommended temperature = 1.0, top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).

Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.

Long context

TypeMoEDense
Active parameters3-4B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
RULER
4k86.991.884.794.696.392.094.889.5
8k83.788.885.093.495.380.892.677.7
16k80.983.787.192.095.270.888.871.3
32k76.372.788.786.493.764.480.970.8
64k72.2–86.882.691.0–––
128k67.9–87.881.089.9–––
256k69.8––72.180.1–––
512k65.5––73.172.2–––
1M63.2––58.557.5–––
HELMET
8k79.174.881.979.187.771.174.864.2
16k77.576.481.980.885.567.173.857.2
32k80.073.582.978.484.160.174.648.1
64k81.2–83.872.778.851.973.1–
128k82.4–85.470.877.3–––

Each RULER row is the mean of its NIAH, VT, WE and QA tasks at that context length, sampled with temperature = 0. A model reports no score (–) at lengths beyond its context window. The models with a 65,536-token window report none at 64k either: its prompts slightly exceed that under some tokenizers. Gemma 4 26B-A4B Base reports none at 256k: with the answer budget, its prompts exceed its 262,144-token window. At 512k and 1M, Kolibri Base and Qwen3.5 35B-A3B Base are served beyond their 262,144-token trained window, Qwen3.5 with static YaRN (factor 4) at every length; Nemotron 3 Nano 30B-A3B Base supports up to 1M.

The HELMET rows are the mean of four HELMET tasks at that length: JSON KV and RAG QA scored by substring exact match, RULER MK3 by recall, and InfBench multiple choice by exact match. Apertus 70B Base and Kolibri Origin Base report none at 64k: some RAG QA prompts exceed their 65,536-token window.

Training Details

Model Dependencies

None. The model was trained from scratch.

Model Architecture

Model detail Value
Vocabulary size 128,000
Number of layers 50
Attention type 4:1 causal sliding-window GQA to full causal GQA
Sliding window 512 preceding tokens plus the current token
Hidden size 2,560
Attention heads 48
Key-value heads 4
Head size 128
QK normalization Per-head RMSNorm
Expert hidden size 512
MLP type SwiGLU
Routed experts per MoE layer 384
Experts selected per token 6
Shared experts per MoE layer 1
Routing Token-choice, top-6 sigmoid routing
Pretraining sequence length 16,384
Maximum context length 262,144 native; 1,048,576 by extrapolation
Position embeddings RoPE, base 10,000, applied only in sliding-window layers

Pre-training

Our tokenizer was trained on our pre-training mix with a vocabulary size of 128,000. With 23.9% of German share in our data, it achieves higher German compression (4.7 bytes/token) than leading models with larger vocabulary, without sacrificing English efficiency (4.2 bytes/token). We trained the tokenizer with a new algorithm, UniBPE, that respects the morphology of languages better than existing approaches, especially the compound structure of German. UniBPE keeps the greedy bottom-up approach of BPE and using the Unigram training objective for selecting which merge to add to the vocabulary.

We randomly initialized all model parameters and pretrained the model with a causal next-token-prediction objective on a large and diverse document corpus described above. Training examples consisted of 16,384-token sequences, with multiple documents potentially packed into a single sequence. The pretraining phase covered 20T tokens over 264,750 optimization steps. We used a global batch size of 4,608 sequences, corresponding to 75.497 million tokens per step, with two gradient-accumulation steps across 768 GPUs.

We linearly warmed up the learning rate over approximately 100 billion tokens, corresponding to 1,325 optimization steps, and then held it constant for the remaining 263,425 steps. AdamW was used for the embedding matrix, one-dimensional backbone parameters, router weights, and language-model head. All other two-dimensional backbone parameters were optimized with Nesterov Muon using a learning rate of 0.001.

The expert-balancing bias was updated separately using a recentered global-quantile target update. We additionally injected load-error gradients into the router scores with a weight of 1e-5 and a soft-clamp value of 1.0.

We applied independent weight decay of 2⁻¹², approximately 0.0002441, to all decay-eligible parameters. Embedding, normalization, and expert-balancing-bias parameters were excluded from weight decay. We used per-head RMS normalization of the query and key representations.

The mid-training phase increased the sequence length to 65,536 tokens and used a global batch size of 1,536 sequences, corresponding to 100.663 million tokens per step. Its 3.44 trillion token budget corresponds to 34,200 optimization steps. Because this phase resumed the pretraining optimizer state, it used no additional learning-rate warmup.

The subsequent long-context phase increased the sequence length to 262,144 tokens and uses a global batch size of 768 sequences, corresponding to 201.327 million tokens per step. Its 201-billion-token budget corresponds to 1000 optimization steps, again without additional warmup.

Training was conducted on 768 GPUs using our PyTorch-based TorchTitan distributed training infrastructure, with bfloat16 parameters, float32 reductions, and a maximum gradient norm of 1.0.

Pre-training data sources

Kolibri was trained on a filtered, bilingual (German/English) corpus combining curated web data, synthetic rephrasings and translations, and high-quality curated sources. We trained on 20T tokens, followed by 3.44T tokens in a mid-training stage and 201B tokens in a long-context adaptation stage.

Pre-training and mid-training data mix

Domain Pre-training Mid-training Long-context extension
English web and documents 43.4% 1.3% 15.8%
German web and documents 23.4% 2.1% 2.6%
English instruction and reasoning 11.9% 31.4% 11.3%
German instruction and reasoning 0.5% 1.9% 0.7%
Code 13.6% 16.3% 13.2%
Agentic code and tool use ~0% 15.4% 5.6%
STEM (documents) 4.5% 5.9% 6.7%
STEM (QA-style) 1.8% 25.4% 8.9%
Specialised Domains (Legal, Public Sector) 1.1% 0.2% 1.9%
OCR'd PDFs 0% 0% 33.3%

Note: Rows do not sum to 100% due to rounding.

Synthetic data

We generated synthetic data using permissively-licensed LLMs.

English pre-training rephrases. Deduplicated English Common Crawl was rephrased with Gemma-4-26B-A4B following a Nemotron-CC-style recipe, using five prompt templates applied to each seed document: diverse QA pairs, distillation, knowledge list, knowledge extraction, and encyclopedia-style rewriting. Generation ran as a background job on 912 GPUs over the substring-deduplicated English corpus. The outputs were post-processed, shuffled, indexed, and included in the pre-training mix.

German pre-training rephrases. German synthetic rephrasings of web data were produced with Mistral-NeMo-12B.

Synthetic annotations. LLM-as-a-judge annotations using Qwen3-32B over randomly sampled English Common Crawl were generated as training data for our text-quality classifiers; see Data curation.

Data curation

The pre-training corpus was built from extracted Common Crawl snapshots. We applied a range of curation techniques, including but not limited to:

  • URL filtering. To prevent training on illegal, harmful, infringing or pirated content, we filtered our data based on a URL filter list.
  • Heuristic filtering and cleanup. Web documents passed a heuristic filter pipeline inspired by the approaches described in Dolma 3, DCLM and RefinedWeb: Gopher document and quality filters (length, symbol ratio, bullet and ellipsis lines, stop words, boilerplate, repeated lines and n-grams), fastText language identification routing documents into an English and a German branch with language-specific thresholds, RefinedWeb line-wise removal of navigation and counter lines, and, for English, the MADLAD-400 questionable-content rules. Unlike these filters, which apply to web data only, a final cleanup pass ran over every dataset in the pool: it collapsed runs of blank lines (at three newlines for code and math sources, two for all others) and capped horizontal whitespace runs, which removed extreme padding such as long runs of spaces. Filter statistics were recorded per filter, language, and source, for development purposes, e.g., understanding how much each step discards and tuning heuristics appropriately.
  • Deduplication and shuffling. The Common Crawl corpus passed through exact/global (cross-dump) deduplication, MinHash fuzzy deduplication, and substring deduplication. When joining the data sources in our pre- and mid-training data pools, we globally shuffled the document indices. For mid-training, we additionally performed exact deduplication per dataset and also globally across all datasets.
  • Quality classification. fastText and Luxical classifiers were trained both on LLM-as-a-judge annotations of sampled Common Crawl and on human-curated datasets. For English, the selected annotators consisted of 5 classifiers which had non-linear interaction terms. For German, we took the mean of the GermanWeb educational quality and grammar fastText model scores. Combined classifier scores were rescaled to [0, 1] in the annotation pipeline, and each language corpus was sorted into 20 equal-token-mass quantile buckets, grouped into five categories (low, medium-low, medium, medium-high, high) to enable quality-aware mixing. Bucket breakpoints were computed from per-shard equal-token-mass quantile documents.
  • PII redaction. We removed personally identifiable information from pre-training data using regular expressions, replacing high-syntax-constraint entity types with special tokens, such as email addresses (<|pii-email-address|>), IP-addresses (<|pii-ip-address|>) and others. Documents in which more than 15% of bytes were replaced were dropped entirely. An ablation comparing redacted and unredacted versions of the same 750M-row sample showed no clear performance difference in either direction.
  • Harmful and unsuitable content. Illegal, harmful, unethical and infringing sources are excluded through the URL-level filtering described above, which also removed known adult and NSFW sites. For English web data, keyword-based filtering additionally removed sexually explicit content. These measures partially address child sexual abuse material (CSAM); no dedicated CSAM detection was applied. Personal data was redacted as described under PII redaction. The training corpus was text-only, so image-based CSAM or non-consensual intimate imagery (NCII) was not ingested.
  • Decontamination. Documents in our mid-training and long-context extension datasets were scored for contamination by n-gram overlap against our evaluation suite, and contaminated documents were dropped in full. Measured contamination rates across the mid-training pool were low, between 0 and roughly 1e-4 of documents per dataset. RULER and MRCR were excluded from the reference set, as their synthetic needle-in-a-haystack construction made overlap checks uninformative.
  • Bias mitigation. German is under-represented in web content, which lowers German performance. We counteracted this by upsampling German data where downstream evaluations showed gains, and by adding translated and synthetically rephrased German data to diversify style and register. Openly available fine-tuning data is often generated by models that carry political bias; we built a dedicated post-training dataset to counteract risk of contamination of our training data with politically biased material, grounded in curated reference material on politically sensitive topics.
  • Mix selection. Pre- and mid-training mixes were chosen empirically using mix-search proxy models trained with different mixes to fit a function which then predicts the best mix based on our eval suite. The predicted mixes were subsequently confirmed using a separate validation proxy setup. In pre-training, this involved training thousands of small dense proxy models (30M params) with different mixes (3B tokens). For mid-training, a dedicated base proxy model was trained to match the tokens-per-parameter of the target scale. We then used the method described in MergeMix (weight-averaging of domain expert continued pre-trainings over nine domains) to find promising mixture candidates.
  • Long-context. Long-context candidates were built by length-bucketing the mid-training pool and blending it with OCR'd PDFs. We included documents up to a sequence length of 256k tokens. We fixed the OCR'd PDFs at 34% of tokens, x% the optimal mid-training mix, and (66-x)% the mid-training pool weighted towards long-context documents. We ablated over multiple candidate mixes for different values of x and found the best mix to be at x=33.

Post-training

Post-training was split into two phases: A supervised fine-tuning stage that instilled reasoning and instruction following capabilities and a reinforcement learning stage that fine-tuned the long-context behavior across a broad set of environments. Across both stages we put particular focus on the key capabilities that shape the strengths of Kolibri.

Capabilities

German. We trained our model to reason in German to make it easier for our users to follow along. For the SFT side we generated translated prompts and generated language consistent reasoning answers and filtered the answers for correctness and language consistency to obtain 10.6B tokens. We continued this training in RL, where we used German environments and included a language consistency reward for both the reasoning trace and the answer.

Reasoning Effort. Kolibri can answer directly (reasoning effort none) or reason at low, medium or high reasoning effort, which the chat template requests through a fixed sentence in the system prompt. In SFT, we assigned reasoning effort labels based on observed dataset-specific reasoning length statistics and judged difficulty, while on RL side we trained each task across all reasoning levels with different length penalties.

Safety. Kolibri was trained in SFT to decline harmful and illegal requests while still providing partial answers where appropriate. The safeguards trained into the model do not replace safeguards at the solution level. When integrating Kolibri into an application, these should be supplemented with additional appropriate measures, such as content filtering and output validation.

Personally Identifiable Information. In addition to the mitigations during pre-training, we incorporated safety data in our SFT mix that provide refusals to PII extraction attacks. Refusals on such behavior were measured as part of our safety evaluations.

Long Context. We trained Kolibri on context windows up to 256k tokens, which enables reasoning over large documents. Since most datasets usually focus on short-context, we curated datasets that contained long-context documents based on publicly available data (e.g., from German legislative documents) with corresponding question/answering prompts. We similarly trained on long-context question-answer during RL.

Agentic Use and Tool Calling. We trained Kolibri to perform well in agentic settings involving tool calls and terminal interactions. For permissively available agentic SFT datasets, we regenerated the reasoning traces. During training, we masked out erroneous tool calls and terminal actions, so the model learned to recover from errors without learning the erroneous actions themselves. In addition to the open data, we generated two software-engineering SFT datasets: one for producing code patches and one for bug fixing as a Terminus 2 agent. During the RL stage, we trained on multiple tool-calling, software-engineering and terminal environments and randomized the harnesses to support generalization.

Retrieval. To make our model suited for customer information retrieval use-cases, we built environments that cover retrieval-augmented question answering over German and English corpora across general knowledge and specialised domains such as legal, electronics, hardware and aviation. The environments were heavily randomised across different harnesses, retrieval parameters, and languages. To warm-start the RL training phase we generate a curated set of 11k high-reward completions for SFT training.

Hallucinations. To address hallucinations, we trained with abstention data. We also constructed RL environments which implement our Merlin-Arthur protocol. The protocol shows the model each question with parts of the context hidden. One player, Merlin, hides parts such that the probability of the ground-truth answer increases; we trained the model to answer Merlin examples correctly. The other, Morgana, hides parts of the context which decreases the probability of the correct answer, and we trained the model to abstain on Morgana samples. We also trained with the original unchanged sample. To warm-start the RL training, we generated Merlin-Arthur rollouts for SFT and filtered out the data obtaining low rewards.

Supervised Fine-Tuning (SFT)

We fine-tuned for 4000 steps with a global batch size of 256 sequences of 262,144 tokens. We kept the optimizer split of pre-training, Muon for the two-dimensional backbone parameters and AdamW for everything else, and configured it at half the pre-training learning rates. The schedule warmed up over 50 steps and decays linearly over the last fifth of the run to a tenth of the peak.

Data Selection and Curation. We used different permissively available datasets as well as in-house generated data. We unified and aligned all datasets and decontaminated w.r.t. evaluation datasets and removed 0.006% of the rows. We filtered all of our data for political bias potentially inherited from generator models. In addition, we generated data centered on human dignity, universal human rights, liberal democracy, the rule of law and pluralism.

Mix Weighting. We adapted MergeMix to determine the relative weighting of our SFT data mix. We first validated the approach on math and code and then scaled the approach to 20 clusters. We trained one specialist per cluster on 4.5B packed positions, merged 69 weightings drawn from Dirichlet distributions around our hand-tuned mix, and scored every merge on 16 benchmarks in seven capabilities. Eight of these weightings we trained as data mixtures at proxy scale. Six beat the hand-tuned baseline, and the four best went into a test at target scale. Its weights were set to reach 1 epoch per cluster at 168B tokens, so the run saw the weighted mix about once.

Model Souping. To produce the final checkpoint we merged two checkpoints trained for the same token horizon on two different data mixtures. Their respective strengths carried over into the final model.

Reinforcement Learning (RL)

We trained in an asynchronous Reinforcement Learning setup, where vLLM inference ran in parallel to training. Updated model weights were synchronized to the inference engine without waiting for ongoing generations to finish. We kept the optimizer consistent with the SFT stage, but adapted the learning rate and additionally used quantization aware training to enable efficient low-precision inference with vLLM.

We trained for a total of 1000 steps with a total batch size of 2048 and a maximum sequence length of 256k tokens across all environment simultaneously. The different environments all contain verifiable rewards for the answer and additional format rewards to encourage general style and consistency. All tasks are heavily randomised over different harnesses, system prompts and languages to encourage generalisation. For hard tasks that the model can't solve, we additionally used on-policy self distillation to hint the model towards the correct solutions.

Responsible Use

Acknowledging the permissive nature of the license under which the weights of Kolibri are released, we nevertheless raise awareness and encourage users to refrain from engaging in unlawful activity. Kolibri should not be used for illegal or unlawful actions of any kind and with any illegal or unlawful content. This includes, in particular, prohibited practices according to Article 5 of Regulation (EU) 2024/1689 (EU AI Act) and other illegal activities such as engaging in terrorism, violence, human trafficking, illegal distribution of materials harmful to minors, sexual solicitation, harassment, discrimination, creating or promoting malicious code, any other criminal activities or activities risking death or harm, including those related to military or nuclear applications, and activities not in compliance with sanction regimes, technology export regulations, and other restrictions that may apply by law. Additionally, we ask users not to engage in any use of Kolibri which could constitute an infringement of any intellectual property rights, especially copyright. Kolibri should be used following ethical standards.

Risks and limitations

Note: The use of language models in high-stake environments, for critical decisions or to support a user's wellbeing should be performed with additional guardrails in place. All use of our model as part of a downstream AI system must meet all applicable regulatory and legal requirements, especially Regulation (EU) 2024/1689 (EU AI Act).

In the following sections, we describe risk categories and provide examples of completions we would consider inappropriate or harmful. We then describe steps to minimize these risks.

Harmful Language

Large language models can sometimes generate undesired outputs that are unsuitable for certain applications. This includes producing content with harmful language, discriminative content, inappropriate tone and style, or systemic biases. Our model has also not been optimized to represent a political opinion or take a specific point of view, and may generate outputs that contradict a user's opinion or expectation, including hateful or violent content. Such outputs can also include incorrect, outdated information, or material that is not suitable for all ages. While we constantly take efforts to reduce the likelihood of such undesired outputs, this possibility can never be fully ruled out. To minimize these issues, the following strategies can be employed:

  • Abide by the guidance on Responsible Use provided for in this Model Card.
  • Crafting prompts carefully to guide the model's output more effectively.
  • Conducting additional validations at the application level to ensure output quality and appropriateness, e.g. via Red-Teaming or classifying the output.

Systemic Biases

Language models obtain world-knowledge from their pre-training data and may therefore exhibit the same systematic biases that are present in the data. Differing deployment scenarios (including differing cultural contexts) can expose systematic biases in different ways. We acknowledge the cultural diversity of communities and users worldwide.

Outdated World Knowledge

Our model's implicit knowledge reflects its training data cutoff (EN/DE June 18, 2026). Pretraining uses a fixed dataset compiled at a fixed point in the past, so the model's world knowledge is limited to what that data contained. This means, more recent events or facts may be unknown to it, or misunderstood if presented as input during live usage. This mainly matters for deployments without internet or tool access, since the model will otherwise use tool calling to work around potentially outdated information.

Risks include:

  • Generation of unintended, irrelevant, or repetitive outputs. This includes the production of incorrect or outdated information.

Risks may be mitigated by:

  • Injecting context, where relevant.
  • Crafting prompts carefully to guide the model's output more effectively.
  • Performing validations on the application layer, e.g., classifying the output.
  • Using a repetition penalty or other parameters available in the API (see vLLM documentation).

Political Bias

We acknowledge the diversity of political contexts our model can be used in. Our model is not optimized to have a consistent political position across a large variety of political topics and may reproduce political biases present in its training data in some contexts.

Our model's training data also contains material generated with Chinese language models which are known to carry bias toward certain political positions (see Blog post). We actively reduced this type of political bias in our model, including through data filtering and dedicated alignment training. (See the Data Curation section for more details.)

Mistaken for a Human

Users may attribute human traits to AI models. This also includes the fact that content generated by the model is not explicitly detectable at this point. It is therefore required to design the system in a way that mitigates the impact of unintended interpretation of the output.

Other Errors

Any AI model can produce errors, even after implementing all legally required and additionally recommended measures. When integrating foundation language models into an application, users should:

  • Be aware of the risk of (harmful) failure cases and implement the use case in a way that mitigates such risks.
  • Be aware that foundation models do not contain application logic, e.g., content filters. Enforcement policies relevant to the use case need to be implemented in the application layer.
  • Avoid unsupervised use in high-stakes environments.
  • Validate output with adequate measures.

Reproducibility

Some inference parameters, e.g., temperature, lead to the random sampling of outputs, which precludes the reproducibility of outputs. Even when such parameters are not in use, outputs may diverge slightly on a numeric level for technical reasons. One may implement the following measures if needed:

  • Logging of past model outputs on the application layer (Aleph Alpha is not storing any data and/or using any data provided in prompts for the training of its LLMs).

This list of risks, biases, and limitations may not be complete, as improving the understanding and behavior of language models is an ongoing research topic in the AI science community.

Public summary

Summary about the content used for training (according to the European Commission's template) aleph-alpha.com/downloads/data-summary.pdf

License and terms

The model weights are published by Aleph Alpha GmbH under Apache 2.0 license. The rights granted thereunder only apply to the weights and configuration files published in this repository. For the avoidance of doubt, any other artifacts not included in the repository are excluded from the license. The license especially does not extend to underlying code, model architecture, parameter settings or any training method. Aleph Alpha retains all rights to its artifacts, code, model architecture, training methods, parameter settings and intellectual property rights.

Point of Contact for Rightsholders

Point of contact for rightsholders and their authorised representatives (Measure 1.5. General-Purpose AI Code of Practice): copyright-compliance@aleph-alpha.com

This model card was auto-generated by Savanna, our Model Factory, at commit 6f2108b924ac96a8e77e2536c4ef79dde0946b12.

Downloads last month
-
Safetensors
Model size
78B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 2 Ask for provider support

Model tree for Aleph-Alpha/Kolibri-1

Quantized
(4)
this model
Quantizations
6 models

Space using Aleph-Alpha/Kolibri-1 1

Collection including Aleph-Alpha/Kolibri-1

Papers for Aleph-Alpha/Kolibri-1