Haiku
Reasoning + tool-call chat model (~655M), preference-tuned with APO
A ~655M hybrid Kimi Delta Attention + gated MLA model: 32k-context SFT on reasoning, tool and chat data, then SmolLM3-style APO
What this is
Haiku is the chat model of the Haiku family, the ~655M successor to Tercet-R.
- Supervised fine-tuning at a 32,768-token sequence length on ~9.5B tokens of reasoning, tool-call, instruction-following and chat data
- Then APO-zero preference optimization on the SmolTalk2 preference split, following the SmolLM3 recipe
- Live demo:
kerzgrr/haiku-demo— Haiku chat plus Haiku-base raw continuation in one Space - Hub weights are the EMA snapshot in bfloat16
This upload is the final checkpoint of the preference run (optimizer step 5,155, one full epoch of 164,948 pairs).
The SFT started from the v2 Haiku pretrain (step 4,800, 2.5B FineWeb-Edu tokens), which is not released separately. The public kerzgrr/Haiku-base is the earlier v1 pretrain of the same architecture and tokenizer (step 8,400, 4.4B tokens).
Chat contract
Thinking
Each assistant turn is prefixed with a zero-loss control token:
| Mode | Prefix | Typical body |
|---|---|---|
| think | <|think|>\n |
<think>…</think> then the answer |
| no-think | <|no_think|>\n |
answer only |
inference.py streams the <think> region live (dim yellow) and hides the control tokens.
Tool calls (SmolTalk JSON)
Tool schemas go in a Hermes-style <tools> block in the system turn. The model calls them with:
<tool_call>
{"name": "web-search", "arguments": {"query": "…"}}
</tool_call>
SFT covered Nemotron Agentic v1/v2 (including web-search), Toucan 1.5M, Hermes reasoning tool-use and Nemotron PTD v1 tool calling. inference.py auto-runs:
| Built-in | Tool name | Observation |
|---|---|---|
--tools web-search |
web-search |
Tavily-shaped JSON (Tavily if TAVILY_API_KEY is set, else DuckDuckGo + Wikipedia) |
--tools python |
stateful_python_code_exec |
Jupyter-style stdout / last value from a restricted interpreter |
--tools calculator |
calculator |
Numeric result of a math expression |
Tool results
Each observation is a tool (or user) turn prefixed with:
<|tool_response|>
{observation}
Install & run
pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Haiku inference.py --local-dir .
python inference.py
python inference.py --prompt "What is the capital of France?"
python inference.py --tools web-search,python,calculator
python inference.py --no-think --prompt "Reply in one sentence."
inference.py auto-downloads weights / tokenizer / tiny_gdn/ and auto-installs pinned flash-linear-attention (plus transformers, which its decode cache imports). Git is required on PATH. A CUDA GPU is required: the Kimi Delta Attention layers run on Triton kernels.
| Flag | Default | Description |
|---|---|---|
--prompt |
— | One-shot user message |
--system |
— | System prompt, used verbatim |
--think / --no-think |
think | Assistant control prefix |
--tools |
— | Built-ins: web-search, python, calculator (comma-separated) |
--temperature |
0.7 |
Sampling temperature |
--max-new-tokens |
4096 |
Max generation length |
--context-length |
32768 |
Conversation tokens kept (oldest dropped first) |
--device |
cuda |
CUDA device, e.g. cuda:1 |
Interactive commands: /think /no_think /system … /reset /exit. Ctrl+C stops the current reply and keeps the session.
Model architecture
Same Haiku hybrid as Haiku-base (655,270,488 parameters):
| Layers | 36 (Kimi Delta Attention ×3 + gated MLA every 4th) |
| Hidden | 1,024 |
| MLP | SiTU-GLU 3,840 |
| Attention | Gated MLA (NoPE), 8 heads, Q LoRA rank 512, KV LoRA rank 256 |
| Linear | Kimi Delta Attention, 8 heads × 128 |
| Residuals | Block attention residuals |
| Vocab | 65,536 BPE |
| Context | 32,768 (SFT sequence length) |
Training
| Stage | Details |
|---|---|
| Pretrain (v2) | FineWeb-Edu, seq 2,048, 4,800 steps = 2.52B tokens, hybrid Muon + AdamW at peak 5×10⁻⁴, 19.2 hours. Val loss (EMA) 2.90 |
| SFT (v3) | Natural-proportion mixture with a full shuffle: Nemotron-Cascade-2 SFT, SmolTalk2 Mid + SFT, the Tercet-R stage-3 tool/agentic mix (Nemotron PTD v1, Agentic v1/v2, IF-Chat v1/v2, Toucan 1.5M, Hermes reasoning tool-use), WildChat-4.8M, SmolTalk, Hermes-3, UltraChat 200k, Tulu 3 IF personas, No Robots. Seq 32,768, hybrid Muon + AdamW 4×10⁻⁴, 18,200 steps ≈ 9.5B tokens, 91.2 hours. Val loss (EMA) 1.100 (ppl 3.01) |
| Preference (APO) | APO-zero, β = 0.05, on the SmolTalk2 Preference split as in SmolLM3: Tulu 3 preference mixture (no-think, 112,706 pairs) + Qwen3-32B vs Qwen3-0.6B (think, 52,567 pairs). One epoch, 32 pairs per step, 5,155 steps. AdamW 1×10⁻⁶ cosine to 1×10⁻⁷, 10% warmup, clip 0.2, max length 24,576. Reference model: the SFT EMA. 3.9 hours |
| Weights | EMA (this repo's model.safetensors) |
| Preference val (EMA) | 1,000 held-out pairs: reward accuracy 81.1%, reward margin 2.95, chosen NLL 1.06 per token |
preference_config.json, validation.json, sft_config.json and sft_validation.json hold the exact settings and held-out results.
Evaluation
lm-eval 0.4.13, EMA weights. SFT is sft-v3 step 18,200 (the APO starting point). Haiku is apo-v1 step 5,155.
ARC-Easy (AI2, 0-shot, 2,376 questions). acc is exact-match over choices; acc_norm length-normalizes the log-likelihood. The think protocol writes one greedy thought of up to 1,024 tokens, closes </think> if the thought hits that cap, then scores the four choices the same way.
| Protocol | SFT 18200 | Haiku (APO 5155) |
|---|---|---|
| Standard prompt, no chat template | 51.09 / 47.64 | 51.56 / 47.22 |
| Chat template, thinking off | 47.39 / 43.22 | 43.64 / 39.60 |
| Chat template, think, then score | 45.12 / 40.87 | 42.09 / 37.54 |
Numbers are acc / acc_norm, in percent. A thought before scoring lowers ARC-Easy. After APO the thoughts get much longer: mean 929 tokens, and 1,674 of 2,371 thoughts hit the 1,024-token cap (SFT: mean 543 tokens, 822 capped). Keep --max-new-tokens high when thinking is on.
IFEval (541 prompts, chat template, thinking off, greedy, 1,280 new tokens):
| SFT 18200 | Haiku (APO 5155) | |
|---|---|---|
| Prompt-level strict | 37.89 ± 2.09 | 48.43 ± 2.15 |
| Instruction-level strict | 49.28 | 59.59 |
| Prompt-level loose | 41.40 ± 2.12 | 53.42 ± 2.15 |
| Instruction-level loose | 52.64 | 64.15 |
242 SFT generations and 173 Haiku generations hit the 1,280-token limit.
Limitations
- Scale: ~655M is a research / edge model, not a frontier system
- On ARC-Easy multiple choice, thinking before the scored answer lowers accuracy, and APO makes those thoughts much longer
- Requires
flash-linear-attentionand a CUDA GPU; GGUF / llama.cpp support is not available today - The Python tool is a restricted interpreter (math-oriented imports only)
Model family
| Model | Stage | Hub |
|---|---|---|
| Haiku-base | Pretrain (v1) | kerzgrr/Haiku-base |
| Haiku | SFT + APO (chat, reasoning, tools) | this repo |
| Demo | ZeroGPU Space (chat + base) | kerzgrr/haiku-demo |
| Tercet-R-1.1 | Previous generation (~502M) | kerzgrr/Tercet-R-1.1 |
Citation
@misc{haiku2026,
title={Haiku: A 655M Hybrid KDA + Gated-MLA Reasoning Model},
author={kerzgrr},
year={2026},
url={https://huggingface.co/kerzgrr/Haiku}
}
Five, seven, five — now with preferences.
- Downloads last month
- 32