Haiku

Reasoning + tool-call chat model (~655M), preference-tuned with APO

Model Stage License Architecture Demo

A ~655M hybrid Kimi Delta Attention + gated MLA model: 32k-context SFT on reasoning, tool and chat data, then SmolLM3-style APO


What this is

Haiku is the chat model of the Haiku family, the ~655M successor to Tercet-R.

  • Supervised fine-tuning at a 32,768-token sequence length on ~9.5B tokens of reasoning, tool-call, instruction-following and chat data
  • Then APO-zero preference optimization on the SmolTalk2 preference split, following the SmolLM3 recipe
  • Live demo: kerzgrr/haiku-demo — Haiku chat plus Haiku-base raw continuation in one Space
  • Hub weights are the EMA snapshot in bfloat16

This upload is the final checkpoint of the preference run (optimizer step 5,155, one full epoch of 164,948 pairs).

The SFT started from the v2 Haiku pretrain (step 4,800, 2.5B FineWeb-Edu tokens), which is not released separately. The public kerzgrr/Haiku-base is the earlier v1 pretrain of the same architecture and tokenizer (step 8,400, 4.4B tokens).


Chat contract

Thinking

Each assistant turn is prefixed with a zero-loss control token:

Mode Prefix Typical body
think <|think|>\n <think>…</think> then the answer
no-think <|no_think|>\n answer only

inference.py streams the <think> region live (dim yellow) and hides the control tokens.

Tool calls (SmolTalk JSON)

Tool schemas go in a Hermes-style <tools> block in the system turn. The model calls them with:

<tool_call>
{"name": "web-search", "arguments": {"query": "…"}}
</tool_call>

SFT covered Nemotron Agentic v1/v2 (including web-search), Toucan 1.5M, Hermes reasoning tool-use and Nemotron PTD v1 tool calling. inference.py auto-runs:

Built-in Tool name Observation
--tools web-search web-search Tavily-shaped JSON (Tavily if TAVILY_API_KEY is set, else DuckDuckGo + Wikipedia)
--tools python stateful_python_code_exec Jupyter-style stdout / last value from a restricted interpreter
--tools calculator calculator Numeric result of a math expression

Tool results

Each observation is a tool (or user) turn prefixed with:

<|tool_response|>
{observation}

Install & run

pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Haiku inference.py --local-dir .
python inference.py
python inference.py --prompt "What is the capital of France?"
python inference.py --tools web-search,python,calculator
python inference.py --no-think --prompt "Reply in one sentence."

inference.py auto-downloads weights / tokenizer / tiny_gdn/ and auto-installs pinned flash-linear-attention (plus transformers, which its decode cache imports). Git is required on PATH. A CUDA GPU is required: the Kimi Delta Attention layers run on Triton kernels.

Flag Default Description
--prompt — One-shot user message
--system — System prompt, used verbatim
--think / --no-think think Assistant control prefix
--tools — Built-ins: web-search, python, calculator (comma-separated)
--temperature 0.7 Sampling temperature
--max-new-tokens 4096 Max generation length
--context-length 32768 Conversation tokens kept (oldest dropped first)
--device cuda CUDA device, e.g. cuda:1

Interactive commands: /think /no_think /system … /reset /exit. Ctrl+C stops the current reply and keeps the session.


Model architecture

Same Haiku hybrid as Haiku-base (655,270,488 parameters):

Layers 36 (Kimi Delta Attention ×3 + gated MLA every 4th)
Hidden 1,024
MLP SiTU-GLU 3,840
Attention Gated MLA (NoPE), 8 heads, Q LoRA rank 512, KV LoRA rank 256
Linear Kimi Delta Attention, 8 heads × 128
Residuals Block attention residuals
Vocab 65,536 BPE
Context 32,768 (SFT sequence length)

Training

Stage Details
Pretrain (v2) FineWeb-Edu, seq 2,048, 4,800 steps = 2.52B tokens, hybrid Muon + AdamW at peak 5×10⁻⁴, 19.2 hours. Val loss (EMA) 2.90
SFT (v3) Natural-proportion mixture with a full shuffle: Nemotron-Cascade-2 SFT, SmolTalk2 Mid + SFT, the Tercet-R stage-3 tool/agentic mix (Nemotron PTD v1, Agentic v1/v2, IF-Chat v1/v2, Toucan 1.5M, Hermes reasoning tool-use), WildChat-4.8M, SmolTalk, Hermes-3, UltraChat 200k, Tulu 3 IF personas, No Robots. Seq 32,768, hybrid Muon + AdamW 4×10⁻⁴, 18,200 steps ≈ 9.5B tokens, 91.2 hours. Val loss (EMA) 1.100 (ppl 3.01)
Preference (APO) APO-zero, β = 0.05, on the SmolTalk2 Preference split as in SmolLM3: Tulu 3 preference mixture (no-think, 112,706 pairs) + Qwen3-32B vs Qwen3-0.6B (think, 52,567 pairs). One epoch, 32 pairs per step, 5,155 steps. AdamW 1×10⁻⁶ cosine to 1×10⁻⁷, 10% warmup, clip 0.2, max length 24,576. Reference model: the SFT EMA. 3.9 hours
Weights EMA (this repo's model.safetensors)
Preference val (EMA) 1,000 held-out pairs: reward accuracy 81.1%, reward margin 2.95, chosen NLL 1.06 per token

preference_config.json, validation.json, sft_config.json and sft_validation.json hold the exact settings and held-out results.


Evaluation

lm-eval 0.4.13, EMA weights. SFT is sft-v3 step 18,200 (the APO starting point). Haiku is apo-v1 step 5,155.

ARC-Easy (AI2, 0-shot, 2,376 questions). acc is exact-match over choices; acc_norm length-normalizes the log-likelihood. The think protocol writes one greedy thought of up to 1,024 tokens, closes </think> if the thought hits that cap, then scores the four choices the same way.

Protocol SFT 18200 Haiku (APO 5155)
Standard prompt, no chat template 51.09 / 47.64 51.56 / 47.22
Chat template, thinking off 47.39 / 43.22 43.64 / 39.60
Chat template, think, then score 45.12 / 40.87 42.09 / 37.54

Numbers are acc / acc_norm, in percent. A thought before scoring lowers ARC-Easy. After APO the thoughts get much longer: mean 929 tokens, and 1,674 of 2,371 thoughts hit the 1,024-token cap (SFT: mean 543 tokens, 822 capped). Keep --max-new-tokens high when thinking is on.

IFEval (541 prompts, chat template, thinking off, greedy, 1,280 new tokens):

SFT 18200 Haiku (APO 5155)
Prompt-level strict 37.89 ± 2.09 48.43 ± 2.15
Instruction-level strict 49.28 59.59
Prompt-level loose 41.40 ± 2.12 53.42 ± 2.15
Instruction-level loose 52.64 64.15

242 SFT generations and 173 Haiku generations hit the 1,280-token limit.


Limitations

  • Scale: ~655M is a research / edge model, not a frontier system
  • On ARC-Easy multiple choice, thinking before the scored answer lowers accuracy, and APO makes those thoughts much longer
  • Requires flash-linear-attention and a CUDA GPU; GGUF / llama.cpp support is not available today
  • The Python tool is a restricted interpreter (math-oriented imports only)

Model family

Model Stage Hub
Haiku-base Pretrain (v1) kerzgrr/Haiku-base
Haiku SFT + APO (chat, reasoning, tools) this repo
Demo ZeroGPU Space (chat + base) kerzgrr/haiku-demo
Tercet-R-1.1 Previous generation (~502M) kerzgrr/Tercet-R-1.1

Citation

@misc{haiku2026,
  title={Haiku: A 655M Hybrid KDA + Gated-MLA Reasoning Model},
  author={kerzgrr},
  year={2026},
  url={https://huggingface.co/kerzgrr/Haiku}
}

Five, seven, five — now with preferences.

Downloads last month
32
Safetensors
Model size
0.7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kerzgrr/Haiku

Finetunes
1 model

Datasets used to train kerzgrr/Haiku

Collection including kerzgrr/Haiku