Haiku-Cyber
Research-only cybersecurity chat model (~655M), fine-tuned from Haiku
Haiku (SFT + APO) fine-tuned for one epoch on defensive cybersecurity chat. Research and non-commercial use only.
License
Research use only. Commercial use is not permitted.
The fine-tune is oi-uae/cyber-security at revision 2ca77817376c289b407cb7f269a38a6ca80aea9b. That dataset's license says models trained on it inherit its non-commercial restriction, including paid APIs and commercial hosted offerings. LICENSE in this repo is that text. The free demo Space is a research demo, not a commercial service.
What this is
Haiku-Cyber is kerzgrr/Haiku (the APO chat model, step 5,155 EMA) fine-tuned on defensive cybersecurity conversations.
- One epoch of supervised fine-tuning at sequence length 8,192
- 123,461 train conversations, 110.3M packed tokens, 68.1M assistant targets
- Hub weights are the EMA snapshot in bfloat16, optimizer step 842
- Live demo:
kerzgrr/haiku-cyber
cve_qa (1,700,781 train rows) and doc_qa (52,967 train rows, AWS documentation chunks) were left out. The model is not a CVE lookup table and was not trained to recite those PDFs.
Train mix that remains:
| Task | Conversations |
|---|---|
| Causal-reasoning essays | 98,876 |
| Community security Q&A | 15,560 |
| Classification | 2,962 |
| MITRE and reference Q&A | 2,193 |
| Code tutoring | 1,819 |
| Instructions | 1,078 |
| Vulnerability fixes | 803 |
| Incident, playbook, and scenario Q&A | 170 |
Chat contract
Same contract as Haiku. Each assistant turn is prefixed with a zero-loss control token: <|think|> or <|no_think|>. Tool schemas go in a <tools> block. Calls look like:
<tool_call>
{"name": "web-search", "arguments": {"query": "…"}}
</tool_call>
inference.py auto-runs web-search, stateful_python_code_exec, and calculator. Other observations are sent as <|tool_response|>.
Install & run
pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Haiku-Cyber inference.py --local-dir .
python inference.py
python inference.py --prompt "Explain the difference between stored and reflected XSS."
python inference.py --no-think --prompt "Reply in one sentence."
inference.py auto-downloads weights, the tokenizer, and tiny_gdn/, and auto-installs pinned flash-linear-attention. Git is required on PATH. A CUDA GPU is required: the Kimi Delta Attention layers run on Triton kernels.
| Flag | Default | Description |
|---|---|---|
--prompt |
— | One-shot user message |
--system |
— | System prompt, used verbatim |
--think / --no-think |
think | Assistant control prefix |
--tools |
— | Built-ins: web-search, python, calculator |
--temperature |
0.7 |
Sampling temperature |
--max-new-tokens |
4096 |
Max generation length |
--context-length |
32768 |
Conversation tokens kept |
--device |
cuda |
CUDA device, e.g. cuda:1 |
This fine-tune's sequence length was 8,192. The architecture still accepts 32,768, which is the default context window above.
Model architecture
Same Haiku hybrid as Haiku (655,270,488 parameters): 36 layers at width 1,024, Kimi Delta Attention in the 27 linear layers, gated MLA every fourth layer, SiTU-GLU feed-forwards, block attention residuals, two MTP heads, vocab 65,536.
Training
| Stage | Details |
|---|---|
| Init | kerzgrr/Haiku APO EMA, apo-v1 step 5,155 |
| Data | oi-uae/cyber-security full config, cve_qa and doc_qa excluded |
| SFT | AdamW 2×10⁻⁵ peak, cosine to 0.1×, 3% warmup, no weight decay, clip 1.0, seq 8,192, 131,072 tokens/step, 842 steps, one epoch, 0.87 hours |
| Weights | EMA (model.safetensors) |
| Validation (EMA) | Loss 1.561, perplexity 4.76, on 85,114 held-out assistant targets |
sft_config.json and validation.json are the checkpoint's own files.
Limitations
- License: research and other non-commercial use only
- Not an authority: on a spot check it described CVE-2021-44228 (Log4Shell) as a Linux-kernel bug, and it did not correctly distinguish stored XSS from reflected XSS. CVE identifiers were removed from training, so it has nothing to recall them from
- Scale: ~655M parameters. Long answers can repeat and run into the token cap
- Requires
flash-linear-attentionand a CUDA GPU. GGUF / llama.cpp is not available - The Python tool is a restricted interpreter
Model family
| Model | Stage | Hub |
|---|---|---|
| Haiku-base | Pretrain (v1) | kerzgrr/Haiku-base |
| Haiku | SFT + APO | kerzgrr/Haiku |
| Haiku-Cyber | Cyber SFT, research-only | this repo |
| Demo | ZeroGPU Space | kerzgrr/haiku-cyber |
Citation
@misc{haikucyber2026,
title={Haiku-Cyber: A 655M Research-Only Cybersecurity Fine-Tune of Haiku},
author={kerzgrr},
year={2026},
url={https://huggingface.co/kerzgrr/Haiku-Cyber}
}
- Downloads last month
- -
Model tree for kerzgrr/Haiku-Cyber
Base model
kerzgrr/Haiku