Haiku-Cyber

Research-only cybersecurity chat model (~655M), fine-tuned from Haiku

Model Stage License Architecture Demo

Haiku (SFT + APO) fine-tuned for one epoch on defensive cybersecurity chat. Research and non-commercial use only.


License

Research use only. Commercial use is not permitted.

The fine-tune is oi-uae/cyber-security at revision 2ca77817376c289b407cb7f269a38a6ca80aea9b. That dataset's license says models trained on it inherit its non-commercial restriction, including paid APIs and commercial hosted offerings. LICENSE in this repo is that text. The free demo Space is a research demo, not a commercial service.


What this is

Haiku-Cyber is kerzgrr/Haiku (the APO chat model, step 5,155 EMA) fine-tuned on defensive cybersecurity conversations.

  • One epoch of supervised fine-tuning at sequence length 8,192
  • 123,461 train conversations, 110.3M packed tokens, 68.1M assistant targets
  • Hub weights are the EMA snapshot in bfloat16, optimizer step 842
  • Live demo: kerzgrr/haiku-cyber

cve_qa (1,700,781 train rows) and doc_qa (52,967 train rows, AWS documentation chunks) were left out. The model is not a CVE lookup table and was not trained to recite those PDFs.

Train mix that remains:

Task Conversations
Causal-reasoning essays 98,876
Community security Q&A 15,560
Classification 2,962
MITRE and reference Q&A 2,193
Code tutoring 1,819
Instructions 1,078
Vulnerability fixes 803
Incident, playbook, and scenario Q&A 170

Chat contract

Same contract as Haiku. Each assistant turn is prefixed with a zero-loss control token: <|think|> or <|no_think|>. Tool schemas go in a <tools> block. Calls look like:

<tool_call>
{"name": "web-search", "arguments": {"query": "…"}}
</tool_call>

inference.py auto-runs web-search, stateful_python_code_exec, and calculator. Other observations are sent as <|tool_response|>.


Install & run

pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Haiku-Cyber inference.py --local-dir .
python inference.py
python inference.py --prompt "Explain the difference between stored and reflected XSS."
python inference.py --no-think --prompt "Reply in one sentence."

inference.py auto-downloads weights, the tokenizer, and tiny_gdn/, and auto-installs pinned flash-linear-attention. Git is required on PATH. A CUDA GPU is required: the Kimi Delta Attention layers run on Triton kernels.

Flag Default Description
--prompt — One-shot user message
--system — System prompt, used verbatim
--think / --no-think think Assistant control prefix
--tools — Built-ins: web-search, python, calculator
--temperature 0.7 Sampling temperature
--max-new-tokens 4096 Max generation length
--context-length 32768 Conversation tokens kept
--device cuda CUDA device, e.g. cuda:1

This fine-tune's sequence length was 8,192. The architecture still accepts 32,768, which is the default context window above.


Model architecture

Same Haiku hybrid as Haiku (655,270,488 parameters): 36 layers at width 1,024, Kimi Delta Attention in the 27 linear layers, gated MLA every fourth layer, SiTU-GLU feed-forwards, block attention residuals, two MTP heads, vocab 65,536.


Training

Stage Details
Init kerzgrr/Haiku APO EMA, apo-v1 step 5,155
Data oi-uae/cyber-security full config, cve_qa and doc_qa excluded
SFT AdamW 2×10⁻⁵ peak, cosine to 0.1×, 3% warmup, no weight decay, clip 1.0, seq 8,192, 131,072 tokens/step, 842 steps, one epoch, 0.87 hours
Weights EMA (model.safetensors)
Validation (EMA) Loss 1.561, perplexity 4.76, on 85,114 held-out assistant targets

sft_config.json and validation.json are the checkpoint's own files.


Limitations

  • License: research and other non-commercial use only
  • Not an authority: on a spot check it described CVE-2021-44228 (Log4Shell) as a Linux-kernel bug, and it did not correctly distinguish stored XSS from reflected XSS. CVE identifiers were removed from training, so it has nothing to recall them from
  • Scale: ~655M parameters. Long answers can repeat and run into the token cap
  • Requires flash-linear-attention and a CUDA GPU. GGUF / llama.cpp is not available
  • The Python tool is a restricted interpreter

Model family

Model Stage Hub
Haiku-base Pretrain (v1) kerzgrr/Haiku-base
Haiku SFT + APO kerzgrr/Haiku
Haiku-Cyber Cyber SFT, research-only this repo
Demo ZeroGPU Space kerzgrr/haiku-cyber

Citation

@misc{haikucyber2026,
  title={Haiku-Cyber: A 655M Research-Only Cybersecurity Fine-Tune of Haiku},
  author={kerzgrr},
  year={2026},
  url={https://huggingface.co/kerzgrr/Haiku-Cyber}
}
Downloads last month
-
Safetensors
Model size
0.7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kerzgrr/Haiku-Cyber

Base model

kerzgrr/Haiku
Finetuned
(1)
this model

Datasets used to train kerzgrr/Haiku-Cyber

Collection including kerzgrr/Haiku-Cyber