Instructions to use srock44/cipher-air with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use srock44/cipher-air with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf srock44/cipher-air:Q4_K_M # Run inference directly in the terminal: llama cli -hf srock44/cipher-air:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf srock44/cipher-air:Q4_K_M # Run inference directly in the terminal: llama cli -hf srock44/cipher-air:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf srock44/cipher-air:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf srock44/cipher-air:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf srock44/cipher-air:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf srock44/cipher-air:Q4_K_M
Use Docker
docker model run hf.co/srock44/cipher-air:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use srock44/cipher-air with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "srock44/cipher-air" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "srock44/cipher-air", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/srock44/cipher-air:Q4_K_M
- Ollama
How to use srock44/cipher-air with Ollama:
ollama run hf.co/srock44/cipher-air:Q4_K_M
- Unsloth Studio
How to use srock44/cipher-air with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for srock44/cipher-air to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for srock44/cipher-air to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for srock44/cipher-air to start chatting
- Pi
How to use srock44/cipher-air with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf srock44/cipher-air:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "srock44/cipher-air:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use srock44/cipher-air with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf srock44/cipher-air:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "srock44/cipher-air:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use srock44/cipher-air with Docker Model Runner:
docker model run hf.co/srock44/cipher-air:Q4_K_M
- Lemonade
How to use srock44/cipher-air with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull srock44/cipher-air:Q4_K_M
Run and chat with the model
lemonade run user.cipher-air-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use srock44/cipher-air with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf srock44/cipher-air:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default srock44/cipher-air:Q4_K_M
Run Hermes
hermes
- Atomic Chat
Cipher Air
Cipher Air is a LoRA fine-tune of Qwen/Qwen2.5-0.5B-Instruct, trained on every LLM-backed feature of a local-first email assistant: email triage (importance/summary/category JSON), chat, daily-summary synthesis, draft reply, compose assist, and voice-command intent parsing โ not just prompted for these tasks, actually trained on them.
Update: added a 6th task, voice-command intent parsing (interpreting a dictated voice
command + a candidate address list into a structured action). The original 5-task version
had never seen this task and would confidently hallucinate a plausible-looking but entirely
fake person/address when it couldn't resolve one โ this retrain fixes that; see Training
below and generate_voice_intent.py for the fix in detail.
Second update: a security-relevant bug was found and fixed โ the previous voice-intent
retrain, done through several rounds of incremental continue-training, left this model
occasionally complying with prompt-injection attempts embedded in email content on the
draft-reply/chat tasks (echoing a fake "security verification" request for banking details
back into a drafted reply). Rather than patch that narrowly, this model was retrained fresh
from the base in one pass on a properly-rebalanced dataset with much richer injection
coverage across every task, not just triage. Verified through repeated adversarial testing
against the actual calling application's code path (not just raw single-shot completions):
no compliance with any credential/wire-transfer/data-exfiltration injection attempt across
multiple test rounds. See Training below.
It's the middle of the three Cipher tiers (cipher-nano / cipher-air / cipher-pro) โ a balanced default at 40% of cipher-pro's disk size and 3x the throughput. Cipher is the local-model engine for an unreleased larger email-assistant project โ that project isn't public yet, but these weights, the training code, the eval script, and all five dataset generators are fully open now, in this repo.
Why this exists
Most email triage today means sending your inbox to a third-party API. Cipher runs entirely on your own hardware via Ollama โ nothing about your email ever leaves your machine.
What's in this repo
cipher-air.Q4_K_M.ggufโ the model weights, ready for OllamaModelfileโ the exact Ollama Modelfile (system prompt + inference params) used in training/evaltrain_cipher_air.py/export_gguf_cipher_air.pyโ the exact scripts used to produce this model (Unsloth LoRA on the base model above)generate2.py,generate_chat.py,generate_daily_summary.py,generate_draft_reply.py,generate_compose.py,generate_voice_intent.pyโ the six task-specific synthetic-data generators (produces the full multi-task training set)eval_voice_intent.pyโ regression harness for the voice-intent task, including the exact hallucinated-address bug case as a required fixtureeval_triage.py/eval_fixtures.jsonโ a standalone benchmark harness (no external dependencies beyondhttpx/pydantic) reproducing the numbers below
Everything needed to reproduce this model from scratch, or fine-tune your own variant, is in this repo โ nothing here depends on an unreleased package.
Benchmark
Evaluated on a 29-fixture triage benchmark on an RTX 5070:
| Model | Disk | Tok/s | JSON-valid | Category acc | Importance-in-band | Injection-safe |
|---|---|---|---|---|---|---|
| cipher-air | 398 MB | 507.8 | 100.0% | 69.0% | 79.3% | 100% |
Honest caveat: cipher-air is the tightest-capacity tier of the three (only 8.8M of 502M
params are trainable via LoRA), and it shows โ of the three tiers it's the one most likely
to occasionally misjudge whether something genuinely needs a reminder/action versus being
routine. cipher-pro and cipher-nano both handle that nuance more reliably. Reproduce
with:
pip install -r requirements.txt
python eval_triage.py --models cipher-air:latest --keep
Usage (Ollama)
ollama create cipher-air -f Modelfile
Query it with grammar-constrained JSON output for reliable parsing:
curl http://localhost:11434/api/chat -d '{
"model": "cipher-air",
"messages": [
{"role": "system", "content": "<system prompt from Modelfile>"},
{"role": "user", "content": "From: alex@acme.com\nSubject: Q3 budget review\n\nBody:\nCan we sync before Friday?"}
],
"format": "json",
"options": {"temperature": 0.1}
}'
Training
- Base:
Qwen/Qwen2.5-0.5B-Instruct, LoRA (r=16, alpha=32, all linear layers), 2 epochs - Data: ~4,800 triage examples (oversampled ~2x to ~60% of the final training mix โ this
size tier needed a stronger triage signal than the other two to hold onto exact JSON
schema output while also learning four other task formats) + ~1,600-2,000 examples each
for chat/daily-summary/draft-reply/compose, all matching production prompts exactly โ
generated by the five
generate_*.pyscripts in this repo - Framework: Unsloth +
trl.SFTTrainer - Sequence packing was tried to speed up training (most examples are well under the 2048-token context window) โ it crashed outright, an Unsloth/trl version incompatibility, not a quality tradeoff. Disabled.
- Reproduce with
train_cipher_air.pyโexport_gguf_cipher_air.py - Voice-intent retrain: added ~1,800 examples from
generate_voice_intent.py, weighted heavily toward the no-match case (a spoken name with no corresponding candidate address โ the model must returnnullrather than inventing one) and matching-with-distractors cases. Oneval_voice_intent.py's 5 fixtures, cipher-air went from 2/5 (including a fabricated address and an out-of-schema action) to 5/5, the cleanest result of the three tiers on this task. - Full retrain (current version): rather than continue-training the voice-intent adapter
further, this version is a fresh LoRA fine-tune from the base model on one consolidated,
properly-balanced dataset covering all 6 tasks in a single pass โ triage (
4,800, oversampled 3x to hold the ~60% mix ratio this tier needs), chat/daily-summary/draft- reply/compose/voice-intent (1,800-2,500 each).generate_draft_reply.pyandgenerate_chat.pyboth gained substantially heavier and more varied injection coverage (credential/wire-transfer/data-exfiltration attempts, not just one generic case) after live testing found the previous incremental-patch version could be induced into complying with an injected "security verification" request for banking details.generate_chat.pyalso gained an explicit "the question is about something with no connection to your email at all (weather, sports, etc.)" scenario category after finding a hallucination regression there. One clean training pass over the properly-balanced result, instead of a chain of narrow continue-trains, avoids the whack-a-mole pattern where each targeted fix risked nudging a different, previously-working case.
A dead end worth knowing about
We tried quantizing this model down further (Q3_K_M, Q2_K) hoping to shrink it toward cipher-nano's size class. It barely helped (355MB / 339MB vs 398MB at Q4_K_M) โ Qwen2.5's 151,936-token vocabulary embedding table dominates disk size and doesn't compress with weight quantization. If you're looking for something genuinely small, use cipher-nano instead (different base model, built specifically to solve this).
License
Apache 2.0, inherited from the base model. Weights, training code, and eval harness are fully open.
- Downloads last month
- 19
4-bit