Instructions to use riposta/CrossbowReviewer-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use riposta/CrossbowReviewer-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="riposta/CrossbowReviewer-9B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("riposta/CrossbowReviewer-9B") model = AutoModelForCausalLM.from_pretrained("riposta/CrossbowReviewer-9B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use riposta/CrossbowReviewer-9B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf riposta/CrossbowReviewer-9B:Q8_0 # Run inference directly in the terminal: llama cli -hf riposta/CrossbowReviewer-9B:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf riposta/CrossbowReviewer-9B:Q8_0 # Run inference directly in the terminal: llama cli -hf riposta/CrossbowReviewer-9B:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf riposta/CrossbowReviewer-9B:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf riposta/CrossbowReviewer-9B:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf riposta/CrossbowReviewer-9B:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf riposta/CrossbowReviewer-9B:Q8_0
Use Docker
docker model run hf.co/riposta/CrossbowReviewer-9B:Q8_0
- LM Studio
- Jan
- vLLM
How to use riposta/CrossbowReviewer-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "riposta/CrossbowReviewer-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "riposta/CrossbowReviewer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/riposta/CrossbowReviewer-9B:Q8_0
- SGLang
How to use riposta/CrossbowReviewer-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "riposta/CrossbowReviewer-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "riposta/CrossbowReviewer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "riposta/CrossbowReviewer-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "riposta/CrossbowReviewer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use riposta/CrossbowReviewer-9B with Ollama:
ollama run hf.co/riposta/CrossbowReviewer-9B:Q8_0
- Unsloth Desktop
- Pi
How to use riposta/CrossbowReviewer-9B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf riposta/CrossbowReviewer-9B:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "riposta/CrossbowReviewer-9B:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use riposta/CrossbowReviewer-9B with Docker Model Runner:
docker model run hf.co/riposta/CrossbowReviewer-9B:Q8_0
- Lemonade
How to use riposta/CrossbowReviewer-9B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull riposta/CrossbowReviewer-9B:Q8_0
Run and chat with the model
lemonade run user.CrossbowReviewer-9B-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use riposta/CrossbowReviewer-9B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf riposta/CrossbowReviewer-9B:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default riposta/CrossbowReviewer-9B:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use riposta/CrossbowReviewer-9B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf riposta/CrossbowReviewer-9B:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "riposta/CrossbowReviewer-9B:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
CrossbowReviewer-9B
Calibrated code review decisions in one forward pass. It asks up to 107 rules per file, runs in ~0.2 s on an A100 and in ~3 s on a Mac with the Q8_0 GGUF.
CrossbowReviewer-9B fine-tunes Qwen3.5-9B-Base for
Jev-style typed decisions about code. It writes no review text. For
each question, such as "is this free of SQL injection?" or "does each class have one reason to change?", it returns
a choice, a calibrated confidence and the full distribution. Its inference server
provides a /v1/systemone endpoint that follows the Jev request schema.
- Typed decisions:
noul(yes/no),choice(up to 26 options) andscore(2–10 levels), asked against the code asstate. - One pass: all questions share one prefill. Option-letter logits are read at each answer slot, so no answer text is generated.
- 120 built-in rules: SOLID and design principles, patterns and anti-patterns, correctness, error handling,
security (OWASP / CWE), performance, readability, clean code, testability, testing, API design and language idioms.
Two aggregate questions are included (
primary_category,severity). - 7 languages: Java, Python, TypeScript, JavaScript, Go, C#, Rust.
- Full weights: no adapter and no conversion step. Safetensors work with transformers; GGUF works with llama.cpp.
Inference server · Quickstart · Benchmarks · Artifact manifest
Downloads
| Files | Format and use | Size |
|---|---|---|
model-0000{1..4}-of-00004.safetensors + configs |
BF16 safetensors. transformers backend (CUDA, Apple MPS, CPU) | 17.9 GB |
gguf/CrossbowReviewer-9B-Q8_0.gguf |
Q8_0 GGUF. llama.cpp backend; fits 12 GB GPUs and Apple Silicon | 9.5 GB |
decision_config.json holds what turns logits into decisions: the fitted temperature (1.044), the option letters and
the prompt format. Both backends read it.
The Q8_0 GGUF gives the same answers as the BF16 safetensors. On a parity set of 79 decisions every choice matched, with a mean probability difference of 0.004.
Quick start
git clone https://github.com/riposta/CrossbowReviewer && cd CrossbowReviewer
uv run uvicorn server:app --port 8000 # transformers backend; downloads this repo on first request
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' \
-d '{"state": {"language": "python", "code": "def add(items=[]):\n items.append(1)\n return items"}}'
Without questions, every rule for state.language is asked. yes means compliant:
{"decisions": {"py_mutable_default": {"choice": "no", "confidence": 0.98, "distribution": {"yes": 0.02, "no": 0.98}}, "...": "..."}}
For llama.cpp, custom questions and all options, see docs/QUICKSTART.md.
Results
| Balanced accuracy (0.5 = chance) | Synthetic | Synthetic, relabeled by independent judges | OWASP Benchmark (human labels) | Real open-source code |
|---|---|---|---|---|
| CrossbowReviewer-9B | 0.824 | 0.841 | 0.654 | 0.783 |
| Jev (jev-1.13.0) | 0.809 | 0.813 | 0.653 | 0.717 |
| Qwen3.5-9B-Base | 0.644 | 0.648 | 0.696 | 0.517 |
| Claude Opus 5.5 (reasoning) | 0.925 | — | 0.979 | 0.831 |
| GPT-6 Sol (reasoning) | 0.911 | — | 0.900 | 0.926 |
- Calibration: expected calibration error is 0.016 on the synthetic test and 0.006 on real code. When the model says 0.9, it is right about 90% of the time.
- Latency: the reasoning models are 3–7 points more accurate, but 25–55× slower (8–10 s per file versus ~0.2 s).
Methodology, per-category results and all metrics are in docs/BENCHMARKS.md and docs/benchmarks.json.
Limitations — read before use
Security taint traps. On OWASP Benchmark the model flags 94% of vulnerable cases. It recognizes only 26% of the safe cases built to look vulnerable, for example when tainted input reaches only a dead branch or is overwritten by a constant. Jev (24%) and the base model (37%) share this weakness. One forward pass cannot trace data flow the way a reasoning model can (Opus 5.5: 94%). Extra training on 500 hard security cases raised this to 42%, which is not enough, so that version is not released.
Recommended use. Use the model directly for design, readability, correctness, testing and idiom rules. For
data-flow security rules (injection, path_traversal, xss, ssrf, open_redirect), treat a no as "needs a
closer look" and escalate it to a reasoning model.
Other limits:
- Synthetic training data. About 1 500 snippets (30–120 lines) were written and labeled by LLMs. They are not human reviews of real pull requests. Scores on synthetic data overstate real-world accuracy.
- One file at a time. The model judges snippets, not diffs, and does not see the rest of the codebase.
- Custom questions outside the 120 trained rules use the same format, but their quality is not measured.
- Prompts are in English.
choicesupports up to 26 options (Jev: 255).
Training
- Data: 1 466 synthetic snippets. Each one is built as language × rule × (compliant / violating), and all ~107 questions are labeled by an LLM. The label agrees with the construction target in 93% of cases. A judge-in-the-loop added 51 samples for the weakest rules.
- Method: LoRA on all linear layers (rank 16, alpha 32; 0.48% of parameters), merged into these full weights.
Cross-entropy over option-letter logits at every answer slot. Each sample shows a random 8–all subset of questions,
always including its target rule. 30% of samples omit
language. - Optimization: 3 rounds × 1 epoch, AdamW, lr 1e-4, effective batch 8, bf16, 1× A100 40 GB (~2 h).
- Calibration: a single temperature fitted on a validation split.
- Not included: the multi-token-prediction head of the base model is not shipped (
mtp_num_hidden_layers: 0), because decisions never generate text.
License
Apache 2.0, the same as the base model. See LICENSE and NOTICE.
- Downloads last month
- 405


