jeeves / README.md
robbie-posthog's picture
Update the serving note: Macs, --precision and the FP8 weights
c754d67 verified
|
Raw History Blame Contribute Delete
5.49 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
language:
- en
pipeline_tag: text-classification
tags:
- decision
- calibration
- reasoning
- speculative-decoding
- qwen3.5
---
# Jeeves-9B
A Jev-like decision model that reasons before it decides. Give it a state (text or JSON) and yes/no, multiple-choice or rating questions; it writes a reasoning chain per question and returns a calibrated probability for every option.
This repository holds the fused weights (Qwen3.5-9B with the LoRA merged), the pointer head, the fitted temperature, and two speculative-decoding drafters. It loads with the Jeeves code, not with `transformers`: the pointer head and the prompt format are part of the model.
## Files
| file | contents |
|---|---|
| `model-*.safetensors`, `config.json` | fused Qwen3.5-9B weights, bf16 |
| `head.pt` | pointer head (query and key projections, 256-dim) |
| `export.json` | temperature (1.859), prompt format, training step and config |
| `tokenizer.json`, `vocab.json`, `merges.txt`, `tokenizer_config.json`, `chat_template.jinja` | Qwen3.5 tokenizer |
| `drafter_k4.safetensors` | block-4 drafter, the serving default, bf16 |
| `drafter_k8.safetensors` | block-8 drafter, faster for single-question requests, bf16 |
| `drafter_k4.config.json`, `drafter_k8.config.json` | drafter training settings |
| `LICENSE` | Apache-2.0, inherited from Qwen3.5-9B |
## Usage
```bash
hf download PostHog/jeeves --local-dir jeeves-weights
git clone https://github.com/PostHog/jeeves && cd jeeves
python -m inference.serve --model ../jeeves-weights --drafter ../jeeves-weights/drafter_k4.safetensors --port 8009
```
```bash
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice for one order.",
"questions": {
"billing": {"type": "noul", "instructions": "Is this about billing?"},
"tone": {"type": "choice", "instructions": "What is the customer'"'"'s tone?", "criteria": {"calm": null, "frustrated": null, "angry": null}}
}}'
```
The request and response follow Jev's `/v1/systemone` format. An optional `options` object sets `think`, `max_think` (truncate chains), `nothink_threshold` (skip thinking when the no-think answer is already confident) and `return_reasoning`. Serving runs on a CUDA GPU or on an Apple Silicon Mac with 48 GB (add `--max-rows 4 --max-len 4096` there). It serves bf16 by default; `--precision fp8` runs the linear layers in FP8 on GPUs with compute capability 8.9 or higher and on Macs. FP8 weights can be downloaded from [PostHog/jeeves-fp8](https://huggingface.co/PostHog/jeeves-fp8).
## Results
Accuracy with thinking, greedy, 2,560-token cap. Kev-9B and Jev numbers are the ones Kev publishes. JevBench uses the same 231 public items for every model (the sealed judge tier is not included); the other rows use the same sources with different items.
| | Kev-9B | Jev | Jeeves-9B |
|---|---|---|---|
| Test overall (out-of-domain and held-out) | 0.822 | 0.857 | **0.889** |
| MMLU-Pro and buried state | 0.579 | **0.800** | 0.746 |
| JevBench, public tiers | 0.715* | 0.866 | **0.935** |
| JevBench hard | 0.451* | 0.730 | **0.865** |
\* Kev-8B (Qwen3); no Kev-9B JevBench result is published.
Without thinking the model scores 0.804 on the Jeeves test split, against 0.840 with it. With the fitted temperature the no-think path has a top-label calibration error of 0.021 and 1.3% confident errors (wrong at p ≥ 0.9).
Serving on one H100, 325 dev questions:
| setting | accuracy | median / p90 latency |
|---|---|---|
| full thinking | 0.825 | 3.3 s / 17.1 s |
| `max_think` 768, `nothink_threshold` 0.9 | 0.806 | 2.0 s / 5.6 s |
| no thinking | 0.775 | about 0.3 s |
## Training
- **SFT.** LoRA r=16 on every projection plus the pointer head, 2 epochs on 19,126 questions from 12 public datasets and synthetic policy data; half the questions carry a reasoning chain sampled from the base model.
- **CISPO.** 9,992 RL questions, 8 rollouts each, up to 2,560 thinking tokens. Reward is the probability of the correct option, discounted by up to 10% for long chains. The released checkpoint is step 402 of a 624-step schedule.
- **Calibration.** One softmax temperature fitted on the dev set.
- **Drafters.** Diffusion views of the frozen model, inspired by Orthrus, distilled by KL divergence on chains sampled from this model. At one question the block-4 drafter decodes 1.6× faster than graphed greedy decoding and the block-8 drafter 1.76×.
A from-scratch reproduction with the released code matched this checkpoint within noise. The training code, dataset build scripts and pinned dataset revisions are in the GitHub repository.
## Limitations
- Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840).
- Thinking is slow at the tail (17 s at p90 with full chains); use `max_think` and `nothink_threshold` when latency matters.
- About a third of full-length chains hit the token cap without closing; accuracy on those items is lower.
- Evaluated on English inputs only.
## License
The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license (`LICENSE`). The Jeeves code is MIT licensed.
## Citation
```bibtex
@software{waltz2026jeeves,
author = {Waltz, Nicholas P.},
title = {Jeeves: Reasoning Improves Jev-like Decisions},
year = {2026},
url = {https://github.com/PostHog/jeeves},
note = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter}
}
```