--- license: apache-2.0 base_model: Qwen/Qwen3.5-9B language: - en pipeline_tag: text-classification tags: - decision - calibration - reasoning - speculative-decoding - qwen3.5 --- # Jeeves-9B A Jev-like decision model that reasons before it decides. Give it a state (text or JSON) and yes/no, multiple-choice or rating questions; it writes a reasoning chain per question and returns a calibrated probability for every option. This repository holds the fused weights (Qwen3.5-9B with the LoRA merged), the pointer head, the fitted temperature, and two speculative-decoding drafters. It loads with the Jeeves code, not with `transformers`: the pointer head and the prompt format are part of the model. ## Files | file | contents | |---|---| | `model-*.safetensors`, `config.json` | fused Qwen3.5-9B weights, bf16 | | `head.pt` | pointer head (query and key projections, 256-dim) | | `export.json` | temperature (1.859), prompt format, training step and config | | `tokenizer.json`, `vocab.json`, `merges.txt`, `tokenizer_config.json`, `chat_template.jinja` | Qwen3.5 tokenizer | | `drafter_k4.safetensors` | block-4 drafter, the serving default, bf16 | | `drafter_k8.safetensors` | block-8 drafter, faster for single-question requests, bf16 | | `drafter_k4.config.json`, `drafter_k8.config.json` | drafter training settings | | `LICENSE` | Apache-2.0, inherited from Qwen3.5-9B | ## Usage ```bash hf download PostHog/jeeves --local-dir jeeves-weights git clone https://github.com/PostHog/jeeves && cd jeeves python -m inference.serve --model ../jeeves-weights --drafter ../jeeves-weights/drafter_k4.safetensors --port 8009 ``` ```bash curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{ "state": "I was charged twice for one order.", "questions": { "billing": {"type": "noul", "instructions": "Is this about billing?"}, "tone": {"type": "choice", "instructions": "What is the customer'"'"'s tone?", "criteria": {"calm": null, "frustrated": null, "angry": null}} }}' ``` The request and response follow Jev's `/v1/systemone` format. An optional `options` object sets `think`, `max_think` (truncate chains), `nothink_threshold` (skip thinking when the no-think answer is already confident) and `return_reasoning`. Serving runs on a CUDA GPU or on an Apple Silicon Mac with 48 GB (add `--max-rows 4 --max-len 4096` there). It serves bf16 by default; `--precision fp8` runs the linear layers in FP8 on GPUs with compute capability 8.9 or higher and on Macs. FP8 weights can be downloaded from [PostHog/jeeves-fp8](https://huggingface.co/PostHog/jeeves-fp8). ## Results Accuracy with thinking, greedy, 2,560-token cap. Kev-9B and Jev numbers are the ones Kev publishes. JevBench uses the same 231 public items for every model (the sealed judge tier is not included); the other rows use the same sources with different items. | | Kev-9B | Jev | Jeeves-9B | |---|---|---|---| | Test overall (out-of-domain and held-out) | 0.822 | 0.857 | **0.889** | | MMLU-Pro and buried state | 0.579 | **0.800** | 0.746 | | JevBench, public tiers | 0.715* | 0.866 | **0.935** | | JevBench hard | 0.451* | 0.730 | **0.865** | \* Kev-8B (Qwen3); no Kev-9B JevBench result is published. Without thinking the model scores 0.804 on the Jeeves test split, against 0.840 with it. With the fitted temperature the no-think path has a top-label calibration error of 0.021 and 1.3% confident errors (wrong at p ≥ 0.9). Serving on one H100, 325 dev questions: | setting | accuracy | median / p90 latency | |---|---|---| | full thinking | 0.825 | 3.3 s / 17.1 s | | `max_think` 768, `nothink_threshold` 0.9 | 0.806 | 2.0 s / 5.6 s | | no thinking | 0.775 | about 0.3 s | ## Training - **SFT.** LoRA r=16 on every projection plus the pointer head, 2 epochs on 19,126 questions from 12 public datasets and synthetic policy data; half the questions carry a reasoning chain sampled from the base model. - **CISPO.** 9,992 RL questions, 8 rollouts each, up to 2,560 thinking tokens. Reward is the probability of the correct option, discounted by up to 10% for long chains. The released checkpoint is step 402 of a 624-step schedule. - **Calibration.** One softmax temperature fitted on the dev set. - **Drafters.** Diffusion views of the frozen model, inspired by Orthrus, distilled by KL divergence on chains sampled from this model. At one question the block-4 drafter decodes 1.6× faster than graphed greedy decoding and the block-8 drafter 1.76×. A from-scratch reproduction with the released code matched this checkpoint within noise. The training code, dataset build scripts and pinned dataset revisions are in the GitHub repository. ## Limitations - Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840). - Thinking is slow at the tail (17 s at p90 with full chains); use `max_think` and `nothink_threshold` when latency matters. - About a third of full-length chains hit the token cap without closing; accuracy on those items is lower. - Evaluated on English inputs only. ## License The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license (`LICENSE`). The Jeeves code is MIT licensed. ## Citation ```bibtex @software{waltz2026jeeves, author = {Waltz, Nicholas P.}, title = {Jeeves: Reasoning Improves Jev-like Decisions}, year = {2026}, url = {https://github.com/PostHog/jeeves}, note = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter} } ```