Jeeves-9B FP8
FP8 weights of PostHog/jeeves, a Jev-like decision model that reasons before it decides. Every large linear layer of the model and of the block-4 drafter is stored in FP8 (e4m3fn), with one float32 scale per output row. The embeddings, the norms, the Gated DeltaNet gates and the pointer head stay as in the bf16 release. The download is 11.6 GB. The bf16 model and its block-4 drafter are 21 GB.
--precision fp8 with the bf16 weights gives the same outputs. These weights load with the Jeeves code, not with transformers.
Files
| file | contents |
|---|---|
model-*.safetensors, config.json |
fused Qwen3.5-9B: each quantized layer's *.weight in F8_E4M3 with a float32 *.weight_scale per output row, other tensors bf16 |
drafter_k4.safetensors |
block-4 drafter, projections in FP8 the same way |
head.pt |
pointer head, unchanged |
export.json |
temperature (1.859), prompt format and training config, with "precision": "fp8" |
tokenizer.json, vocab.json, merges.txt, tokenizer_config.json, chat_template.jinja |
Qwen3.5 tokenizer |
LICENSE |
Apache-2.0, inherited from Qwen3.5-9B |
Usage
hf download PostHog/jeeves-fp8 --local-dir jeeves-fp8
git clone https://github.com/PostHog/jeeves && cd jeeves
python -m inference.serve --model ../jeeves-fp8 --drafter ../jeeves-fp8/drafter_k4.safetensors --precision fp8 --port 8009
The engine serves these weights on CUDA GPUs with compute capability 8.9 or higher and on Apple Silicon Macs. On a Mac with 48 GB, add --max-rows 4 --max-len 4096. These weights only run with --precision fp8.
The request and response follow Jev's /v1/systemone format. See the bf16 model card and the GitHub README for the API, the options and the results.
Accuracy and speed
FP8 changes the outputs slightly against bf16. On dev questions, accuracy and NLL did not change measurably. The serving results in the bf16 model card (325 dev questions, accuracy 0.825 with full thinking) come from --precision fp8 on an H100.
| bf16 | FP8 | |
|---|---|---|
| GPU memory for the weights | 21 GB | 11.5 GB |
speed.py on an L40S (12 dev requests) |
9.5 s | 7.9 s |
| one question thinking on an M4 Pro (48 GB) | about 20 tokens/s | about 38 tokens/s |
How they were made
hf download PostHog/jeeves --revision 8622b7d1652a9dcb8629486b84dce9e8d690c5cd --local-dir jeeves-weights
python export_fp8.py jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --out jeeves-fp8
with export_fp8.py from the GitHub repository. Each row's scale is the row's largest absolute weight divided by 448, the largest e4m3fn value. Each weight is divided by its row's scale and rounded to the nearest e4m3fn value.
License
The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license (LICENSE). The Jeeves code is MIT licensed.
- Downloads last month
- 9