| --- |
| license: mit |
| base_model: 0xSero/DeepSeek-V4-Flash-0731-REAP |
| tags: |
| - speculative-decoding |
| - draft-head |
| - dspark |
| - mtp |
| - jetson |
| - edge-inference |
| library_name: safetensors |
| pipeline_tag: text-generation |
| --- |
| |
| # DeepSeek-V4-Flash-0731-REAP + DSpark MTP draft head — complete checkpoint |
|
|
| A **drop-in replacement** for [`0xSero/DeepSeek-V4-Flash-0731-REAP`](https://huggingface.co/0xSero/DeepSeek-V4-Flash-0731-REAP) |
| with a fine-tuned DSpark MTP draft head merged in. Download it and you have the whole model: no |
| patching step, no separate head file to graft on. |
|
|
| **Every emitted token is bit-identical to base autoregressive decode.** The speed-up is lossless by |
| construction — speculation changes how tokens are *produced*, never which tokens they are — and that |
| invariant is checked on every run by a gate that compares the speculative path against plain AR. |
|
|
| ## What is actually different from the base checkpoint |
|
|
| Three files. Shards **46, 47 and 48** carry the MTP draft-head tensors; shards 1–45, the tokenizer, |
| and the config are byte-identical to the base. If you already have the base checkpoint, the only new |
| weights here are ~6.7 GB of head shards. |
|
|
| | | | |
| |---|---| |
| | provenance | `deepseek-ai/DeepSeek-V4-Flash-0731` @ `9e165c30` → REAP prune (256 → **160** experts/layer) → this head fine-tune | |
| | head trained on | 1,472 self-generated sequences, deficit-weighted, `a_ce` 0.1 / `a_tv` 0.9 | |
| | draft block width | **5** (the checkpoint's own `dspark_block_size`) | |
| | quantisation | native MXFP4 experts, FP8 attention — unchanged from base | |
|
|
| ## Measured speculative decode |
|
|
| Measured on a **Jetson AGX Thor (`sm_110a`, 122.8 GiB unified)** with a from-scratch pure-CUDA |
| inference server. Every figure is from a logged run, not an estimate. |
| |
| | | measured | |
| |---|---| |
| | suite mean `tau` | **3.8413** of a 5-wide draft (77 % of the ceiling) | |
| | suite mean throughput | **28.38 tok/s** | |
| | base AR decode | 14.61 tok/s → speculation is **1.94×** | |
| | vs the stock shipped head | **22.66 → 28.38 tok/s, +25.3 %** | |
| | on copy-heavy agentic prompts | up to **34.97 tok/s, 2.51×** | |
| |
| ### Acceptance by workload, and where this head gives ground |
| |
| `tau` is tokens committed per target forward. The head was selected on the **mean**, and the mean |
| hides a real trade — published here rather than buried: |
| |
| | category | stock head | **this head** | | |
| |---|---:|---:|---| |
| | long_context | 5.00 | 4.53 | **−0.47** | |
| | agentic_format | 4.41 | 4.16 | **−0.25** | |
| | multi_turn | 4.06 | 4.83 | +0.77 | |
| | code_edit | 4.08 | 4.18 | +0.10 | |
| | short_factual | 3.11 | 3.87 | +0.76 | |
| | reasoning | 1.82 | 3.57 | +1.75 | |
| | explanation | 1.70 | 3.00 | +1.30 | |
| | code_gen | 1.86 | 2.59 | +0.73 | |
| | **mean** | 3.2550 | **3.8413** | **+18.0 %** | |
| |
| The gain is concentrated in the *constructive* categories — reasoning nearly doubles — and it is |
| paid for on the two *reconstructive* ones the stock head was already best at. Under this project's |
| own release rule that is a **failure of 3 of 6 category floors**, and the rule is stated so that a |
| downstream user can predict their own workload instead of inheriting a mean. |
| |
| **If your workload is dominated by long-context reconstruction, the stock head may serve you better.** |
| For mixed agentic work — tool calls, multi-turn, code edits, reasoning — this head is a clear win. |
| |
| ## Task accuracy |
| |
| Measured through the same server, `effort=low`, 24k token budget, temperature 1.0 / top-p 0.95. |
| |
| | benchmark | score | n | |
| |---|---:|---:| |
| | AIME 2024 | **91.7 %** [80.0, 100.0] | 60 | |
| | GPQA-Diamond | **83.8 %** [78.1, 88.3] | 198 | |
| | MMLU-Pro | **74.7 %** [67.2, 81.0] | 150 | |
| | BFCL (tool calling) | **86.2 %** | 240 | |
| | BFCL-Live | **78.7 %** | 508 | |
| | LiveCodeBench | 46.9 % [39.6, 54.2] *(8k budget, 59 % truncated — not quotable)* | 175 | |
| |
| **A caveat that matters more than any single number.** GPQA at a 24k budget was *pre-registered* at |
| 87–93 % before the run, with a stated falsification threshold of 85 %. It landed at **83.8 %** — the |
| prediction was wrong, and the reason is a finding about this pruned model: **19 of 51 items that |
| exhaust an 8k budget never terminate even at 24k**, and the ones that do finish score 78.1 %, not the |
| 93.9 % that items terminating inside 8k reach. Prompts needing long chains are *qualitatively* |
| harder here, not merely longer. Budget your context accordingly. |
| |
| ## Context |
| |
| | context | KV format | fits | |
| |---:|---|---| |
| | 32,768 | fp32 rows | yes, 119.9 / 122.8 GiB | |
| | **131,072** | packed fp8 + ue8m0 | **yes, 117.0 / 122.8 GiB** | |
| | 262,144 | packed | allocates at 121.8 / 122.8 — **too little headroom to serve** | |
| |
| KV packing is **bit-exact** (720 B/row vs 2048; the fp32 path already stored e4m3-representable |
| values in four bytes) and costs ~12 % of prefill throughput, so it is worth enabling only when the |
| context needs it. |
| |
| ## Limitations, stated plainly |
| |
| - **Measured on one machine.** Every number above comes from a single Jetson AGX Thor. Throughput |
| will differ on other hardware; accuracy should not. |
| - **Single-stream.** The server these numbers came from serves one request at a time by design. This |
| is not a multi-tenant serving checkpoint. |
| - **The eval battery is partial.** The suite is complete; the 24k extension covers 4 of 8 tasks. |
| LiveCodeBench, MATH-500, HumanEval and SciCode are at the 8k budget and several are truncation- |
| limited, which is why LCB is marked not quotable above. |
| - **This is a REAP-pruned model.** 160 of 256 experts per layer. Comparisons against unpruned |
| DeepSeek-V4-Flash are not like-for-like. |
| |
| ## Licence |
| |
| MIT, following the base checkpoint. The DeepSeek copyright notice is retained in `LICENSE`. This |
| repository adds only the fine-tuned draft-head tensors in shards 46–48. |
| |