File size: 5,759 Bytes
1b985a2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
---
license: mit
base_model: 0xSero/DeepSeek-V4-Flash-0731-REAP
tags:
  - speculative-decoding
  - draft-head
  - dspark
  - mtp
  - jetson
  - edge-inference
library_name: safetensors
pipeline_tag: text-generation
---

# DeepSeek-V4-Flash-0731-REAP + DSpark MTP draft head — complete checkpoint

A **drop-in replacement** for [`0xSero/DeepSeek-V4-Flash-0731-REAP`](https://huggingface.co/0xSero/DeepSeek-V4-Flash-0731-REAP)
with a fine-tuned DSpark MTP draft head merged in. Download it and you have the whole model: no
patching step, no separate head file to graft on.

**Every emitted token is bit-identical to base autoregressive decode.** The speed-up is lossless by
construction — speculation changes how tokens are *produced*, never which tokens they are — and that
invariant is checked on every run by a gate that compares the speculative path against plain AR.

## What is actually different from the base checkpoint

Three files. Shards **46, 47 and 48** carry the MTP draft-head tensors; shards 1–45, the tokenizer,
and the config are byte-identical to the base. If you already have the base checkpoint, the only new
weights here are ~6.7 GB of head shards.

| | |
|---|---|
| provenance | `deepseek-ai/DeepSeek-V4-Flash-0731` @ `9e165c30` → REAP prune (256 → **160** experts/layer) → this head fine-tune |
| head trained on | 1,472 self-generated sequences, deficit-weighted, `a_ce` 0.1 / `a_tv` 0.9 |
| draft block width | **5** (the checkpoint's own `dspark_block_size`) |
| quantisation | native MXFP4 experts, FP8 attention — unchanged from base |

## Measured speculative decode

Measured on a **Jetson AGX Thor (`sm_110a`, 122.8 GiB unified)** with a from-scratch pure-CUDA
inference server. Every figure is from a logged run, not an estimate.

| | measured |
|---|---|
| suite mean `tau` | **3.8413** of a 5-wide draft (77 % of the ceiling) |
| suite mean throughput | **28.38 tok/s** |
| base AR decode | 14.61 tok/s → speculation is **1.94×** |
| vs the stock shipped head | **22.66 → 28.38 tok/s, +25.3 %** |
| on copy-heavy agentic prompts | up to **34.97 tok/s, 2.51×** |

### Acceptance by workload, and where this head gives ground

`tau` is tokens committed per target forward. The head was selected on the **mean**, and the mean
hides a real trade — published here rather than buried:

| category | stock head | **this head** | |
|---|---:|---:|---|
| long_context | 5.00 | 4.53 | **−0.47** |
| agentic_format | 4.41 | 4.16 | **−0.25** |
| multi_turn | 4.06 | 4.83 | +0.77 |
| code_edit | 4.08 | 4.18 | +0.10 |
| short_factual | 3.11 | 3.87 | +0.76 |
| reasoning | 1.82 | 3.57 | +1.75 |
| explanation | 1.70 | 3.00 | +1.30 |
| code_gen | 1.86 | 2.59 | +0.73 |
| **mean** | 3.2550 | **3.8413** | **+18.0 %** |

The gain is concentrated in the *constructive* categories — reasoning nearly doubles — and it is
paid for on the two *reconstructive* ones the stock head was already best at. Under this project's
own release rule that is a **failure of 3 of 6 category floors**, and the rule is stated so that a
downstream user can predict their own workload instead of inheriting a mean.

**If your workload is dominated by long-context reconstruction, the stock head may serve you better.**
For mixed agentic work — tool calls, multi-turn, code edits, reasoning — this head is a clear win.

## Task accuracy

Measured through the same server, `effort=low`, 24k token budget, temperature 1.0 / top-p 0.95.

| benchmark | score | n |
|---|---:|---:|
| AIME 2024 | **91.7 %** [80.0, 100.0] | 60 |
| GPQA-Diamond | **83.8 %** [78.1, 88.3] | 198 |
| MMLU-Pro | **74.7 %** [67.2, 81.0] | 150 |
| BFCL (tool calling) | **86.2 %** | 240 |
| BFCL-Live | **78.7 %** | 508 |
| LiveCodeBench | 46.9 % [39.6, 54.2] *(8k budget, 59 % truncated — not quotable)* | 175 |

**A caveat that matters more than any single number.** GPQA at a 24k budget was *pre-registered* at
87–93 % before the run, with a stated falsification threshold of 85 %. It landed at **83.8 %** — the
prediction was wrong, and the reason is a finding about this pruned model: **19 of 51 items that
exhaust an 8k budget never terminate even at 24k**, and the ones that do finish score 78.1 %, not the
93.9 % that items terminating inside 8k reach. Prompts needing long chains are *qualitatively*
harder here, not merely longer. Budget your context accordingly.

## Context

| context | KV format | fits |
|---:|---|---|
| 32,768 | fp32 rows | yes, 119.9 / 122.8 GiB |
| **131,072** | packed fp8 + ue8m0 | **yes, 117.0 / 122.8 GiB** |
| 262,144 | packed | allocates at 121.8 / 122.8 — **too little headroom to serve** |

KV packing is **bit-exact** (720 B/row vs 2048; the fp32 path already stored e4m3-representable
values in four bytes) and costs ~12 % of prefill throughput, so it is worth enabling only when the
context needs it.

## Limitations, stated plainly

- **Measured on one machine.** Every number above comes from a single Jetson AGX Thor. Throughput
  will differ on other hardware; accuracy should not.
- **Single-stream.** The server these numbers came from serves one request at a time by design. This
  is not a multi-tenant serving checkpoint.
- **The eval battery is partial.** The suite is complete; the 24k extension covers 4 of 8 tasks.
  LiveCodeBench, MATH-500, HumanEval and SciCode are at the 8k budget and several are truncation-
  limited, which is why LCB is marked not quotable above.
- **This is a REAP-pruned model.** 160 of 256 experts per layer. Comparisons against unpruned
  DeepSeek-V4-Flash are not like-for-like.

## Licence

MIT, following the base checkpoint. The DeepSeek copyright notice is retained in `LICENSE`. This
repository adds only the fine-tuned draft-head tensors in shards 46–48.