dspark-mtp-draft-head-0731reap-v1 β€” DSpark MTP draft head for 0xSero/DeepSeek-V4-Flash-0731-REAP

This is not a standalone language model. It is a fine-tuned set of DSpark MTP draft-head tensors for 0xSero/DeepSeek-V4-Flash-0731-REAP, used as the speculator in self-speculative decoding. It only does anything in combination with the base checkpoint.

What it is for

Speculative decoding with greedy verification: the draft head proposes several tokens, the base model verifies them in one pass, and the longest correct prefix is accepted. Because verification is greedy and exact, the emitted token sequence is identical to what the base model would have produced on its own β€” this is a latency optimisation, not a quality trade. Every head in this line is gated on that property before it is measured (LOSSLESS GATE: first 8 tokens match base AR), and a head that changes the output is rejected regardless of how fast it is.

Measured

Hardware: Jetson AGX Thor (sm_110a, 20 SMs, 122.8 GiB unified LPDDR5X, CUDA 13.0), custom CUDA inference engine, block size 6.

tok/s
base autoregressive decode (no speculation) 13.8
stock DSpark head shipped with the checkpoint 22.66
this head 25.5312

Acceptance tau (tokens per verify, 8-prompt suite mean): 3.8438 out of a maximum of 7.

Protocol, which is part of the number. 8-prompt suite spanning agentic_format, code_edit, code_gen, explanation, long_context, multi_turn, reasoning and short_factual; NGEN0 >= 200 generated tokens (acceptance is a transient below ~128 tokens and a short-generation figure is not comparable to anything); clocks pinned; page cache dropped; no profiling instruments in the binary; adaptK 1.50. Measurements on this box have a residual sd of ~0.5 % per run with a discarded warm-up and shuffled run order β€” differences below ~1 % are not meaningful without replicates.

Limitations, measured rather than assumed

  1. It regresses on held-out continuation drafting. Against a true paired control β€” the stock head over the same prompts, budget and threshold β€” the fine-tuned heads are slower at drafting the model's own long continuations (-0.40 tau for the first session's head, -0.65 for the second). They win the mixed suite and lose this. Both numbers are real; which one matters depends on the workload.
  2. Training helps weak categories and hurts strong ones. Rank-ordered across the suite: the categories the stock head was worst at improved most, and the ones it was best at regressed. Balancing the training corpus repaired some of that (short_factual, half of code_edit) and left long_context and agentic_format untouched.
  3. Tuned for one engine and one checkpoint. The draft head taps specific backbone layers and is fine-tuned against activations captured from this exact checkpoint. It is not expected to transfer to another quantisation, another REAP revision, or another serving stack.

Files

  • config.json (0 MB)
  • eval.log (0 MB) β€” the raw measurement log behind the numbers above
  • head_card.json (0 MB) β€” the archive record written when this head was measured
  • mtp_trained.safetensors (1035 MB) β€” the trained MTP draft-head tensors β€” the artifact
  • train_metrics.json (0 MB) β€” per-step training loss history

How to use it

The loadable head is the base checkpoint with these tensors substituted in, re-quantised to the formats the checkpoint uses (fp8 e4m3 with F8_E8M0 128x128 block scales for the projections, bf16 and fp32 for the rest). Materialise it locally:

git clone https://github.com/<your-org>/deepseek-v4-flash-0731-cuda
python3 tools/build_trained_head.py \
    --base /path/to/DeepSeek-V4-Flash-0731-REAP \
    --trained mtp_trained.safetensors \
    --out /path/to/head_out

The build is deterministic and self-checked: it recomputes each block scale from the trained values (reusing the old scale against new weights is the silent way to get a plausible tensor with the wrong magnitude), snaps that scale to an exact power of two before quantising because E8M0 stores a bare exponent, and refuses to write any tensor whose round-trip error exceeds 0.10. Worst observed round-trip error for this head: see provenance.json.

Provenance

Base model 0xSero/DeepSeek-V4-Flash-0731-REAP, used as shipped β€” no re-pruning and no additional quantisation. Trained on activations captured from that checkpoint over the session's prompt continuations the base model itself generated. Full training and measurement history, including the rejected candidates and the corrections, is in LOOP_LOG.md and S5_PROGRESSION.md in the engine repository.

Engine revision at measurement time: 85dbea6cf.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for patrickbdevaney/deepseek-v4-flash-0731-REAP-calibrated-draft-head

Quantized
(2)
this model