File size: 2,688 Bytes
a202601
 
 
134557f
a202601
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
134557f
 
 
 
a202601
 
 
 
 
134557f
a202601
 
 
134557f
 
 
 
 
 
 
 
 
 
 
 
a202601
134557f
 
 
 
 
 
 
 
 
 
a202601
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
134557f
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
---
pipeline_tag: text-generation
base_model: moonshotai/Kimi-K3
checkpoint_step: 5780
tags:
  - speculative-decoding
  - vllm
  - dspark
  - kimi-k3
---

# Kimi-K3 DSpark — Block 5

A DSpark draft model for speculative decoding with
[`moonshotai/Kimi-K3`](https://huggingface.co/moonshotai/Kimi-K3). Three full-causal MLA
layers, conditioned on hidden states extracted from three of the target's layers and
trained at `block_size=5`: one forward pass drafts the full five-token block.

## Training

The draft completed a four-epoch run with
[**speculators**](https://github.com/vllm-project/speculators), the speculative-decoding
library from the [vLLM project](https://github.com/vllm-project), against Kimi-K3 hidden
states streamed live from **vLLM**.

## Acceptance

Evaluated against Kimi-K3 on 8×B300 (TP8), drafting 5 tokens with probabilistic
sampling and block rejection at temperature 1.0 and top-p 0.95. All 1,604 requests
succeeded. The aggregate is computed from summed counters, not an average of arm rates.

| Workload | Requests | Acceptance | Mean accepted length (max 6) |
| --- | ---: | ---: | ---: |
| GSM8K | 256 | 71.62% | 4.581 |
| HumanEval | 164 | 63.74% | 4.187 |
| MBPP | 256 | 56.38% | 3.819 |
| MATH-500 | 500 | 47.48% | 3.374 |
| SWE-bench Pro | 128 | 38.09% | 2.904 |
| MT-Bench | 80 | 33.19% | 2.659 |
| AIME26 | 30 | 29.78% | 2.489 |
| AA-LCR (~100K) | 100 | 41.04% | 3.052 |
| BEAM 100K | 20 | 37.17% | 2.858 |
| BEAM 500K | 35 | 30.26% | 2.513 |
| BEAM 1M | 35 | 31.37% | 2.569 |
| **Aggregate** | **1,604** | **46.10%** | **3.305** |

The separate five-shot GSM8K accuracy check scored **96.51% exact match** on all 1,319
examples with zero request errors.

**The full 1M context window is supported through YaRN**, evaluated on every row of the
[BEAM](https://huggingface.co/datasets/Mohammadta/BEAM) 100K, 500K, and 1M splits above.
The fixed 1M acceptance corpus left-truncates oldest turns in 14 of 35 over-limit source
conversations to keep prompt plus output within 1,048,576 tokens; its longest measured
prompt is 1,039,871 tokens. Two successful completions were flagged as degenerate by the
client (BEAM 500K conversation 4 and BEAM 1M conversation 18). Full counters,
per-position rates, and quality flags are in `benchmark_results.json`.

## Serving

```bash
vllm serve moonshotai/Kimi-K3 \
  --trust-remote-code \
  --tensor-parallel-size 8 \
  --max-model-len 1048576 \
  --speculative-config '{
    "method": "dspark",
    "model": "Inferact/Kimi-K3-DSpark-Block5",
    "num_speculative_tokens": 5,
    "attention_backend": "FLASHINFER_MLA",
    "draft_sample_method": "probabilistic",
    "rejection_sample_method": "block"
  }'
```