| --- |
| pipeline_tag: text-generation |
| base_model: moonshotai/Kimi-K3 |
| checkpoint_step: 5780 |
| tags: |
| - speculative-decoding |
| - vllm |
| - dspark |
| - kimi-k3 |
| --- |
| |
| # Kimi-K3 DSpark — Block 5 |
|
|
| A DSpark draft model for speculative decoding with |
| [`moonshotai/Kimi-K3`](https://huggingface.co/moonshotai/Kimi-K3). Three full-causal MLA |
| layers, conditioned on hidden states extracted from three of the target's layers and |
| trained at `block_size=5`: one forward pass drafts the full five-token block. |
|
|
| ## Training |
|
|
| The draft completed a four-epoch run with |
| [**speculators**](https://github.com/vllm-project/speculators), the speculative-decoding |
| library from the [vLLM project](https://github.com/vllm-project), against Kimi-K3 hidden |
| states streamed live from **vLLM**. |
|
|
| ## Acceptance |
|
|
| Evaluated against Kimi-K3 on 8×B300 (TP8), drafting 5 tokens with probabilistic |
| sampling and block rejection at temperature 1.0 and top-p 0.95. All 1,604 requests |
| succeeded. The aggregate is computed from summed counters, not an average of arm rates. |
|
|
| | Workload | Requests | Acceptance | Mean accepted length (max 6) | |
| | --- | ---: | ---: | ---: | |
| | GSM8K | 256 | 71.62% | 4.581 | |
| | HumanEval | 164 | 63.74% | 4.187 | |
| | MBPP | 256 | 56.38% | 3.819 | |
| | MATH-500 | 500 | 47.48% | 3.374 | |
| | SWE-bench Pro | 128 | 38.09% | 2.904 | |
| | MT-Bench | 80 | 33.19% | 2.659 | |
| | AIME26 | 30 | 29.78% | 2.489 | |
| | AA-LCR (~100K) | 100 | 41.04% | 3.052 | |
| | BEAM 100K | 20 | 37.17% | 2.858 | |
| | BEAM 500K | 35 | 30.26% | 2.513 | |
| | BEAM 1M | 35 | 31.37% | 2.569 | |
| | **Aggregate** | **1,604** | **46.10%** | **3.305** | |
|
|
| The separate five-shot GSM8K accuracy check scored **96.51% exact match** on all 1,319 |
| examples with zero request errors. |
|
|
| **The full 1M context window is supported through YaRN**, evaluated on every row of the |
| [BEAM](https://huggingface.co/datasets/Mohammadta/BEAM) 100K, 500K, and 1M splits above. |
| The fixed 1M acceptance corpus left-truncates oldest turns in 14 of 35 over-limit source |
| conversations to keep prompt plus output within 1,048,576 tokens; its longest measured |
| prompt is 1,039,871 tokens. Two successful completions were flagged as degenerate by the |
| client (BEAM 500K conversation 4 and BEAM 1M conversation 18). Full counters, |
| per-position rates, and quality flags are in `benchmark_results.json`. |
|
|
| ## Serving |
|
|
| ```bash |
| vllm serve moonshotai/Kimi-K3 \ |
| --trust-remote-code \ |
| --tensor-parallel-size 8 \ |
| --max-model-len 1048576 \ |
| --speculative-config '{ |
| "method": "dspark", |
| "model": "Inferact/Kimi-K3-DSpark-Block5", |
| "num_speculative_tokens": 5, |
| "attention_backend": "FLASHINFER_MLA", |
| "draft_sample_method": "probabilistic", |
| "rejection_sample_method": "block" |
| }' |
| ``` |
|
|