| --- |
| license: other |
| base_model: k-l-lambda/kimi-k2.6-eagle3-mla |
| base_model_relation: quantized |
| tags: |
| - text-generation |
| - speculative-decoding |
| - eagle3 |
| - kimi-k2.6 |
| - mla |
| - torchspec |
| - fp8 |
| - compressed-tensors |
| --- |
| |
| # kimi-k2.6-eagle3-mla-fp8 |
|
|
| FP8 (W8A8 dynamic) quantization of |
| [k-l-lambda/kimi-k2.6-eagle3-mla](https://huggingface.co/k-l-lambda/kimi-k2.6-eagle3-mla), |
| an Eagle3 MLA draft model for speculative decoding of |
| [Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6). |
|
|
| Format is `compressed-tensors` / `float-quantized`: per-output-channel static |
| FP8 weight scales plus dynamic per-token activation scales — the same layout |
| llm-compressor's `FP8_DYNAMIC` recipe produces. Served directly by vLLM via |
| `CompressedTensorsW8A8Fp8`. |
|
|
| ## Read this before using it |
|
|
| **FP8 saves only ~11% of file size here, and buys no measured speed.** This is |
| a 3.02 B-parameter draft in which `embed_tokens` and `lm_head` (163840x7168 |
| each) are **2.35 B params — 78% of the model** — and neither is quantizable. |
| Only the 22% in linear layers converts: |
|
|
| | | size | |
| |---|---:| |
| | bf16 original | 5.617 GiB | |
| | this FP8 model | 5.004 GiB (**10.9%** smaller) | |
|
|
| Measured per-user decode throughput is **indistinguishable** from the bf16 |
| original (see below). Use this model for **kernel-path compatibility with an |
| FP8 serving stack**, not to save memory or gain speed. |
|
|
| ## What is quantized |
|
|
| Quantized to `float8_e4m3fn` with per-channel scales (8 tensors): |
| `fc`, `mlp.{gate,up,down}_proj`, |
| `self_attn.{q_a_proj,q_b_proj,kv_a_proj_with_mqa,o_proj}`. |
|
|
| Left in bf16, bit-identical to the source (9 tensors): |
|
|
| - `embed_tokens`, `lm_head` — vLLM builds the Eagle3 `ParallelLMHead` with no |
| quant_config, so a quantized head would never be read. |
| - `kv_b_proj` — MLA absorbs it into W_UK/W_UV at load time via |
| `get_and_maybe_dequant_weights()`; for compressed-tensors that takes a |
| generic O(N^3) identity-matmul dequant path. It is 8.4 M params (0.3%), so |
| bf16 costs nothing and avoids that path. |
| - all RMSNorm weights. |
| |
| Expressed in the config as `ignore: ["lm_head", "re:.*kv_b_proj.*"]`. |
|
|
| `strategy: channel` is deliberate, not incidental: vLLM fuses |
| `q_a_proj`+`kv_a_proj_with_mqa` into `fused_qkv_a_proj` and |
| `gate_proj`+`up_proj` into `gate_up_proj`. Channel scales load through |
| `ChannelQuantScaleParameter` (a `_ColumnvLLMParameter`) and concatenate per |
| shard along dim 0. A per-tensor scheme would require the fused shards to share |
| one scale. |
|
|
| ## Weight-level fidelity |
|
|
| 512 tokens of N(0,1) activations, dequantized-FP8 matmul vs bf16, fp32 |
| reference, worst case across all 8 quantized layers: |
|
|
| | metric | worst | |
| |---|---:| |
| | cosine similarity | 0.999648 | |
| | output relative L2 error | 2.654% | |
|
|
| Error is uniform (2.644–2.654%) across every layer — the E4M3 floor for |
| Gaussian weights under per-channel scaling, not a per-layer defect. |
|
|
| ## Evaluation |
|
|
| vLLM 0.20.0, verifier `Kimi-K2.6`, 8xH200, TP=1 / EP=8 / DP=8, `kv-cache-dtype |
| fp8`, `max-model-len 32768`, `num_speculative_tokens=3`, greedy. Four held-out |
| test sets derived from real Kimi-K2.6 production traffic (1000 conversations |
| each; accept_len over a 128-prompt subsample). |
| |
| **accept_length** (higher is better): |
| |
| | model | random | general-chat | multimodal | agentic | |
| |---|---:|---:|---:|---:| |
| | lightseekorg/kimi-k2.6-eagle3-mla (init) | 2.349 | 2.385 | 2.469 | 2.363 | |
| | k-l-lambda/kimi-k2.6-eagle3-mla (bf16) | 2.380 | 2.383 | 2.468 | 2.409 | |
| | **this model (FP8)** | 2.400 | 2.393 | 2.496 | 2.403 | |
| |
| Deltas vs bf16: **+0.020 / +0.010 / +0.028 / −0.006**. Mixed signs, all within |
| ±0.03 — read this as *quantization costs no measurable accept_length*, **not** |
| as FP8 being better. Run-to-run variance at this subsample size has not been |
| measured, so ±0.03 is not an established error bar. |
| |
| **per-position acceptance** (pos0/pos1/pos2 %): |
| |
| | model | random | general-chat | multimodal | agentic | |
| |---|---|---|---|---| |
| | bf16 | 66/43/29 | 66/43/29 | 68/46/32 | 67/44/30 | |
| | FP8 | 67/44/29 | 66/44/29 | 69/48/33 | 67/44/29 | |
| |
| **per-user decode TPS** (speedup vs no-spec baseline): |
| |
| | c | model | random | general-chat | multimodal | agentic | |
| |---|---|---|---|---|---| |
| | 128 | no-spec baseline | 7.84 | 7.93 | 10.02 | 10.69 | |
| | 128 | bf16 | 9.74 (1.24x) | 8.79 (1.11x) | 10.88 (1.09x) | 10.02 (0.94x) | |
| | 128 | FP8 | 9.62 (1.23x) | 8.83 (1.11x) | 10.89 (1.09x) | 9.90 (0.93x) | |
| | 192 | no-spec baseline | 6.54 | 6.46 | 8.59 | 8.43 | |
| | 192 | bf16 | 7.69 (1.18x) | 7.17 (1.11x) | 8.93 (1.04x) | 8.57 (1.02x) | |
| | 192 | FP8 | 7.83 (1.20x) | 7.27 (1.13x) | 8.86 (1.03x) | 8.16 (0.97x) | |
| | 256 | no-spec baseline | 5.57 | 5.62 | 7.23 | 7.31 | |
| | 256 | bf16 | 6.76 (1.21x) | 6.32 (1.12x) | 7.93 (1.10x) | 7.37 (1.01x) | |
| | 256 | FP8 | 6.97 (1.25x) | 6.36 (1.13x) | 8.08 (1.12x) | 7.39 (1.01x) | |
| |
| FP8-vs-bf16 TPS differences scatter in both directions (max |Δ| 0.41 tok/s) with |
| no consistent sign by dataset or concurrency. Expected: the draft forward pass |
| is a small fraction of each speculation step, and only 22% of its params are |
| quantized, so the kernel gain does not surface. |
| |
| Note on the long-conversation set: **speculative decoding is a net loss on |
| `agentic` at c=128** (0.93–0.94x) for *all three* drafts, breaking even only at |
| higher concurrency. That workload averages 28.5 turns with 100% tool use and |
| has the highest no-spec baseline, leaving least headroom for drafting to win. |
| This is a property of the workload, not of quantization. |
|
|
| ## Requirements |
|
|
| - Compute capability **>= 89** (Ada / Hopper). `CompressedTensorsW8A8Fp8` |
| reports `min_capability 89`; H200 is 90. |
| - vLLM with draft-model quantization support. Verified on **0.20.0** |
| (`compressed_tensors` 0.15.0.1), where `deepseek_eagle3.py` calls |
| `get_draft_quant_config()` so the draft's own quant config is honored rather |
| than the verifier's. |
|
|
| ## Usage |
|
|
| ```bash |
| vllm serve /path/to/Kimi-K2.6 \ |
| --speculative-config '{"method":"eagle3","model":"k-l-lambda/kimi-k2.6-eagle3-mla-fp8","num_speculative_tokens":3}' |
| ``` |
|
|
| ## Inherited caveat |
|
|
| The bf16 parent is a K2.6 fine-tune selected by validation loss on the K2.6 |
| teacher distribution, and its card notes it can over-specialize: on |
| K2.7-Code production traffic the official `lightseekorg` init showed higher |
| accept-length. The evaluation above is K2.6 traffic, where the fine-tune is |
| ahead or level — it does not test, and so does not refute, the K2.7 claim. |
| Benchmark against your own traffic before choosing a draft. |
|
|