--- license: other base_model: k-l-lambda/kimi-k2.6-eagle3-mla base_model_relation: quantized tags: - text-generation - speculative-decoding - eagle3 - kimi-k2.6 - mla - torchspec - fp8 - compressed-tensors --- # kimi-k2.6-eagle3-mla-fp8 FP8 (W8A8 dynamic) quantization of [k-l-lambda/kimi-k2.6-eagle3-mla](https://huggingface.co/k-l-lambda/kimi-k2.6-eagle3-mla), an Eagle3 MLA draft model for speculative decoding of [Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6). Format is `compressed-tensors` / `float-quantized`: per-output-channel static FP8 weight scales plus dynamic per-token activation scales — the same layout llm-compressor's `FP8_DYNAMIC` recipe produces. Served directly by vLLM via `CompressedTensorsW8A8Fp8`. ## Read this before using it **FP8 saves only ~11% of file size here, and buys no measured speed.** This is a 3.02 B-parameter draft in which `embed_tokens` and `lm_head` (163840x7168 each) are **2.35 B params — 78% of the model** — and neither is quantizable. Only the 22% in linear layers converts: | | size | |---|---:| | bf16 original | 5.617 GiB | | this FP8 model | 5.004 GiB (**10.9%** smaller) | Measured per-user decode throughput is **indistinguishable** from the bf16 original (see below). Use this model for **kernel-path compatibility with an FP8 serving stack**, not to save memory or gain speed. ## What is quantized Quantized to `float8_e4m3fn` with per-channel scales (8 tensors): `fc`, `mlp.{gate,up,down}_proj`, `self_attn.{q_a_proj,q_b_proj,kv_a_proj_with_mqa,o_proj}`. Left in bf16, bit-identical to the source (9 tensors): - `embed_tokens`, `lm_head` — vLLM builds the Eagle3 `ParallelLMHead` with no quant_config, so a quantized head would never be read. - `kv_b_proj` — MLA absorbs it into W_UK/W_UV at load time via `get_and_maybe_dequant_weights()`; for compressed-tensors that takes a generic O(N^3) identity-matmul dequant path. It is 8.4 M params (0.3%), so bf16 costs nothing and avoids that path. - all RMSNorm weights. Expressed in the config as `ignore: ["lm_head", "re:.*kv_b_proj.*"]`. `strategy: channel` is deliberate, not incidental: vLLM fuses `q_a_proj`+`kv_a_proj_with_mqa` into `fused_qkv_a_proj` and `gate_proj`+`up_proj` into `gate_up_proj`. Channel scales load through `ChannelQuantScaleParameter` (a `_ColumnvLLMParameter`) and concatenate per shard along dim 0. A per-tensor scheme would require the fused shards to share one scale. ## Weight-level fidelity 512 tokens of N(0,1) activations, dequantized-FP8 matmul vs bf16, fp32 reference, worst case across all 8 quantized layers: | metric | worst | |---|---:| | cosine similarity | 0.999648 | | output relative L2 error | 2.654% | Error is uniform (2.644–2.654%) across every layer — the E4M3 floor for Gaussian weights under per-channel scaling, not a per-layer defect. ## Evaluation vLLM 0.20.0, verifier `Kimi-K2.6`, 8xH200, TP=1 / EP=8 / DP=8, `kv-cache-dtype fp8`, `max-model-len 32768`, `num_speculative_tokens=3`, greedy. Four held-out test sets derived from real Kimi-K2.6 production traffic (1000 conversations each; accept_len over a 128-prompt subsample). **accept_length** (higher is better): | model | random | general-chat | multimodal | agentic | |---|---:|---:|---:|---:| | lightseekorg/kimi-k2.6-eagle3-mla (init) | 2.349 | 2.385 | 2.469 | 2.363 | | k-l-lambda/kimi-k2.6-eagle3-mla (bf16) | 2.380 | 2.383 | 2.468 | 2.409 | | **this model (FP8)** | 2.400 | 2.393 | 2.496 | 2.403 | Deltas vs bf16: **+0.020 / +0.010 / +0.028 / −0.006**. Mixed signs, all within ±0.03 — read this as *quantization costs no measurable accept_length*, **not** as FP8 being better. Run-to-run variance at this subsample size has not been measured, so ±0.03 is not an established error bar. **per-position acceptance** (pos0/pos1/pos2 %): | model | random | general-chat | multimodal | agentic | |---|---|---|---|---| | bf16 | 66/43/29 | 66/43/29 | 68/46/32 | 67/44/30 | | FP8 | 67/44/29 | 66/44/29 | 69/48/33 | 67/44/29 | **per-user decode TPS** (speedup vs no-spec baseline): | c | model | random | general-chat | multimodal | agentic | |---|---|---|---|---|---| | 128 | no-spec baseline | 7.84 | 7.93 | 10.02 | 10.69 | | 128 | bf16 | 9.74 (1.24x) | 8.79 (1.11x) | 10.88 (1.09x) | 10.02 (0.94x) | | 128 | FP8 | 9.62 (1.23x) | 8.83 (1.11x) | 10.89 (1.09x) | 9.90 (0.93x) | | 192 | no-spec baseline | 6.54 | 6.46 | 8.59 | 8.43 | | 192 | bf16 | 7.69 (1.18x) | 7.17 (1.11x) | 8.93 (1.04x) | 8.57 (1.02x) | | 192 | FP8 | 7.83 (1.20x) | 7.27 (1.13x) | 8.86 (1.03x) | 8.16 (0.97x) | | 256 | no-spec baseline | 5.57 | 5.62 | 7.23 | 7.31 | | 256 | bf16 | 6.76 (1.21x) | 6.32 (1.12x) | 7.93 (1.10x) | 7.37 (1.01x) | | 256 | FP8 | 6.97 (1.25x) | 6.36 (1.13x) | 8.08 (1.12x) | 7.39 (1.01x) | FP8-vs-bf16 TPS differences scatter in both directions (max |Δ| 0.41 tok/s) with no consistent sign by dataset or concurrency. Expected: the draft forward pass is a small fraction of each speculation step, and only 22% of its params are quantized, so the kernel gain does not surface. Note on the long-conversation set: **speculative decoding is a net loss on `agentic` at c=128** (0.93–0.94x) for *all three* drafts, breaking even only at higher concurrency. That workload averages 28.5 turns with 100% tool use and has the highest no-spec baseline, leaving least headroom for drafting to win. This is a property of the workload, not of quantization. ## Requirements - Compute capability **>= 89** (Ada / Hopper). `CompressedTensorsW8A8Fp8` reports `min_capability 89`; H200 is 90. - vLLM with draft-model quantization support. Verified on **0.20.0** (`compressed_tensors` 0.15.0.1), where `deepseek_eagle3.py` calls `get_draft_quant_config()` so the draft's own quant config is honored rather than the verifier's. ## Usage ```bash vllm serve /path/to/Kimi-K2.6 \ --speculative-config '{"method":"eagle3","model":"k-l-lambda/kimi-k2.6-eagle3-mla-fp8","num_speculative_tokens":3}' ``` ## Inherited caveat The bf16 parent is a K2.6 fine-tune selected by validation loss on the K2.6 teacher distribution, and its card notes it can over-specialize: on K2.7-Code production traffic the official `lightseekorg` init showed higher accept-length. The evaluation above is K2.6 traffic, where the fine-tune is ahead or level — it does not test, and so does not refute, the K2.7 claim. Benchmark against your own traffic before choosing a draft.