k-l-lambda's picture
Add files using upload-large-folder tool
a71d353 verified
|
Raw
History Blame Contribute Delete
6.42 kB
---
license: other
base_model: k-l-lambda/kimi-k2.6-eagle3-mla
base_model_relation: quantized
tags:
- text-generation
- speculative-decoding
- eagle3
- kimi-k2.6
- mla
- torchspec
- fp8
- compressed-tensors
---
# kimi-k2.6-eagle3-mla-fp8
FP8 (W8A8 dynamic) quantization of
[k-l-lambda/kimi-k2.6-eagle3-mla](https://huggingface.co/k-l-lambda/kimi-k2.6-eagle3-mla),
an Eagle3 MLA draft model for speculative decoding of
[Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6).
Format is `compressed-tensors` / `float-quantized`: per-output-channel static
FP8 weight scales plus dynamic per-token activation scales — the same layout
llm-compressor's `FP8_DYNAMIC` recipe produces. Served directly by vLLM via
`CompressedTensorsW8A8Fp8`.
## Read this before using it
**FP8 saves only ~11% of file size here, and buys no measured speed.** This is
a 3.02 B-parameter draft in which `embed_tokens` and `lm_head` (163840x7168
each) are **2.35 B params — 78% of the model** — and neither is quantizable.
Only the 22% in linear layers converts:
| | size |
|---|---:|
| bf16 original | 5.617 GiB |
| this FP8 model | 5.004 GiB (**10.9%** smaller) |
Measured per-user decode throughput is **indistinguishable** from the bf16
original (see below). Use this model for **kernel-path compatibility with an
FP8 serving stack**, not to save memory or gain speed.
## What is quantized
Quantized to `float8_e4m3fn` with per-channel scales (8 tensors):
`fc`, `mlp.{gate,up,down}_proj`,
`self_attn.{q_a_proj,q_b_proj,kv_a_proj_with_mqa,o_proj}`.
Left in bf16, bit-identical to the source (9 tensors):
- `embed_tokens`, `lm_head` — vLLM builds the Eagle3 `ParallelLMHead` with no
quant_config, so a quantized head would never be read.
- `kv_b_proj` — MLA absorbs it into W_UK/W_UV at load time via
`get_and_maybe_dequant_weights()`; for compressed-tensors that takes a
generic O(N^3) identity-matmul dequant path. It is 8.4 M params (0.3%), so
bf16 costs nothing and avoids that path.
- all RMSNorm weights.
Expressed in the config as `ignore: ["lm_head", "re:.*kv_b_proj.*"]`.
`strategy: channel` is deliberate, not incidental: vLLM fuses
`q_a_proj`+`kv_a_proj_with_mqa` into `fused_qkv_a_proj` and
`gate_proj`+`up_proj` into `gate_up_proj`. Channel scales load through
`ChannelQuantScaleParameter` (a `_ColumnvLLMParameter`) and concatenate per
shard along dim 0. A per-tensor scheme would require the fused shards to share
one scale.
## Weight-level fidelity
512 tokens of N(0,1) activations, dequantized-FP8 matmul vs bf16, fp32
reference, worst case across all 8 quantized layers:
| metric | worst |
|---|---:|
| cosine similarity | 0.999648 |
| output relative L2 error | 2.654% |
Error is uniform (2.644–2.654%) across every layer — the E4M3 floor for
Gaussian weights under per-channel scaling, not a per-layer defect.
## Evaluation
vLLM 0.20.0, verifier `Kimi-K2.6`, 8xH200, TP=1 / EP=8 / DP=8, `kv-cache-dtype
fp8`, `max-model-len 32768`, `num_speculative_tokens=3`, greedy. Four held-out
test sets derived from real Kimi-K2.6 production traffic (1000 conversations
each; accept_len over a 128-prompt subsample).
**accept_length** (higher is better):
| model | random | general-chat | multimodal | agentic |
|---|---:|---:|---:|---:|
| lightseekorg/kimi-k2.6-eagle3-mla (init) | 2.349 | 2.385 | 2.469 | 2.363 |
| k-l-lambda/kimi-k2.6-eagle3-mla (bf16) | 2.380 | 2.383 | 2.468 | 2.409 |
| **this model (FP8)** | 2.400 | 2.393 | 2.496 | 2.403 |
Deltas vs bf16: **+0.020 / +0.010 / +0.028 / −0.006**. Mixed signs, all within
±0.03 — read this as *quantization costs no measurable accept_length*, **not**
as FP8 being better. Run-to-run variance at this subsample size has not been
measured, so ±0.03 is not an established error bar.
**per-position acceptance** (pos0/pos1/pos2 %):
| model | random | general-chat | multimodal | agentic |
|---|---|---|---|---|
| bf16 | 66/43/29 | 66/43/29 | 68/46/32 | 67/44/30 |
| FP8 | 67/44/29 | 66/44/29 | 69/48/33 | 67/44/29 |
**per-user decode TPS** (speedup vs no-spec baseline):
| c | model | random | general-chat | multimodal | agentic |
|---|---|---|---|---|---|
| 128 | no-spec baseline | 7.84 | 7.93 | 10.02 | 10.69 |
| 128 | bf16 | 9.74 (1.24x) | 8.79 (1.11x) | 10.88 (1.09x) | 10.02 (0.94x) |
| 128 | FP8 | 9.62 (1.23x) | 8.83 (1.11x) | 10.89 (1.09x) | 9.90 (0.93x) |
| 192 | no-spec baseline | 6.54 | 6.46 | 8.59 | 8.43 |
| 192 | bf16 | 7.69 (1.18x) | 7.17 (1.11x) | 8.93 (1.04x) | 8.57 (1.02x) |
| 192 | FP8 | 7.83 (1.20x) | 7.27 (1.13x) | 8.86 (1.03x) | 8.16 (0.97x) |
| 256 | no-spec baseline | 5.57 | 5.62 | 7.23 | 7.31 |
| 256 | bf16 | 6.76 (1.21x) | 6.32 (1.12x) | 7.93 (1.10x) | 7.37 (1.01x) |
| 256 | FP8 | 6.97 (1.25x) | 6.36 (1.13x) | 8.08 (1.12x) | 7.39 (1.01x) |
FP8-vs-bf16 TPS differences scatter in both directions (max |Δ| 0.41 tok/s) with
no consistent sign by dataset or concurrency. Expected: the draft forward pass
is a small fraction of each speculation step, and only 22% of its params are
quantized, so the kernel gain does not surface.
Note on the long-conversation set: **speculative decoding is a net loss on
`agentic` at c=128** (0.93–0.94x) for *all three* drafts, breaking even only at
higher concurrency. That workload averages 28.5 turns with 100% tool use and
has the highest no-spec baseline, leaving least headroom for drafting to win.
This is a property of the workload, not of quantization.
## Requirements
- Compute capability **>= 89** (Ada / Hopper). `CompressedTensorsW8A8Fp8`
reports `min_capability 89`; H200 is 90.
- vLLM with draft-model quantization support. Verified on **0.20.0**
(`compressed_tensors` 0.15.0.1), where `deepseek_eagle3.py` calls
`get_draft_quant_config()` so the draft's own quant config is honored rather
than the verifier's.
## Usage
```bash
vllm serve /path/to/Kimi-K2.6 \
--speculative-config '{"method":"eagle3","model":"k-l-lambda/kimi-k2.6-eagle3-mla-fp8","num_speculative_tokens":3}'
```
## Inherited caveat
The bf16 parent is a K2.6 fine-tune selected by validation loss on the K2.6
teacher distribution, and its card notes it can over-specialize: on
K2.7-Code production traffic the official `lightseekorg` init showed higher
accept-length. The evaluation above is K2.6 traffic, where the fine-tune is
ahead or level — it does not test, and so does not refute, the K2.7 claim.
Benchmark against your own traffic before choosing a draft.