Myric's picture
add measured draft-depth x temperature surface; document v2 builds
e4ae997 verified
|
Raw
History Blame Contribute Delete
13 kB
---
license: apache-2.0
base_model: Kwaipilot/KAT-Coder-V2.5-Dev
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- moe
- apex
- quantized
- imatrix
- torch-imatrix
- mtp
- speculative-decoding
- qwen3_5_moe
- coding
- agentic-coding
- llama.cpp
---
# KAT-Coder-V2.5-Dev β€” APEX GGUF with MTP head
MoE-aware, mixed-precision **APEX** quantization of
[Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev), an
agentic-coding fine-tune on the Qwen3.6-35B-A3B backbone (`qwen3_5_moe`: 40 layers, 256 routed
+ 1 shared expert, hybrid GatedDeltaNet linear-attention with periodic full attention).
Kwaipilot reports SWE-bench Verified 69.40 for this size class, against Qwen3.6-35B-A3B's 64.40.
**What makes this build different: it has a working MTP head.** KAT-Coder ships
`mtp_num_hidden_layers: 0` β€” no multi-token-prediction head at all, so no speculative decoding
is possible out of the box, and that is true of the vendor release and of every other quant of
this model I am aware of. This repo transplants Qwen3.6-35B-A3B's own trained MTP head onto
KAT-Coder and quantizes the result. Measured below: **2.03x on a hard agentic-coding suite,
with correctness unchanged.**
## Recommended speculative settings
```
llama-server -m KAT-Coder-V2.5-Dev-MTP-APEX-i-quality-v2.gguf \
--spec-type draft-mtp --spec-draft-n-max 3
```
Full draft-depth x temperature surface, measured on `-v2`. One long-context prompt (~82 KB),
`max_tokens=400`, 3 seeds per cell, median tok/s, `top_p 0.95 top_k 20`, ctx 32768, build
`6d05498`. Speedup is against the no-drafter baseline **at the same temperature**:
| depth | T=0 (greedy) | T=0.25 | T=0.5 | T=0.75 | T=1.0 |
|---|---:|---:|---:|---:|---:|
| no drafter | 59.3 | 59.4 | 59.4 | 59.4 | 59.2 |
| `n-max 1` | 71.3 (1.20x) | 75.4 (1.27x) | 75.4 (1.27x) | 73.7 (1.24x) | 71.9 (1.21x) |
| `n-max 2` | 76.5 (1.29x) | 76.2 (1.28x) | 77.7 (1.31x) | 75.2 (1.27x) | 71.9 (1.21x) |
| **`n-max 3`** | **85.7 (1.44x)** | **78.7 (1.33x)** | **78.8 (1.33x)** | **77.0 (1.30x)** | **74.9 (1.26x)** |
| `n-max 5` | 68.4 (1.15x) | 61.9 (1.04x) | 66.5 (1.12x) | 66.4 (1.12x) | 62.0 (1.05x) |
| `n-max 8` | 51.2 (0.86x) | 55.9 (0.94x) | 53.8 (0.91x) | 54.5 (0.92x) | 48.6 (0.82x) |
Draft acceptance for the same cells:
| depth | T=0 | T=0.25 | T=0.5 | T=0.75 | T=1.0 |
|---|---:|---:|---:|---:|---:|
| `n-max 1` | 0.814 | 0.820 | 0.842 | 0.791 | 0.736 |
| `n-max 2` | 0.685 | 0.692 | 0.696 | 0.631 | 0.620 |
| `n-max 3` | 0.692 | 0.588 | 0.605 | 0.586 | 0.552 |
| `n-max 5` | 0.455 | 0.415 | 0.441 | 0.414 | 0.388 |
| `n-max 8` | 0.290 | 0.297 | 0.308 | 0.305 | 0.260 |
**Depth 3 wins at every temperature**, and the falloff past it is steep: acceptance decays roughly
geometrically with depth while draft compute grows linearly, so at `n-max 8` the wasted work
exceeds the gain and it runs **slower than no drafter at all** (0.82-0.94x). If you tune one knob,
tune this one.
The no-drafter row is flat across temperature (59.2-59.4), which is the control: all the variation
above it is acceptance, not decode speed.
Absolute tok/s here is lower than the from-scratch figures further down (59.3 vs ~70.8 baseline)
because this sweep uses a long-context prompt β€” decode slows as context grows. The *ratios* are
the transferable part.
Every number in both tables is throughput. Correctness is unaffected by construction: the target
model verifies each drafted token against its own computation, so a rejected draft never reaches
the output. Speculative decoding cannot change what the model says, only how fast it says it.
## Files
| file | size | what it is |
|---|---:|---|
| `KAT-Coder-V2.5-Dev-MTP-APEX-i-quality-v2.gguf` | 19.24 GiB | **recommended default** β€” supersedes the file below. Same tier, 56 MiB smaller, better PPL. See "v2" |
| `KAT-Coder-V2.5-Dev-APEX-dynamic-v2.gguf` | 11.91 GiB | supersedes the dynamic file below; better PPL on both eval corpora |
| `KAT-Coder-V2.5-Dev-MTP-APEX-i-quality.gguf` | 20.72 GB | superseded by `-v2`; kept so earlier results stay reproducible |
| `KAT-Coder-V2.5-Dev-APEX-dynamic.gguf` | 11.86 GiB | superseded by `-v2`; role-aware allocation targeting a 16GB card, no MTP head |
| `kat-coder.imatrix` | 192 MB | importance matrix used to build both; standard GGUF format, works with any `llama-quantize --imatrix` |
| `model-00014-of-mtp.safetensors` | 1.69 GB | the bf16 MTP head shard, if you want to do the transplant yourself |
| `config.json`, `model.safetensors.index.json` | β€” | already merged to reference all 14 shards |
| `TOOL_CALLING.md` | β€” | reproducible chained tool-calling transcript |
Pick `MTP-APEX-i-quality` unless you are VRAM-constrained, in which case `APEX-dynamic` is
the best quality available at its size. To run either you need only the `.gguf`. The safetensors shard and index are there for
anyone who wants to redo the splice against the original checkpoint: drop those three files
into a local copy of
[Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)
alongside its own 13 `model-*-of-00013.safetensors` shards. KAT-Coder's original shards are
untouched and byte-identical; this is a drop-in addition, not a re-upload of the 65 GB
checkpoint.
## The MTP head β€” honest framing
This is the **naive, un-fine-tuned transplant**. The head's weights are copied straight from
stock Qwen3.6 with zero training on KAT-Coder's own hidden-state distribution. It is
architecturally clean rather than a hack β€” hidden_size, expert count, and attention/MoE shapes
match exactly, because KAT-Coder is itself built on the Qwen3.6-35B-A3B backbone β€” but a
fine-tuning pass (freezing KAT-Coder's 40 layers, training only the ~850M-parameter head) is a
planned follow-up. These numbers are the baseline before that.
### Draft-acceptance rate
`llama-server --spec-type draft-mtp --spec-draft-n-max 2`, 5 real from-scratch coding prompts,
`n_predict=400` each:
| prompt | draft tokens | accepted | rate | tok/s |
|---|---:|---:|---:|---:|
| binary search tree (insert/delete/traverse) | 286 | 255 | 89.2% | 105.5 |
| CSV parser (stdlib only) | 114 | 84 | 73.7% | 94.6 |
| retry decorator w/ exponential backoff | 188 | 152 | 80.9% | 100.6 |
| merge two sorted linked lists | 74 | 58 | 78.4% | 97.6 |
| thread-safe LRU cache | 360 | 218 | 60.6% | 85.9 |
| **overall** | **1022** | **767** | **75.1%** | β€” |
The same quant tier without any MTP head runs ~70.8 tok/s solo. Acceptance tracks
predictability: BST traversal boilerplate hits 89%, the LRU cache with genuine locking and
eviction logic is hardest at 61%.
### Agentic coding benchmarks
Two task suites run through `opencode`'s real tool-calling loop, MTP build vs the same quant
without MTP. **Correctness was identical β€” 100% on both suites either way** (59/59 and 28/28
test cases), which is what the theory predicts: the MTP head only drafts, the main model always
verifies against its own computation, so a bad draft is rejected and never reaches the output.
Only speed moves.
| hard-tier task | no MTP | with MTP | speedup |
|---|---:|---:|---:|
| batch_scheduler | 22.5s | 20.8s | 1.08x |
| buddy_allocator | 123.5s | 63.5s | 1.94x |
| circuit_breaker | 50.7s | 14.0s | 3.62x |
| expr_parser_advanced | 54.7s | 21.1s | 2.59x |
| inventory_fix_and_extend | 16.4s | 12.3s | 1.33x |
| **total** | **267.8s** | **131.7s** | **2.03x** |
| tier-4 task | no MTP | with MTP | speedup |
|---|---:|---:|---:|
| lru_ttl_cache | 20.5s | 25.2s | 0.81x |
| rate_limiter_fix_and_extend | 19.7s | 12.1s | 1.63x |
| weighted_interval_scheduling | 22.5s | 17.4s | 1.29x |
| **total** | **62.7s** | **54.7s** | **1.15x** |
The single slower result (`lru_ttl_cache`, 0.81x) is most likely run-to-run variance in how
much iteration that attempt needed rather than a systematic MTP cost β€” every other task
improved.
## v2
Two allocation defects were found and fixed. Both were reported or surfaced after the original
upload; the `-v2` files are the corrected builds and the originals are retained so earlier numbers
stay reproducible.
**1. State coefficients were quantised.** `blk.N.ssm_alpha` and `blk.N.ssm_beta` β€” 60 tensors,
[2048,32], 3.93 M params total β€” were absent from the allocation recipe and silently inherited the
base quant type (Q4_K here, Q3_K in the dynamic tier). They now ship **F32**. These parameterise a
gate: `beta = sigmoid(ssm_beta . x)`, `alpha = softplus(ssm_alpha . x + dt)`, and the result
multiplies the state path, so error there compounds along the sequence rather than adding to one
layer's output. Cost of pinning them: ~14 MB on a 19 GiB file. Reported by
[@anjeysapkovski](https://huggingface.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF/discussions/2).
**2. The MTP block was never allocated.** `blk.40` β€” the transplanted head, 805 M params of
experts plus `nextn.eh_proj` β€” was also missing from the recipe and took the base type by accident,
landing four expert tensors on Q6_K. A drafter's precision buys acceptance rate, never output
quality, so those were bits spent where the return is zero by construction. `blk.40` is now pinned
Q4_K throughout. Measured effect on acceptance: none (within run-to-run spread at every depth),
and the file is 56 MiB smaller.
| tier | published | v2 | code-heavy PPL | wiki.test PPL |
|---|---:|---:|---:|---:|
| MTP i-quality | 19,765 MiB | **19,709 MiB** | 2.1247 -> **2.1231** | 7.1169 -> **7.1128** |
| dynamic | 12,141 MiB | 12,193 MiB | 2.4479 -> **2.4415** | 7.5435 -> **7.5160** |
Both v2 files pass a recipe-vs-artifact audit: every requested tensor type honoured, and zero
tensors quantised that the recipe never asked about (the original files had 71 such tensors).
## The imatrix
`llama-imatrix` is impractical on this architecture: GatedDeltaNet linear-attention is a serial
state-space recurrence, and the imatrix collection callback breaks the GPU path, falling back to
a single CPU thread. `kat-coder.imatrix` was instead generated with a band-serialized PyTorch
forward pass over a general text/code corpus. It is a standard GGUF-format importance matrix
(`in_sum2` + `counts` per tensor), bit-compatible with `llama-quantize --imatrix`, so it is
directly reusable if you want to build your own tiers.
One checkpoint quirk worth knowing if you do: KAT-Coder stores MoE experts as fully separate
per-expert tensors (`mlp.experts.{e}.{gate,up,down}_proj.weight` Γ— 256) rather than the fused
`gate_up_proj`/`down_proj` layout the HF module expects. Both imatrix generation and GGUF
conversion have to fuse them.
## Allocation
APEX structural mode: routed experts are bucketed into edge / near / middle depth bands (edge =
first and last layers, generally most load-bearing) with decreasing precision toward the middle.
For i-quality that is routed edge β†’ Q6_K, near β†’ Q4_K, middle β†’ Q3_K. Attention, shared-expert,
and router tensors stay at higher fixed precision (Q6_K–Q8_0).
### Measured quality
Wikitext perplexity across the tiers that were built. The two marked *hosted* are in this repo;
the other two were measured for comparison and are not uploaded:
| tier | file size | bits/weight | wikitext PPL | vs bf16 | |
|---|---:|---:|---:|---:|---|
| bf16 (reference) | 64.61 GiB | 16.01 | 7.028 | β€” | |
| APEX-i-quality | 19.23 GiB | 4.65 | 7.117 | +1.27% | recipe used by the hosted MTP build |
| **APEX-dynamic** | 12.13 GiB | 2.94 | 7.543 | +7.33% | **hosted** |
| IQ3_XXS (stock, imatrix-guided) | 12.68 GB | ~2.98 | 7.881 | +12.13% | stock baseline, for comparison |
Two things worth drawing out. i-quality lands within 1.3% of full bf16 at under a third of the
size. And **APEX-dynamic beats the stock IQ3_XXS baseline on both axes at once** β€” smaller
(12.13 vs 12.68 GiB) *and* lower perplexity (7.543 vs 7.881) β€” which is why it is the one
hosted alongside the flagship rather than the stock tier.
That result is cleaner than the same method got on Laguna-XS, where it traded size for a small
PPL cost. Whether role/depth-aware allocation beats a well-tuned stock quant appears to depend
on the architecture and calibration, and is not something to assume without measuring.
The hosted MTP build uses the i-quality recipe; its file is slightly larger than the 19.23 GiB
row because it carries the MTP head as well.
## Attribution
- KAT-Coder base model: **Kwaipilot** β€” [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) (Apache-2.0).
- MTP head weights: **Qwen** β€” [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0), transplanted unmodified.
- APEX recipe & toolkit: **LocalAI** β€” [localai-org/apex-quant](https://github.com/localai-org/apex-quant).
- Quantization engine: **llama.cpp** ([ggml-org](https://github.com/ggml-org/llama.cpp)).
Unofficial community quantization and experiment; not affiliated with or endorsed by Kwaipilot
or Qwen.