File size: 9,385 Bytes
6ec1471
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b3ac135
 
 
 
6ec1471
 
 
 
b3ac135
 
 
 
6ec1471
b3ac135
 
 
 
 
 
 
 
 
 
6ec1471
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b3ac135
 
 
6ec1471
 
 
 
b3ac135
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6ec1471
 
b3ac135
 
 
 
6ec1471
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b3ac135
6ec1471
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b3ac135
6ec1471
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
---
license: apache-2.0
base_model: Kwaipilot/KAT-Coder-V2.5-Dev
base_model_relation: quantized
pipeline_tag: text-generation
language:
- en
- zh
quantized_by: 1337Hero
tags:
  - gguf
  - kat-coder
  - quantized
  - rocm
  - amd
  - rdna4
  - gfx1201
  - vulkan
  - moe
  - code
  - experimental
---

# KAT-Coder-V2.5-Dev — ROCmFP4 GGUF (experimental, AMD RDNA4 / gfx1201)

Two experimental 4-bit quantizations of
[Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev),
a 34.66B-parameter MoE coding model (256 experts, 8 active, 256K context,
`qwen35moe` architecture). Converted and quantized on a Radeon AI PRO R9700
(`gfx1201`, RDNA4).

> [!IMPORTANT]
> These files do **not** run on upstream llama.cpp, Ollama, LM Studio, or
> vLLM. They use the custom `Q4_0_ROCMFP4` tensor layouts and require a
> [ROCmFPX](https://github.com/charlie12345/ROCmFPX) build as described below.
> Unsupported runtimes should reject the files; if a tool appears to load one
> anyway, do not trust the output.

> [!WARNING]
> Validation was performed on RDNA4 `gfx1201` only: both files load, generate
> coherent output, were throughput-benchmarked, and were measured against the
> BF16 source for wikitext-2 perplexity. No Strix Halo testing and **no
> code-specific or agentic evaluation** — see [What was not
> measured](#what-was-not-measured) before relying on either file.

## Which file?

| File | Size | Effective BPW | Wikitext-2 PPL | Pick it if |
| --- | ---: | ---: | ---: | --- |
| `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf` | 21.18 GiB | 5.25 | 6.9182 (+1.38%) | You care about output quality. **Recommended for coding.** |
| `KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4_STRIX_LEAN.gguf` | 17.32 GiB | 4.29 | 7.1079 (+4.16%) | You need the smaller file or the extra decode speed. |

This is a **real tradeoff, not a clean win for either file.** `STRIX_LEAN` is
18% smaller and 13% faster at decode, but gives up three times as much
perplexity against the BF16 source. For a coding model — where a single wrong
token breaks a program — the plain `Q4_0_ROCMFP4` is the safer default, and
21.18 GiB still fits a 32 GB card comfortably.

Take `STRIX_LEAN` if you are memory-constrained (24 GB cards), or if you are
throughput-bound and have validated that the quality holds on your own tasks.
Its recipe was tuned on `gfx1151`; nothing about the file format is
Strix-specific.

## Why the sizes differ from the nominal BPW

Both presets apply tensor-aware routing, and on a 256-expert MoE the expert
tensors dominate the file. The routing difference between the two is almost
entirely one tensor family:

| Tensor | `Q4_0_ROCMFP4` | `Q4_0_ROCMFP4_STRIX_LEAN` |
| --- | --- | --- |
| `ffn_down_exps` | `q6_K` | `q4_0_rocmfp4_fast` |
| `ffn_gate_exps` / `ffn_up_exps` | `q4_0_rocmfp4` | `q4_0_rocmfp4_fast` |
| `attn_qkv` | `q5_K` | Strix attn K/V recipe |
| `token_embd` | `q6_K` | `q5_K` |

That is why `Q4_0_ROCMFP4` lands at 5.25 BPW rather than its nominal 4.50 —
`ffn_down_exps` alone is roughly a third of the model's parameters.

## Measured throughput — Radeon AI PRO R9700, `gfx1201`

`llama-bench`, `pp512` + `tg128`, 3 repetitions, full offload, FlashAttention
on, one model resident at a time on an otherwise idle GPU.

| Backend | Quant | Prompt fill `pp512` t/s | Decode `tg128` t/s |
| --- | --- | ---: | ---: |
| Vulkan0 | **`STRIX_LEAN`** | **3278.14 ± 50.90** | **122.27 ± 1.27** |
| Vulkan0 | `Q4_0_ROCMFP4` | 3120.99 ± 17.61 | 107.80 ± 1.02 |
| ROCm0 | `STRIX_LEAN` | 2598.84 ± 5.06 | 59.39 ± 0.19 |
| ROCm0 | `Q4_0_ROCMFP4` | 1787.51 ± 74.26 | 52.51 ± 0.26 |

Two results worth acting on:

- **Use Vulkan on this hardware.** Vulkan decodes roughly **2× faster** than
  HIP/ROCm for both files (122 vs 59 t/s on `STRIX_LEAN`) and also leads on
  prompt fill. This matches ROCmFPX's own Strix Halo findings.
- **`STRIX_LEAN` is the faster file** — +13% decode and +5% prefill on Vulkan,
  +13% decode and +45% prefill on ROCm — but see the quality section below
  before choosing it on speed alone.

No control quant (Q4_K_M or similar) was benchmarked, so these numbers compare
the two ROCmFP4 files against each other, not against ordinary GGUF quants.

## Measured quality — wikitext-2 perplexity

`llama-perplexity`, full wikitext-2 test set (580 chunks), `-c 512 -b 512`,
FlashAttention on, Vulkan. The BF16 source GGUF was measured on the same host
with the same settings, split across three GPUs.

| File | BPW | PPL | Δ vs BF16 |
| --- | ---: | ---: | ---: |
| `KAT-Coder-V2.5-Dev-BF16.gguf` (source) | 16.01 | 6.8237 ± 0.04537 | — |
| `Q4_0_ROCMFP4` | 5.25 | 6.9182 ± 0.04607 | **+1.38%** |
| `Q4_0_ROCMFP4_STRIX_LEAN` | 4.29 | 7.1079 ± 0.04762 | **+4.16%** |

Both quants land where you would expect for their bit budgets, and neither is
degenerate. The gap between them is larger than the error bars, so it is a
real difference and not measurement noise: `STRIX_LEAN` buys its 18% size
reduction with roughly 3× the perplexity cost.

Perplexity is a weak proxy for coding ability. It measures next-token
prediction on English Wikipedia, not code correctness or tool-call formatting.
Treat it as a floor check — it rules out a broken quantization, it does not
establish that either file codes as well as the source.

## What was not measured

- **Coding ability.** No HumanEval, MBPP, or any code benchmark. Wikitext-2
  perplexity was measured (see above), but it does not measure code
  correctness.
- **KL-divergence** against the BF16 source. Perplexity only.
- **Agentic and tool-calling behavior**, which is the point of a coding model.
  Untested.
- **Any hardware other than `gfx1201`.** Not tested on Strix Halo, RDNA3,
  RDNA2, or CPU.
- **Long context.** Benchmarked at `pp512`/`tg128`; the model claims 262144.
  Deep-context behavior and KV-cache pressure are unmeasured.
- **Batch > 1 / concurrent requests.** Single-stream only.
- **A non-ROCmFPX control quant.** The two files were compared to each other,
  not to Q4_K_M.

## Required runtime

```bash
git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX && git checkout main   # built and quantized at commit db6844d

env JOBS=16 scripts/build-rdna4.sh    # -> build-rdna4/ (gfx1201 auto-detected)
```

On a Navi 48 card (RX 9070, 9070 XT, AI PRO R9700) the script builds `gfx1201`
automatically. `gfx1200` builds are **not** interchangeable on these cards.

## Example run

```bash
./build-rdna4/bin/llama-server \
  -m KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -c 32768 \
  -b 512 -ub 512 \
  --jinja
```

`-dev Vulkan0`, not `ROCm0` — see the benchmark table above. Swap in `-dev
ROCm0` only if Vulkan is unavailable on your system.

`--jinja` is required — the model ships a chat template with `<think>`
reasoning blocks.

The model has **no MTP/NextN head** (`mtp_num_hidden_layers = 0` in the source
config), so ROCmFPX's self-speculative decoding is not available here.

## Artifacts

| Field | `STRIX_LEAN` | `Q4_0_ROCMFP4` |
| --- | --- | --- |
| Size | 18,597,337,248 bytes | 22,741,457,056 bytes |
| Effective BPW | 4.29 | 5.25 |
| SHA-256 | `857d39a696d448a9349a000ffdd5811c88761df59ec7d5a67e7ada3ac46a8161` | `f87c3f509c487876dc76d8a7606583faa1a3cbd04407b288c36dd2031b2cd92f` |
| Quantization | `Q4_0_ROCMFP4_STRIX_LEAN` | `Q4_0_ROCMFP4` |
| Importance matrix | none | none |

Source: `KAT-Coder-V2.5-Dev-BF16.gguf`, 69,376,637,408 bytes, converted from
the upstream `safetensors` release with ROCmFPX's `convert_hf_to_gguf.py` at
commit `db6844d`:

```bash
python convert_hf_to_gguf.py /path/to/KAT-Coder-V2.5-Dev \
  --outtype bf16 --outfile KAT-Coder-V2.5-Dev-BF16.gguf
```

Quantization (same commit):

```bash
./build-rdna4/bin/llama-quantize \
  KAT-Coder-V2.5-Dev-BF16.gguf \
  KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  Q4_0_ROCMFP4_STRIX_LEAN 16

./build-rdna4/bin/llama-quantize \
  KAT-Coder-V2.5-Dev-BF16.gguf \
  KAT-Coder-V2.5-Dev-Q4_0_ROCMFP4.gguf \
  Q4_0_ROCMFP4 16
```

Verify after download:

```bash
sha256sum -c SHA256SUMS
```

## Notes on the source model

The upstream open-weight release ships **language-model weights only** — the
vision and multimodal components described in the model card are not included,
and the converted GGUFs contain no multimodal projector. Despite the
`Qwen3_5MoeForConditionalGeneration` class name, these are text-only files.

## Limitations

- Requires the ROCmFPX fork; no upstream llama.cpp compatibility.
- Validated on exactly one `gfx1201` host, batch 1, shallow context.
- Quality evidence is wikitext-2 perplexity only; no code or agentic evals.
- 34.66B MoE: needs ~18–22 GB for weights plus KV cache. Comfortable on a
  32 GB card, tight on 24 GB with meaningful context.

## License and attribution

- **Base model:** KAT-Coder-V2.5-Dev, Kwaipilot, Apache-2.0. This repository
  redistributes a converted and quantized derivative under the same license.
- **Format and execution path:** the `Q4_0_ROCMFP4` representations and
  kernels are the work of the
  [ROCmFPX](https://github.com/charlie12345/ROCmFPX) project, which builds on
  [llama.cpp](https://github.com/ggml-org/llama.cpp).
- **This repository:** the quantized artifacts only.

KAT-Coder and related marks belong to their owners. This community
quantization is not affiliated with or endorsed by Kwaipilot, AMD, ROCmFPX, or
llama.cpp.