Qwen3.8-27B — ROCmFP4 + MTP drafter ladder (Strix Halo)

ROCmFP4 builds of Qwen3.8-27B, quantised from ggml-org/Qwen3.8-27B-GGUF's BF16 (sha256 5a3eedc837bcbd13…, verified), plus an MTP drafter at five precisions so the speculative-decoding numbers below can be reproduced rather than taken on trust.

Other ROCmFP4 builds of this model already exist. What is here and (at publication) not elsewhere: draft-acceptance rates, a per-backend n-max sweep, a -ub sweep, and a ROCm 7.2 vs 10.1 comparison that reverses the preset ranking.

Read this before choosing FP4. On the same machine, mainline llama.cpp on Vulkan with a plain Q4_K_M and the official MTP drafter is faster — 330 vs 227 prefill (+43%), decode tied. These FP4 files are worth it for the 1.7 GiB smaller footprint (15.6 vs 17.3 GiB resident), which matters when co-residing two models. They are not the throughput winner.

Files

file preset size
Qwen3.8-27B-ROCMFP4-STRIX.gguf Q4_0_ROCMFP4_STRIX 13.75 GiB
Qwen3.8-27B-ROCMFP4-COHERENT.gguf Q4_0_ROCMFP4_COHERENT 14.41 GiB
mtp-Qwen3.8-27B-ROCMFP4-STRIX.gguf FP4 drafter 1.85 GiB
mtp-Qwen3.8-27B-ROCMFP3.gguf FP3 drafter 1.55 GiB
mtp-Qwen3.8-27B-ROCMFP6.gguf FP6 drafter 2.27 GiB
mtp-Qwen3.8-27B-ROCMFP8.gguf FP8 drafter 2.86 GiB
mtp-Qwen3.8-27B-ROCMFP2.gguf FP2 drafter — broken, see below 1.48 GiB

Requires a ROCmFPX build; mainline llama.cpp does not know the Q4_0_ROCMFP4_* tensor types.

Hardware / method

AMD Ryzen AI MAX+ 395, Radeon 8060S (gfx1151, RDNA 3.5), 124 GiB GTT, Debian sid, kernel 7.2. Server-measured (llama-server + probe), ~8K-token prompt, temperature 0, one job at a time. ROCm nightly pinned to therock-dist-linux-gfx1151-10.1.0a20260815.

1. Draft depth (--spec-draft-n-max) is per-backend

n-max Vulkan Q4_K_M decode acc FP4 decode acc
3 33.42 100.0%
4 35.80 88.1% 35.16 98.7%
5 38.94 91.6% 38.67 98.1%
6 38.47 86.5% 38.04 97.5%
7 37.84 82.0% 39.26 94.7%
8 28.56 78.1% 32.36 95.2%
10 25.47 59.4%

Acceptance decays monotonically with depth; the knee is where verifying rejected drafts costs more than the accepted ones save. FP4 holds higher acceptance, so its knee sits deeper and flatter. This is not DeepSeek-V4's n=2 — draft depth does not transfer between models.

2. Drafter precision is a bandwidth lever, not a quality one

Target fixed, drafter varied, Vulkan, n=5:

drafter size decode acceptance
Q4_K_M 1.89 GiB 39.16 91.6%
Q6_K 2.28 GiB 38.28 92.1%
Q5_K_M 2.08 GiB 36.93 89.0%
Q8_0 2.95 GiB 34.74 89.0%

Acceptance is flat (89–92%) while decode spans 13% — so shrinking the drafter buys bandwidth and costs nothing in draft quality. Advice to keep drafters at ≥Q8 does not hold here.

3. FP2 destroys a drafter

FPX ladder, STRIX target, ROCm 10.1, n=5:

drafter decode acceptance
FP4-STRIX 37.03 97.4%
FP3 36.07 98.1%
FP6 30.92 96.6%
FP8 29.76 96.6%
FP2 22.07 64.0%

FP2's codebook has no exact zero. This drafter is BF16-sourced — the case usually assumed safe — and acceptance still collapses. Do not use FP2 for a draft model.

4. The preset ranking flips with the ROCm version

llama-bench, pp2048:

preset ROCm 7.2.4 ROCm 10.1 nightly
COHERENT 205.7 208.6 (+1%)
STRIX 151.8 272.0 (+79%)

COHERENT wins on 7.2; STRIX wins on 10.1. Any "preset X is best" claim — including ones in other repos — is conditional on a ROCm version that usually goes unstated.

5. -ub 256, not the default

-ub Vulkan pp2048 FP4 pp2048 (10.1)
256 370.6 300.3
512 360.9 267.3
1024 343.7 236.5
2048 332.0 235.8

Both backends prefer a small micro-batch; -ub 2048 costs FP4 27% of its prefill. Batch size (-b 512→4096) changes nothing.

Usage

llama-server \
  -m Qwen3.8-27B-ROCMFP4-STRIX.gguf \
  -md mtp-Qwen3.8-27B-ROCMFP4-STRIX.gguf \
  --spec-type draft-mtp --spec-draft-n-max 5 \
  -ngl 99 -ngld 99 -fa on -ub 256

Provenance

Every file derives from ggml-org/Qwen3.8-27B-GGUF, sha256-verified before quantisation: Qwen3.8-27B-BF16.gguf = 5a3eedc837bcbd13…, mtp-Qwen3.8-27B-BF16.gguf = 5723e551c4ee2b8c….

Downloads last month
-
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for singulared/Qwen3.8-27B-ROCmFP4-MTP-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(367)
this model