| --- |
| base_model: Mia-AiLab/Qwable-3.6-27b-MTP |
| base_model_relation: quantized |
| license: mit |
| library_name: gguf |
| tags: |
| - gguf |
| - rocmfp4 |
| - qwen3.6 |
| - fable |
| - fable5 |
| - reasoning |
| - mtp |
| - speculative-decoding |
| - strix-halo |
| - amd |
| - rocm |
| - vulkan |
| --- |
| |
| <div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;"> |
| <div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div> |
| <div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;"> |
| <pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;"> |
| ▗▇▇▇▇▇▇▇▖ |
| ▗█▘▝██████▖ |
| ▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅ |
| ▟▛ ▗█████████████████▙▖ |
| ▄▄▄▄▄▟▛ ▟████████████████████▖ |
| ▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘ |
| ▗████▖ ▜▖ ▗█▘ |
| ▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙ |
| ▜█████▙ ▝████████████▛ ▜▙ |
| ▜█████▙ ▝██████████▛ ▃ ▜▙ |
| ▀█████▙▖ ▝████████▘ ▟█▙ ▀▙ |
| ▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█ |
| ▟███████▖ ▜███▘ ▗███████████▛ |
| ▟█████████▄ ▜▛ ▗███████████▀ |
| ▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘ |
| ▜██▘ ▗▛ ▟█████▛▘ |
| ▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛ |
| ▝█▖ ▟█████▛ |
| ▝███████▀ |
| </pre> |
| <div style="flex:0 1 auto; max-width:100%; text-align:center;"> |
| <div style="font-size:23px; font-weight:800; letter-spacing:1px;">QWABLE-3.6-27B-MTP</div> |
| <div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">imatrix + f16 EMBEDDINGS</span> · <span style="white-space:nowrap;">Q6_K HEAD</span> · <span style="white-space:nowrap;">MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div> |
| <div style="font-size:11px; opacity:0.7; margin-top:6px;">a ROCmFP4 quant of <b>Mia-AiLab/Qwable-3.6-27b-MTP</b> — a Fable-5 reasoning fine-tune of Qwen3.6-27B</div> |
| </div> |
| </div> |
| <table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;"> |
| <tr> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT · STRIX</div></td> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">4.82 BPW</div></td> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">16.9 GB</div></td> |
| <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td> |
| </tr> |
| <tr> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5</div></td> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">EMBEDDINGS</div><div style="font-weight:700;">f16 · HEAD Q6_K</div></td> |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">Vulkan0 · ROCm</div></td> |
| <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CALIBRATION</div><div style="font-weight:700;">imatrix (froggeric)</div></td> |
| </tr> |
| </table> |
| </div> |
| |
| <div style="border:2px solid #2563eb; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;"> |
| <b style="color:#2563eb; letter-spacing:1px;">★ ORIGINAL MODEL — ALL CREDIT TO Mia-AiLab</b><br> |
| This is a quantization. The model itself is <b><a href="https://huggingface.co/Mia-AiLab/Qwable-3.6-27b-MTP">Mia-AiLab/Qwable-3.6-27b-MTP</a></b> — a full fine-tune of <code>qwen/Qwen3.6-27B</code> on a cleaned Fable-5 reasoning/instruction dataset, with native MTP. Please ⭐ and follow the author: <a href="https://x.com/MiaAI_lab">@MiaAI_lab</a>. License inherited: <b>MIT</b>. |
| </div> |
| |
| <div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;"> |
| <b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br> |
| The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, Ollama, Jan, or koboldcpp</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>: |
| <br><br> |
| <code>git clone https://github.com/charlie12345/ROCmFPX</code><br> |
| <code>cd ROCmFPX</code><br> |
| <code>env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh</code> |
| </div> |
|
|
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;"> |
| <b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" badge — its parser only knows standard GGUF quant types, can't read ROCmFP4, and "sees" only the genuinely-f16 token embeddings. This is a <b>~4.82 bpw 4-bit</b> file; ~16.9 GB. |
| </div> |
|
|
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · WHAT THIS IS</div> |
|
|
| A 4-bit **ROCmFP4** quant of **[Mia-AiLab/Qwable-3.6-27b-MTP](https://huggingface.co/Mia-AiLab/Qwable-3.6-27b-MTP)** (a Fable-5 reasoning/instruction fine-tune of `qwen/Qwen3.6-27B`), built on the PlunderStruck **STRIX** daily-driver recipe for AMD Strix Halo (Ryzen AI Max+ 395, gfx1151): |
|
|
| - **Body** — `Q4_0_ROCMFP4_STRIX`: FP4 (E2M1) weights + UE4M3 microscale; the Strix quality/speed recipe (dual-scale ROCmFP4 on attention K/V, fast single-scale path on the large tensors). |
| - **Token embeddings** — kept at **f16** (the coherence lever that's actually felt). |
| - **Output head** — **Q6_K**. |
| - **imatrix** — importance matrix over the froggeric general corpus (`groups_merged` + `technical` + `code`), applied to all quantizable tensors. |
| - **MTP** — Mia's **native** multi-token-prediction head (`blk.64.nextn.*`) is preserved → llama.cpp **self-speculative decode**, no separate draft model. |
|
|
| Bundled in this repo: **`chat_template.jinja`** (froggeric's unified Qwen3.6 template — tool calls + inline `<|think_off|>`/`<|think_on|>`), **`mmproj-F32.gguf`** (shared Qwen3-VL projector, `projection_dim 5120` — vision, optional), and **`qwable-3.6-27b.imatrix`** for exact reproduction. Runs entirely on a single AMD APU — Vulkan/RADV or ROCm/HIP (gfx1151). |
|
|
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div> |
|
|
| Run from the folder holding the `.gguf` + `chat_template.jinja`: |
|
|
| ```bash |
| env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \ |
| llama-server \ |
| -m Qwable-3.6-27B-MTP-ROCmFP4-STRIX-imatrix-embF16-headQ6.gguf \ |
| --alias qwable-mtp \ |
| --host 0.0.0.0 \ |
| --port 8080 \ |
| -dev Vulkan0 \ |
| -ngl 999 \ |
| -fa on \ |
| -c 262144 \ |
| -b 2048 \ |
| -ub 256 \ |
| -t 16 \ |
| -tb 16 \ |
| -ctk f16 \ |
| -ctv f16 \ |
| -cpent 256 \ |
| -ctxcp 32 \ |
| --cache-reuse 256 \ |
| --cache-ram 65536 \ |
| --temp 0.6 \ |
| --top-p 0.95 \ |
| --top-k 20 \ |
| --min-p 0.0 \ |
| --spec-type draft-mtp \ |
| --spec-draft-device Vulkan0 \ |
| --spec-draft-ngl all \ |
| --spec-draft-type-k f16 \ |
| --spec-draft-type-v f16 \ |
| --spec-draft-n-max 5 \ |
| --spec-draft-n-min 0 \ |
| --spec-draft-p-min 0.0 \ |
| --spec-draft-p-split 0.10 \ |
| --chat-template-file chat_template.jinja \ |
| --reasoning on \ |
| --reasoning-format deepseek \ |
| --chat-template-kwargs '{"preserve_thinking": true}' \ |
| --jinja \ |
| --parallel 1 \ |
| --metrics \ |
| --no-mmap \ |
| --mmproj mmproj-F32.gguf \ |
| --image-min-tokens 1024 |
| ``` |
|
|
| The **last two lines enable vision** via the bundled `mmproj-F32.gguf`; omit them for text-only. **`--image-min-tokens 1024` is required whenever `--mmproj` is set.** |
|
|
| <div style="overflow:hidden;"> |
| <table style="width:100%; border-collapse:collapse; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;"> |
| <thead><tr> |
| <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th> |
| <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th> |
| </tr></thead> |
| <tbody> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">use the full 128 GB unified memory</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0 · -ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Vulkan (KHR_coopmat) beats ROCm/HIP here · all layers · flash attention</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch (256 is the optimum here) · CPU threads</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache; drop to <code>q8_0</code>/<code>q4_0</code> for less memory</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing (every 256 tok, keep 32, reuse ≥256-tok prefix) + 64 GB resident reuse cache</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3.6 "precise" sampling (temp 1.0 for general/creative)</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5</code></td><td style="border:1px solid currentColor; padding:6px 10px;">built-in MTP head, self-speculative; draft depth 5</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 draft KV</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">bundled froggeric template (tool calls + think-toggle)</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning on --reasoning-format deepseek + kwargs {preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">clean <code>content</code> + <code>reasoning_content</code>; keep <code><think></code> across turns so cross-turn cache survives</td></tr> |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--mmproj mmproj-F32.gguf · --image-min-tokens 1024</code></td><td style="border:1px solid currentColor; padding:6px 10px;">vision via the bundled Qwen3-VL projector (omit for text-only)</td></tr> |
| </tbody> |
| </table> |
| </div> |
| |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · CODING AGENT / CROSS-TURN CACHE</div> |
| |
| **Multi-turn prompt-cache reuse is what makes a 27B usable on one APU.** Qwen3.6's recurrent (SSM) state can't be partially rewound, so multi-turn reuse needs a context *checkpoint* at/before the divergence point. Two defaults otherwise force a full re-prefill every turn — both fixed by the flags above: |
| |
| 1. **Checkpoint cadence.** Default `-cpent` is 8192, so prompts under 8K never get a usable checkpoint. Fix: **`-cpent 256 -ctxcp 32 --cache-reuse 256`**. |
| 2. **Thinking text breaking the prefix match.** Use **`--reasoning-format deepseek`** + **`--chat-template-kwargs '{"preserve_thinking": true}'`** so the template keeps `<think>` across all turns and reuse holds. |
| |
| `--jinja` is required so the chat template (and `preserve_thinking`) apply. Point any OpenAI-compatible client (OpenCode, etc.) at the server; in single-model mode llama-server ignores the request's `model` field. |
|
|
| ## Credits |
|
|
| - **Model:** [Mia-AiLab/Qwable-3.6-27b-MTP](https://huggingface.co/Mia-AiLab/Qwable-3.6-27b-MTP) by **Mia-AiLab** ([@MiaAI_lab](https://x.com/MiaAI_lab)) — Fable-5 reasoning fine-tune of `qwen/Qwen3.6-27B`. **All model credit is theirs.** |
| - **Base:** [qwen/Qwen3.6-27B](https://huggingface.co/qwen/Qwen3.6-27B). |
| - **ROCmFP4 kernels/format:** [charlie12345/ROCmFPX](https://github.com/charlie12345/ROCmFPX). |
| - **Quant + packaging:** PlunderStruck. License: MIT (inherited). |
|
|