File size: 11,286 Bytes
97785d7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 | # ROCmFPX Quantization Guide β q4_0_rocmfp4_fast experts + q8_0_rocmfpx everything else
Self-contained playbook for producing "hybrid" ROCmFPX GGUF quants from a Hugging Face repo link:
**routed MoE experts (up/down/gate) β `q4_0_rocmfp4_fast`, every other tensor β `q8_0_rocmfpx`**.
Runs fully **CPU-only** (no GPU needed; quantization is a CPU reference path).
Validated end-to-end on:
| Model | Source | Recipe applied to | Result |
|---|---|---|---|
| Laguna-S-2.1 (118B-A10B, poolside, arch `laguna`) | `unsloth/Laguna-S-2.1-GGUF/BF16` (5 shards) | `ffn_{gate,up,down}_exps` | 224 GB BF16 β 61.6 GB (4.39 bpw) |
| Qwen3.8-27B (dense) | `unsloth/Qwen3.8-27B-GGUF/BF16` (2 shards) | n/a (dense) | 52 GB BF16 β 26.9 GB (8.25 bpw) |
| G4-MeroMero-26B-A4B (Gemma4 MoE) | raw HF weights `llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic` | `ffn_gate_up_exps`, `ffn_down_exps` | converted to BF16 GGUF (50.5 GB) β 14.0 GB (4.64 bpw) |
| Qwen3.8-27B hybrid (dense, `qwen35` arch) | `unsloth/Qwen3.8-27B-GGUF/BF16` | sensitive tensors per Unsloth UD-Q4_K_XL (see Β§8) | 16.4 GB (5.15 bpw), targets ~16 GB VRAM-fit file |
---
## 0. One-time setup (CPU-only)
```bash
# System deps: gcc, cmake, git, python3 + pip
git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX
# CPU-only build (no HIP/CUDA/Vulkan/Metal)
cmake -B build-cpu -DGGML_HIP=OFF -DGGML_CUDA=OFF -DGGML_VULKAN=OFF \
-DGGML_METAL=OFF -DCMAKE_BUILD_TYPE=Release
cmake --build build-cpu -j "$(nproc)" --target llama-quantize llama-gguf-split llama-gguf
# Python deps for HFβGGUF conversion (only needed for Workflow B)
pip install -r requirements/requirements-convert_hf_to_gguf.txt
pip install -e gguf-py # install the repo's gguf lib (older pip gguf can't read ROCmFPX types!)
```
Binaries land in `build-cpu/bin/`: `llama-quantize`, `llama-gguf-split`, `llama-gguf`.
> **Note:** the ROCmFPX quant types (`q4_0_rocmfp4_fast` = 101, `q8_0_rocmfpx` = 103, etc.)
> are only understood by this fork's tooling and recent-ish llama.cpp. Use the binaries built here.
---
## 1. Decide which workflow applies
Look at the HF repo (`https://huggingface.co/<repo>`):
- **Workflow A** β repo already contains GGUF files, e.g. `BF16/*-00001-of-0000N.gguf`
(typical for `unsloth/*-GGUF` repos). Skip straight to quantizing.
- **Workflow B** β repo contains raw weights (`model*.safetensors` + `config.json`).
Convert to BF16 GGUF first, then quantize.
Check sizes before downloading:
```bash
curl -s https://huggingface.co/api/models/<ORG>/<REPO>/tree/main/BF16 | \
python3 -c "import json,sys; [print(s['path'], round(s['size']/1e9,2),'GB') for s in json.load(sys.stdin)]"
```
Download with `hf` CLI (shards download in parallel; pass the first shard to llama.cpp later β
remaining shards are discovered automatically):
```bash
pip install -q huggingface_hub
hf download <ORG>/<REPO> --repo-type model --include "BF16/*" --local-dir /workspace/models/<NAME>
```
---
## 2. Find the routed-expert tensor names (critical step)
Tensor naming depends on architecture. Inspect the first GGUF shard (header only β works as soon as
the file starts existing, and only needs a few MB):
```bash
python3 - <<'EOF'
from gguf import GGUFReader
import re, collections
r = GGUFReader("<first-shard>.gguf", "r")
print("arch:", bytes(r.fields['general.architecture'].parts[-1]).decode())
names = sorted(set(re.sub(r'^blk\.\d+', 'blk.N', t.name) for t in r.tensors))
for n in names: print(n)
print(collections.Counter(t.tensor_type.name for t in r.tensors))
EOF
```
Build the expert regex from the patterns you see:
| What you see in the tensor list | Expert regex for `--tensor-type` |
|---|---|
| `blk.N.ffn_gate_exps.weight`, `ffn_up_exps`, `ffn_down_exps` (packed, e.g. `laguna`) | `ffn_(gate\|up\|down)_exps` |
| `blk.N.ffn_gate_up_exps.weight`, `ffn_down_exps` (fused gate+up, e.g. `gemma4`) | `ffn_(gate_up\|down)_exps\.weight` |
| `blk.N.ffn_gate.M`, `ffn_up.M`, `ffn_down.M` (per-expert tensors) | `ffn_(gate\|up\|down)\.\d+\.` |
Patterns are **regexes matched against lowercased tensor names** (regex_search, i.e. partial match).
Norms, biases, router scales (`*.scale`, `*.bias`, 1-D tensors) are never quantized regardless, so
matching a little too broadly is harmless β but still verify with a dry run.
---
## 3. Dry run (always do this first)
`--output-tensor-type` only affects `output.weight`; to get "everything else = q8_0_rocmfpx" use the
**base ftype `Q8_0_ROCMFPX`** and override experts:
```bash
./build-cpu/bin/llama-quantize --dry-run \
--tensor-type "<EXPERT_REGEX>=q4_0_rocmfp4_fast" \
<input-BF16-shard-1.gguf> <output-name.gguf> Q8_0_ROCMFPX
```
Confirm in the output:
- `applying manual override: q8_0_rocmfp4_fast...` lines appear for **every** expert tensor
(count = layers Γ {gate,up,down}, e.g. 47 MoE layers Γ 3 = 141 for Laguna),
- attention / shared experts / embeddings β `(q8_0_rocmfpx)`,
- norms/biases/routers stay `f32`,
- the final `quant size = ... (X.XX BPW)` line looks sane (~4.4 bpw for mostly-expert MoE,
~8.25 bpw for dense).
For a dense model, drop `--tensor-type` entirely.
---
## 4. Quantize
```bash
./build-cpu/bin/llama-quantize \
--tensor-type "<EXPERT_REGEX>=q4_0_rocmfp4_fast" \
<input-BF16-shard-1.gguf> <output-name.gguf> Q8_0_ROCMFPX
```
Flags worth knowing:
- `--keep-split` β output keeps the input's shard layout.
- Default (no flag) β auto-splits output at 50 GiB.
- `--tensor-type` can be repeated, or use `--tensor-type-file list.txt` (one `name=type` per line).
- Add trailing `nthreads` positional arg to cap thread count.
Speed reference (64-core server): 118B MoE β 8 min, 27B dense β 1 min, 26B MoE β 3 min.
Memory use is modest (mmap-based); RAM β source size is plenty.
---
## 5. Verify + merge shards
Verify the result (works on merged or shard-1 file):
```bash
./build-cpu/bin/llama-quantize --dry-run <output.gguf> /tmp/dummy.gguf COPY | grep "type .*tensors"
# e.g.: f32: 287 | q4_0_rocmfp4_fast: 141 | q8_0_rocmfpx: 386
```
Merge shards into one file (only pass the first shard; the rest are auto-discovered):
```bash
./build-cpu/bin/llama-gguf-split --merge \
<output>-00001-of-0000N.gguf <output-merged.gguf>
```
(`--split` / `--split-max-size` do the reverse. Note: the `llama-gguf <file> r` example tool aborts
on ROCmFPX types even for valid files β use the quantize-dry-run check above instead.)
---
## 6. Workflow B only: HF weights β BF16 GGUF
For repos with `*.safetensors` + `config.json` (check `architectures` in config.json is registered in
`convert_hf_to_gguf.py` β e.g. `Gemma4ForConditionalGeneration` is supported by this fork):
```bash
hf download <ORG>/<REPO> --local-dir /workspace/models/<NAME> # all files incl. tokenizer
python3 convert_hf_to_gguf.py /workspace/models/<NAME> \
--outfile /workspace/models/<NAME>-BF16.gguf --outtype bf16
# for multimodal models, add --mmproj to also export the vision projector (separate file)
```
Then continue at **Step 2** with the produced BF16 GGUF.
---
## 7. Cheat sheet of the runs documented here
```bash
# Laguna-S-2.1 (laguna arch, packed experts)
./build-cpu/bin/llama-quantize \
--tensor-type "ffn_(gate|up|down)_exps=q4_0_rocmfp4_fast" \
--keep-split \
Laguna-S-2.1-BF16-00001-of-00005.gguf Laguna-S-2.1-Q8_0_ROCMFPX-Q4FAST-experts.gguf Q8_0_ROCMFPX
# Qwen3.8-27B (dense β pure q8_0_rocmfpx)
./build-cpu/bin/llama-quantize \
Qwen3.8-27B-BF16-00001-of-00002.gguf Qwen3.8-27B-Q8_0_ROCMFPX.gguf Q8_0_ROCMFPX
# Gemma4 MoE (fused gate_up experts), after convert_hf_to_gguf.py
./build-cpu/bin/llama-quantize \
--tensor-type "ffn_(gate_up|down)_exps.weight=q4_0_rocmfp4_fast" \
G4-MeroMero-26B-A4B-BF16.gguf G4-MeroMero-26B-A4B-Q8_0_ROCMFPX-Q4FAST-experts.gguf Q8_0_ROCMFPX
```
### Other ROCmFPX ftypes you may be asked for
`Q4_0_ROCMFP4`, `Q4_0_ROCMFP4_FAST` (4.25 bpw, speed-first), `Q4_0_ROCMFP4_STRIX(_LEAN)`,
`Q4_0_ROCMFP4_COHERENT`, `Q8_0_ROCMFPX`, `Q6_0_ROCMFPX`, `Q3_0_ROCMFPX`, `Q2_0_ROCMFPX` (2.5 bpw).
List all with `./build-cpu/bin/llama-quantize --help`.
---
## 8. VRAM-fit hybrid recipe (q4 bulk + q8 sensitive), e.g. ~16 GB target
When the file must fit a smaller GPU (e.g. 24 GB laptop VRAM minus KV cache β ~16 GB file),
invert the recipe: **base ftype `Q4_0_ROCMFP4_FAST`**, then override *sensitive* tensors to
`q8_0_rocmfpx`. To pick the sensitive set, mirror the tiers of Unsloth's Dynamic recipe
(e.g. `UD-Q4_K_XL`) for the same model:
1. Download the reference quant and tabulate per-tensor types:
```python
from gguf import GGUFReader
import re, collections
r = GGUFReader("<UD-Q4_K_XL>.gguf", "r")
g = collections.defaultdict(set)
for t in r.tensors:
g[re.sub(r'^blk\.\d+', 'blk.N', t.name)].add(t.tensor_type.name)
for k in sorted(g): print(k, sorted(g[k]))
```
2. Map tiers: Unsloth's **Q6_K tensors = most sensitive** (typically `attn_v`, `output.weight`,
MTP/nextn projections), **Q5_K = next** (attn q/k/qkv/gate/output, `ffn_up/down`, SSM/`ssm_out`),
**IQ4_XS/Q4_K = bulk** (`ffn_gate`, `token_embd`).
3. Budget math: each tensor bumped q4βq8 costs `params Γ 0.5` bytes (8.25 vs 4.25 bpw).
Count instances per group from the BF16 shards, then greedily add from most sensitive
until file size β target (dry-run reports exact `quant size`).
Worked example β Qwen3.8-27B (`qwen35` arch, 65 layers: 17 full-attn + 48 linear-attn/SSM,
dense FFN, MTP `nextn` head), 16 GB target:
```bash
./build-cpu/bin/llama-quantize \
--tensor-type "attn_q.weight=q8_0_rocmfpx" \
--tensor-type "attn_k.weight=q8_0_rocmfpx" \
--tensor-type "attn_v.weight=q8_0_rocmfpx" \
--tensor-type "attn_output.weight=q8_0_rocmfpx" \
--tensor-type "attn_gate.weight=q8_0_rocmfpx" \
--tensor-type "ssm_out.weight=q8_0_rocmfpx" \
--tensor-type "^output.weight=q8_0_rocmfpx" \
Qwen3.8-27B-BF16-00001-of-00002.gguf Qwen3.8-27B-Q4FAST-Q8-sensitive.gguf Q4_0_ROCMFP4_FAST
# β 16773 MiB (5.15 bpw): 340Γ q4_0_rocmfp4_fast + 165Γ q8_0_rocmfpx + 360Γ f32
```
Notes for this recipe:
- Use `^output.weight` (anchored regex) β plain `output` would also match `attn_output`.
- With base `Q4_0_ROCMFP4_FAST`, the quantizer **auto-bumps MTP/draft-sensitive projections**
(`nextn.eh_proj` etc.) to plain `q8_0` β expected and desirable.
- `attn_qkv` (fused QKV of linear-attn layers) was left at q4 here to stay near budget;
it's the first thing to bump if you have ~1.2 GiB more headroom.
- `token_embd` stays q4 β Unsloth's own XL recipe keeps it at Q4_K.
- For MoE models, apply both sections: experts to q4 (Β§3/Β§7) *and* sensitive non-expert
tensors to q8 β the regexes compose freely.
### Gotchas
- Install the **repo's** `gguf-py` (`pip install -e gguf-py`); the stock pip `gguf` raises
`ValueError: ... is not a valid GGMLQuantizationType` on these types.
- `--output-tensor-type` only touches `output.weight` β the base ftype defines "everything else".
- Dense leading layers (e.g. Laguna blk.0 `ffn_gate/up/down`) are not routed experts and correctly
stay `q8_0_rocmfpx`.
- Pass the **first** shard as input; llama.cpp picks up `-0000N-of-0000M` siblings automatically.
- Quantize from BF16/F16 sources for quality; never requantize an already-quantized GGUF.
|