File size: 11,286 Bytes
97785d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
# ROCmFPX Quantization Guide β€” q4_0_rocmfp4_fast experts + q8_0_rocmfpx everything else

Self-contained playbook for producing "hybrid" ROCmFPX GGUF quants from a Hugging Face repo link:
**routed MoE experts (up/down/gate) β†’ `q4_0_rocmfp4_fast`, every other tensor β†’ `q8_0_rocmfpx`**.
Runs fully **CPU-only** (no GPU needed; quantization is a CPU reference path).

Validated end-to-end on:
| Model | Source | Recipe applied to | Result |
|---|---|---|---|
| Laguna-S-2.1 (118B-A10B, poolside, arch `laguna`) | `unsloth/Laguna-S-2.1-GGUF/BF16` (5 shards) | `ffn_{gate,up,down}_exps` | 224 GB BF16 β†’ 61.6 GB (4.39 bpw) |
| Qwen3.8-27B (dense) | `unsloth/Qwen3.8-27B-GGUF/BF16` (2 shards) | n/a (dense) | 52 GB BF16 β†’ 26.9 GB (8.25 bpw) |
| G4-MeroMero-26B-A4B (Gemma4 MoE) | raw HF weights `llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic` | `ffn_gate_up_exps`, `ffn_down_exps` | converted to BF16 GGUF (50.5 GB) β†’ 14.0 GB (4.64 bpw) |
| Qwen3.8-27B hybrid (dense, `qwen35` arch) | `unsloth/Qwen3.8-27B-GGUF/BF16` | sensitive tensors per Unsloth UD-Q4_K_XL (see Β§8) | 16.4 GB (5.15 bpw), targets ~16 GB VRAM-fit file |

---

## 0. One-time setup (CPU-only)

```bash
# System deps: gcc, cmake, git, python3 + pip
git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX

# CPU-only build (no HIP/CUDA/Vulkan/Metal)
cmake -B build-cpu -DGGML_HIP=OFF -DGGML_CUDA=OFF -DGGML_VULKAN=OFF \
      -DGGML_METAL=OFF -DCMAKE_BUILD_TYPE=Release
cmake --build build-cpu -j "$(nproc)" --target llama-quantize llama-gguf-split llama-gguf

# Python deps for HF→GGUF conversion (only needed for Workflow B)
pip install -r requirements/requirements-convert_hf_to_gguf.txt
pip install -e gguf-py          # install the repo's gguf lib (older pip gguf can't read ROCmFPX types!)
```

Binaries land in `build-cpu/bin/`: `llama-quantize`, `llama-gguf-split`, `llama-gguf`.

> **Note:** the ROCmFPX quant types (`q4_0_rocmfp4_fast` = 101, `q8_0_rocmfpx` = 103, etc.)
> are only understood by this fork's tooling and recent-ish llama.cpp. Use the binaries built here.

---

## 1. Decide which workflow applies

Look at the HF repo (`https://huggingface.co/<repo>`):

- **Workflow A** β€” repo already contains GGUF files, e.g. `BF16/*-00001-of-0000N.gguf`
  (typical for `unsloth/*-GGUF` repos). Skip straight to quantizing.
- **Workflow B** β€” repo contains raw weights (`model*.safetensors` + `config.json`).
  Convert to BF16 GGUF first, then quantize.

Check sizes before downloading:

```bash
curl -s https://huggingface.co/api/models/<ORG>/<REPO>/tree/main/BF16 | \
  python3 -c "import json,sys; [print(s['path'], round(s['size']/1e9,2),'GB') for s in json.load(sys.stdin)]"
```

Download with `hf` CLI (shards download in parallel; pass the first shard to llama.cpp later β€”
remaining shards are discovered automatically):

```bash
pip install -q huggingface_hub
hf download <ORG>/<REPO> --repo-type model --include "BF16/*" --local-dir /workspace/models/<NAME>
```

---

## 2. Find the routed-expert tensor names (critical step)

Tensor naming depends on architecture. Inspect the first GGUF shard (header only β€” works as soon as
the file starts existing, and only needs a few MB):

```bash
python3 - <<'EOF'
from gguf import GGUFReader
import re, collections
r = GGUFReader("<first-shard>.gguf", "r")
print("arch:", bytes(r.fields['general.architecture'].parts[-1]).decode())
names = sorted(set(re.sub(r'^blk\.\d+', 'blk.N', t.name) for t in r.tensors))
for n in names: print(n)
print(collections.Counter(t.tensor_type.name for t in r.tensors))
EOF
```

Build the expert regex from the patterns you see:

| What you see in the tensor list | Expert regex for `--tensor-type` |
|---|---|
| `blk.N.ffn_gate_exps.weight`, `ffn_up_exps`, `ffn_down_exps` (packed, e.g. `laguna`) | `ffn_(gate\|up\|down)_exps` |
| `blk.N.ffn_gate_up_exps.weight`, `ffn_down_exps` (fused gate+up, e.g. `gemma4`) | `ffn_(gate_up\|down)_exps\.weight` |
| `blk.N.ffn_gate.M`, `ffn_up.M`, `ffn_down.M` (per-expert tensors) | `ffn_(gate\|up\|down)\.\d+\.` |

Patterns are **regexes matched against lowercased tensor names** (regex_search, i.e. partial match).
Norms, biases, router scales (`*.scale`, `*.bias`, 1-D tensors) are never quantized regardless, so
matching a little too broadly is harmless β€” but still verify with a dry run.

---

## 3. Dry run (always do this first)

`--output-tensor-type` only affects `output.weight`; to get "everything else = q8_0_rocmfpx" use the
**base ftype `Q8_0_ROCMFPX`** and override experts:

```bash
./build-cpu/bin/llama-quantize --dry-run \
  --tensor-type "<EXPERT_REGEX>=q4_0_rocmfp4_fast" \
  <input-BF16-shard-1.gguf> <output-name.gguf> Q8_0_ROCMFPX
```

Confirm in the output:
- `applying manual override: q8_0_rocmfp4_fast...` lines appear for **every** expert tensor
  (count = layers Γ— {gate,up,down}, e.g. 47 MoE layers Γ— 3 = 141 for Laguna),
- attention / shared experts / embeddings β†’ `(q8_0_rocmfpx)`,
- norms/biases/routers stay `f32`,
- the final `quant size = ... (X.XX BPW)` line looks sane (~4.4 bpw for mostly-expert MoE,
  ~8.25 bpw for dense).

For a dense model, drop `--tensor-type` entirely.

---

## 4. Quantize

```bash
./build-cpu/bin/llama-quantize \
  --tensor-type "<EXPERT_REGEX>=q4_0_rocmfp4_fast" \
  <input-BF16-shard-1.gguf> <output-name.gguf> Q8_0_ROCMFPX
```

Flags worth knowing:
- `--keep-split` β€” output keeps the input's shard layout.
- Default (no flag) β€” auto-splits output at 50 GiB.
- `--tensor-type` can be repeated, or use `--tensor-type-file list.txt` (one `name=type` per line).
- Add trailing `nthreads` positional arg to cap thread count.

Speed reference (64-core server): 118B MoE β‰ˆ 8 min, 27B dense β‰ˆ 1 min, 26B MoE β‰ˆ 3 min.
Memory use is modest (mmap-based); RAM β‰ˆ source size is plenty.

---

## 5. Verify + merge shards

Verify the result (works on merged or shard-1 file):

```bash
./build-cpu/bin/llama-quantize --dry-run <output.gguf> /tmp/dummy.gguf COPY | grep "type .*tensors"
# e.g.:  f32: 287 | q4_0_rocmfp4_fast: 141 | q8_0_rocmfpx: 386
```

Merge shards into one file (only pass the first shard; the rest are auto-discovered):

```bash
./build-cpu/bin/llama-gguf-split --merge \
  <output>-00001-of-0000N.gguf <output-merged.gguf>
```

(`--split` / `--split-max-size` do the reverse. Note: the `llama-gguf <file> r` example tool aborts
on ROCmFPX types even for valid files β€” use the quantize-dry-run check above instead.)

---

## 6. Workflow B only: HF weights β†’ BF16 GGUF

For repos with `*.safetensors` + `config.json` (check `architectures` in config.json is registered in
`convert_hf_to_gguf.py` β€” e.g. `Gemma4ForConditionalGeneration` is supported by this fork):

```bash
hf download <ORG>/<REPO> --local-dir /workspace/models/<NAME>   # all files incl. tokenizer

python3 convert_hf_to_gguf.py /workspace/models/<NAME> \
  --outfile /workspace/models/<NAME>-BF16.gguf --outtype bf16
# for multimodal models, add --mmproj to also export the vision projector (separate file)
```

Then continue at **Step 2** with the produced BF16 GGUF.

---

## 7. Cheat sheet of the runs documented here

```bash
# Laguna-S-2.1 (laguna arch, packed experts)
./build-cpu/bin/llama-quantize \
  --tensor-type "ffn_(gate|up|down)_exps=q4_0_rocmfp4_fast" \
  --keep-split \
  Laguna-S-2.1-BF16-00001-of-00005.gguf Laguna-S-2.1-Q8_0_ROCMFPX-Q4FAST-experts.gguf Q8_0_ROCMFPX

# Qwen3.8-27B (dense β€” pure q8_0_rocmfpx)
./build-cpu/bin/llama-quantize \
  Qwen3.8-27B-BF16-00001-of-00002.gguf Qwen3.8-27B-Q8_0_ROCMFPX.gguf Q8_0_ROCMFPX

# Gemma4 MoE (fused gate_up experts), after convert_hf_to_gguf.py
./build-cpu/bin/llama-quantize \
  --tensor-type "ffn_(gate_up|down)_exps.weight=q4_0_rocmfp4_fast" \
  G4-MeroMero-26B-A4B-BF16.gguf G4-MeroMero-26B-A4B-Q8_0_ROCMFPX-Q4FAST-experts.gguf Q8_0_ROCMFPX
```

### Other ROCmFPX ftypes you may be asked for
`Q4_0_ROCMFP4`, `Q4_0_ROCMFP4_FAST` (4.25 bpw, speed-first), `Q4_0_ROCMFP4_STRIX(_LEAN)`,
`Q4_0_ROCMFP4_COHERENT`, `Q8_0_ROCMFPX`, `Q6_0_ROCMFPX`, `Q3_0_ROCMFPX`, `Q2_0_ROCMFPX` (2.5 bpw).
List all with `./build-cpu/bin/llama-quantize --help`.

---

## 8. VRAM-fit hybrid recipe (q4 bulk + q8 sensitive), e.g. ~16 GB target

When the file must fit a smaller GPU (e.g. 24 GB laptop VRAM minus KV cache β‡’ ~16 GB file),
invert the recipe: **base ftype `Q4_0_ROCMFP4_FAST`**, then override *sensitive* tensors to
`q8_0_rocmfpx`. To pick the sensitive set, mirror the tiers of Unsloth's Dynamic recipe
(e.g. `UD-Q4_K_XL`) for the same model:

1. Download the reference quant and tabulate per-tensor types:
   ```python
   from gguf import GGUFReader
   import re, collections
   r = GGUFReader("<UD-Q4_K_XL>.gguf", "r")
   g = collections.defaultdict(set)
   for t in r.tensors:
       g[re.sub(r'^blk\.\d+', 'blk.N', t.name)].add(t.tensor_type.name)
   for k in sorted(g): print(k, sorted(g[k]))
   ```
2. Map tiers: Unsloth's **Q6_K tensors = most sensitive** (typically `attn_v`, `output.weight`,
   MTP/nextn projections), **Q5_K = next** (attn q/k/qkv/gate/output, `ffn_up/down`, SSM/`ssm_out`),
   **IQ4_XS/Q4_K = bulk** (`ffn_gate`, `token_embd`).
3. Budget math: each tensor bumped q4β†’q8 costs `params Γ— 0.5` bytes (8.25 vs 4.25 bpw).
   Count instances per group from the BF16 shards, then greedily add from most sensitive
   until file size β‰ˆ target (dry-run reports exact `quant size`).

Worked example β€” Qwen3.8-27B (`qwen35` arch, 65 layers: 17 full-attn + 48 linear-attn/SSM,
dense FFN, MTP `nextn` head), 16 GB target:

```bash
./build-cpu/bin/llama-quantize \
  --tensor-type "attn_q.weight=q8_0_rocmfpx" \
  --tensor-type "attn_k.weight=q8_0_rocmfpx" \
  --tensor-type "attn_v.weight=q8_0_rocmfpx" \
  --tensor-type "attn_output.weight=q8_0_rocmfpx" \
  --tensor-type "attn_gate.weight=q8_0_rocmfpx" \
  --tensor-type "ssm_out.weight=q8_0_rocmfpx" \
  --tensor-type "^output.weight=q8_0_rocmfpx" \
  Qwen3.8-27B-BF16-00001-of-00002.gguf Qwen3.8-27B-Q4FAST-Q8-sensitive.gguf Q4_0_ROCMFP4_FAST
# β†’ 16773 MiB (5.15 bpw): 340Γ— q4_0_rocmfp4_fast + 165Γ— q8_0_rocmfpx + 360Γ— f32
```

Notes for this recipe:
- Use `^output.weight` (anchored regex) β€” plain `output` would also match `attn_output`.
- With base `Q4_0_ROCMFP4_FAST`, the quantizer **auto-bumps MTP/draft-sensitive projections**
  (`nextn.eh_proj` etc.) to plain `q8_0` β€” expected and desirable.
- `attn_qkv` (fused QKV of linear-attn layers) was left at q4 here to stay near budget;
  it's the first thing to bump if you have ~1.2 GiB more headroom.
- `token_embd` stays q4 β€” Unsloth's own XL recipe keeps it at Q4_K.
- For MoE models, apply both sections: experts to q4 (Β§3/Β§7) *and* sensitive non-expert
  tensors to q8 β€” the regexes compose freely.

### Gotchas
- Install the **repo's** `gguf-py` (`pip install -e gguf-py`); the stock pip `gguf` raises
  `ValueError: ... is not a valid GGMLQuantizationType` on these types.
- `--output-tensor-type` only touches `output.weight` β€” the base ftype defines "everything else".
- Dense leading layers (e.g. Laguna blk.0 `ffn_gate/up/down`) are not routed experts and correctly
  stay `q8_0_rocmfpx`.
- Pass the **first** shard as input; llama.cpp picks up `-0000N-of-0000M` siblings automatically.
- Quantize from BF16/F16 sources for quality; never requantize an already-quantized GGUF.