ROCmFPX Quantize Playbook

A field-tested, agent-ready playbook for producing ROCmFPX hybrid GGUF quantizations (q4_0_rocmfp4_fast / q8_0_rocmfpx) with the ROCmFPX fork of llama.cpp โ€” fully CPU-only, from either pre-quantized GGUF repos (e.g. Unsloth BF16) or raw HF safetensors.

The complete guide is in rocmfpx-quantization-guide.md. Hand this repo (or just that file) to a coding agent together with a Hugging Face link and a recipe, and it can reproduce every step below.

What it covers

  • One-time CPU-only build of llama-quantize / llama-gguf-split + Python conversion deps
  • Workflow A: source repo already has BF16 GGUF shards (e.g. unsloth/*-GGUF)
  • Workflow B: raw safetensors โ†’ convert_hf_to_gguf.py โ†’ BF16 GGUF
  • How to discover routed-expert tensor names per architecture (regex lookup table)
  • Dry-run verification, shard merging, and output validation
  • VRAM-fit hybrid recipe (q4 bulk + q8 sensitive) derived from Unsloth Dynamic tiers
  • Cheatsheets, timings, and gotchas (regex anchoring, MTP auto-protection, gguf-py version trapโ€ฆ)

Validated runs

Model Recipe Result
Laguna-S-2.1 118B-A10B q4 experts / q8 rest 224 GB โ†’ 61.6 GB (4.39 bpw)
Qwen3.8-27B pure q8 and 16 GB hybrid from UD-Q4_K_XL tiers 26.9 GB (8.25 bpw) / 16.4 GB (5.15 bpw)
G4-MeroMero-26B-A4B q4 experts / q8 rest + mmproj 50.5 GB โ†’ 14.0 GB (4.64 bpw)

Timings on a 64-core CPU box: 118B MoE โ‰ˆ 8 min, 27B dense โ‰ˆ 1 min, 26B MoE โ‰ˆ 3 min.

Related

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including JackBinary/ROCmFPX-Quantize-Playbook