GPT-OSS-120B β€” ROCmFP4 STRIX_LEAN for AMD Strix Halo

ROCmFP4 GGUF requantization of openai/gpt-oss-120b (117B MoE, 5.1B active), produced for AMD Strix Halo (Ryzen AI Max, gfx1151).

To our knowledge the first ROCmFP4-family quant of gpt-oss. +21% decode over the native MXFP4 release on Strix Halo at essentially identical size β€” the gain is pure kernel efficiency (the ROCmFP4 layout streams at near-peak bandwidth on gfx1151).

⚠️ Stock llama.cpp will reject this file (invalid ggml type 106). It runs on:

Companion repos: Qwen3.6-35B-A3B Β· Nemotron-3-Nano-30B Β· Ornith-1.0-35B Β· Ornith-1.0-9B

Files

File Quant BPW Size
gpt-oss-120b-Q4_0_ROCMFP4_STRIX_LEAN.gguf Q4_0_ROCMFP4_STRIX_LEAN ~4.25 62.4 GB

Measured performance

AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X, ROCm, llama-bench -fa 1 --mmap 0):

Quant Size pp512 tg128
Q4_0_ROCMFP4_STRIX_LEAN 58.1 GiB 593 61.7
native MXFP4 (ggml-org) 59.0 GiB 587 50.9

Provenance & quality disclosure

Requantized from the native MXFP4 release (no BF16 master exists for gpt-oss) with llama-quantize --allow-requantize, calibrated with bartowski's imatrix. FP4β†’FP4 regridding adds rounding noise; on an 8-prompt quality battery vs the native release (temp 0), 7/8 answers were byte-equivalent-or-equal-quality and one contained a minor arithmetic slip in an illustrative example (the substantive answer remained correct). Judge that trade for your workload β€” the native MXFP4 remains available from ggml-org.

Serving

llama-server -m gpt-oss-120b-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  -ngl 999 -fa on --jinja -c 65536

Credits

  • Base model: OpenAI β€” gpt-oss-120b (Apache-2.0)
  • imatrix: bartowski
  • ROCmFP4 quant formats: Hal0ai; GPU quantizer build: kyuz0
Downloads last month
198
GGUF
Model size
117B params
Architecture
gpt-oss
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for gsrunion/gpt-oss-120b-ROCmFP4-STRIX_LEAN-GGUF

Quantized
(121)
this model