โš ๏ธ STOCK llama.cpp WILL NOT LOAD THIS MODEL

๐Ÿš€ 39.90 tok/s โ€” ~13% faster than Q4_K_M, ranges disjoint by a wide margin.

Granite-4.1-8B โ€” ROCmFP4 (tier 102 COHERENT) GGUF

A 4-bit ROCmFP4 quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 GGUF โ€” a lossless source, not a requantization of a lower-bit build.

File granite-4.1-8b-Q4_0_ROCMFP4_COHERENT.gguf
Size 5.162 GB
BPW 4.69
ftype Q4_0_ROCMFP4_COHERENT (102)

โ›” Requires a llama.cpp with the ROCmFP4 quant types

Q4_0_ROCMFP4_COHERENT (ftype 102) exists only in charlie12345/ROCmFPX, not upstream llama.cpp. Ignore the auto-generated "Use this model" commands above.


Measured

Ryzen AI MAX+ 395 (gfx1151, 128 GB unified, ROCm 7.2.4). Median of 3+, warm-up discarded, otherwise-idle box. Correctness at the model's official sampling.

build size decode (median) range
this build 5.162 GB 39.90 [39.87 โ€“ 39.93]
Q4_K_M 5.348 GB 35.27 [35.22 โ€“ 35.28]

~13% faster, ranges disjoint (39.87 min vs 35.28 max). Also 3.5% smaller.

Correctness: 17ร—23 โ‡’ โœ… 391 ยท capital of Japan โ‡’ โœ… Tokyo ยท days in 2024 โ‡’ โœ… 366

Per-tensor types (audited in the finished file)

token_embd Q6_K ยท norms F32 ยท bulk Q4_0_ROCMFP4 (type 100)

granite is dense โ€” no SSM/MoE state to protect.


What was NOT measured

  • No perplexity run, and no quality A/B against the baseline or the source. The checks above are memorized-fact prompts โ€” necessary but not sufficient; a damaged model can pass them.
  • No long-context testing.
  • No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
112
GGUF
Model size
9B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kingjones777/Granite-4.1-8B-ROCmFP4-GGUF

Quantized
(78)
this model