โš ๏ธ STOCK llama.cpp WILL NOT LOAD THIS MODEL

๐Ÿ“ฆ 1.985 GB โ€” 5.5% smaller than Q4_K_M, decode within noise.

Granite-4.1-3B โ€” ROCmFP4 (tier 102 COHERENT) GGUF

A 4-bit ROCmFP4 quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 GGUF (6.81 GB) โ€” a lossless source, not a requantization of a lower-bit build.

File granite-4.1-3b-Q4_0_ROCMFP4_COHERENT.gguf
Size 1.985 GB
BPW 4.66
ftype Q4_0_ROCMFP4_COHERENT (102)

โ›” Requires a llama.cpp with the ROCmFP4 quant types

Q4_0_ROCMFP4_COHERENT (ftype 102) exists only in charlie12345/ROCmFPX, not upstream llama.cpp. Ignore the auto-generated "Use this model" commands above.


Measured

Ryzen AI MAX+ 395 (gfx1151, 128 GB unified, ROCm 7.2.4). Median of 3+, warm-up discarded, otherwise-idle box. Correctness at the model's official sampling.

build size decode (median) range
this build 1.985 GB 62.08 [61.17 โ€“ 63.84]
Q4_K_M 2.100 GB 59.37 [58.18 โ€“ 61.70]

โš ๏ธ Decode ranges OVERLAP โ€” we make no speed claim. The medians differ but the intervals cross, so this is within noise. The size saving is the real result.

Correctness: 17ร—23 โ‡’ โœ… 391 ยท capital of Japan โ‡’ โœ… Tokyo ยท days in 2024 โ‡’ โœ… 366

Per-tensor types (audited in the finished file)

280ร— Q4_0_ROCMFP4 (type 100) ยท token_embd Q6_K ยท 81 norms F32

granite is a dense architecture here โ€” no SSM/MoE state to protect.


What was NOT measured

  • No perplexity run, and no quality A/B against the baseline or the source. The checks above are memorized-fact prompts โ€” necessary but not sufficient; a damaged model can pass them.
  • No long-context testing.
  • No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
-
GGUF
Model size
3B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kingjones777/Granite-4.1-3B-ROCmFP4-GGUF

Quantized
(57)
this model