How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
# Run inference directly in the terminal:
llama cli -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
# Run inference directly in the terminal:
llama cli -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
# Run inference directly in the terminal:
./llama-cli -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
# Run inference directly in the terminal:
./build/bin/llama-cli -hf kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
Use Docker
docker model run hf.co/kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF:Q4_0_ROCMFP
Quick Links

⚠️ STOCK llama.cpp WILL NOT LOAD THIS MODEL

🚀 37.86 tok/s on a 119B MoE — 63.07 GiB, 7 GiB smaller than UD-Q4_K_XL.

Mistral-Small-4-119B-A6.5B — ROCmFP4 (tier 102 COHERENT) GGUF

A 4-bit ROCmFP4 quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 (222 GiB) — a lossless source, not a requantization of a lower-bit build.

File Mistral-Small-4-119B-2603-Q4_0_ROCMFP4_COHERENT.gguf (sharded)
Size 63.07 GiB
BPW 4.55
ftype Q4_0_ROCMFP4_COHERENT (102)

⛔ Requires a llama.cpp with the ROCmFP4 quant types

Q4_0_ROCMFP4_COHERENT (ftype 102) exists only in charlie12345/ROCmFPX, not upstream llama.cpp. Ignore the auto-generated "Use this model" commands above.


Files — where each build lives

build ftype size where
4-bit COHERENT (2 shards) 102 63.07 GiB in this repo
8-bit plain 111 114.38 GiB in this repo · also at …-ROCmFPX-Q8_0-GGUF
8-bit AGENT 115 116.15 GiB ➡️ …-ROCmFPX-Q8_0-AGENT-GGUF

The AGENT build is hosted in its own repo rather than mirrored here — at ~116 GiB the duplication is not worth it, and (see the warning above) neither 8-bit build fits in GPU memory on a 128 GB Strix Halo. For a single Strix Halo, use the 4-bit build in this repo.

The separate per-quant repos exist so a user searching HF for a specific quant finds it directly.

All quant variants

variant ftype size bpw GPU on 128 GB Strix Halo decode
4-bit COHERENT 102 63.07 GiB 4.55 ✅ fits 37.86 tok/s
8-bit AGENT 115 116.15 GiB 8.39 ⛔ does not fit not measurable on this box
8-bit plain 111 114.38 GiB 8.26 ⛔ does not fit not measurable on this box

Repos: 4-bit · 8-bit AGENT · 8-bit plain

Measured

Ryzen AI MAX+ 395 (gfx1151, 128 GB unified, ROCm 7.2.4). Median of 3+, warm-up discarded, otherwise-idle box. Correctness at the model's official sampling.

build size decode (median) range
this build 63.07 GiB 37.86 [37.83 – 38.65] (tight)
UD-Q4_K_XL 70 GiB 18.75 [15.99 – 34.30] (wide)

Ranges disjoint (34.30 < 37.83). The median ratio looks like +102%, but the baseline's own variance is large — we report the win without leaning on that headline number. Note the baseline is UD-Q4_K_XL, not Q4_K_M.

Correctness: 17×23 ⇒ ✅ 391 · capital of Japan ⇒ ✅ Tokyo · days in 2024 ⇒ ✅ 366

Per-tensor types (audited in the finished file)

output.weight Q6_K · token_embd Q6_K · shexp 108× Q8_0 (shared-expert protection) · router 36× F32 · norms 145× F32 · bulk TYPE_100 (288)

119B total / 6.5B active. Built with --tensor-type shexp=q8_0; shared experts are dense (they see every token) so their error is systematic, not averaged.


What was NOT measured

  • No perplexity run, and no quality A/B against the baseline or the source. The checks above are memorized-fact prompts — necessary but not sufficient; a damaged model can pass them.
  • No long-context testing.
  • No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
50
GGUF
Model size
119B params
Architecture
mistral4
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/Mistral-Small-4-119B-ROCmFP4-GGUF

Quantized
(34)
this model