Instructions to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: llama cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: llama cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: ./llama-cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: ./build/bin/llama-cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Use Docker
docker model run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- LM Studio
- Jan
- Ollama
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Ollama:
ollama run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- Unsloth Studio
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
- Docker Model Runner
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Docker Model Runner:
docker model run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- Lemonade
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-0731-ROCmFP3-MIX-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
I've bit of trial on Spreadsheet benchmark2, Following are the Claude conclusion, hope this help.
Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX — end-to-end evaluation
Environment
Hardware: AMD Ryzen AI MAX+ 395 (Strix Halo), gfx1151, 128 GB unified memory.
Host: Fedora, kernel 6.19.10, in-tree amdgpu, ROCm 7.2.4 userspace.
Build: GeometricAGI/lucebox-hub @ 119739b, branch feat/ds4-adaptive-on-upstream, via Dockerfile.rocm with DFLASH_HIP_ARCHES=gfx1151.
Server: dflash_server --ds4-fused-decode --m
Artifact: ds4-0731-opt1.gguf, 102,321,006,592 bytes, 3.5 bits/weight (14 bytes per 32 weights).
qtype-105 confirmed active:
using codebooks embedded in the GGUF (375320 bytes, no sidecar file needed)
registered 43 qtype-105 down-expert layer(s)
placement: hot=11008 (91.38 GiB) cold=0
ds4_fused = ON
Sustained decode 19–22 tok/s — the fastest Dun on this hardware. --ds4-fused-decode wasworth about 33%.
Models compared
All are quants of the same base checkpoint, DeepSeek-V4-Flash-0731.
ROCmFP3-MIX — Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX, file ds4-0731-opt1.gguf, 95.3 GiB, 3.5 bpw. Served by
lucebox dflash_server from this fork.
Layers37-42-Q4KExperts — DeepSeek-V4-Flash-LertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed-0731.gguf, 90.9 GiB, mixed precision. Served by antirez ds4-server.
IQ2XXS-chat-v2 — DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf, 80.8 GiB, ~2.1 bpw.
Served by antirez ds4-server.
Both baselines ran under kyuz0/strix-halo-ds.4.
Benchmark — SpreadsheetBench-2 Template-16, thinking enabled
task ROCmFP3-MIX Layers37-42-Q4KExpert
02_01 0.0% 97.1% 92.8%
03_01 0.0% 88.1
04_03 77.4% 95.2% 95.2%
05_01 51.5% * 98.5
- truncated by our token budget, not compara
The headline: ROCmFP3-MIX at 3.5 bpw loses t.1 bpw. This is not bit starvation.
Two failure modes
A — misreads tokens present in its own conte
wb = openpyxl.load_workbox(wb_path) # Attr
The prompt contains the literal line wb = o). Reproduced three times identically — 662output tokens, about 62 s each, temperature 0, under two different token budgets. The same reply also carries an off-by-one column range (it writes a B row c would fail even with the identifiercorrected.
B — never terminates reasoning. On 03_01 the model never emits . With --think-max-tokens 24576 it produced exactly 24,577 output tokens; with --think-mxactly 40,961. It consumes whatever ceilingis set and then force-closes with no parseable answer. Layers37-42-Q4KExperts completes the same task in 2,015 s at 88.1%.
Three defects in the release
The published sha256 is wrong.
model card : accabb4cb83cf180ce18e4c5c5bbc3941a6
git-LFS oid: a431f955c639107d3f34ab4cc65edc806038588a0c258abd0bc3e9074b908096
x-linked-etag agrees with the LFS oid, and our byte-verified file matches it. A correct download fails the card's check.No model card sidecar, and no chat templaV4-Flash-0731-ROCMFPX, which matches nothingin share/model_cards/ — only gemma, laguna and qwen ship — so the server falls back to hard defaults. The GGUF also carries no tokenizer.chat_template, so we aumplate to get correct prompting. Shippingdeepseek-v4-flash-0731-rocmfpx.json would remove a real setup barrier.
ROCM_VERSION=6.4.1, the Dockerfile's own default, builds but cannot load. Exit 139 (SIGSEGV) in the hybrid expert load path, reproducible. 7.2.4 works. Your onst a ROCm 7 host driver; it deserves to bethe default, or a build-time check.
Reading
Both failure modes are precision failures rather than reasoning failures — the output stays fluent and well-structured while exact-token discrimination degrades. Toks fitted to minimise average divergence:the reported KL −57% and PPL −5.6% on wikitext2 and c4 are plausibly real, and simply don't transfer to code generation, where one wrong identifier scorefraction of a nat. Your own caveat —"fake-quant research measurements, not end-to-end serving benchmarks" — is exactly right, and this is the end-to-end number. The 96.4% expert coverage may compous sit in the uncovered 3.6% that fell back to the uniform rung, they are strictly worse than baseline while everything else improved.
Caveats
Four tasks, one seed — a small sample, and 04_03 did complete cleanly at 77.4%.
Different inference engines: ROCmFP3-MIX ran on lucebox dflash_server, both baselines on antirez ds4-server. Engine effects cannot be fully excluded, though a m from its own prompt is hard to attribute toa runtime.
The uniform DeepSeek-V4-Flash-0731-ROCMFPX baseline was not tested — that repo is gated: manual and our access request is outstanding. So we cannot separate "adaptm "ROCmFP family behaviour in general".Ungating it, or publishing a paired A/B on a code benchmark, would settle that question.
Suggested next step: code-domain calibration, or a code-weighted expert-coverage pass, re-measured on HumanEval or similar.
thanks, investigating and running more tests.
https://huggingface.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCMFPX should now be ungated now.
Thank you — this is an unusually careful report, and two of the three defects were ours and are now fixed. Still investigating the third, which is the one that matters most.
Your failure mode B is our bug, not the quant
With
--think-max-tokens 24576it produced exactly 24,577 output tokens; with 40960, exactly 40,961.
That ceiling + 1 is the signature of a defect in our decode loop, and your two data points pinned it. The level-2 force-close reserves hard_limit_reply_budget tokens so a model that exhausts its thinking budget still writes a visible answer. On the DeepSeek4 backend it pushed the close sequence into the output and then break-ed out of the decode loop — so the reserve was computed, logged, and thrown away.
We reproduced it independently before seeing your post, with close sequences of 1, 3 and 23 tokens, landing on exactly 8193 / 8195 / 8215 output tokens. Zero answer tokens in every case, for every sequence length, and finish_reason still reported "stop" — which is why it looks like a wrong answer rather than a truncation. One item returned 22,706 characters of reasoning against 1 character of content.
Fixed in f4b212e: override the sampled token with the close sequence one token per step and keep decoding, so the injected tokens go through a forward pass and the reserve is actually spendable. (qwen35_backend already did this; DeepSeek4 was the outlier.) Verified on hardware just now — same item, same reasoning length, answer goes from 1 character to 10,060:
before: completion_tokens=8193 reasoning=22,706 chars content=1 char
after: completion_tokens=12288 reasoning=22,707 chars content=10,060 chars
Your 02_01 and 03_01 zeros are almost certainly this. Your baselines ran on antirez ds4-server, which doesn't share our decode loop — which is exactly why only ours produced them. We have not re-run SpreadsheetBench yet, so we're not claiming your scores are fixed, only that the mechanism you documented is.
Defect 1: the published sha256 was wrong. Corrected.
You were right, and the diagnosis was exact. The card carried accabb4c…, which was a pre-release build's hash — a correct download failed our own check. The real value is a431f955…, matching the LFS oid and x-linked-etag, as you found. The card is now corrected with a dated note so anyone who checked earlier knows why it failed. Sorry for the wasted verification time.
Defect 2: sidecar and chat template
Confirmed on both counts. The GGUF carries no tokenizer.chat_template — we verified directly, you were not missing anything. general.name is DeepSeek-V4-Flash-0731-ROCMFPX, which normalises to exactly the filename you asked for: share/model_cards/deepseek-v4-flash-0731-rocmfpx.json. That file now exists and is in review, along with a deepseek4 family fallback so the arch stops landing on hard defaults even without a sidecar, and a fix for share/model_cards/ not being found from a CMake build tree at all (it was looked for one directory too shallow, so shipped cards silently never applied). Embedding the chat template in the GGUF is the right fix and is on the list.
Defect 3: ROCm 6.4.1 default
Accepted. A default that builds and then SIGSEGVs is worse than no default; it should be ROCm 7 or a build-time check.
The paired A/B you asked for
You noted the uniform baseline was gated so you couldn't separate "adaptive is worse" from "ROCmFP family behaviour". That repo is now ungated, and we ran exactly that comparison overnight — same base checkpoint, ds4-0731-uniform.gguf vs ds4-0731-opt1.gguf, identical items, temperature 0, paired McNemar on discordant pairs:
| suite | n | uniform | adaptive | Δ | McNemar |
|---|---|---|---|---|---|
| HumanEval | 164 | 65.9% | 72.6% | +6.7 | χ²=1.82, ns |
| GSM8K | 200 | 86.5% | 86.5% | 0.0 | χ²=0.03, ns |
| MATH-500 | 120 | 75.0% | 73.3% | −1.7 | χ²=0.04, ns |
| pooled | 484 | 76.7% | 78.5% | +1.9 | χ²=0.54, ns |
No significant difference in either direction, with the sign flipping by suite. Two caveats that matter: this ran with thinking off, precisely because of the bug above, so it does not exercise the long-reasoning path where you saw failures; and it is our-quant-vs-our-quant, so it structurally cannot detect a defect shared by both. Your comparison against IQ2XXS is the absolute reference we didn't have, which is why your report is more informative than ours on this point.
Your failure mode A is the open question
load_workbook → load_workbox, with the correct string sitting in the prompt, reproduced three times at temperature 0, is not explained by the decode bug and survives the fix. We take it seriously. It is also invisible to everything in the table above: those suites score a final answer, and a relative A/B cancels any damage common to both arms.
Your reading is sharper than our own framing here — KL and PPL are fitted to minimise average divergence, and one wrong identifier costs a fraction of a nat while failing the task outright. So "KL −57%" and "correct code" are not the same claim, and we shouldn't have let the former stand in for the latter.
We're running an exact-copy fidelity probe now — identifiers placed in the prompt and checked for verbatim reproduction, scored absolutely rather than arm-vs-arm, since copying a string that's in context is not a matter of model skill. If that reproduces what you saw, code-domain calibration or a code-weighted coverage pass is the right next step, as you suggest.
Investigation is ongoing and we'll post results either way, including negative ones. If you're willing to re-run 02_01 / 03_01 against f4b212e once it lands, that would be the cleanest confirmation that B is closed — no obligation, you've already done more than your share of our QA.
Update, and it is not good news for the artifact: chasing your failure mode A we found a correctness bug in the adaptive decode path, not the precision loss we were both reasoning about. Tracking it down is now our top priority ahead of everything else.
The finding
We built an absolute probe rather than another A/B — put an identifier in the prompt and ask for it back verbatim, temperature 0, since copying a string that is present in context is not a matter of model skill. 20 identifiers, 2 repeats, both arms:
| run | result |
|---|---|
| uniform ROCMFPX, fused, ctx 4096 | 40/40 clean |
| uniform ROCMFPX, fused, ctx 8192 | 24/24 clean |
| ROCmFP3-MIX, fused, ctx 4096 | 24/40 — 16 failures |
| ROCmFP3-MIX, fused, ctx 8192 | 24/40 — the same 16, deterministic |
| ROCmFP3-MIX, fusion disabled | 5 requests fine, then the server died |
The failures are not near-miss identifiers. On "repeat the following line of Python exactly" we get:
value = openpyxl.load_workbook(argument, flag=True)
-> " 1. **Bootstrap** (or [open]S) 2. **PIPE** (or [open]S) 3. **PIPE** ..."
value = train_test_split(argument, flag=True)
-> " 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 ..."
value = normalize_whitespace_inplace(argument, flag=True)
-> "" (empty)
while other cases in the same run come back perfectly. Deterministic across repeats and reproducible across two server configurations, on the same eight identifiers. Garbage on one code path and a hard crash on the other, only for the artifact that uses qtype-105/106, reads as a memory-safety or indexing fault in the adaptive expert decode, not as bit starvation. That is a hypothesis about the mechanism — the reproducibility is not.
What this changes about our earlier reply
Two corrections we owe you.
We told you your 02_01 / 03_01 zeros were "almost certainly" the budget-hook bug. That was too confident. The budget-hook defect is real and fixed — your 24,577 / 40,961 numbers pinned it exactly — but this new bug is input-dependent and produces exactly the kind of unparseable output that scores 0.0%. Your zeros may be this, or both. We should not have given you a single confident cause.
We agreed too readily that both your failure modes were "precision failures". Your framing on KL was correct and remains correct — averaged divergence cannot see a single wrong identifier. But mode A now looks structural rather than precision-limited, which means code-domain calibration, the fix we were preparing on your suggestion, would not have addressed it. We have parked that work until this is resolved.
Practical advice
Treat ds4-0731-opt1.gguf as defective for now. We would not want you spending more bench time on it. The uniform DeepSeek-V4-Flash-0731-ROCMFPX artifact is clean on every probe we have run and is now ungated, so it is the better basis for any comparison in the meantime.
Worth being explicit about why we did not catch this ourselves: we ran 484 paired items across HumanEval, GSM8K and MATH-500 overnight, and this artifact scored 86.5% on GSM8K. The bug is input-dependent, so aggregate benchmarks average straight over it — and both arms of our A/B were our own, so any defect the comparison shares is invisible by construction. Your comparison against a differently-produced quant is what made the problem visible at all. That is a lesson about our methodology, not just about this bug.
We will post the root cause and a fix here, including if it turns out to be worse than we currently think.
Root cause found, and it is a serving flag, not the artifact. Please disregard my previous message telling you to treat ds4-0731-opt1.gguf as defective — that was wrong, and I am sorry for the wasted signal.
--ds4-expert-top-k 4 is the cause
That flag keeps four of the model's six routed experts and renormalizes. Exact-copy fidelity, 20 identifiers x 3 repeats, temperature 0, all four cells measured back to back on the same box and binary:
| artifact | top-k 6 (model default) | top-k 4 |
|---|---|---|
| ROCmFP3-MIX (adaptive) | 95.0% (57/60) | 60.0% (36/60) |
| ROCMFPX (uniform) | 100% (60/60) | 100% (60/60) |
The uniform artifact is completely indifferent to the approximation. The adaptive one loses 35 points, and the failures are the degenerate output I showed you last time. So top-4 leaves no error margin, and the adaptive formats' different error profile crosses a threshold the uniform formats' does not. At the model default the adaptive artifact is fine.
If your runs passed --ds4-expert-top-k 4, that is very likely your failure mode A, and possibly your 02_01 / 03_01 zeros too. Omit the flag, or pass 0, and the model uses its six routed experts. Worth a re-run before you spend anything else on this.
Where the flag came from: us
server/docs/DS4.md shipped --ds4-expert-top-k 4 in five copy-paste example commands, including the block labelled "the validated single-device Strix Halo profile". Elsewhere the same document warned it was "an approximate inference policy [that] must be quality-validated" and that you should "omit it to retain the model's default six routed experts" — but a caveat in paragraph nine loses to a working command in paragraph three every time. That is on us, not on anyone who copied it.
Fixed now: the examples no longer include it, the flag's description carries the table above, and this model card has a serving section saying the same. The throughput note that compared top-4 against the default (32.12 vs 28.26 tok/s) is annotated, because it was measured on a uniform artifact and that trade is not available on an adaptive one.
Three corrections in this thread, which is the real lesson
I have now given you three different causes for the same symptom: the thinking-budget bug, then a "correctness bug in the adaptive decode path", now the top-k flag. The budget-hook defect was real and is fixed — your 24,577 / 40,961 numbers nailed it. The second was me over-reading a 40% failure rate before running the obvious control, which was to vary the serving flags I had been holding constant all along. Both of my wrong calls came from the same habit: reaching for an explanation about the quantization before eliminating the configuration.
Your original diagnosis deserves better credit than I gave it. You wrote that both failure modes were "precision failures rather than reasoning failures", and I agreed and then went looking for precision damage. What you had actually identified was that the output stays fluent while exact-token discrimination collapses — which is exactly what an under-provisioned expert mixture does, and is why aggregate benchmarks miss it. We ran 484 paired items overnight and saw nothing, because both arms had the same bad flag.
Two things still open on our side, and we will post either way: re-running that A/B at the model default so we have an honest quality comparison, and the KL framing you criticised, which remains a fair hit.