Instructions to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: llama cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: llama cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: ./llama-cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX # Run inference directly in the terminal: ./build/bin/llama-cli -hf Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Use Docker
docker model run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- LM Studio
- Jan
- Ollama
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Ollama:
ollama run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- Unsloth Studio
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX to start chatting
- Docker Model Runner
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Docker Model Runner:
docker model run hf.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
- Lemonade
How to use Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-0731-ROCmFP3-MIX-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
File size: 5,325 Bytes
d2be8b2 1287d04 d2be8b2 1287d04 b05323d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | ---
license: mit
base_model: deepseek-ai/DeepSeek-V4-Flash-0731
tags:
- gguf
- deepseek
- rocm
- strix-halo
- quantized
---
# DeepSeek-V4-Flash-0731 — Option 1: adaptive down-experts (ROCmFP3-MIX)
The same bytes as the uniform ROCmFPX baseline, holding a measurably better model.
The 43 down-projection expert tensors move from **qtype 104** (`Q3_0_ROCMFPX`, uniform) to
**qtype 105** (`Q3_1_ROCMFP3_MIX`, adaptive). Both are **14 bytes per 32 weights**, so the
file is byte-for-byte the same size, reads the same bytes per token, and costs the kernels
the same work. Nothing else in the model changes: attention, dense, gate/up, norms and
embeddings are identical to the baseline.
## What "adaptive" means here
A uniform format compiles one interpretation into the format itself — the same codebook for
every block in the model. An adaptive format moves that interpretation into the artifact:
each expert carries its own small learned table, chosen by a per-expert search against
Hessian-weighted sensitivity rather than by a rule applied to the whole tensor.
Those tables are ~375 KB against ~102 GB. That asymmetry is the whole trick — a per-expert
choice costs almost nothing to record and buys a much better fit, because expert weight
distributions genuinely differ.
## Measured quality
From the quality-transfer ledger that selected this assignment (fake-quant, against the
unquantized model, 2 seeds x 2 corpora x 5 policies):
| | uniform (P0) | adaptive |
|---|---|---|
| wikitext2 PPL | 3.8573 | **3.6760** |
| wikitext2 KL | 0.2157 | **0.0953** |
| c4 PPL | 12.5342 | **11.5245** |
| c4 KL | 0.2064 | **0.0677** |
**Aggregate: KL -57.0%, perplexity -5.6%, at identical footprint.**
KL is the number that matters most: it measures divergence from the *unquantized* model, so
it says this artifact reproduces the original's output distribution roughly twice as
faithfully as the uniform recipe does — using the same bytes.
Two honest caveats:
- **96.4% weighted expert coverage.** The calibration corpus never routed to some experts;
those fall back to the conservative uniform rung rather than being guessed at. The gains
above are earned on the covered majority.
- These are **fake-quant research measurements** against the native model, not end-to-end
serving benchmarks. Throughput and speculative-decode figures are reported separately,
under one protocol with the prompt mix stated.
## Single file — no sidecar
Earlier adaptive artifacts shipped as a model **plus a required `.p4mix.bin` companion**:
separate the two and the tensors are undecodable. The learned tables are now carried inside
the GGUF, as a verbatim byte blob in the KV block:
```
deepseek4.p4mix.sidecar 375,320 bytes
```
Loaders prefer the embedded copy and fall back to a loose file, so existing two-file
artifacts keep working unchanged.
## Requirements
Needs a server build with the qtype-105 decode kernels — upstream `lucebox` plus the
adaptive series, on branch
[`feat/ds4-adaptive-on-upstream`](https://github.com/GeometricAGI/lucebox-hub/tree/feat/ds4-adaptive-on-upstream)
(upstream main + 20 reviewable commits, full suite green). A stock build without those
kernels will not decode qtype 105.
Pair with
[DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF](https://huggingface.co/Geometric-AI/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF)
for speculative decode, and compare against
[DeepSeek-V4-Flash-0731-ROCMFPX](https://huggingface.co/Geometric-AI/DeepSeek-V4-Flash-0731-ROCMFPX),
the uniform baseline this is measured against.
## Integrity
```
sha256 a431f955c639107d3f34ab4cc65edc806038588a0c258abd0bc3e9074b908096
file ds4-0731-opt1.gguf (102,321,006,592 bytes)
```
That is **+375,392 bytes** over the uniform baseline — the embedded codebooks and their KV
overhead, and the entire size cost of the adaptive encoding.
## License
MIT, inherited from the base model.
> **Checksum correction (2026-08-05).** An earlier revision of this card listed
> `accabb4c…`, which was the hash of a pre-release build, not of the published file.
> A correct download failed that check. The value above (`a431f955…`) is the real one
> and matches the git-LFS oid and `x-linked-etag`. Thanks to @hycris for catching it.
## Serving: keep the model's six routed experts
**Do not pass `--ds4-expert-top-k 4` with this artifact.** That flag keeps only four of the
six routed experts and renormalizes. On the uniform ROCMFPX artifact it is harmless; on this
adaptive one it produces degenerate output.
Measured 2026-08-05, exact-copy fidelity over 20 identifiers x 3 repeats at temperature 0
(the model is asked to echo a line of Python already present in its prompt, so anything but a
verbatim copy is a defect):
| | top-k 6 (model default) | top-k 4 |
|---|---|---|
| this artifact (adaptive) | 95.0% | **60.0%** |
| uniform ROCMFPX | 100% | 100% |
The failures are not degraded paraphrases -- they are `" 0 0 0 0 0 ..."`, repeated markdown
fragments and empty strings, on prompts as trivial as "repeat this line". Deterministic, and
reproducible at two context sizes and with fused decode on or off.
Some published example commands (including in this project's own `server/docs/DS4.md` until
today) included `--ds4-expert-top-k 4`. Omit it, or pass `0`, to use the model default.
|