Running Qwen3.8-Flash-Next NVFP4 on 2x RTX A6000 (Ampere) — GPTQ-requantized pack, Linux, full numbers
I built the best-accuracy Flash-Next quant I could on 94 GB RAM / 2x A6000 48 GB and benchmarked it properly. First NVFP4-fork numbers I've seen on Ampere (sm_86, no FP4 tensor cores), and the first with a user-requantized GPTQ pack. Full reproducible pipeline + pack hashes in the files tab.
What's running
- Engine: sergqwer/strata-nvfp4 fork @ 84fe9ed (0.1.39-nvfp4.2), source build, Linux, CUDA 13.2, sm_86. (Pitfall: system nvcc must be >= 12.4 — Ubuntu 24.04's 12.0 fails the VMM shim. CUDA 13.2 builds clean for Ampere.)
- Pack: NVIDIA NVFP4 checkpoint -> nvfp4_convert.py (verified bit-for-bit) -> requant.py GPTQ re-quantization calibrated on this box (4 prompts, all 48 layers), with Q8_0 down-projections in the 27 worst layers per requant_plan.py at a 74 GiB budget. Measured expert error: 33.5% of plain rounding with GPTQ alone, 15.7% with the Q8_0-down plan.
- Aux tensors at checkpoint precision: PLE n-gram table FP8, embedding BF16, MTP draft head q2_0, int8 KV (Hadamard-rotated), float64 RoPE table.
- Flags: --max-context 262144 --kv int8 --kv-resident 32768 --layer-split 16 --spec 4 --spec-min-p 0.60 --expert-cache auto --prefill auto
Numbers (cold engine, median [min-max] of 3, salted cold prompts, temp 0)
| Prompt | Prefill tok/s | Decode tok/s | Needle |
|---|---|---|---|
| 16K | 2,328 [2,322-362] | 99.0 [89.7-104.2] | 3/3 |
| 64K | 2,839 [2,791-876] | 102.3 [93.8-110.3] | 3/3 |
| 128K | 2,971 [2,952-010] | 94.4 [67.0-104.3] | 3/3 |
| 250K | 2,897 [2,841-928] | 94.2 [94.0-97.4] | 3/3 |
12/12 recall. First 128K touch on an empty expert cache costs speed (67.0), then 94-104 every time after.
The finding worth knowing: agent sessions eat your long-context decode
Re-ran the identical matrix with one live ~180K-token Hermes agent session sharing the engine: 250K decode dropped to 75.8 [71.7-95.3] — ~20% median hit, 3x the spread — while 16K-128K barely moved. The resident session's KV competes with the request's KV streaming ring, and it only bites at 250K. If your long-context numbers look noisy, measure what else is resident.
Tuning A/B (measured, zero accuracy cost — verified decoding)
--spec-min-p 0.70 -> 0.60: suffix-draft acceptance ~50% -> 74%, real-turn decode 92.9 -> 104-112 tok/s.
Why bother vs stock IQ3_S
IQ3_S was faster (104-112 decode) but the GPTQ+Q8_0 pack is in a different accuracy class: expert error 15.7% of plain rounding, ~half the KL of stock NVIDIA NVFP4, with every auxiliary tensor at checkpoint precision. The ~10% decode cost buys the most accurate Flash-Next representation this hardware can hold.
Hardware: 2x A6000 48 GB (PCIe Gen4 x16, NVLink present/unused), Ryzen 9 7900X (AVX-512+VNNI), 94 GB RAM, NVMe. RAM fully committed: the 73.83 GiB expert arena is pinned resident; no SSD streaming during decode.
Pack SHA-256 (experts.bin 73.83 GiB): ed5c0ff9763442fee857737892d12551eeaa11a5b5e811b3be30a885b13a5900 Full hashes + method + per-run JSON in the files.
Credits
None of this is my architecture — I ran the pipeline and measured the result.
- sergqwer/strata-nvfp4 — the NVFP4 fork of Strata; its kernels are what make NVFP4 run on Ampere without FP4 tensor cores. docs/NVFP4.md was the build reference. https://github.com/sergqwer/strata-nvfp4
- Niko1221/Strata — upstream engine: expert cache, layer-split pipeline, the bench harness used here. https://github.com/Niko1221/Strata
- NVIDIA — the NVFP4 checkpoint and ModelOpt recipe. https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4
- Qwen — the base model. https://huggingface.co/Qwen/Qwen3.8-Flash-Next
If this report is useful, the credit flows upstream to those four.
Model tree for Lafonse/strata-nvfp4-a6000-bench
Base model
Qwen/Qwen3.8-Flash-Next