Qwen3.8-Flash-Next-NVFP4 β€” see instead

I opened this name intending to publish a plain NVFP4 quantization of the base Qwen3.8-Flash-Next. I am not going to, because there is no reason to ship a second one.

If you want base Flash-Next in NVFP4, use RadixArk/Qwen3.8-Flash-Next-NVFP4. It is a complete NVFP4 (W4A4) build of the base model β€” its own card lists SGLang support and marks it a candidate release. Duplicating a base-model NVFP4 here would add nothing.

What I actually made is the thing that did not exist: a refusal-removed Flash-Next, and an NVFP4 build of that.

  • Qwen3.8-Flash-Next-Abliterated β€” the bf16 abliterated model. AdvBench refusal 99.42 % β†’ 0.96 %, MMLU βˆ’1.6 pp (paired), GSM8K unchanged, byte-level verified. Ships ARCHITECTURE.md, the first written account of the qwen4_exp internals.
  • Qwen3.8-Flash-Next-Abliterated-NVFP4 β€” that model in NVFP4. W4A16 rather than W4A4 on purpose: abliteration moves the activation distribution, so a static-activation-scale W4A4 build cannot reuse calibration and W4A16 sidesteps the question entirely. Vision tower and MTP head grafted back and verified present.
  • Qwen3.8-Flash-Next-Abliterated-GGUF β€” for llama.cpp (against PR #27742). All five quants are live β€” Q4_K_M / Q5_K_M / Q6_K / Q8_0 / BF16 β€” each sha256-verified against its local build (5/5, 0 problems).

This repo is kept only as a signpost. Nothing will be published here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for windowsxp811203/Qwen3.8-Flash-Next-NVFP4

Finetuned
(20)
this model