Qwythos-9B-v2-NVFP4 / README.md
artificial-citizen's picture
card: quality-gate PASS table + honest MTP-0%-accept status
b615a01 verified
|
Raw
History Blame Contribute Delete
4.84 kB
---
license: apache-2.0
base_model: empero-ai/Qwythos-9B-v2
pipeline_tag: text-generation
library_name: vllm
tags:
- nvfp4
- w4a4
- compressed-tensors
- vllm
- blackwell
- qwen3.5
- quantized
- mtp
language:
- en
---
# Qwythos-9B-v2-NVFP4
**NVFP4 (W4A4) build of [`empero-ai/Qwythos-9B-v2`](https://huggingface.co/empero-ai/Qwythos-9B-v2) for vLLM on NVIDIA Blackwell** β€” the native 4-bit serving path nobody ships for this checkpoint. Empero already publishes the MTP-GGUF for llama.cpp; this is the piece that was missing: a compressed-tensors NVFP4 artifact that runs the model's GEMMs directly on the Blackwell FP4 tensor cores under vLLM.
Quantized by [protoLabs](https://huggingface.co/protoLabsAI). Base model, its capabilities, and its uncensored research posture are Empero AI's β€” see [their card](https://huggingface.co/empero-ai/Qwythos-9B-v2).
## What this is
- **NVFP4 W4A4** on the 128 transformer linears (attention + MLP). The **hybrid DeltaNet (`linear_attn`) layers, the vision tower, `lm_head`, and the MTP head are kept BF16** β€” DeltaNet corrupts under 4-bit, and the vision path stays lossless.
- **11.2 GB** on disk (vs 19.3 GB BF16 source; the BF16-preserved vision tower is most of the remainder).
- **MTP sidecar included** (`model-mtp.safetensors`) β€” see the status note below.
- Inherits the base wholesale: **1M-token context (YaRN), Qwen3.5 multimodal stack, FTPO looping-fix (greedy loop rate 6.7%β†’0%), uncensored** for research/red-team/bio-chem-clinical work.
## Speed
Single RTX PRO 6000 Blackwell (sm120), vLLM 0.24.0, NVFP4 `--linear-backend marlin`, BF16 KV:
| Regime | tok/s | TPOT |
|---|---|---|
| single-stream (real text) | **138** | β€” |
| chat 1k/1k Β· C1 | 152 | 6.5 ms |
| chat 1k/1k Β· C8 | 1051 | 7.1 ms |
| chat 1k/1k Β· C32 | **2523** | 11.4 ms |
*(Client-side `vllm bench serve`, random dataset, fixed seed. `--ignore-eos`. Single trial.)*
## Serving (vLLM β‰₯ 0.24.0, Blackwell / sm120)
```bash
vllm serve protoLabsAI/Qwythos-9B-v2-NVFP4 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--linear-backend marlin \
--max-model-len 65536 \
--trust-remote-code
```
sm120 requires the FlashInfer/CUDA-13 env (or you hit the "no CUDA arch for major 12" crash):
`VLLM_USE_FLASHINFER_SAMPLER=0`, `CUDA_HOME=<cu13>`, `FLASHINFER_CUDA_ARCH_LIST=12.0f`,
`NVCC_APPEND_FLAGS=-DCCCL_DISABLE_CTK_COMPATIBILITY_CHECK`. `--linear-backend marlin` is the
proven NVFP4 kernel on sm120 (fastest at the single-stream / low-batch regime that matters here).
## Recipe
`llm-compressor` NVFP4, calibrated (512 samples, seq 2048). Ignore list: `lm_head`, `re:.*visual.*`,
`re:.*linear_attn.*`, `re:.*mtp.*`. VL keys canonicalized post-quant. Reproducible from the BF16 source.
## Quality gate β€” NVFP4 vs BF16 base (paired)
Every suite run against **both** the NVFP4 build and the BF16 source, same prompts/harness, on this rig.
Deterministic suites are judge-free (solver-verified reasoning, execution-graded code); claw uses an
independent LLM grader. Single trial (coding Γ—3).
| Suite | NVFP4 | BF16 base | Ξ” |
|---|---|---|---|
| Function-calling | 89% (48/54) | 94% (51/54) | βˆ’5 pp |
| Reasoning-v2 (solver-verified) | 0.759 | 0.789 | βˆ’0.03 |
| Coding (exec-graded, Γ—3) | 0.518 | 0.553 | βˆ’0.03 |
| Claw (agentic, 10 tasks) | **0.651** | 0.614 | +0.04 |
| **Coherence @ depth 4K–60K** | **clean** | β€” | needle βœ“, no repetition/degradation |
**Near-parity.** The only real cost is a few function-calling cases (single-trial, 54-task suite); reasoning
and coding are within noise, agentic is flat-to-better, and there is **no quant-rot at long context** (verified
to 60K). Expected NVFP4 behaviour for a 9B.
## Status (2026-07-11)
- βœ… Serves + coherent output on vLLM 0.24.0/sm120; **quality gate passed** (table above).
- ⚠️ **MTP sidecar is included but does not accelerate vLLM decoding for this checkpoint** (0% draft
acceptance). We isolated the cause: the sidecar is byte-identical in layout to a head that drafts at ~67%
on the *same* NVFP4+MTP stack for another Qwen3.5-9B, so this is **the head's weights, not our packaging or
the runtime** β€” the restored MTP head doesn't predict this FTPO-tuned checkpoint's distribution. The sidecar
ships for completeness/llama.cpp; for speculative decoding today use Empero's
[GGUF-MTP](https://huggingface.co/empero-ai/Qwythos-9B-v2-GGUF). A distilled-head fix may follow.
## Requests
Need a different size or format (GGUF, other quant, longer-context config)? Open a **Community discussion** β€”
we usually ship within 48h.
## License
apache-2.0, inherited from the base. **Uncensored** β€” this is Empero's research/red-team model; the quant
changes numerics, not alignment. Use responsibly and within the base model's terms.