File size: 12,786 Bytes
f5372a2 ecf0dc8 f5372a2 c69b611 f5372a2 c69b611 f5372a2 05ad718 f5372a2 3a66e59 05ad718 f5372a2 e5bb100 f5372a2 52ac4aa e6a53c8 f5372a2 05ad718 df241f8 05ad718 df241f8 05ad718 f5372a2 05ad718 f5372a2 05ad718 e6a53c8 f5372a2 05ad718 f5372a2 b82c4fd 34b2fc2 05ad718 df241f8 05ad718 df241f8 5623cda 05ad718 e6a53c8 4fdd2fd e6a53c8 3158193 e6a53c8 b58acbf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 | ---
license: apache-2.0
base_model: Nanbeige/Nanbeige4.2-3B
base_model_relation: quantized
pipeline_tag: text-generation
language: [en, zh]
tags: [fp8, nvfp4, quantized, compressed-tensors, vllm, looped-transformer, nanbeige]
model-index:
- name: Nanbeige4.2-3B-NVFP4-FP8-LoopShield
results:
- task:
type: text-generation
dataset:
name: GSM8K
type: gsm8k
metrics:
- name: GSM8K (strict, thinking, n=100)
type: exact_match
value: 89.0
verified: false
- task:
type: text-generation
dataset:
name: GSM8K
type: gsm8k
metrics:
- name: GSM8K (flexible)
type: exact_match
value: 96.0
verified: false
- task:
type: text-generation
dataset:
name: IFEval
type: ifeval
metrics:
- name: IFEval prompt-strict (non-thinking, n=250)
type: exact_match
value: 76.0
verified: false
- task:
type: text-generation
dataset:
name: MMLU-Pro
type: mmlu_pro
metrics:
- name: MMLU-Pro (non-thinking, 25/category)
type: exact_match
value: 64.9
verified: false
- task:
type: text-generation
dataset:
name: BBH
type: bbh
metrics:
- name: BBH CoT few-shot (non-thinking)
type: exact_match
value: 59.5
verified: false
- task:
type: text-generation
dataset:
name: MultiHop-RAG
type: multihop_rag
metrics:
- name: MultiHop-RAG generator-only, gold evidence (n=248)
type: exact_match
value: 74.6
verified: false
---
# Nanbeige4.2-3B-NVFP4-FP8-LoopShield
Mixed-precision quantization of
[Nanbeige/Nanbeige4.2-3B](https://huggingface.co/Nanbeige/Nanbeige4.2-3B) (looped
transformer, 22 layers × 2 passes/token): **FP8-dynamic on attention, `down_proj`, and
the first/last 3 layers' MLP; NVFP4 (weight-only, group-16) on middle-layer
`gate/up_proj`**. 4.5 GB. `format: mixed-precision` (compressed-tensors).
## The nitty gritty
Naive uniform NVFP4 on this looped architecture loses **8 points of GSM8K strict**
(quantization error compounds across the two loop passes — consistent with LoopQ,
arXiv:2605.16343, the only prior looped-LLM PTQ study, which tested INT only; these
are, as far as I can tell, the first FP4-family numbers for a looped LLM). The
protection recipe that llama.cpp's quant mixes, Unsloth's ablations, and
llmcompressor's own non-uniform example all converge on recovered it completely:
| step | GSM8K strict (n=100) |
|---|---|
| uniform NVFP4A16 | 81 |
| + FP8 attention & down_proj (all layers) | 84 |
| **+ FP8 gate/up in layers 0-2 & 19-21 (this repo)** | **89 — bf16/FP8 parity** |
The edge-band ratio follows APEX's ablation (~12.5% of depth per side); the tensor
priority (down_proj > attention > gate/up) matches llama.cpp's quant mixes, Unsloth's
sensitivity ablations, and llmcompressor's own non-uniform example.
**Super-weight verification (2026-07-23)**: this model's super weight (the
single most load-bearing scalar, [Apple 2411.07191](https://arxiv.org/abs/2411.07191))
sits at `layers.1.mlp.down_proj.weight[1252, 6883]` (largest weight in its tensor,
11.4x p99.99; drives a 26,752-magnitude activation spike, 1,300x the median). This
recipe protects both the weight (all-layer FP8 down_proj) and its production path
(L1 gate/up in the FP8 edge band) — verified by direct scan, not assumed. One
looped-arch novelty from the scan: the spike is pass-asymmetric (26,752 on loop
pass 1 of 2; 1,352 on pass 2).
## Quality evals (RTX 5090, vLLM v0.25.1, 64k ctx, fp8 KV; champion = my FP8-Dynamic quant)
| bench | mode | bf16 original | FP8-Dynamic | **this repo** |
|---|---|---|---|---|
| GSM8K strict (n=100) | thinking | not measured | 89 | **89** |
| GSM8K flexible | thinking | not measured | 93 | **96** |
| IFEval prompt-strict (n=250) | non-thinking | not measured | 76.4 | 76.0 |
| IFEval inst-strict | non-thinking | not measured | 83.5 | 83.3 |
| MMLU-Pro (25/category) | non-thinking | not measured | 62.6 | **64.9** |
| BBH CoT few-shot | non-thinking | not measured | 64.9 | 59.5 |
| MultiHop-RAG (gold evidence, n=248) | non-thinking | not measured | 74.2 | **74.6** |
| Blind-judge summarization (closed, same-judge pair, n=46) | non-thinking | n/c (judged in a separate pass; scores only comparable within a pass) | 4.57 | 4.41 (coverage −0.24, ~1.5σ) |
| Judged faithfulness (closed) | non-thinking | n/c | 4.90 | 4.87 |
| Judged fabrication / leaks (closed) | non-thinking | 0% / 0% | 2% / 0% | 2% / 0% |
| Dictation-rewrite taxonomy (closed) | non-thinking | 18/20 | 18/20 | 17/20 |
| JSON parse rate (closed, /48) | non-thinking | 44 | 42 | **46** |
The trade: **reasoning at full parity, best-in-family structured-output
reliability, small summarization-coverage cost.** If you want maximum quality use my
FP8-Dynamic; if you want the smallest artifact that keeps reasoning intact on this
architecture, use this one.
## Speed (RTX 5090, batch-1, vLLM v0.25.1, 64k ctx, fp8 KV)
Decode tok/s, single stream, per-workload best speculative config:
| workload | spec decode | bf16 | FP8-Dynamic | **this repo** |
|---|---|---|---|---|
| freeform / chat / agent | off | 96 | 154 | **159** |
| summarize / RAG (2k+ ctx prompts) | ngram, 8 tok | 136 | 206 | 202 |
Two workload anchors, not an ISL sweep; decode speed shifts with context
length, batch size, and attention backend. The mid-MLP NVFP4 weights serve
via the Marlin W4A16 kernel; the FP8 tensors via the FP8 path; the small
freeform edge over FP8-Dynamic comes from the lighter weight reads, the small
summarize deficit from mixed-kernel overhead under the ngram verify batch.
## Serving (vLLM)
Arch not yet upstream ([vLLM PR #49433](https://github.com/vllm-project/vllm/pull/49433));
install the bundled plugin first:
```bash
pip install --no-deps ./vllm_plugin
vllm serve <this-repo> --trust-remote-code \
--max-model-len 65536 --kv-cache-dtype fp8 \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml
```
Thinking/sampling notes as in the base model: thinking ON by default
(`chat_template_kwargs.enable_thinking=false` to disable), T=0.6/top_p=.95/top_k=20
defaults ship in `generation_config.json`, T=1.0 for agentic use.
Download just this artifact:
```bash
hf download NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield --local-dir Nanbeige4.2-3B-NVFP4-FP8-LoopShield
```
## Creation (llmcompressor, data-free)
```python
g_fp8 = dict(FP8_DYNAMIC)
g_fp8["targets"] = ["re:.*self_attn\\.q_proj.*", "re:.*self_attn\\.k_proj.*",
"re:.*self_attn\\.v_proj.*", "re:.*self_attn\\.o_proj.*",
"re:.*down_proj.*"] + \
[f"re:model\\.layers\\.{i}\\.mlp\\.(gate|up)_proj.*" for i in (0,1,2,19,20,21)]
g_fp4 = dict(NVFP4A16)
g_fp4["targets"] = [f"re:model\\.layers\\.{i}\\.mlp\\.(gate|up)_proj.*" for i in range(3,19)]
QuantizationModifier(config_groups={"group_0": g_fp8, "group_1": g_fp4}, ignore=["lm_head"])
```
GPTQ/AWQ were NOT usable on this architecture in llmcompressor 0.12: GPTQ's Hessian
inversion fails on all 154 modules (root cause unknown — the double-fire
hook was checked and is not the cause); AWQ lacks arch mappings. Reported upstream:
[#2952](https://github.com/vllm-project/llm-compressor/issues/2952), [#2953](https://github.com/vllm-project/llm-compressor/issues/2953).
## Limitations
- vLLM-only until PR #49433 merges. Bundled `modeling_nanbeige.py` carries two
one-line transformers-5 compat patches (rope key, tied-weights type).
- W4A4 config-variant possible (same bytes) but NOT recommended: batch-1 decode is
faster via the Marlin W4A16 path on consumer Blackwell, and activation quantization
is the looped arch's worst failure mode.
- English-only evals; n=100/46 — directional, not leaderboard-grade.
## Which artifact should I choose?



| artifact | size | tok/s (freeform / summ) | pick when |
|---|---|---|---|
| [FP8-Dynamic](https://huggingface.co/NullSense/Nanbeige4.2-3B-FP8-Dynamic) | 4.9 GB | 154 / 206 | default: no measured quality loss on any gate. Recommended. |
| [NVFP4-FP8-LoopShield](https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield) (this repo) | 4.5 GB | 159 / 202 | smallest artifact that keeps reasoning at FP8 parity; best JSON reliability. Recommended for tight VRAM. |
| [NVFP4A16](https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4A16) | 3.6 GB | 175 / 220 | fastest; summarization/extraction only (reasoning drops 8 GSM8K points). Not for math/agentic. |
| [EAGLE3 draft](https://huggingface.co/NullSense/Nanbeige4.2-3B-EAGLE3) | +1.5 GB | +12-41% decode | add-on speculator for any of the above; thinking-aware retrain (2026-07-24), thinking-mode acceptance 0.41, lossless. Serve with TRITON_ATTN. |
All three serve identically (same plugin, same flags); only the checkpoint differs.
Comparison chain: the columns here use my FP8-Dynamic quant as reference; the FP8 card carries the [bf16-original matrix](https://huggingface.co/NullSense/Nanbeige4.2-3B-FP8-Dynamic) linking the chain back to the unquantized model.
## Links & provenance
- Base model: [Nanbeige/Nanbeige4.2-3B](https://huggingface.co/Nanbeige/Nanbeige4.2-3B)
- Family: [FP8-Dynamic](https://huggingface.co/NullSense/Nanbeige4.2-3B-FP8-Dynamic) · [NVFP4-FP8-LoopShield](https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield) · [NVFP4A16](https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4A16) · [EAGLE3 draft](https://huggingface.co/NullSense/Nanbeige4.2-3B-EAGLE3)
- Arch serving support: [vLLM PR #49433](https://github.com/vllm-project/vllm/pull/49433) (until merged, the bundled `vllm_plugin/` registers it out-of-tree)
- Upstream tooling reports I filed from this work: [llm-compressor#2952](https://github.com/vllm-project/llm-compressor/issues/2952) (GPTQ Hessian inversion fails on all modules of this looped arch — root cause unknown; silent RTN fallback) · [llm-compressor#2953](https://github.com/vllm-project/llm-compressor/issues/2953) (AWQ lacks arch mappings)
- Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) 0.12.0 (compressed-tensors format)
## Benchmark provenance (ordered: public-harness first, then my closed harnesses)
**Public, reproducible** (lm-eval-harness `local-chat-completions`, exact configs in each row's annotation):
1. [GSM8K](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/gsm8k) — grade-school math, the reasoning gate (thinking mode, n=100, max_tokens 8192)
2. [IFEval](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/ifeval) — verifiable instruction following (non-thinking, n=250)
3. [MMLU-Pro](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/mmlu_pro) — 10-choice knowledge/reasoning (non-thinking, 25/category)
4. [BBH](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/bbh) — hard reasoning suite, CoT few-shot (non-thinking, 15/subtask)
5. [MultiHop-RAG](https://huggingface.co/datasets/yixuantt/MultiHopRAG) — multi-doc news QA; I run generator-only with gold evidence (custom harness, dataset public)
**Closed/personal harnesses** (not publicly reproducible — my own serving-workload gates; treat as relative signals between artifacts in THIS family, not cross-model scores):
- **Blind-judge summarization** — 48 stratified real articles, per-article anonymized+shuffled candidates, single LLM judge scoring faithfulness/coverage/fabrication. Tests: does the quant change long-form grounded generation quality?
- **Rewrite taxonomy** — 20 dictation-cleanup cases from a production ASR pipeline. Tests: instruction-constrained short-form editing.
- **JSON parse rate** — structured-output emission over the summarization set. Tests: format discipline under quantization.
## Citation
```bibtex
@misc{peciukonis2026nanbeige42loopshield,
author = {Pe{\v{c}}iukonis, Matas (NullSense)},
title = {NVFP4-FP8-LoopShield: loop-aware mixed-precision NVFP4 quantization of Nanbeige4.2-3B},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield},
note = {Placement recipe recovering GSM8K 81->89 on a looped/weight-shared LLM; first FP4-family results for the architecture class.}
}
```
|