fafstmobel
A locally built 27B derivative of the Swift abliterated Qwen3.8-27B checkpoint: Huihui's observed abliteration delta transplanted onto Swift, quantized to an NVFP4/FP8 text allocation, and exported to a single NInfer v3 artifact with the published DFlash2 draft component bundled alongside vision and MTP.
Built end to end on one workstation (RTX 5090, CUDA 13.1, Python 3.11) by a reproducible pipeline. Every number below was measured on that machine.
This revision stores the DFlash2 drafter's projections as Q4 (see Revision note). The target model is byte-identical to the previous revision.
Artifact
| File | fafstmobel.ninfer — 22,884,131,844 bytes |
| SHA-256 | 828e5dffc1e023c2032e901cf060197d5c282cd2d9f4566297aad4b8f155a16b |
| Container | NInfer v3, 1246 objects / 1240 tensors / 6 resources / 1513 bindings / 844 uses, 1 file |
| Components | text, vision, mtp, dflash2, plus the indexed proposal head |
| Text model | 64 layers, hidden 5120, vocab 248320, 16 full-attention + 48 linear-attention |
| Draft | DFlash2, 5 layers, target taps [5,19,33,47,61], block 8, selector top-k 16; projections Q4 (group 64), fused QKV Q8 (group 32) |
| Metadata name | fafstmobel |
What was changed relative to upstream
- Delta transplant onto Swift (
scripts/abliterate.py). Huihui's abliteration was not re-derived; the base→Huihui weight delta was applied to Swift:result = (swift.float() + (huihui.float() - base.float())).to(bfloat16). Exactly 70 tensors changed, in language-model layers 17–51:mlp.down_projplusself_attn.o_proj(full-attention layers) orlinear_attn.out_proj(linear-attention layers). Max |Δ| = 2.83e-2, mean |Δ| = 2.02e-4, 4,220,518,400 elements rewritten. The other 1129 tensors — all vision, MTP, embeddings, norms and the output head — are byte-identical to the Swift checkpoint. - Text quantization to the published NVFP4/FP8 allocation: 168 NVFP4 matrices (MLP gate/up/down, group size 16, 4-bit weights, local-dynamic 4-bit activations with calibrated input scales) and 233 FP8 per-channel matrices (attention q/k/v/o, GDN in/out projections, output head, MLP layers 56–63). 32 UltraChat-200k calibration samples × 2048 tokens. Vision, MTP and GDN gate projections stay unquantized.
- Export to NInfer v3 with the
tools.convertof Cinference (revision8364ed75cb5c7edb282e761cd3d151bfa3b10df8, a fork of NInfer) and itsqwen3_8_27b_nvfp4recipe, including the maintained Qwen3.8 chat template. Target weights are imported unchanged; the recipe quantizes the DFlash2 drafter's projections to Q4 and its fused QKV projection to Q8. - Naming normalization of the text intermediate (
model.language_model.*→model.*) so the checkpoint is consistent with its flattened text-only config; only safetensors headers were rewritten, verified against per-shard payload digests.
Revision note
The previous revision (54202e174c5f05945fbb873d1c2d8384e2643bd3, 23,719,715,844 bytes) was
exported with NInfer 98dada0e03cb073fd07f905400b5904bc6e82759, whose recipe stores the drafter
projections as Q8. This revision changes only the 16 drafter projection objects; every target,
vision, MTP and proposal-head object is byte-identical.
The drafter only proposes tokens and the target verifies them exactly, so outputs are unchanged.
With Cinference 8364ed7 on the RTX 5090:
| Check | Q8 drafter (previous) | Q4 drafter (this revision) |
|---|---|---|
| Greedy decoding, BF16 KV, 6 prompts × 256 tokens, DFlash2 | reference | identical token IDs on all 6 prompts |
| Accepted draft tokens per round, 32 thinking-mode coding requests, verify trees, default sampling | 4.008 | 4.016 |
| Drafter projection kernels per 8K verification round | 1,447.5 µs | 1,118.2 µs |
| Resident weights | 21.7 GiB | 20.9 GiB |
| Full Cinference profile (262,144-token K8V4 context, vision, trees) with 1.5 GB of desktop VRAM in use | does not start (275 MB short) | starts with 993 MiB free |
Measured results (this revision; RTX 5090, 32 GB; Cinference c59d0da; context 262,144; K8V4 KV; CUDA graphs on)
| Check | Result |
|---|---|
Arithmetic (17 × 19) |
323; same-seed replay identical |
| Code prompt | valid slicing def, 30 tokens |
| Cinference DFlash2 integration test (draft 15, verify trees, K8V4, vision) | exit 0 |
| Serving | /health ok, /v1/models id fafstmobel (262,144 tokens), tools.smoke.serve_contract pass |
| Multimodal | 256×256 red/blue image → "left is red, right is blue" |
| Speculation (seed 42, greedy, 256 tokens, batch 1) | autoregressive 54.1 tok/s · MTP-3 155.3 tok/s (2.87×, acceptance 74.0%) · DFlash2-15 with verify trees 293.7 tok/s (5.43×) |
| Serving profile | DFlash2-15 with verify trees and K8V4 at 262,144 tokens vs DFlash2-7 with fp8 KV at 16,384, alternating on the same prompts: median 241.4 → 294.4 tok/s (+22%) |
KV cache quality (ninfer-perplexity, 261,167 tokens, context 65,536) |
BF16 4.08915 · fp8 4.08843 · K8V4 4.08607 |
A fullscreen screensaver shared the GPU during these runs, so absolute rates are depressed;
speedups and acceptance counts reproduce exactly across repetitions. The previous revision
measured autoregressive 68.5–77.1, MTP-3 210.0–215.3 and DFlash2-7 304.1–312.9 tok/s (NInfer
98dada0, 16,384 context, fp8 KV) on a quieter desktop.
Usage
DFlash2 with this revision's Q4 drafter requires Cinference
8364ed7 or later (its installer pins a
matching runtime). Older Cinference builds and upstream NInfer load the target, MTP and vision but
reject this drafter; use revision 54202e174c5f05945fbb873d1c2d8384e2643bd3 with them.
ninfer-serve fafstmobel.ninfer --host 127.0.0.1 --port 8088 --model-id fafstmobel \
--spec dflash2 --draft-tokens 15 --lm-head-draft --verify-tree --vision \
--kv-dtype k8v4 --max-context 262144 --kv-capacity 262144 --max-concurrency 1
ninfer fafstmobel.ninfer --prompt "What is 17 multiplied by 19?" --greedy --max-new 32 \
--no-thinking --vision --kv-dtype k8v4 --max-context 16384 \
--spec dflash2 --draft-tokens 15 --lm-head-draft --verify-tree
Licensing and notices
This artifact contains three upstream contributions, each retained under its own terms:
- Swift contribution (
ukisai/Swift-Qwen3.8-27b) — Swift Open License v1.0, bundled asLICENSE.swift. That licence grants reproduction and distribution of derivative works, including converted weights, subject to its conditions: recipients receive a copy of the licence, modified files carry prominent change notices, attribution notices are retained, and Commercial Use is not licensed for a Legal Entity exceeding the one-million-USD Threshold (Section 5). A separate Swift Enterprise License is required above that threshold. This derivative is distributed under the same terms for the Swift contribution. - Base model (
Qwen/Qwen3.8-27B) — Apache-2.0, bundled asLICENSE-APACHE-2.0. Qwen3.8-27B portions remain under the Apache License 2.0. - Huihui abliteration delta (
huihui-ai/Huihui-Qwen3.8-27B-abliterated) — Apache-2.0, bundled asLICENSE.huihui. Only the observed weight difference was transplanted. - DFlash2 draft (
z-lab/Qwen3.8-27B-DFlash2) — Apache-2.0; bundled by conversion with its projection weights quantized to Q4 (fused QKV Q8), otherwise unmodified. - Cinference (https://github.com/satellitedown/cinference, Apache-2.0, a fork of NInfer, https://github.com/Neroued/ninfer) — the converter used for this revision; NInfer converted the previous revision.
NOTICE carries the upstream notice plus this derivative's transformation notice;
PROVENANCE.md records the exact repository revisions.
Claim boundaries
- This is a delta transplant, not a newly fitted refusal direction, and it is not proof that all refusals are removed. Behavioural claims about refusal removal are not made here.
- The DFlash2 draft is pretrained upstream and bundled by conversion; no drafter was trained, and Swift was not retrained.
- The quantization is a reproduction of the published allocation, not a bit-exact replication of any other derivative; no other derivative's weights were downloaded or compared.
- Not suitable for safety-critical use. Outputs are unmoderated model text.
- Downloads last month
- 109
Model tree for satellitedown/fafstmobel
Base model
Qwen/Qwen3.8-27B