fafstmobel

A locally built 27B derivative of the Swift abliterated Qwen3.8-27B checkpoint: Huihui's observed abliteration delta transplanted onto Swift, quantized to an NVFP4/FP8 text allocation, and exported to a single NInfer v3 artifact with the published DFlash2 draft component bundled alongside vision and MTP.

Built end to end on one workstation (RTX 5090, CUDA 13.1, Python 3.11) by a reproducible pipeline. Every number below was measured on that machine.

This revision stores the DFlash2 drafter's projections as Q4 (see Revision note). The target model is byte-identical to the previous revision.

Artifact

File fafstmobel.ninfer — 22,884,131,844 bytes
SHA-256 828e5dffc1e023c2032e901cf060197d5c282cd2d9f4566297aad4b8f155a16b
Container NInfer v3, 1246 objects / 1240 tensors / 6 resources / 1513 bindings / 844 uses, 1 file
Components text, vision, mtp, dflash2, plus the indexed proposal head
Text model 64 layers, hidden 5120, vocab 248320, 16 full-attention + 48 linear-attention
Draft DFlash2, 5 layers, target taps [5,19,33,47,61], block 8, selector top-k 16; projections Q4 (group 64), fused QKV Q8 (group 32)
Metadata name fafstmobel

What was changed relative to upstream

  1. Delta transplant onto Swift (scripts/abliterate.py). Huihui's abliteration was not re-derived; the base→Huihui weight delta was applied to Swift: result = (swift.float() + (huihui.float() - base.float())).to(bfloat16). Exactly 70 tensors changed, in language-model layers 17–51: mlp.down_proj plus self_attn.o_proj (full-attention layers) or linear_attn.out_proj (linear-attention layers). Max |Δ| = 2.83e-2, mean |Δ| = 2.02e-4, 4,220,518,400 elements rewritten. The other 1129 tensors — all vision, MTP, embeddings, norms and the output head — are byte-identical to the Swift checkpoint.
  2. Text quantization to the published NVFP4/FP8 allocation: 168 NVFP4 matrices (MLP gate/up/down, group size 16, 4-bit weights, local-dynamic 4-bit activations with calibrated input scales) and 233 FP8 per-channel matrices (attention q/k/v/o, GDN in/out projections, output head, MLP layers 56–63). 32 UltraChat-200k calibration samples × 2048 tokens. Vision, MTP and GDN gate projections stay unquantized.
  3. Export to NInfer v3 with the tools.convert of Cinference (revision 8364ed75cb5c7edb282e761cd3d151bfa3b10df8, a fork of NInfer) and its qwen3_8_27b_nvfp4 recipe, including the maintained Qwen3.8 chat template. Target weights are imported unchanged; the recipe quantizes the DFlash2 drafter's projections to Q4 and its fused QKV projection to Q8.
  4. Naming normalization of the text intermediate (model.language_model.* → model.*) so the checkpoint is consistent with its flattened text-only config; only safetensors headers were rewritten, verified against per-shard payload digests.

Revision note

The previous revision (54202e174c5f05945fbb873d1c2d8384e2643bd3, 23,719,715,844 bytes) was exported with NInfer 98dada0e03cb073fd07f905400b5904bc6e82759, whose recipe stores the drafter projections as Q8. This revision changes only the 16 drafter projection objects; every target, vision, MTP and proposal-head object is byte-identical.

The drafter only proposes tokens and the target verifies them exactly, so outputs are unchanged. With Cinference 8364ed7 on the RTX 5090:

Check Q8 drafter (previous) Q4 drafter (this revision)
Greedy decoding, BF16 KV, 6 prompts × 256 tokens, DFlash2 reference identical token IDs on all 6 prompts
Accepted draft tokens per round, 32 thinking-mode coding requests, verify trees, default sampling 4.008 4.016
Drafter projection kernels per 8K verification round 1,447.5 µs 1,118.2 µs
Resident weights 21.7 GiB 20.9 GiB
Full Cinference profile (262,144-token K8V4 context, vision, trees) with 1.5 GB of desktop VRAM in use does not start (275 MB short) starts with 993 MiB free

Measured results (this revision; RTX 5090, 32 GB; Cinference c59d0da; context 262,144; K8V4 KV; CUDA graphs on)

Check Result
Arithmetic (17 × 19) 323; same-seed replay identical
Code prompt valid slicing def, 30 tokens
Cinference DFlash2 integration test (draft 15, verify trees, K8V4, vision) exit 0
Serving /health ok, /v1/models id fafstmobel (262,144 tokens), tools.smoke.serve_contract pass
Multimodal 256×256 red/blue image → "left is red, right is blue"
Speculation (seed 42, greedy, 256 tokens, batch 1) autoregressive 54.1 tok/s · MTP-3 155.3 tok/s (2.87×, acceptance 74.0%) · DFlash2-15 with verify trees 293.7 tok/s (5.43×)
Serving profile DFlash2-15 with verify trees and K8V4 at 262,144 tokens vs DFlash2-7 with fp8 KV at 16,384, alternating on the same prompts: median 241.4 → 294.4 tok/s (+22%)
KV cache quality (ninfer-perplexity, 261,167 tokens, context 65,536) BF16 4.08915 · fp8 4.08843 · K8V4 4.08607

A fullscreen screensaver shared the GPU during these runs, so absolute rates are depressed; speedups and acceptance counts reproduce exactly across repetitions. The previous revision measured autoregressive 68.5–77.1, MTP-3 210.0–215.3 and DFlash2-7 304.1–312.9 tok/s (NInfer 98dada0, 16,384 context, fp8 KV) on a quieter desktop.

Usage

DFlash2 with this revision's Q4 drafter requires Cinference 8364ed7 or later (its installer pins a matching runtime). Older Cinference builds and upstream NInfer load the target, MTP and vision but reject this drafter; use revision 54202e174c5f05945fbb873d1c2d8384e2643bd3 with them.

ninfer-serve fafstmobel.ninfer --host 127.0.0.1 --port 8088 --model-id fafstmobel \
  --spec dflash2 --draft-tokens 15 --lm-head-draft --verify-tree --vision \
  --kv-dtype k8v4 --max-context 262144 --kv-capacity 262144 --max-concurrency 1
ninfer fafstmobel.ninfer --prompt "What is 17 multiplied by 19?" --greedy --max-new 32 \
  --no-thinking --vision --kv-dtype k8v4 --max-context 16384 \
  --spec dflash2 --draft-tokens 15 --lm-head-draft --verify-tree

Licensing and notices

This artifact contains three upstream contributions, each retained under its own terms:

  • Swift contribution (ukisai/Swift-Qwen3.8-27b) — Swift Open License v1.0, bundled as LICENSE.swift. That licence grants reproduction and distribution of derivative works, including converted weights, subject to its conditions: recipients receive a copy of the licence, modified files carry prominent change notices, attribution notices are retained, and Commercial Use is not licensed for a Legal Entity exceeding the one-million-USD Threshold (Section 5). A separate Swift Enterprise License is required above that threshold. This derivative is distributed under the same terms for the Swift contribution.
  • Base model (Qwen/Qwen3.8-27B) — Apache-2.0, bundled as LICENSE-APACHE-2.0. Qwen3.8-27B portions remain under the Apache License 2.0.
  • Huihui abliteration delta (huihui-ai/Huihui-Qwen3.8-27B-abliterated) — Apache-2.0, bundled as LICENSE.huihui. Only the observed weight difference was transplanted.
  • DFlash2 draft (z-lab/Qwen3.8-27B-DFlash2) — Apache-2.0; bundled by conversion with its projection weights quantized to Q4 (fused QKV Q8), otherwise unmodified.
  • Cinference (https://github.com/satellitedown/cinference, Apache-2.0, a fork of NInfer, https://github.com/Neroued/ninfer) — the converter used for this revision; NInfer converted the previous revision.

NOTICE carries the upstream notice plus this derivative's transformation notice; PROVENANCE.md records the exact repository revisions.

Claim boundaries

  • This is a delta transplant, not a newly fitted refusal direction, and it is not proof that all refusals are removed. Behavioural claims about refusal removal are not made here.
  • The DFlash2 draft is pretrained upstream and bundled by conversion; no drafter was trained, and Swift was not retrained.
  • The quantization is a reproduction of the published allocation, not a bit-exact replication of any other derivative; no other derivative's weights were downloaded or compared.
  • Not suitable for safety-critical use. Outputs are unmoderated model text.
Downloads last month
109
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for satellitedown/fafstmobel

Base model

Qwen/Qwen3.8-27B
Finetuned
(450)
this model