PixelML/Qwen3.8-Flash-Next-KVA-Projector

LLKVApprox KVA projector (v2) for Qwen3.8-Flash-Next (qwen4_exp hybrid: 48 layers = 36 gated-delta-net + 12 full attention, 512-expert MoE, hyper-connections, lightning indexer). Port of kishida's LLKVApprox "Encoder-Decoder" prefill to the Flash-Next hybrid; follows the v1 Qwen3.8-27B playbook (PixelML issue #134/#139).

What it does

Prefill runs layers 0..23 over all prompt tokens; this projector predicts the approximated layers' (24..47) state fills from the boundary hyper-connection state โ€” 18 GDN recurrence-input sets (qkv/b/a) and 6 FA (K pre-RoPE + V + lightning-indexer raw keys); layers 24..47 then run only for the last prompt position(s); decode is exact.

Training

2000 steps, cached-target distillation (zero teacher kernels in the loop), AdamW lr 1e-3, 16 x 2048-token sequences of cached teacher targets. Final held-out cosines: FA k 0.945-0.953, GDN k 0.978-0.982. Ridge ceilings (single global linear): FA k 0.62, GDN qkv 0.95 โ€” the per-layer own-weights init + MLP heads close the FA gap.

Honest limitations

  • Quality is measured as teacher-forced/greedy token agreement vs the full baseline โ€” a diagnostic, not task accuracy.
  • On this 512-expert model, bf16 GEMM differences between the CED suffix path (M=1) and the full prefill (M=T) flip near-tie router decisions; divergence is bounded but non-zero by construction (see issue #139 gate note).
  • v1's suffix ladder (exact replay of the last N tokens) remains available as the quality knob.

Usage

Requires the qwen4_exp modeling code (transformers >= 5.18-dev) and the LLKVApprox engine (shipping in the club-170hx PR). Load projector.safetensors into KVAProjector(model); the frozen w0 buffers are included so no weight surgery is needed at load time.

Downloads last month
92
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for PixelML/Qwen3.8-Flash-Next-KVA-Projector

Finetuned
(70)
this model