Qwen3.8-27B MXFP4 + DFlash2 β€” radiance container

amd/Qwen3.8-27B-Quark-AWQ-MXFP4 β€” Qwen3.8-27B with AWQ-calibrated OCP-MXFP4 weights from AMD Quark β€” as a single .rad container for the radiance inference engine (AMD RDNA4, ROCm), with its vision tower and the z-lab/Qwen3.8-27B-DFlash2 block-diffusion drafter merged in for speculative decoding.

Needs a radiance build newer than 1.1.1: the MXFP4 GEMMs this file is served with arrive in the release after it, and 1.1.1 cannot serve it.

File qwen3.8-27b-mxfp4.rad β€” 18.80 GiB
Weights AMD's MXFP4 trunk exactly as released β€” e2m1 codes, an E8M0 exponent per 32 β€” served at 4-bit weights against E4M3 activations; the lm_head made block FP8 by the recipe
Speculator DFlash2, FP8 per 128Γ—128 block from its bf16 release; its vocabulary head as 2-bit codes. Drafts are verified, so they change the speed and never the output
Vision the 27-block vision tower, bf16: images and video in chat requests
Context 262,144 tokens trained; 200K tested

Speed

2Γ— Radeon AI PRO R9700 (gfx1201), --tp 2, greedy, against the FP8 Qwen3.8-27B container with the same drafter:

this file FP8 27B
one stream, prose (tok/s) 108 78
one stream, code (tok/s) 251 192
four streams, each (tok/s) 79–96 63–71
prefill, 23.6K-token prompt (tok/s) 4,439 4,148

AMD reports the quality of these weights on its model card (GSM8K 5-shot, 99–102% of bf16); this file serves the same weights.

Serve

radiance --model qwen3.8-27b-mxfp4.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \
    --max-num-seqs 8 --host 0.0.0.0 --port 8000

The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and structured output, and image_url / video_url parts in chat messages. The drafter's depth is chosen automatically (--num-speculative-tokens N states one, 0 turns speculation off).

How this file was made

AMD's release stores each quantized linear as packed e2m1 codes and an E8M0 exponent plane. quark_repack (quark_repack.cpp in this repository) rewrites every such pair as the bf16 values it encodes β€” exactly, since every e2m1 code times a power of two is a bf16 value β€” and states the checkpoint as a block-FP8-family one. The recipe then re-derives AMD's MXFP4 codes from those values with the OCP shared-exponent rule; the converter reports zero error on all 288 trunk weights.

g++ -std=c++20 -O2 -I<nlohmann/json include dir> quark_repack.cpp -o quark_repack
quark_repack amd/Qwen3.8-27B-Quark-AWQ-MXFP4 Qwen3.8-27B-Quark-AWQ-MXFP4-repack
rad-convert Qwen3.8-27B-Quark-AWQ-MXFP4-repack --draft-model z-lab/Qwen3.8-27B-DFlash2 \
    --recipe q38-27b-quark-mxfp4-df2.recipe -o qwen3.8-27b-mxfp4.rad

The recipe (q38-27b-quark-mxfp4-df2.recipe in this repository); everything it does not name is the checkpoint's own:

blk.*.attn_qg.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.attn_k.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.attn_v.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.attn_output.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.ssm_inz.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.ssm_out.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.ffn_gate_up.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
blk.*.ffn_down.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
dflash.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16

The delta net's a/b projection is declared bf16 and keeps AMD's values that way.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for StillDeadcode/qwen3.8-27b-mxfp4

Base model

Qwen/Qwen3.8-27B
Finetuned
(1)
this model