huginnfork/ThinkingCap-Qwen3.6-27B-FP8
Image-Text-to-Text • 28B • Updated • 498
FP8 quants of bottlecapai/ThinkingCap-Qwen3.6-27B. bf16 SSM/attention layout, and unlike the upstream FP8 it loads in transformers on Blackwell.
Note FP8 attnbf16 — MLPs only, the whole self_attn path stays bf16. KLD 0.0132 nats, ΔPPL +0.96 % vs the bf16 parent, and a much lighter worst-case token KLD (5.04) than the upstream FP8 (13.47). 75.0 % vLLM MTP draft acceptance.