rdtand/Qwen3.6-27B-prismaquant-gridbook-5.5bit-vllm
Image-Text-to-Text • 20B • Updated • 2.09k • 5
Product-VQ codebook weights (NVFP4-CB / FP8-CB) served by stock vLLM on Blackwell tensor cores. Spec + plugin: github.com/RobTand/gridbook
Note 23 GB · vision-language + MTP · the quality-validated one: -77% held-out KL vs the matched-size NVFP4+FP8 artifact, ToolEvalBench 87
Note 89 GB · 117B sparse MoE coding model · full 256k context on one DGX Spark, 14.9 tok/s decode · no quality claims published
Note 106 GB · 295B-A21B MoE at 2.9 bpp on one 128 GB Spark, MTP drafter included · no quality claims published