Remove stale fork patch: superseded by xianbaoqian/vllm#1 (fix-spec-decode), which the card now points to 2c6a808 verified cloudnathan5 commited on 7 days ago
Fix install step: plain editable install, not --no-deps (matches the validated setup) c8f757c verified cloudnathan5 commited on 7 days ago
Add SGLang instructions via cloudnathan5/sglang W4A16 fork: 111-207 tok/s with DFlash ad1be1d verified cloudnathan5 commited on 7 days ago
Make ignore list portable: name the unquantized vision stack explicitly (SGLang compatibility) 3fdc15d verified cloudnathan5 commited on 7 days ago
Document engine support (SGLang cannot serve W4A16 yet) and the related Preyazz checkpoint 6a79bf2 verified cloudnathan5 commited on 7 days ago
Record that the quant is verified on the upstream PR + fix-spec-decode route 398b534 verified cloudnathan5 commited on 8 days ago
Point install instructions at upstream PR #51655 + xianbaoqian/vllm#1 rather than this fork 75161ab verified cloudnathan5 commited on 8 days ago
GPTQ-rounded NVFP4A16: PPL 5.5631 (was 5.5965 RTN), 139.3 tok/s with DFlash 11f2fa4 verified cloudnathan5 commited on 8 days ago
Replace W4A4 with NVFP4A16: faster (74.7 vs 54.3 tok/s) and better PPL (5.5965 vs 5.7660) 05fb253 verified cloudnathan5 commited on 8 days ago
Update vllm-muse-glimmer-fork.patch: point at live fork branch, refresh patch 406d3c0 verified cloudnathan5 commited on 8 days ago
Update README.md: point at live fork branch, refresh patch 7bf9ebd verified cloudnathan5 commited on 8 days ago