Commit History

Remove stale fork patch: superseded by xianbaoqian/vllm#1 (fix-spec-decode), which the card now points to
2c6a808
verified

cloudnathan5 commited on

Fix install step: plain editable install, not --no-deps (matches the validated setup)
c8f757c
verified

cloudnathan5 commited on

Add SGLang instructions via cloudnathan5/sglang W4A16 fork: 111-207 tok/s with DFlash
ad1be1d
verified

cloudnathan5 commited on

Make ignore list portable: name the unquantized vision stack explicitly (SGLang compatibility)
3fdc15d
verified

cloudnathan5 commited on

Document engine support (SGLang cannot serve W4A16 yet) and the related Preyazz checkpoint
6a79bf2
verified

cloudnathan5 commited on

Record that the quant is verified on the upstream PR + fix-spec-decode route
398b534
verified

cloudnathan5 commited on

Point install instructions at upstream PR #51655 + xianbaoqian/vllm#1 rather than this fork
75161ab
verified

cloudnathan5 commited on

GPTQ-rounded NVFP4A16: PPL 5.5631 (was 5.5965 RTN), 139.3 tok/s with DFlash
11f2fa4
verified

cloudnathan5 commited on

Update README.md
f962df5
verified

cloudnathan5 commited on

Replace W4A4 with NVFP4A16: faster (74.7 vs 54.3 tok/s) and better PPL (5.5965 vs 5.7660)
05fb253
verified

cloudnathan5 commited on

Update vllm-muse-glimmer-fork.patch: point at live fork branch, refresh patch
406d3c0
verified

cloudnathan5 commited on

Update README.md: point at live fork branch, refresh patch
7bf9ebd
verified

cloudnathan5 commited on

Upload folder using huggingface_hub
3392d1d
verified

cloudnathan5 commited on

initial commit
b2dba82
verified

cloudnathan5 commited on