Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
cloudnathan5
/
Muse-Glimmer-30B-NVFP4
like
2
Image-Text-to-Text
Safetensors
vllm
muse_glimmer
nvfp4
fp4
compressed-tensors
llm-compressor
quantized
blackwell
multimodal
conversational
8-bit precision
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
Muse-Glimmer-30B-NVFP4
23.4 GB
Ctrl+K
Ctrl+K
1 contributor
History:
14 commits
cloudnathan5
Remove stale fork patch: superseded by xianbaoqian/vllm#1 (fix-spec-decode), which the card now points to
2c6a808
verified
about 8 hours ago
.gitattributes
Safe
1.57 kB
Upload folder using huggingface_hub
1 day ago
LICENSE
Safe
11.4 kB
Upload folder using huggingface_hub
1 day ago
README.md
14.6 kB
Fix install step: plain editable install, not --no-deps (matches the validated setup)
about 9 hours ago
USAGE_POLICY.md
5.23 kB
Upload folder using huggingface_hub
1 day ago
chat_template.jinja
7.17 kB
Upload folder using huggingface_hub
1 day ago
config.json
6.3 kB
Make ignore list portable: name the unquantized vision stack explicitly (SGLang compatibility)
about 17 hours ago
generation_config.json
185 Bytes
Upload folder using huggingface_hub
1 day ago
model-00001-of-00002.safetensors
20 GB
xet
GPTQ-rounded NVFP4A16: PPL 5.5631 (was 5.5965 RTN), 139.3 tok/s with DFlash
about 23 hours ago
model-00002-of-00002.safetensors
3.38 GB
xet
GPTQ-rounded NVFP4A16: PPL 5.5631 (was 5.5965 RTN), 139.3 tok/s with DFlash
about 23 hours ago
model.safetensors.index.json
224 kB
Replace W4A4 with NVFP4A16: faster (74.7 vs 54.3 tok/s) and better PPL (5.5965 vs 5.7660)
1 day ago
processor_config.json
1.08 kB
Upload folder using huggingface_hub
1 day ago
recipe.yaml
579 Bytes
GPTQ-rounded NVFP4A16: PPL 5.5631 (was 5.5965 RTN), 139.3 tok/s with DFlash
about 23 hours ago
tokenizer.json
28.1 MB
xet
Replace W4A4 with NVFP4A16: faster (74.7 vs 54.3 tok/s) and better PPL (5.5965 vs 5.7660)
1 day ago
tokenizer_config.json
80 kB
Upload folder using huggingface_hub
1 day ago