| --- |
| title: Muse Glimmer 30B |
| emoji: ✨ |
| colorFrom: purple |
| colorTo: blue |
| sdk: gradio |
| sdk_version: 6.22.0 |
| python_version: 3.10.13 |
| app_file: app.py |
| startup_duration_timeout: 1h |
| fullWidth: true |
| pinned: false |
| license: apache-2.0 |
| models: |
| - meta-models/Muse-Glimmer-30B |
| - meta-models/Muse-Glimmer-30B-assistant |
| --- |
| |
| # Muse Glimmer 30B · private inference |
|
|
| A private text-and-image chat for both: |
| - [`meta-models/Muse-Glimmer-30B`](https://huggingface.co/meta-models/Muse-Glimmer-30B), full BF16 |
| - [`meta-models/Muse-Glimmer-30B-assistant`](https://huggingface.co/meta-models/Muse-Glimmer-30B-assistant), assistant checkpoint |
|
|
| `meta-models/Muse-Glimmer-30B-assistant` is a companion DFlash drafter and does not expose standalone |
| `.generate(...)` in this release of Transformers. In this Space, assistant selection runs the full 30B |
| checkpoint (same prompt/history behavior) while keeping the model selector and UI behavior intact. |
|
|
| They are served at pinned revisions: |
| `f84ecc3a0ea984a4c04542a84269e3d065350a6e` (full BF16, 29,776,626,688 parameters) |
| and `2c86316d689027b91123638739743fef1d425233` (assistant BF16, 2,555,985,152 parameters). |
|
|
| ## Runtime |
|
|
| - Full, unquantized BF16 checkpoint (29,776,626,688 parameters) |
| - ZeroGPU allocation through the `@spaces.GPU(size="xlarge")` request. |
| The full 30B checkpoint reserves up to `120s` for its cold start on first load by default. |
| Assistant selection follows the same warm-path budget rules: |
| - if a full model is already loaded in RAM, assistant generation reuses that active model and uses the normal |
| per-turn estimate, |
| - if assistant must warm an alternative path, it also uses the `120s` cold-start cap. |
| - Pinned read-only model storage is mounted from Hugging Face model volumes into |
| - `/models/muse-glimmer` |
| - `/models/muse-glimmer-assistant` |
| - Runtime defaults can be overridden with `MUSE_MODEL_PATH`, `MUSE_ASSISTANT_MODEL_PATH`, |
| and `MUSE_DEFAULT_MODEL_ID`. |
| - `transformers==5.15.0`, the first stable release with native Muse Glimmer support |
| - Sequential safetensors `pread` loading, avoiding whole-shard host-RAM copies from the volume |
| - One inference request at a time |
| - Optional bucket-backed checkpoints are supported by setting `MUSE_MODEL_PATH` |
| and `MUSE_ASSISTANT_MODEL_PATH` to mounted bucket locations (for example |
| `/mnt/muse-glimmer` when mounted via `hf://buckets/<namespace>/<bucket>:/mnt`). |
| - The private Space API path is locked down (`api_visibility: private`) so all chat requests flow |
| through the same Gradio queue and preflight validation. |
| - Model loading is deferred until the first request so the Space can start even when a mount |
| is temporarily unavailable. |
| To keep both checkpoints selectable, mount both model volumes in Space settings: |
| `/models/muse-glimmer` and `/models/muse-glimmer-assistant`. |
| |
| ### Optional Hugging Face bucket storage |
|
|
| If you need storage that persists beyond the Space runtime, mount an HF bucket: |
|
|
| ```bash |
| hf spaces volumes set ssdataanalysis/muse-glimmer-30b \\ |
| -v hf://buckets/ssdataanalysis/muse-glimmer-30b-store:/mnt/muse-glimmer-store \\ |
| -v hf://models/meta-models/Muse-Glimmer-30B:/models/muse-glimmer:ro \\ |
| -v hf://models/meta-models/Muse-Glimmer-30B-assistant:/models/muse-glimmer-assistant:ro |
| ``` |
|
|
| Then point the app to bucket copies without changing code: |
|
|
| ```bash |
| export MUSE_MODEL_PATH=/mnt/muse-glimmer-store/Muse-Glimmer-30B |
| export MUSE_ASSISTANT_MODEL_PATH=/mnt/muse-glimmer-store/Muse-Glimmer-30B-assistant |
| ``` |
|
|
| To inspect configured mounts: |
|
|
| ```bash |
| hf spaces volumes list ssdataanalysis/muse-glimmer-30b |
| ``` |
|
|
| The app loads the selected checkpoint on demand; a selected model switch swaps checkpoints as needed. |
| Loaded checkpoints are placed on `cuda`, as required by ZeroGPU. The app does not download full |
| checkpoints into the Space's ephemeral disk. |
|
|
| ZeroGPU status check: |
|
|
| - `hf spaces info ssdataanalysis/muse-glimmer-30b --json` should show: |
| - `requested_hardware: zero-a10g` |
| - `hardware: zero-a10g` |
| - This app requests `@spaces.GPU(size="xlarge")` and should run on ZeroGPU's 96GB path for BF16. |
|
|
| ## Runtime access and quota checks |
|
|
| - The Space is private. If you open `https://huggingface.co/spaces/ssdataanalysis/muse-glimmer-30b` or |
| `https://ssdataanalysis-muse-glimmer-30b.hf.space` without being authenticated as the owner, you will see |
| an authentication/error page. |
| - API calls must include an authenticated Hugging Face token with access to this private Space. |
| - You can verify the endpoint is wired with: |
|
|
| ```bash |
| HF_TOKEN=<your-token> |
| curl -H "Authorization: Bearer $HF_TOKEN" \ |
| -H "Content-Type: application/json" \ |
| -d '{"data":["hi",null,"meta-models/Muse-Glimmer-30B",[],[], "", "high", false, 32, 1.0, 0.95, 64, 1.0, 42, false, true]}' \ |
| https://ssdataanalysis-muse-glimmer-30b.hf.space/gradio_api/call/chat |
| ``` |
|
|
| - If you receive a ZeroGPU quota response such as |
| `You have exceeded your Pro ZeroGPU quota ...`, the container/startup is healthy and the failure is account-level |
| quota, not a model loading/runtime regression. Add credits here if you need immediate inference: |
| <https://huggingface.co/settings/billing?add-credits=true> |
|
|
| ## Defaults and controls |
|
|
| The checkpoint's `generation_config.json` uses greedy decoding (`do_sample: false`), so the UI |
| opens in **Native greedy** mode. The separate **Meta recommended sampling** preset enables the |
| model-card values: |
|
|
| - temperature: `1.0` |
| - top-p: `0.95` |
| - top-k: `64` |
|
|
| Reasoning strength defaults to `high`, matching the supplied chat template. It can be changed to |
| `low`, `medium`, `high`, or `xhigh`. The app also exposes the response token budget, sampling |
| toggle, repetition penalty, seed, optional system instruction, reasoning visibility, and optional |
| image input. The 512-token response budget is an app safety default; the model repository does not |
| define a practical per-response budget. |
|
|
| The app uses the model's native response parser to keep reasoning and final-answer regions |
| separate. No tools are registered or executed. |
|
|
| ## Privacy and safety |
|
|
| Access is enforced by the private Space setting. This application does not write prompts, |
| responses, or uploaded images to persistent storage and does not print them to application logs. |
| Inference still runs on Hugging Face-hosted infrastructure; it is not local-only. Image metadata is |
| discarded before inference. |
|
|
| Use is subject to the model's |
| [`USAGE_POLICY.md`](https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/f84ecc3a0ea984a4c04542a84269e3d065350a6e/USAGE_POLICY.md) or [`USAGE_POLICY.md`](https://huggingface.co/meta-models/Muse-Glimmer-30B-assistant/blob/2c86316d689027b91123638739743fef1d425233/USAGE_POLICY.md). |
| Model outputs may be inaccurate and should be independently verified for consequential use. |
|
|