--- base_model: - darkc0de/Muse-Glimmer-30B-heretic - meta-models/Muse-Glimmer-30B-assistant base_model_relation: quantized library_name: llama.cpp pipeline_tag: image-text-to-text license: apache-2.0 tags: - gguf - llama.cpp - rocm - rocmfpx - amd - strix-halo - gfx1151 - imatrix - multimodal - muse-glimmer - conversational - heretic - uncensored - decensored - abliterated - dflash - speculative-decoding --- # Muse-Glimmer-30B Heretic ROCmFPX GGUF ROCmFP4 and ROCmFP8 builds of [darkc0de/Muse-Glimmer-30B-heretic](https://huggingface.co/darkc0de/Muse-Glimmer-30B-heretic), targeted and tested on AMD Strix Halo (`gfx1151`). The source is a reproducible Heretic v1.4.0 abliteration of Meta's Muse Glimmer 30B. Rawr. šŸ¦– These are custom ROCmFPX formats, not ordinary llama.cpp Q4/Q8 files. Read the compatibility section before downloading. > **Experimental runtime required:** these GGUFs need a patched build of > [charlie12345/ROCmFPX](https://github.com/charlie12345/ROCmFPX). Stock > llama.cpp does not implement the ROCmFP4/ROCmFP8 tensor layouts, while the > pinned ROCmFPX base predates Muse Glimmer support. Apply the included > [`ROCmFPX-Muse-Glimmer.patch`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/blob/main/ROCmFPX-Muse-Glimmer.patch) > to ROCmFPX commit `00d54526e…`, then build that checkout. ## Files Click a filename to download it directly from the Hub. | File | ROCmFPX preset | Size | BPW | iMatrix | | --- | --- | ---: | ---: | --- | | [`Muse-Glimmer-30B-heretic-ROCmFP4.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/Muse-Glimmer-30B-heretic-ROCmFP4.gguf?download=true) | `Q4_0_ROCMFP4_STRIX` | 14.17 GiB | 4.36 | Yes | | [`Muse-Glimmer-30B-heretic-ROCmFP4-Q6-QUALITY.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/Muse-Glimmer-30B-heretic-ROCmFP4-Q6-QUALITY.gguf?download=true) | `Q4_0_ROCMFP4_COHERENT` | 14.93 GiB | 4.60 | Yes | | [`Muse-Glimmer-30B-heretic-ROCmFP8.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/Muse-Glimmer-30B-heretic-ROCmFP8.gguf?download=true) | `Q8_0_ROCMFPX` | 26.77 GiB | 8.25 | No | | [`mmproj-Muse-Glimmer-30B-heretic-BF16.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/mmproj-Muse-Glimmer-30B-heretic-BF16.gguf?download=true) | BF16 vision projector | 3.58 GiB | — | — | | [`Muse-Glimmer-30B-DFlash-ROCmFP4.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/Muse-Glimmer-30B-DFlash-ROCmFP4.gguf?download=true) | `Q4_0_ROCMFP4_STRIX` drafter | 1.39 GiB | 4.63 | No | | [`Muse-Glimmer-30B-DFlash-ROCmFP8.gguf`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/resolve/main/Muse-Glimmer-30B-DFlash-ROCmFP8.gguf?download=true) | `Q8_0_ROCMFPX` drafter | 2.47 GiB | 8.25 | No | Suggested choices: - **ROCmFP4:** default Strix Halo speed/quality build. Fast FP4 body, dual-scale FP4 attention K/V, and Q6_K token embeddings. - **ROCmFP4-Q6-QUALITY:** coherence-biased build. Dual-scale FP4 throughout the body with Q6_K token embeddings. - **ROCmFP8:** high-fidelity 8.25-bpw reference build. - **DFlash ROCmFP8:** recommended drafter for this Heretic target on the tested host. It produced the best measured six-token speculative speed. - **DFlash ROCmFP4:** smaller drafter alternative. The fresh Heretic-labelled BF16 projector works with all three text models. ## About the abliteration The source model was produced with [Heretic](https://heretic-project.org) v1.4.0 and retains Muse Glimmer's architecture, tokenizer, chat template, perception encoder, and native 131,072-token context. The source author reports KL divergence `0.0743` from the original and 11 refusals out of 100, compared with 59/100 for the original. Published abliteration parameters: | Parameter | Value | | --- | ---: | | `direction_index` | 38.49 | | `attn.o_proj.max_weight` | 1.46 | | `attn.o_proj.max_weight_position` | 30.98 | | `attn.o_proj.min_weight` | 0.75 | | `attn.o_proj.min_weight_distance` | 29.88 | | `mlp.down_proj.max_weight` | 1.49 | | `mlp.down_proj.max_weight_position` | 31.58 | | `mlp.down_proj.min_weight` | 0.32 | | `mlp.down_proj.min_weight_distance` | 26.40 | See the [source model card](https://huggingface.co/darkc0de/Muse-Glimmer-30B-heretic) for its reproduction recipe and the original author's description. These figures describe the BF16 source; this repository did not rerun that evaluation on each quantization. ## iMatrix Both FP4 targets use the same GGUF importance matrix: - 500 chunks Ɨ 512 tokens (approximately 256k calibration tokens) - 416 tensor importance entries consumed by each quantizer - varied narrative/general-language calibration corpus - checkpoints saved every 100 chunks `Q8_0_ROCMFPX` does not consume importance weights, so the FP8 reference was intentionally built without an iMatrix. The two DFlash files are quantizations of Meta's official [Muse-Glimmer-30B-assistant](https://huggingface.co/meta-models/Muse-Glimmer-30B-assistant), not a separately trained Heretic assistant. Every proposal is still verified by the Heretic target. Compatibility and speed were measured rather than assumed; see Validation below. ## Compatibility These files use experimental ROCmFPX tensor types and **will not load in stock llama.cpp**. The validated runtime was built from: - [ROCmFPX](https://github.com/charlie12345/ROCmFPX) base commit `00d54526e24e3aba4c76474e3147cbf9c7cc034c` - upstream llama.cpp Muse converter commit `d2f83055d6e3b379b5d34c4837122a918cf402c2` - the included Muse text, vision, DFlash, sparse-attention, and quantized-cache compatibility patch The runtime was built with ROCm and Vulkan backends. Reported generation tests used `ROCm0` on `gfx1151`. Minimal runtime setup: ```bash git clone https://github.com/charlie12345/ROCmFPX.git cd ROCmFPX git checkout 00d54526e24e3aba4c76474e3147cbf9c7cc034c hf download vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF \ ROCmFPX-Muse-Glimmer.patch \ --local-dir /tmp/muse-glimmer-heretic-rocmfpx git apply /tmp/muse-glimmer-heretic-rocmfpx/ROCmFPX-Muse-Glimmer.patch BUILD_DIR=build-muse-rocmfpx \ JOBS=16 \ CMAKE_HIP_COMPILER=/opt/rocm-7.2.0/lib/llvm/bin/clang++ \ GGML_HIP_ROCWMMA_FATTN=OFF \ ./scripts/build-strix-rocmfp4-mtp.sh ``` Adjust `CMAKE_HIP_COMPILER` for the installed ROCm version. The patch must be applied to the exact pinned commit. ## Download and run ```bash hf download vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF \ Muse-Glimmer-30B-heretic-ROCmFP4.gguf \ Muse-Glimmer-30B-DFlash-ROCmFP8.gguf \ mmproj-Muse-Glimmer-30B-heretic-BF16.gguf \ --local-dir ./Muse-Glimmer-30B-heretic-ROCmFPX ``` Text: ```bash ./llama-cli \ -m ./Muse-Glimmer-30B-heretic-ROCmFPX/Muse-Glimmer-30B-heretic-ROCmFP4.gguf \ -dev ROCm0 -ngl all -c 8192 -cnv ``` Vision: ```bash ./llama-cli \ -m ./Muse-Glimmer-30B-heretic-ROCmFPX/Muse-Glimmer-30B-heretic-ROCmFP4.gguf \ -mm ./Muse-Glimmer-30B-heretic-ROCmFPX/mmproj-Muse-Glimmer-30B-heretic-BF16.gguf \ --image ./image.png \ -p "Describe this image." \ -dev ROCm0 -ngl all -c 8192 -cnv -st ``` DFlash speculative decoding (recommended starting point): ```bash ./llama-cli \ -m ./Muse-Glimmer-30B-heretic-ROCmFPX/Muse-Glimmer-30B-heretic-ROCmFP4.gguf \ --spec-draft-model ./Muse-Glimmer-30B-heretic-ROCmFPX/Muse-Glimmer-30B-DFlash-ROCmFP8.gguf \ --spec-type draft-dflash \ -dev ROCm0 -ngl all \ --spec-draft-device ROCm0 --spec-draft-ngl all \ -ctk q4_0 -ctv q4_0 \ --spec-draft-type-k q4_0 --spec-draft-type-v q4_0 \ --spec-draft-n-max 6 --spec-draft-n-min 0 \ --spec-draft-p-min 0.0 --spec-draft-p-split 0.10 \ --no-spec-draft-backend-sampling \ -c 8192 -cnv ``` ## Validation All three target quantizations loaded on `ROCm0` with every layer offloaded and generated tokens. The default FP4 also passed an end-to-end image test using the fresh projector; it correctly identified both the Hugging Face site and the repository shown in the test screenshot. Three deterministic 256-token runs compared native ROCmFP4 decoding with the official FP8 DFlash. Both used Q4_0 target KV cache, seed 42, temperature 0, a 4,096-token context, six-token drafts, and the same prompts. | Prompt | Native tok/s | DFlash tok/s | Accepted / proposed | | --- | ---: | ---: | ---: | | Technical explanation | 13.8 | 22.4 | 155 / 595 | | Backup strategy | 13.7 | 24.1 | 163 / 543 | | Fiction opening | 13.8 | 21.0 | 149 / 628 | | **Mean / aggregate** | **13.77** | **22.50** | **467 / 1,766 (26.44%)** | The measured mean speedup was **1.63Ɨ**. At least one draft token was accepted in 199 of 297 verification rounds (67.0%), and the mean accepted span including the target token was 2.57 tokens. In a technical-prompt comparison, the smaller FP4 DFlash reached 20.3 tok/s with 140/684 proposals accepted, so FP8 is the recommended drafter on this host. Native and speculative greedy outputs were not byte-identical. With custom ROCmFP4 kernels, speculative verification changes target batch shapes; small floating-point differences can redirect the model's reasoning trace. Treat this cross-checkpoint pairing as a measured throughput option, not a claim of bit-exact decoding. Additional verification: - all 13 downloaded Heretic Safetensors shards passed the publisher's manifest - both FP4 targets consumed all 416 iMatrix entries - all six release files pass the published `SHA256SUMS` - source, intermediate, calibration, and patch hashes are retained in [`PROVENANCE_SHA256SUMS`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/blob/main/PROVENANCE_SHA256SUMS) - detailed build and validation report: [`BUILD_RESULTS.md`](https://huggingface.co/vmlinux/Muse-Glimmer-30B-heretic-ROCmFPX-GGUF/blob/main/BUILD_RESULTS.md) ## DFlash, not MTP The companion files are converted from Meta's official five-layer `MuseGlimmerAssistantModel`, which uses DFlash block diffusion with a trained block size of 16. Run it with `--spec-type draft-dflash`; it is not an MTP checkpoint and should not be run with `draft-mtp`. ## Provenance - Heretic source revision: `64a36ddcb9745b521bd9eb114465c93f860a594f` - Original target revision: `f84ecc3a0ea984a4c04542a84269e3d065350a6e` - DFlash source revision: `2c86316d689027b91123638739743fef1d425233` - ROCmFPX base: `00d54526e24e3aba4c76474e3147cbf9c7cc034c` - Conversion: upstream llama.cpp `d2f83055d6e3b379b5d34c4837122a918cf402c2` The source model's Apache 2.0 license and Muse Glimmer usage policy apply. This abliterated model may be more likely to produce unsafe, objectionable, or unreliable output; deployers should perform their own evaluation and add guardrails appropriate to their application.