Instructions to use unsloth/MiniMax-H3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Studio
How to use unsloth/MiniMax-H3-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="unsloth/MiniMax-H3-GGUF", max_seq_length=2048, )
Is the arch labeled as minimax?
I noticed people commenting on my posts about these ggufs and they said the arch doesnt pass through the gguf loader and is causing errors for them.
When i qusnted the model i passed it as "wan" arch and keys to avoid this error and alot of the time you can trick the quantizer to believe its wan then swap architecture back later but i left it at wan because it passes checks in the gguf loader without needing to code for it and PR it on github!
I do have a gguf loader as well that supports the model passing as wan for if you want to do so:
(It says w3a8 loader but thats just the main use, it does gguf and w4a8 if need be but thats native now)
https://github.com/RealRebelAI/Rebels_w3a8_Loader
I noticed people commenting on my posts about these ggufs and they said the arch doesnt pass through the gguf loader and is causing errors for them.
When i qusnted the model i passed it as "wan" arch and keys to avoid this error and alot of the time you can trick the quantizer to believe its wan then swap architecture back later but i left it at wan because it passes checks in the gguf loader without needing to code for it and PR it on github!
I do have a gguf loader as well that supports the model passing as wan for if you want to do so:
(It says w3a8 loader but thats just the main use, it does gguf and w4a8 if need be but thats native now)
https://github.com/RealRebelAI/Rebels_w3a8_Loader
I still can't get these pruned GGUFs to load. I installed your Rebels_w3a8_Loader node, loaded the GGUF Unet Loader + model_type (Rebels), and tried setting various values for model_type, including “minimax,” but it doesn't work—it always gives the error “This model is not currently supported - (Unknown model architecture!)”. How do I load these pruned GGUFs into ComfyUI?
The pruned ggufs are not mine, i cant help you there boss you gotta ask the contributor who made them
The pruned ggufs are not mine, i cant help you there boss you gotta ask the contributor who made them
I found the solution! Right at the bottom. https://github.com/city96/ComfyUI-GGUF/issues/471 His GGUF fork loads them.