Instructions to use WiktorMatuszek/smaug-mini-mxfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WiktorMatuszek/smaug-mini-mxfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="WiktorMatuszek/smaug-mini-mxfp4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("WiktorMatuszek/smaug-mini-mxfp4") model = AutoModelForMultimodalLM.from_pretrained("WiktorMatuszek/smaug-mini-mxfp4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WiktorMatuszek/smaug-mini-mxfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WiktorMatuszek/smaug-mini-mxfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-mxfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/WiktorMatuszek/smaug-mini-mxfp4
- SGLang
How to use WiktorMatuszek/smaug-mini-mxfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-mxfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-mxfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-mxfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-mxfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use WiktorMatuszek/smaug-mini-mxfp4 with Docker Model Runner:
docker model run hf.co/WiktorMatuszek/smaug-mini-mxfp4
Configuration Parsing Warning:In UNKNOWN_FILENAME: "quantization_config.config_groups.group_0.format" must be a string
Smaug-Mini MXFP4
Community MXFP4 quantization of abacusai/Smaug-Mini. The vision tower remains in BF16.
The original NVIDIA ModelOpt MXFP4 export was repacked into the compressed-tensors mxfp4-pack-quantized format. The conversion preserves the packed FP4 values and E8M0 group scales rather than requantizing the model.
Quantization
- NVIDIA ModelOpt:
0.47.0 - Format: MXFP4 (E2M1 weights, group size 32, E8M0 scales)
- Calibration prompts: 1,024
- Calibration sequence length: 4,096
- Vision tower: BF16
- Storage format:
compressed-tensors/mxfp4-pack-quantized - Recommended serving KV cache: FP8 (
--kv-cache-dtype fp8); the repacked checkpoint does not encode a KV-cache quantization scheme.
The included repack_report.json records conversion checks over representative layers.
Hugging Face's automated safetensors metadata may display an 8-bit tag and a lower parameter total for this checkpoint because packed FP4 tensors are stored in uint8 containers. The model architecture remains the 27B Smaug-Mini architecture.
Serving
Example vLLM invocation:
vllm serve WiktorMatuszek/smaug-mini-mxfp4 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
vLLM supports the mxfp4-pack-quantized compressed-tensors format. On supported SM100+ hardware it can use a true W4A4 path; otherwise the format can fall back to W4A16 through Marlin.
Evaluation
The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.
License and attribution
Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.
- Downloads last month
- 41