Configuration Parsing Warning:In UNKNOWN_FILENAME: "quantization_config.config_groups.group_0.format" must be a string

Smaug-Mini MXFP4

Community MXFP4 quantization of abacusai/Smaug-Mini. The vision tower remains in BF16.

The original NVIDIA ModelOpt MXFP4 export was repacked into the compressed-tensors mxfp4-pack-quantized format. The conversion preserves the packed FP4 values and E8M0 group scales rather than requantizing the model.

Quantization

  • NVIDIA ModelOpt: 0.47.0
  • Format: MXFP4 (E2M1 weights, group size 32, E8M0 scales)
  • Calibration prompts: 1,024
  • Calibration sequence length: 4,096
  • Vision tower: BF16
  • Storage format: compressed-tensors / mxfp4-pack-quantized
  • Recommended serving KV cache: FP8 (--kv-cache-dtype fp8); the repacked checkpoint does not encode a KV-cache quantization scheme.

The included repack_report.json records conversion checks over representative layers.

Hugging Face's automated safetensors metadata may display an 8-bit tag and a lower parameter total for this checkpoint because packed FP4 tensors are stored in uint8 containers. The model architecture remains the 27B Smaug-Mini architecture.

Serving

Example vLLM invocation:

vllm serve WiktorMatuszek/smaug-mini-mxfp4 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

vLLM supports the mxfp4-pack-quantized compressed-tensors format. On supported SM100+ hardware it can use a true W4A4 path; otherwise the format can fall back to W4A16 through Marlin.

Evaluation

The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.

License and attribution

Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.

Downloads last month
41
Safetensors
Model size
15B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WiktorMatuszek/smaug-mini-mxfp4

Base model

Qwen/Qwen3.8-27B
Quantized
(6)
this model