MiniMax-H3 10Eros-Max beta2 (FL2VA) β€” W4A8

W4A8 (asym_w4a8_int8) quantization of 10Eros_Max_h3_fl2va_beta2_pruned (bf16 source from TenStrip/10Eros-Max).

  • Format: asym_w4a8_int8, group_size=16, convrot_groupsize=256, per-tensor Lloyd-Max codebook + fp8 group scales (Kijai / comfy-kitchen W4A8 layout).
  • Size: ~11.7 GB (68.8% smaller than the 40 GB bf16 source) β€” fits a 24 GB card with headroom.
  • Layers: 200 2D-linear layers quantized (96% of policy-targeted bytes); norms / first / last kept higher-precision.
  • Quantizer: comfyui-mixed-quantizer --format w4a8 --group-size 16 --codebook-mode fit.
  • Reconstruction: relL2 β‰ˆ 0.073 (bound 0.25), SNR β‰ˆ 22.8 dB, cos β‰ˆ 0.9973.

Requirements

  • ComfyUI β‰₯ v0.31.0 (native W4A8 loader) or the comfyui_w4a8_loader.patch.
  • comfy-kitchen with AsymW4A8Int8Layout (PR #90).
  • CUDA SM β‰₯ 8.0 (verified on an RTX 3090 / SM 8.6).

Verification

Loads via ComfyUI's diffusion-model loader and generates end-to-end on an RTX 3090 (image still executed in ~28 s).

Community experiment; inherits the source model's license.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for berryber09/MiniMax-H3-10Eros-beta2-fl2va-w4a8

Finetuned
(4)
this model