MiniMax-H3 · SDNQ UINT4, group size 64

Verified generation configurations: 2 × NVIDIA CMP 170HX without CPU offload, and 1 × NVIDIA CMP 170HX (64 GiB) with host-RAM/CPU offload. Both completed text-to-audio/video (t2va) generation. The single-card test took approximately 259 seconds; this is an observed run, not a general performance guarantee.

Проверено на одной карте: генерация видео со звуком по тексту работает на 1 × CMP 170HX 64 GiB с выгрузкой компонентов в оперативную память (CPU offload). Работа целиком в видеопамяти без ОЗУ этим тестом не подтверждена.

This is a quantized derivative of MiniMaxAI/MiniMax-H3, not an official MiniMax release. Source revision: 42ed227ee7df40d41602854ae760620d6eb651fe.

transformer, transformer_ref, and text_encoder were converted using SDNQ 0.2.5+ai.manager.5, recipe uint4-g64-fp32-preserve-v1. Selected layers retain their original precision; UINT4 does not mean that every tensor is four-bit. The visual and audio VAEs are retained in their source representation. See ai_manager_quantization.json and conversion_manifest.json for details.

Download size and memory

Component GB GiB
transformer 20.768 19.342
transformer_ref 20.768 19.342
text_encoder 22.983 21.405
vae 10.416 9.700
audio_vae 0.605 0.564
tokenizer 0.011 0.011
processor 0.011 0.011
scheduler 0.000 0.000
audio_scheduler 0.000 0.000
Core artifact (including conversion metadata) 75.563 70.374

Sizes use decimal GB and binary GiB explicitly. The download includes both transformer variants. A single generation workflow does not load both variants simultaneously, so download size is not a VRAM requirement.

The artifact contains weights, component configurations, tokenizer/processor files, and conversion metadata. Server-local catalog backups, locks, download state, credentials, and private application source are not included.

Tested on two NVIDIA CMP 170HX cards

Verified on 2026-10-07, using AI Server Manager's SDNQ runtime:

Item Observed configuration/result
GPUs 2 × NVIDIA CMP 170HX, 64 GiB per card, Ampere SM80
Placement Generator + VAEs on GPU 0; text encoder on GPU 1
Workflow t2va (text to audio/video)
Compute precision BF16
Attention FlashAttention 2 with a guarded SDPA fallback
CPU offload Disabled
Resident reserve 8 GiB per selected GPU
Generation Completed; progress reached 100%; result endpoint returned HTTP 200
Observed denoising loop 15 iterations
Job elapsed time Approximately 123.5 seconds, including orchestration and encoding

The completed job was 1791389692176-i8wwmrq3lhc. The runtime's final health snapshot counted 10,164 FlashAttention calls and zero SDPA calls during that loaded session. This is a session counter, not a throughput benchmark.

GPU memory was approximately 30 GiB + 22 GiB after loading and approximately 34 GiB + 23 GiB after the completed test. These are observed snapshots, not peak-memory measurements or guarantees for other resolutions and lengths.

The exact request dimensions/prompt and independent audio/visual quality assessment were not retained in this report. No quality-parity claim or speedup relative to BF16 is made. fl2va and ref2va passed component/workflow load validation during conversion; end-to-end generation in those modes has not been established by this test. See validation_report.json.

Tested on one NVIDIA CMP 170HX + host RAM

Single-card resident status: not yet verified. The current manager's disk-size-based preflight rejects the full artifact's resident single-card startup at an estimated 81 GB. This is not proof that the active workflow needs 81 GB.

Single GPU + host-RAM startup: passed on 2026-10-07 using a separate t2va/fl2va test view (51.032 GiB of logical files), which excludes the unused transformer_ref from that view only. The complete reference weights remain in this repository and in the source artifact. The view uses hard links, not a second copy of the weights. With CUDA_VISIBLE_DEVICES=0, H3 auto CPU offload, FA2, and an 8 GiB reserve, /ready returned HTTP 200 with model_loaded=true. Idle GPU allocation was approximately 0.383 GiB; process RSS was approximately 9.493 GiB with zero swapped process memory. Offload brings components to the GPU on demand; these idle numbers do not describe generation-time memory. Single-GPU generation completed: the user's t2va run 1791391519005-khwx9gmsghs completed on the same single-card/offload setup. The MP4 was retrievable (HTTP 206 range response, 592,301 bytes). Server timings: 259.109 seconds total, 257.372 seconds inference, no additional model load. The request's exact prompt/dimensions and peak memory were not retained for this report; visual/audio quality has not been reviewed. This does not establish parity with original weights. The test view does not support the audio/image- reference ref2va workflow.

Runtime compatibility

These are SDNQ-packed weights, not GGUF, AWQ, ordinary BF16 weights, or a drop-in vLLM model. Use a compatible SDNQ-aware Diffusers MiniMax-H3 modular pipeline runtime. The tested environment used PyTorch 2.13.0+cu129, CUDA 12.9, Diffusers 0.41.0.dev0 with MiniMax-H3 support, and SDNQ 0.2.5+ai.manager.5. The version strings alone are not a guarantee that an unpatched upstream installation has the same behavior. In particular, the tested two-card adapter preserves CPU token metadata while moving accelerator conditioning tensors between devices.

For the tested resident two-card configuration:

Selected GPUs: two CMP170HX cards
SDNQ_MINIMAX_H3_GPU_MODE=dual
SDNQ_MINIMAX_H3_WORKFLOW=t2va
SDNQ_MINIMAX_H3_CPU_OFFLOAD=0
SDNQ_MINIMAX_H3_RESERVE_GIB=8
SDNQ_DEMO_TORCH_DTYPE=bfloat16
SDNQ_DEMO_ATTENTION_BACKEND=flash_attention_2
SDNQ_DEMO_USE_QUANTIZED_MATMUL=0
SDNQ_FORCE_DISABLE_QUANTIZED_MATMUL=1
SDNQ_ALLOW_FP8_MM=0
SDNQ_USE_TORCH_COMPILE=0

Disabling quantized matmul does not expand all stored weights to BF16; it selects the compatible execution path for the packed SDNQ representation. Download the complete repository, including shard indices and quantization configs. The published modular indices use this Hub repository ID instead of the original conversion machine's absolute /results/... path.

License and changes

The weights remain subject to the MiniMax H3 Community License Agreement; quantization does not relicense them as MIT or Apache-2.0. Read LICENSE and NOTICE before use or distribution. The Qwen3-VL encoder also carries its upstream Apache-2.0 licensing obligations; preserve applicable notices.

The MiniMax agreement excludes the EU, UK, Republic of Korea, and USA from its Applicable Territory, and restricts redistribution accordingly. Commercial products/services with yearly revenue above USD 20 million require separate prior written authorization under that agreement. Additional use restrictions also apply. Private hosting is not a substitute for complying with the license.

Modified artifacts: packed UINT4 weights and SDNQ configs for the three listed components; shard indices; pipeline indices normalized for this repository; conversion metadata and this quantization-specific model card. The source checkpoint is attributed above. No endorsement by MiniMax is implied.

Downloads last month
-
Safetensors
Model size
18B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai-server-mgr/MiniMax-H3-SDNQ-uint4

Finetuned
(166)
this model