Instructions to use ai-server-mgr/MiniMax-H3-SDNQ-uint4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ai-server-mgr/MiniMax-H3-SDNQ-uint4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("ai-server-mgr/MiniMax-H3-SDNQ-uint4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("ai-server-mgr/MiniMax-H3-SDNQ-uint4", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]MiniMax-H3 · SDNQ UINT4, group size 64
Verified generation configurations: 2 × NVIDIA CMP 170HX without CPU offload,
and 1 × NVIDIA CMP 170HX (64 GiB) with host-RAM/CPU offload. Both completed
text-to-audio/video (t2va) generation. The single-card test took approximately
259 seconds; this is an observed run, not a general performance guarantee.
Проверено на одной карте: генерация видео со звуком по тексту работает на 1 × CMP 170HX 64 GiB с выгрузкой компонентов в оперативную память (CPU offload). Работа целиком в видеопамяти без ОЗУ этим тестом не подтверждена.
This is a quantized derivative of MiniMaxAI/MiniMax-H3, not an official MiniMax release.
Source revision: 42ed227ee7df40d41602854ae760620d6eb651fe.
transformer, transformer_ref, and text_encoder were converted using SDNQ
0.2.5+ai.manager.5, recipe uint4-g64-fp32-preserve-v1. Selected layers retain
their original precision; UINT4 does not mean that every tensor is four-bit.
The visual and audio VAEs are retained in their source representation.
See ai_manager_quantization.json and conversion_manifest.json for details.
Download size and memory
| Component | GB | GiB |
|---|---|---|
transformer |
20.768 | 19.342 |
transformer_ref |
20.768 | 19.342 |
text_encoder |
22.983 | 21.405 |
vae |
10.416 | 9.700 |
audio_vae |
0.605 | 0.564 |
tokenizer |
0.011 | 0.011 |
processor |
0.011 | 0.011 |
scheduler |
0.000 | 0.000 |
audio_scheduler |
0.000 | 0.000 |
| Core artifact (including conversion metadata) | 75.563 | 70.374 |
Sizes use decimal GB and binary GiB explicitly. The download includes both transformer variants. A single generation workflow does not load both variants simultaneously, so download size is not a VRAM requirement.
The artifact contains weights, component configurations, tokenizer/processor files, and conversion metadata. Server-local catalog backups, locks, download state, credentials, and private application source are not included.
Tested on two NVIDIA CMP 170HX cards
Verified on 2026-10-07, using AI Server Manager's SDNQ runtime:
| Item | Observed configuration/result |
|---|---|
| GPUs | 2 × NVIDIA CMP 170HX, 64 GiB per card, Ampere SM80 |
| Placement | Generator + VAEs on GPU 0; text encoder on GPU 1 |
| Workflow | t2va (text to audio/video) |
| Compute precision | BF16 |
| Attention | FlashAttention 2 with a guarded SDPA fallback |
| CPU offload | Disabled |
| Resident reserve | 8 GiB per selected GPU |
| Generation | Completed; progress reached 100%; result endpoint returned HTTP 200 |
| Observed denoising loop | 15 iterations |
| Job elapsed time | Approximately 123.5 seconds, including orchestration and encoding |
The completed job was 1791389692176-i8wwmrq3lhc. The runtime's final health
snapshot counted 10,164 FlashAttention calls and zero SDPA calls during that
loaded session. This is a session counter, not a throughput benchmark.
GPU memory was approximately 30 GiB + 22 GiB after loading and approximately 34 GiB + 23 GiB after the completed test. These are observed snapshots, not peak-memory measurements or guarantees for other resolutions and lengths.
The exact request dimensions/prompt and independent audio/visual quality
assessment were not retained in this report. No quality-parity claim or speedup
relative to BF16 is made. fl2va and ref2va passed component/workflow load
validation during conversion; end-to-end generation in those modes has not
been established by this test. See validation_report.json.
Tested on one NVIDIA CMP 170HX + host RAM
Single-card resident status: not yet verified. The current manager's disk-size-based preflight rejects the full artifact's resident single-card startup at an estimated 81 GB. This is not proof that the active workflow needs 81 GB.
Single GPU + host-RAM startup: passed on 2026-10-07 using a separate
t2va/fl2va test view (51.032 GiB of logical files), which excludes the unused
transformer_ref from that view only. The complete reference weights remain in
this repository and in the source artifact. The view uses hard links, not a
second copy of the weights. With CUDA_VISIBLE_DEVICES=0, H3 auto CPU offload,
FA2, and an 8 GiB reserve, /ready returned HTTP 200 with model_loaded=true.
Idle GPU allocation was approximately 0.383 GiB; process RSS was approximately
9.493 GiB with zero swapped process memory. Offload brings components to the
GPU on demand; these idle numbers do not describe generation-time memory.
Single-GPU generation completed: the user's t2va run
1791391519005-khwx9gmsghs completed on the same single-card/offload setup.
The MP4 was retrievable (HTTP 206 range response, 592,301 bytes). Server timings:
259.109 seconds total, 257.372 seconds inference, no additional model load.
The request's exact prompt/dimensions and peak memory were not retained for this
report; visual/audio quality has not been reviewed. This does not establish
parity with original weights. The test view does not support the audio/image-
reference ref2va workflow.
Runtime compatibility
These are SDNQ-packed weights, not GGUF, AWQ, ordinary BF16 weights, or a
drop-in vLLM model. Use a compatible SDNQ-aware Diffusers MiniMax-H3 modular
pipeline runtime. The tested environment used PyTorch 2.13.0+cu129, CUDA 12.9,
Diffusers 0.41.0.dev0 with MiniMax-H3 support, and SDNQ 0.2.5+ai.manager.5.
The version strings alone are not a guarantee that an unpatched upstream
installation has the same behavior. In particular, the tested two-card adapter
preserves CPU token metadata while moving accelerator conditioning tensors
between devices.
For the tested resident two-card configuration:
Selected GPUs: two CMP170HX cards
SDNQ_MINIMAX_H3_GPU_MODE=dual
SDNQ_MINIMAX_H3_WORKFLOW=t2va
SDNQ_MINIMAX_H3_CPU_OFFLOAD=0
SDNQ_MINIMAX_H3_RESERVE_GIB=8
SDNQ_DEMO_TORCH_DTYPE=bfloat16
SDNQ_DEMO_ATTENTION_BACKEND=flash_attention_2
SDNQ_DEMO_USE_QUANTIZED_MATMUL=0
SDNQ_FORCE_DISABLE_QUANTIZED_MATMUL=1
SDNQ_ALLOW_FP8_MM=0
SDNQ_USE_TORCH_COMPILE=0
Disabling quantized matmul does not expand all stored weights to BF16; it
selects the compatible execution path for the packed SDNQ representation.
Download the complete repository, including shard indices and quantization
configs. The published modular indices use this Hub repository ID instead of
the original conversion machine's absolute /results/... path.
License and changes
The weights remain subject to the MiniMax H3 Community License Agreement;
quantization does not relicense them as MIT or Apache-2.0. Read LICENSE and
NOTICE before use or distribution. The Qwen3-VL encoder also carries its
upstream Apache-2.0 licensing obligations; preserve applicable notices.
The MiniMax agreement excludes the EU, UK, Republic of Korea, and USA from its Applicable Territory, and restricts redistribution accordingly. Commercial products/services with yearly revenue above USD 20 million require separate prior written authorization under that agreement. Additional use restrictions also apply. Private hosting is not a substitute for complying with the license.
Modified artifacts: packed UINT4 weights and SDNQ configs for the three listed components; shard indices; pipeline indices normalized for this repository; conversion metadata and this quantization-specific model card. The source checkpoint is attributed above. No endorsement by MiniMax is implied.
- Downloads last month
- -
Model tree for ai-server-mgr/MiniMax-H3-SDNQ-uint4
Base model
MiniMaxAI/MiniMax-H3