Text-to-Video
Diffusers
Safetensors
MiniMax H3
MiniMaxH3ModularPipeline
image-to-video
audio-video-generation
sparse-attention
block-sparse
inference-acceleration
triton
comfyui
Instructions to use Aazeus/Spark-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Aazeus/Spark-H3 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Aazeus/Spark-H3", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
|
Download docs/kernel_performance_matrix.md from Aazeus/Spark-H3: direct link, hf CLI and curl.
- Browser
- Download file 2.97 kB
-
https://huggingface.co/Aazeus/Spark-H3/resolve/main/docs/kernel_performance_matrix.md
- Command line
-
hf download hf://Aazeus/Spark-H3/docs/kernel_performance_matrix.md
-
curl -L -o kernel_performance_matrix.md https://huggingface.co/Aazeus/Spark-H3/resolve/main/docs/kernel_performance_matrix.md
2.97 kB
| # Kernel performance matrix | |
| This table tracks comparable Dense-versus-Spark kernel measurements across GPU | |
| architectures and inference profiles. Each duration cell reports **speedup** | |
| followed by `(mean Dense DiT time -> mean Spark DiT time)`. Speedups are rounded | |
| to two decimal places and times to one decimal place after aggregation. | |
| | GPU | SM | Model / profile | DiT evals | Warmup steps | Dense layers | Kernel / route | Commit | 5s / 120f | 10s / 240f | 14.4s / 345f | | |
| |---|---:|---|---:|---:|---:|---|---|---:|---:|---:| | |
| | A800-SXM4-80GB | 80 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `1185195` | **1.41x** (326.3s -> 231.4s) | **1.70x** (988.0s -> 581.5s) | **1.84x** (1841.9s -> 1003.0s) | | |
| | RTX PRO 6000 Blackwell Server Edition | 120 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `57d401b` + dynamic-cache working tree | **1.42x** (197.8s -> 139.5s) | **1.72x** (583.9s -> 338.7s) | **1.87x** (1070.1s -> 571.8s) | | |
| ## Benchmark profile | |
| - Resolution: 1344x768; durations: 120, 240, and 345 frames. | |
| - Precision and seed: BF16, seed 42. | |
| - MiniMax-H3 native requests 20 inference steps and executes 19 transformer/DiT | |
| evaluations. The table records actual DiT evaluations so that profiles whose | |
| requested and executed step counts match can be added without ambiguity. | |
| - Spark configuration: 10% Top-K, legacy midpoint, threshold routing, four | |
| dense warmup evaluations, and one forced dense layer per sparse evaluation. | |
| - Runtime warmup is architecture-row specific. The A800 row uses a discarded | |
| pass through the four configured dense evaluations plus the first sparse | |
| evaluation. The PRO 6000 row uses the validated short pass with three | |
| requested steps: exactly one dense and one sparse evaluation. In both cases, | |
| the measured pass reuses the compiled callable and excludes first-sparse | |
| compilation time. | |
| - Timing covers only CUDA-synchronized DiT/denoising. Pipeline loading, VAE | |
| decode, and other end-to-end work are excluded. | |
| - Each displayed value aggregates four independent Dense measurements and four | |
| independent Spark measurements. Speedup is calculated from the unrounded | |
| means: `mean(Dense) / mean(Spark)`. | |
| - All measured Spark runs completed 19/19 evaluations with finite output. The | |
| A800 row used `sm80_fused_virtual_query`; no fallback backend was used. | |
| - The PRO 6000 measurements use four prompts (cases 2, 13, 31, and 33). Raw | |
| records and the aggregate are under | |
| `${H3_EXPERIMENTS_ROOT}/pro6000_task1_matrix_20261003`; its four | |
| formal Spark runs at every duration reported zero new fused-kernel compile | |
| calls after runtime warmup. | |
| ## Adding results | |
| Add one row for each GPU, model/inference profile, kernel backend, and tested | |
| commit. Keep actual DiT evaluations explicit. Results from different profiles | |
| may share this matrix, but absolute times should only be compared when their | |
| profile and evaluation count match. | |