Spark-H3 / docs /kernel_performance_matrix.md
Aazeus's picture
Publish Spark-H3 code and model card
a506693 verified
|
Raw History Blame Contribute Delete
2.97 kB
# Kernel performance matrix
This table tracks comparable Dense-versus-Spark kernel measurements across GPU
architectures and inference profiles. Each duration cell reports **speedup**
followed by `(mean Dense DiT time -> mean Spark DiT time)`. Speedups are rounded
to two decimal places and times to one decimal place after aggregation.
| GPU | SM | Model / profile | DiT evals | Warmup steps | Dense layers | Kernel / route | Commit | 5s / 120f | 10s / 240f | 14.4s / 345f |
|---|---:|---|---:|---:|---:|---|---|---:|---:|---:|
| A800-SXM4-80GB | 80 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `1185195` | **1.41x** (326.3s -> 231.4s) | **1.70x** (988.0s -> 581.5s) | **1.84x** (1841.9s -> 1003.0s) |
| RTX PRO 6000 Blackwell Server Edition | 120 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `57d401b` + dynamic-cache working tree | **1.42x** (197.8s -> 139.5s) | **1.72x** (583.9s -> 338.7s) | **1.87x** (1070.1s -> 571.8s) |
## Benchmark profile
- Resolution: 1344x768; durations: 120, 240, and 345 frames.
- Precision and seed: BF16, seed 42.
- MiniMax-H3 native requests 20 inference steps and executes 19 transformer/DiT
evaluations. The table records actual DiT evaluations so that profiles whose
requested and executed step counts match can be added without ambiguity.
- Spark configuration: 10% Top-K, legacy midpoint, threshold routing, four
dense warmup evaluations, and one forced dense layer per sparse evaluation.
- Runtime warmup is architecture-row specific. The A800 row uses a discarded
pass through the four configured dense evaluations plus the first sparse
evaluation. The PRO 6000 row uses the validated short pass with three
requested steps: exactly one dense and one sparse evaluation. In both cases,
the measured pass reuses the compiled callable and excludes first-sparse
compilation time.
- Timing covers only CUDA-synchronized DiT/denoising. Pipeline loading, VAE
decode, and other end-to-end work are excluded.
- Each displayed value aggregates four independent Dense measurements and four
independent Spark measurements. Speedup is calculated from the unrounded
means: `mean(Dense) / mean(Spark)`.
- All measured Spark runs completed 19/19 evaluations with finite output. The
A800 row used `sm80_fused_virtual_query`; no fallback backend was used.
- The PRO 6000 measurements use four prompts (cases 2, 13, 31, and 33). Raw
records and the aggregate are under
`${H3_EXPERIMENTS_ROOT}/pro6000_task1_matrix_20261003`; its four
formal Spark runs at every duration reported zero new fused-kernel compile
calls after runtime warmup.
## Adding results
Add one row for each GPU, model/inference profile, kernel backend, and tested
commit. Keep actual DiT evaluations explicit. Results from different profiles
may share this matrix, but absolute times should only be compared when their
profile and evaluation count match.