File size: 2,970 Bytes
a506693
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
# Kernel performance matrix

This table tracks comparable Dense-versus-Spark kernel measurements across GPU
architectures and inference profiles. Each duration cell reports **speedup**
followed by `(mean Dense DiT time -> mean Spark DiT time)`. Speedups are rounded
to two decimal places and times to one decimal place after aggregation.

| GPU | SM | Model / profile | DiT evals | Warmup steps | Dense layers | Kernel / route | Commit | 5s / 120f | 10s / 240f | 14.4s / 345f |
|---|---:|---|---:|---:|---:|---|---|---:|---:|---:|
| A800-SXM4-80GB | 80 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `1185195` | **1.41x** (326.3s -> 231.4s) | **1.70x** (988.0s -> 581.5s) | **1.84x** (1841.9s -> 1003.0s) |
| RTX PRO 6000 Blackwell Server Edition | 120 | MiniMax-H3 native | 19 | 4 | 1 | fused virtual-query / legacy+threshold / Top-K 10% | `57d401b` + dynamic-cache working tree | **1.42x** (197.8s -> 139.5s) | **1.72x** (583.9s -> 338.7s) | **1.87x** (1070.1s -> 571.8s) |

## Benchmark profile

- Resolution: 1344x768; durations: 120, 240, and 345 frames.
- Precision and seed: BF16, seed 42.
- MiniMax-H3 native requests 20 inference steps and executes 19 transformer/DiT
  evaluations. The table records actual DiT evaluations so that profiles whose
  requested and executed step counts match can be added without ambiguity.
- Spark configuration: 10% Top-K, legacy midpoint, threshold routing, four
  dense warmup evaluations, and one forced dense layer per sparse evaluation.
- Runtime warmup is architecture-row specific. The A800 row uses a discarded
  pass through the four configured dense evaluations plus the first sparse
  evaluation. The PRO 6000 row uses the validated short pass with three
  requested steps: exactly one dense and one sparse evaluation. In both cases,
  the measured pass reuses the compiled callable and excludes first-sparse
  compilation time.
- Timing covers only CUDA-synchronized DiT/denoising. Pipeline loading, VAE
  decode, and other end-to-end work are excluded.
- Each displayed value aggregates four independent Dense measurements and four
  independent Spark measurements. Speedup is calculated from the unrounded
  means: `mean(Dense) / mean(Spark)`.
- All measured Spark runs completed 19/19 evaluations with finite output. The
  A800 row used `sm80_fused_virtual_query`; no fallback backend was used.
- The PRO 6000 measurements use four prompts (cases 2, 13, 31, and 33). Raw
  records and the aggregate are under
  `${H3_EXPERIMENTS_ROOT}/pro6000_task1_matrix_20261003`; its four
  formal Spark runs at every duration reported zero new fused-kernel compile
  calls after runtime warmup.

## Adding results

Add one row for each GPU, model/inference profile, kernel backend, and tested
commit. Keep actual DiT evaluations explicit. Results from different profiles
may share this matrix, but absolute times should only be compared when their
profile and evaluation count match.