RTX 5090 built-artifact results
Environment: NVIDIA GeForce RTX 5090 (SM120), driver 580.159.03, PyTorch
2.11.0+cu128, CUDA runtime 12.8. Artifact variant:
torch211-cxx11-cu128-x86_64-linux, source commit 4baeadf. Measurements use
CUDA events after warmup.
| Workload | M x top-k | N x K | W4A4 region us | W4A4 kernel us | W4A16 us | Per-pair loop us | Region speedup |
|---|---|---|---|---|---|---|---|
| gate_up decode | 1 x 8 | 1024 x 2048 | 12.286 | 8.177 | 6.204 | 65.612 | 5.34x |
| gate_up verify | 7 x 8 | 1024 x 2048 | 32.785 | 28.691 | 24.605 | 459.022 | 14.00x |
| down decode | 8 x 1 | 2048 x 512 | 10.259 | 6.162 | 6.169 | 64.847 | 6.32x |
| down verify | 56 x 1 | 2048 x 512 | 22.540 | 18.449 | 24.602 | 451.777 | 20.04x |
W4A4 region includes one batched activation quantization plus one grouped
compute launch. Per-pair loop quantizes and launches each routed pair
separately. On this cu128 artifact W4A16 wins gate-up, the two kernel-only paths
tie for down decode, and W4A4 wins down verify. The complete W4A4 region still
removes the legacy launch storm, but lower precision is not presented as a
universal per-kernel winner.
Full source correctness: 21/21 checks passed. Worst dequantized-contract result from the random sweep: max abs 0.004084, p99 abs 0.002907, mean abs 0.000507, cosine 0.9999988. Grouped-vs-native-loop and graph replay checks are bitwise.