liangsu9988's picture
Add native SM110 backend and validation
4a45a53 verified
|
Raw
History Blame
1.4 kB

RTX 5090 built-artifact results

Environment: NVIDIA GeForce RTX 5090 (SM120), driver 580.159.03, PyTorch 2.11.0+cu128, CUDA runtime 12.8. Artifact variant: torch211-cxx11-cu128-x86_64-linux, source commit 4baeadf. Measurements use CUDA events after warmup.

Workload M x top-k N x K W4A4 region us W4A4 kernel us W4A16 us Per-pair loop us Region speedup
gate_up decode 1 x 8 1024 x 2048 12.286 8.177 6.204 65.612 5.34x
gate_up verify 7 x 8 1024 x 2048 32.785 28.691 24.605 459.022 14.00x
down decode 8 x 1 2048 x 512 10.259 6.162 6.169 64.847 6.32x
down verify 56 x 1 2048 x 512 22.540 18.449 24.602 451.777 20.04x

W4A4 region includes one batched activation quantization plus one grouped compute launch. Per-pair loop quantizes and launches each routed pair separately. On this cu128 artifact W4A16 wins gate-up, the two kernel-only paths tie for down decode, and W4A4 wins down verify. The complete W4A4 region still removes the legacy launch storm, but lower precision is not presented as a universal per-kernel winner.

Full source correctness: 21/21 checks passed. Worst dequantized-contract result from the random sweep: max abs 0.004084, p99 abs 0.002907, mean abs 0.000507, cosine 0.9999988. Grouped-vs-native-loop and graph replay checks are bitwise.