Buckets:

abayb's picture
|
download
raw
2.18 kB
# int3 g128 MLP on A10G — sub-4-bit feasibility study (v0-v3), lane parked
Goal: spend the idle PPL budget (2.027 vs 2.42 cap) on the verify path's
biggest line item — MLP gate_up+down = 1.70 GB of the 2.44 GB step (70%).
int3 bit-plane packing (32 weights -> 3 int32 + bf16 g128 scale) = 1.29 GB.
## What works (reusable)
- In-boot requantization from the served int4 checkpoint via identity-probe
extraction (no quant-format unpacking), 42 layers in ~6s
- MSE scale search: relerr 0.232 -> 0.189
- Bit-plane pack/unpack with exact roundtrip; composed selftests pass at
rel 0.005 (decode) / 0.004 (prefill)
- fullgraph-safe integration via torch.library.custom_op (dynamo.disable is
FATAL under vLLM fullgraph compile — v0 finding)
- Single-numerics discipline: decode kernels + prefill dequant-scratch serve
the SAME values; PPL measures the real model. Partial mode (per-component
int3/int4) keeps that property.
- Gate architecture: microbench >= threshold else stock fallback — 5 runs,
zero leaderboard damage.
## The kernel-efficiency ladder (the actual finding)
v1 row-major packing: 55.7 GB/s (uncoalesced lane loads)
v1 tiled-coalesced + stages: 55.7 -> no change (NOT a load problem)
v2 BLOCK_K=128, trans-free unpack: 179 GB/s (3.2x: per-iteration
tl.trans/barrier latency was the wall)
v3 per-kernel layout+config autotune: gateup 224.8 (bo=64,nw=4,ns=3)
down 125.4 (config-INSENSITIVE)
Full per-config tables in kernel_ladder.txt.
## Conclusion
Beating int4-Marlin time (needs ~>=400 GB/s on packed bytes at M=8) is not
reachable with single-CTA-per-tile Triton GEMV on sm_86: the down projection
(K=10240, 80 CTAs x 80 chained dot-iterations) is dependency-latency-bound
and insensitive to launch configs. The remaining 2x needs split-K partial
accumulation, cp.async double-buffering, and warp specialization — a
Marlin-class CUDA kernel. "No sub-4-bit Ampere kernel" is now a measured
engineering bill, not folklore. PPL cost of int3 RTN (relerr 0.189) remains
unmeasured — no run got past the speed gate, by design.

Xet Storage Details

Size:
2.18 kB
·
Xet hash:
32d628d6226900cd22e1e8fdab7aae7e318b438f6b498326539c95d4c6df37cc

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.