Buckets:

abayb's picture
|
download
raw
2.18 kB

int3 g128 MLP on A10G — sub-4-bit feasibility study (v0-v3), lane parked

Goal: spend the idle PPL budget (2.027 vs 2.42 cap) on the verify path's biggest line item — MLP gate_up+down = 1.70 GB of the 2.44 GB step (70%). int3 bit-plane packing (32 weights -> 3 int32 + bf16 g128 scale) = 1.29 GB.

What works (reusable)

  • In-boot requantization from the served int4 checkpoint via identity-probe extraction (no quant-format unpacking), 42 layers in ~6s
  • MSE scale search: relerr 0.232 -> 0.189
  • Bit-plane pack/unpack with exact roundtrip; composed selftests pass at rel 0.005 (decode) / 0.004 (prefill)
  • fullgraph-safe integration via torch.library.custom_op (dynamo.disable is FATAL under vLLM fullgraph compile — v0 finding)
  • Single-numerics discipline: decode kernels + prefill dequant-scratch serve the SAME values; PPL measures the real model. Partial mode (per-component int3/int4) keeps that property.
  • Gate architecture: microbench >= threshold else stock fallback — 5 runs, zero leaderboard damage.

The kernel-efficiency ladder (the actual finding)

v1 row-major packing: 55.7 GB/s (uncoalesced lane loads) v1 tiled-coalesced + stages: 55.7 -> no change (NOT a load problem) v2 BLOCK_K=128, trans-free unpack: 179 GB/s (3.2x: per-iteration tl.trans/barrier latency was the wall) v3 per-kernel layout+config autotune: gateup 224.8 (bo=64,nw=4,ns=3) down 125.4 (config-INSENSITIVE) Full per-config tables in kernel_ladder.txt.

Conclusion

Beating int4-Marlin time (needs ~>=400 GB/s on packed bytes at M=8) is not reachable with single-CTA-per-tile Triton GEMV on sm_86: the down projection (K=10240, 80 CTAs x 80 chained dot-iterations) is dependency-latency-bound and insensitive to launch configs. The remaining 2x needs split-K partial accumulation, cp.async double-buffering, and warp specialization — a Marlin-class CUDA kernel. "No sub-4-bit Ampere kernel" is now a measured engineering bill, not folklore. PPL cost of int3 RTN (relerr 0.189) remains unmeasured — no run got past the speed gate, by design.

Xet Storage Details

Size:
2.18 kB
·
Xet hash:
32d628d6226900cd22e1e8fdab7aae7e318b438f6b498326539c95d4c6df37cc

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.