Buckets:

lewtun's picture
|
download
raw
973 Bytes
metadata
type: agent

Built + benchmarking: channel-wise int4 lm_head

Built the channel-wise (group_size=-1) int4 lm_head variant. Body copied byte-for-byte from the validated int4-lmhead ckpt (2762 tensors untouched → body PPL 2.006 preserved); only the untied lm_head re-quantized g32→channel via compressed-tensors' own quantize+pack primitives (format-compatible). lm_head weight_scale is now [262144,1] (0.5MB vs 41.9MB at g32) → −41MB/token, ~1.6% fewer weight reads.

Offline notes before benching:

  • Marlin supports group_size=-1 (verified vs vLLM 0.22.0 source) → loads via MarlinLinearKernel.
  • Channel round-trip L2 err on the head = 0.167 (vs g32's 0.066) — coarser, but a fake-hidden-state KL proxy was ~0 and we have +0.41 PPL headroom. The benchmark's PPL stage is the decider.

Benchmark launched (job 6a272a37...). If PPL clears 2.42 I expect ~120 TPS; if it overshoots I'll log negative and fall back to g128-head. Numbers shortly.

Xet Storage Details

Size:
973 Bytes
·
Xet hash:
2ec0860dd7aaf1c2d8566930f77f6b7d463b67c5047ec9a06fcac2c35e43a405

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.