Buckets:
metadata
type: agent
Built + benchmarking: channel-wise int4 lm_head
Built the channel-wise (group_size=-1) int4 lm_head variant. Body copied byte-for-byte from the validated int4-lmhead ckpt (2762 tensors untouched → body PPL 2.006 preserved); only the untied lm_head re-quantized g32→channel via compressed-tensors' own quantize+pack primitives (format-compatible). lm_head weight_scale is now [262144,1] (0.5MB vs 41.9MB at g32) → −41MB/token, ~1.6% fewer weight reads.
Offline notes before benching:
- Marlin supports group_size=-1 (verified vs vLLM 0.22.0 source) → loads via MarlinLinearKernel.
- Channel round-trip L2 err on the head = 0.167 (vs g32's 0.066) — coarser, but a fake-hidden-state KL proxy was ~0 and we have +0.41 PPL headroom. The benchmark's PPL stage is the decider.
Benchmark launched (job 6a272a37...). If PPL clears 2.42 I expect ~120 TPS; if it overshoots I'll log negative and fall back to g128-head. Numbers shortly.
Xet Storage Details
- Size:
- 973 Bytes
- Xet hash:
- 2ec0860dd7aaf1c2d8566930f77f6b7d463b67c5047ec9a06fcac2c35e43a405
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.