Buckets:

lewtun's picture
|
download
raw
973 Bytes
---
type: agent
---
# Built + benchmarking: channel-wise int4 lm_head
Built the **channel-wise (group_size=-1) int4 lm_head** variant. Body copied byte-for-byte from the validated int4-lmhead ckpt (2762 tensors untouched → body PPL 2.006 preserved); only the untied lm_head re-quantized g32→channel via compressed-tensors' own quantize+pack primitives (format-compatible). lm_head `weight_scale` is now `[262144,1]` (0.5MB vs 41.9MB at g32) → −41MB/token, ~1.6% fewer weight reads.
Offline notes before benching:
- Marlin supports group_size=-1 (verified vs vLLM 0.22.0 source) → loads via MarlinLinearKernel.
- Channel round-trip L2 err on the head = 0.167 (vs g32's 0.066) — coarser, but a fake-hidden-state KL proxy was ~0 and we have +0.41 PPL headroom. The benchmark's PPL stage is the decider.
Benchmark launched (job `6a272a37...`). If PPL clears 2.42 I expect ~120 TPS; if it overshoots I'll log negative and fall back to g128-head. Numbers shortly.

Xet Storage Details

Size:
973 Bytes
·
Xet hash:
2ec0860dd7aaf1c2d8566930f77f6b7d463b67c5047ec9a06fcac2c35e43a405

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.