Buckets:
| type: agent | |
| # Built + benchmarking: channel-wise int4 lm_head | |
| Built the **channel-wise (group_size=-1) int4 lm_head** variant. Body copied byte-for-byte from the validated int4-lmhead ckpt (2762 tensors untouched → body PPL 2.006 preserved); only the untied lm_head re-quantized g32→channel via compressed-tensors' own quantize+pack primitives (format-compatible). lm_head `weight_scale` is now `[262144,1]` (0.5MB vs 41.9MB at g32) → −41MB/token, ~1.6% fewer weight reads. | |
| Offline notes before benching: | |
| - Marlin supports group_size=-1 (verified vs vLLM 0.22.0 source) → loads via MarlinLinearKernel. | |
| - Channel round-trip L2 err on the head = 0.167 (vs g32's 0.066) — coarser, but a fake-hidden-state KL proxy was ~0 and we have +0.41 PPL headroom. The benchmark's PPL stage is the decider. | |
| Benchmark launched (job `6a272a37...`). If PPL clears 2.42 I expect ~120 TPS; if it overshoots I'll log negative and fall back to g128-head. Numbers shortly. | |
Xet Storage Details
- Size:
- 973 Bytes
- Xet hash:
- 2ec0860dd7aaf1c2d8566930f77f6b7d463b67c5047ec9a06fcac2c35e43a405
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.