File size: 17,132 Bytes
fdc6474
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
# GLM-5.2 AQLM convergence plan (executable runbook)

Goal: (1) finish the lite activation-aware convergence, rebuild + validate +
upload all three checkpoints β€” take the win; (2) run full AQLM (full-Hessian
beam encoding + blockwise PV-tuning against a streamed BF16 teacher), then
rebuild + validate + upload again. Sanity gates between every step; a mixed
NVFP4+AQLM checkpoint must LOAD and INFER before anything ships.

This runbook assumes no context beyond this file. Follow it top to bottom.
When a gate FAILS: stop, do not delete anything, record the failing output
in /data/glm52-RUNLOG.md, and fix or escalate before proceeding.

--------------------------------------------------------------------------
## 0. Ground rules and environment

Every shell that touches vLLM must run this preamble:

```bash
cd /home/coder/git/glm52/vllm && source ../.venv/bin/activate
export CUDA_HOME=/home/coder/git/glm52/.venv/lib/python3.12/site-packages/nvidia/cu13
```

Standing serve env (the ONLY blessed recipe; 8xB200 emulating 4x96GB):

```bash
export VLLM_PP_LAYER_PARTITION="21,19,19,19"
export NCCL_MAX_NCHANNELS=4 NCCL_BUFFSIZE=1048576
export VLLM_SPARSE_INDEXER_MAX_LOGITS_MB=256
SERVE_FLAGS="--pipeline-parallel-size 4 --gpu-memory-utilization 0.509 \
  --kv-cache-dtype fp8_ds_mla --max-num-seqs 2 \
  --max-num-batched-tokens 2048 --enforce-eager --port 8199"
```

Hard-won footguns β€” DO NOT repeat these mistakes:
- NEVER `pkill -f <pattern>` where <pattern> appears in your own command
  line or in a launch you make in the same Bash call (it kills itself /
  the new launch). Kill in one call, verify GPUs free with nvidia-smi in
  the same call, launch in a SEPARATE call.
- /tmp is volatile on this box (was wiped once). Teacher downloads live in
  /tmp/glm52-hot-dl2 β€” if missing, regenerate via
  `tools/make_hot_manifest2.py` + `tools/range_download.py` (manifest
  source /data/glm52-need-experts.json).
- `hf upload-large-folder` only targets a repo's main branch. It is
  resumable; rerun on failure.
- A vLLM serve that dies with "Engine core initialization failed" has the
  real error higher up in the log: grep for `ValueError|KeyError` first.
- Long jobs: run_in_background + `until <condition>; do sleep 30; done`
  waiters. Never a bare long sleep.

Checkpoint inventory (do not delete any of these):

| path | what |
|---|---|
| /data/glm52 | LIVE 1M two-tier (v5): 30% hot NVFP4 / 70% cold 2-bpw AQLM |
| /data/glm52-500k, /data/glm52-250k | variants: 48% / 57% hot |
| /data/glm52-v4-uniform | pre-demotion two-tier (hot-array superset source) |
| /data/glm52-v3-3tier, /data/glm52-old-layerwise | older tiers; old-layerwise also = teacher for layers 3,4,5,8,74-77 (all-256 per-expert NVFP4) |
| /data/glm52-aqlm-parts | init-grade AQLM parts, ALL layers, 2-book w2 |
| /data/glm52-aqlm-conv | converger output (this plan, phase 1) |
| /data/glm52-acts | calibration activations, 24k routed tokens/layer |
| /data/glm52-expert-stats-v2.npz, /data/glm52-expert-assignment*.json | routing stats + assignments |
| /data/glm52-sm120-golden | golden bundle (must be RE-CAPTURED whenever weights change) |
| /tmp/glm52-hot-dl2 | teacher NVFP4 regions (cold experts, 67 layers) |

HF repos (public, overwrite in place):
jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid {,-500k,-250k}

--------------------------------------------------------------------------
## 1. Sanity-check toolbox (reusable gates)

Run gates in this order; each assumes the previous passed. "CKPT" = the
checkpoint directory under test.

### SC-1 Schema check (seconds, no GPU)
```bash
python tools/sanity/sc1_schema.py CKPT
```
Verifies: index total_size == sum of shard tensor bytes; every layer 3..77
has the full two-tier tensor set with shapes consistent with
config.json aqlm_layer_books (n_nvfp4+n_cold=256, n_base=0); hyb_kind
counts match; no NaN/Inf in any fp16/fp32 tensor sampled per layer;
codebooks are fp16 [1,65536,8]; codes int16. PASS = prints `SC1 PASS`.

### SC-2 Dequant statistics (1 GPU, ~2 min)
```bash
python tools/sanity/sc2_dequant_stats.py CKPT --layers 3,21,40,60,77
```
For each listed layer: dequantize 4 cold experts (AQLM) and 4 hot experts
(NVFP4) to fp16; check per-tensor RMS in [1e-3, 1.0], zero-fraction < 30%,
no NaN; cosine similarity of AQLM dequant vs the pure-torch reference == 1.
PASS = `SC2 PASS`.

### SC-3 Kernel tests (1 GPU, ~1 min)
```bash
CUDA_VISIBLE_DEVICES=7 python -m pytest tests/kernels/quantization/test_aqlm_moe.py -q
```
PASS = 32 passed.

### SC-4 Load + short-context serve smoke (4 GPUs, ~6 min)
```bash
timeout 900 python -m vllm.entrypoints.cli.main serve CKPT $SERVE_FLAGS \
  --max-model-len 8192 > /tmp/sc4.log 2>&1 &
until grep -qE "startup complete|initialization failed" /tmp/sc4.log; do sleep 15; done
```
PASS = "Application startup complete" and no Traceback. Leave running for
SC-5.

### SC-5 Coherence probes (against SC-4 server, ~1 min)
```bash
python tools/validate_serve.py --port 8199
```
PASS = all three completions coherent (Paris / correct fibonacci or
quicksort / H2O), decode >= 10 tok/s. A model with broken cold experts
produces repetitive garbage here β€” this is the primary "weights kaput"
detector.

### SC-6 Perplexity delta (against SC-4 server, ~5 min)
```bash
python tools/sanity/sc6_ppl.py --port 8199 --ref /data/glm52-heldout.txt
```
Teacher-forced logprob over ~50k held-out tokens (code+prose+medical,
NOT in the calibration set) via the completions API with echo/logprobs.
Record ppl in RUNLOG. PASS rule: after any requantization, ppl must be
<= previous shipped ppl + 1% (phase-1) / must IMPROVE (phase-2 gates).

### SC-7 Full-context revalidation (4 GPUs, ~15 min) β€” 1M checkpoint only
Serve with --max-model-len 1048576 (full SERVE_FLAGS recipe), then:
```bash
python tools/validate_serve.py --port 8199 --long 200000
```
PASS = KV >= 1,048,576 tokens; needle answer contains BLUEBERRY42;
worst GPU <= 97,887 MiB.

### SC-8 Golden e2e comparison (against SC-4 server)
```bash
python tools/verify_sm120.py /data/glm52-sm120-golden --port 8199 --stages ""
```
Compares greedy generations vs stored goldens. NOTE: after INTENTIONAL
weight changes text may legitimately differ; the check is that outputs are
coherent and >= 40/50 top-50 logprob overlap on early steps. After each
shipped rebuild, RE-CAPTURE goldens (tools/capture_golden.py + 
make_kernel_vectors.py) so the bundle matches shipped weights.

### Gate bundles
- GATE-A (any rebuilt checkpoint): SC-1, SC-2, SC-4, SC-5
- GATE-B (before any upload): GATE-A + SC-6 recorded + (1M only) SC-7
- GATE-C (after any fitting phase, per ~8 layers): spot SC-2 on parts +
  fitting-metric monotonicity (err_after < err_before on every layer;
  any layer where err_after > err_before*0.999 -> refit that layer)

TODO(first task): tools/sanity/sc1_schema.py, sc2_dequant_stats.py,
sc6_ppl.py and the held-out set /data/glm52-heldout.txt do not exist yet.
Build them exactly to the contracts above (~200 lines total; reuse
_dequant_reference from nvfp4_aqlm_hybrid.py and the FP4 LUT from
tools/aqlm_quantize.py). Held-out: take 25 files from vllm docs/ + 15
MedQA paragraphs + 10 code files NOT matched by the corpus builder's
random.Random(42) selection; ~50k tokens total.

--------------------------------------------------------------------------
## 2. PHASE 1 β€” finish lite convergence, ship it (est. 6-9 h wall)

### 1.1 Confirm smoke, then launch the full run
Smoke (layer 40, GPU 4) is running; on completion check
`/data/glm52-aqlm-conv/smoke.log` shows `saved` and err_after < err_before
for both w13 and w2. Then:
```bash
# GPUs must be idle first (nvidia-smi). Launch (resumable; skips done layers):
python tools/aqlm_converge.py > /data/glm52-aqlm-conv/run.log 2>&1 &
# progress: grep -c "saved" /data/glm52-aqlm-conv/run.log  (target 75)
```
~10-20 min/layer/GPU => 75 layers on 8 GPUs β‰ˆ 2-4 h. Monitor with a
tail|grep on "saved|Error|Traceback|OutOfMemory". If a worker OOMs:
rerun with that layer alone on an idle GPU (memory fragmentation clears).

### 1.2 GATE-C on the parts
All 75 layer files present; every layer's `*_err_after < *_err_before`;
`python tools/sanity/sc2_dequant_stats.py --parts /data/glm52-aqlm-conv`
(parts mode: reference-dequant a few experts per layer; RMS/NaN checks).

### 1.3 Rebuild the three checkpoints (cold arrays only)
Write tools/build_checkpoint_v7.py (clone of v6's writer): for each target
in {/data/glm52, /data/glm52-500k, /data/glm52-250k}: stream every shard;
copy all tensors EXCEPT layers' `w13_codes|w13_codebooks|w13_scales|
w2c_codes|w2c_codebooks|w2c_scales`, which are replaced by slicing the
conv parts: `sel = positions of the target's cold ids (from its hyb_kind)
within parts.expert_ids` (parts cover the 1M cold set = superset; assert
every target cold id is found). Write to <target>-conv, then GATE-A it,
then swap: `mv <target> <target>-preconv && mv <target>-conv <target>`.
Order: 1M first, then 500k, then 250k.

### 1.4 GATE-B per checkpoint
1M: full GATE-B incl SC-7. 500k/250k: GATE-A + SC-6 (serve at
--max-model-len 8192 is fine for the gate; partitions unchanged).
Record all ppl values in RUNLOG. Abort rule: if 1M ppl regresses > 1%
vs pre-conv, STOP β€” swap back (`mv` reversal) and investigate.

### 1.5 Re-capture goldens (weights changed)
```bash
python tools/capture_golden.py && python tools/make_kernel_vectors.py
```
(armed-marker flow is automatic; outputs overwrite /data/glm52-sm120-golden.)

### 1.6 Upload β€” DEFERRED (revised 2026-07-07)
Do NOT upload here. The gated conv checkpoints are the fallback artifacts;
proceed directly to phase 1.5 (REAP re-tier), gate it, then upload ONCE:
the REAP version if its gates pass (ppl <= conv ppl), else the conv
version. This halves upload traffic; the win still ships strictly better
than what is currently on HF. Original upload commands (run after 1.5):
```bash
hf upload-large-folder jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid      /data/glm52      --repo-type model --num-workers 12
hf upload-large-folder jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid-500k /data/glm52-500k --repo-type model --num-workers 12
hf upload-large-folder jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid-250k /data/glm52-250k --repo-type model --num-workers 12
```
Sequential; each ends with `committed: N/N`. Also `hf upload ... /data/glm52/code code`
if tools changed, and update model-card READMEs to note the
activation-aware convergence. Verify each repo's file count via HfApi.

--------------------------------------------------------------------------
## 2.5 PHASE 1.5 β€” REAP-based expert re-selection (est. 4-6 h, after the
phase-1 upload so the win ships first)

Motivation: the shipped hot/cold split ranks experts by routing FREQUENCY
(bincount of topk_ids). That demotes rarely-fired experts whose outputs are
large and decisive when they do fire. Adopt REAP's router-weighted
activation saliency, adapted for precision demotion (not pruning):

    score_e = SUM_t  g_{t,e} * ||f_e(x_t)||_2   *   relerr_e
      g       = router weight (topk_weights, captured in /data/glm52-acts)
      f_e(x)  = down(silu(gate x) * up x) with TEACHER weights
      relerr_e= h-weighted 2-bit reconstruction error of expert e
                (from conv/full parts; near-uniform, catches fragile tails)

Steps:
a. tools/score_experts_reap.py: per layer load acts + teachers for ALL 256
   experts (hot: dequant the live checkpoint's nvfp4_* arrays; cold:
   /tmp/glm52-hot-dl2 regions / old-layerwise), compute score_e ->
   /data/glm52-reap-scores.npz [75,256]. 8-GPU layer-parallel, ~2 h.
   Sanity: per layer, spearman(score, old counts) printed β€” expect 0.4-0.8
   (correlated but meaningfully different); investigate if < 0.2.
b. Re-solve assignments for all three budgets with the same byte-exact
   solver, ranking by score (floor 8/cap 176). Print overlap vs current
   hot sets; expect 60-85%.
c. Fit AQLM (converger, warm start from full codebook of that layer) for
   newly-cold experts only; promote newly-hot from teacher bytes already
   on disk (cold-teacher superset β€” no downloads needed).
d. Rebuild all three (build_checkpoint_v7 slicing by new assignment),
   GATE-A each, GATE-B the 1M; ppl must improve or match (this changes
   WHICH experts are 2-bit, same byte budget β€” if ppl regresses, the REAP
   scores or fit are wrong: stop and investigate, keep -prereap dirs).
e. This assignment becomes phase 2's fit universe (its cold sets).
f. UPLOAD NOW (single combined upload of phase 1 + 1.5): the three repos,
   using the commands in Β§1.6, plus the code bundle; verify file counts;
   record commit URLs in RUNLOG.

--------------------------------------------------------------------------
## 3. PHASE 2 β€” full AQLM (est. 2.5-3.5 days wall)

### 2.0 Staging (CPU/network only β€” run DURING phase 1, no GPU contention)
a. Calibration corpus v3, ~15M tokens, saved as token-id shards
   /data/glm52-calib-v3/*.npy: reuse corpus builder mix + download
   supplements (HF datasets: a code sample, an instruction set, keep the
   MedQA textbooks; ~30-50 GB). Hold out 50k tokens -> /data/glm52-heldout.txt.
b. BF16 teacher streamer: tools/bf16_stream.py with contract
   `get_expert(layer, expert, proj) -> bf16 tensor`, backed by ranged
   HTTP reads of zai-org/GLM-5.2 (index + shard headers cached once;
   LRU on-disk cache /data/bf16-cache, cap 400 GB). Verify vs NVFP4
   teacher: cosine > 0.98 on 10 random experts.
c. Baseline evals: SC-6 ppl of the phase-1-shipped 1M checkpoint +
   teacher NVFP4 model (serve lukealonso config on 8 GPUs PP=8 briefly);
   plus 20-problem HumanEval subset + 50-problem GSM8K subset through the
   server (tools/sanity/bench_small.py; greedy; record scores in RUNLOG).

### 2.1 Full-Hessian beam encoding (days 1-2 of phase 2)
tools/aqlm_full.py, upgrade of aqlm_converge.py per (layer, projection):
- Hessians: per-expert FULL H = X_e^T X_e (w13: 6144^2 fp32 = 151 MB
  transient per expert; w2: 2048^2). Accumulate from calib-v3 activations
  (re-capture acts at 128k tokens/layer with the existing hook β€” one
  ~30-min PP=4 pass over calib-v3). Damping: H += 1e-2*mean(diag)*I.
- Teacher: BF16 via streamer (fallback NVFP4 regions if a fetch fails).
- Encode: GPTQ-order sequential CD with error feedback within each row
  (process groups in descending diag(H) order; after fixing a group,
  propagate residual via H off-diagonal to remaining groups), candidates
  by beam-4: top-4 codebook entries under the diag metric (GEMM+topk),
  exact H-scored selection among the 4.
- Alternate with weighted codebook update + scale refit (as lite), 3 outer
  iterations. Early-stop per layer when H-weighted err improves < 0.1%.
- Output: /data/glm52-aqlm-full/layer_N.pt (same schema as conv parts).
- GATE-C every 8 layers; on any layer regression vs conv parts, keep the
  conv version for that layer (per-layer best-of).
Budget check: must average <= 3.5 h/layer on one GPU (75 layers/8 GPUs/
2 days); if the first 4 layers exceed it, reduce beam to 2 and/or
subsample H tokens; record the decision.

### 2.2 Blockwise PV-tuning (days 2-3.5 of phase 2)
tools/pv_tune.py per transformer block (layer-parallel, 1 block/GPU):
- Student block: BF16 non-expert weights + hot NVFP4 dequantized frozen +
  cold experts as differentiable dequant (codebook gather x scales);
  trainable: codebooks, scales (cold), NOTHING else. Teacher block: same
  block with BF16 streamed experts everywhere.
- Data: block inputs recorded once per phase (run calib-v3 through the
  phase-2.1 model with a hidden-state capture hook at each block boundary,
  save 256k tokens per block boundary, bf16, ~3 GB/block).
- Loss: MSE(student_out, teacher_out) token-weighted by router prob mass;
  Adam lr 1e-4 (codebooks) / 1e-3 (scales), bs 4096 tokens, ~600 steps;
  every 200 steps: straight-through re-encode (beam-1 full-H) and reset
  optimizer state for reassigned entries. Early-stop on plateau.
- Output: /data/glm52-aqlm-pv/layer_N.pt. GATE-C per 8 blocks + one
  mid-phase GATE-A+SC-6 rebuild of the 1M checkpoint after ~half the
  blocks (catch systemic drift early; expect ppl improvement already).

### 2.3 Final rebuild + gates + upload
- Rebuild all three from pv parts (fallback per layer: pv > full > conv,
  choose best by held-out block-output error; record table in RUNLOG).
- GATE-B all three (1M incl SC-7 needle at 200k). Run bench_small.py:
  ppl must improve vs phase-1; HumanEval/GSM8K must not regress > 1 item.
- Re-capture goldens; update model cards ("PV-tuned AQLM"); upload all
  three repos + code bundle; verify file counts; final RUNLOG summary.

--------------------------------------------------------------------------
## 4. Rollback
Every swap keeps the previous directory as <target>-preconv / -prefull /
-prepv. HF keeps full commit history β€” revert = re-upload the kept dir or
`huggingface-cli` revert to a commit. Never delete a -pre* dir until the
next phase's GATE-B passes.

## 5. RUNLOG
Append every gate result, metric, decision, and anomaly to
/data/glm52-RUNLOG.md with a timestamp. The uploads' commit URLs go there
too. If context is lost, this file + PLAN.md are sufficient to resume.