--- license: other tags: - minimax-h3 - svdquant - nvfp4 - quantization --- # MiniMax-H3 NVFP4 — what the smoothing scale should be calibrated on Five end-to-end 50-step generations, one BF16 reference, same prompt and same seed (1101), all on one RTX 5090. They differ in **one thing**: which rows of the packed `[text | video | audio]` sequence were used to derive the per-input-channel smoothing scale `lambda`. The question they answer: H3 runs full self-attention over a packed sequence of three modalities, but a linear layer has one weight, so `W * lambda` is shared by all three. Per-channel activation profiles differ per modality — measured correlations between them are near zero or negative — so a single `lambda` cannot be right for all of them. **Is favouring video worth it?** Answer: **no.** `lambda` calibrated on video alone is the worst of the four smoothed configs end-to-end. Calibrating on all rows is the right default. ## The runs | # | file | lambda from | LoRA | lambda spread (p50 max/min) | | --- | --- | --- | --- | --- | | 0 | `0_bf16.mp4` | — | — | — | | 1 | `1_plain_w4a4.mp4` | none (`lambda = 1`) | none | 1.0 | | 2 | `2_lambda_none.mp4` | none (`lambda = 1`) | rank 32 | 1.0 | | 3 | `3_lambda_all.mp4` | all rows | rank 32 | 3.1 | | 4 | `4_lambda_video.mp4` | video rows | rank 32 | 5.9 | | 5 | `5_lambda_text.mp4` | text rows | rank 32 | 1.6 | `grid.mp4` is all of them on one timeline, captioned. ## Read the numbers with care `RESULTS.md` has the table. One column in it is easy to misread, so it is worth saying here: **Mean RGB is not a quality metric between two working quantizations.** Every run shares BF16's seed and therefore its initial noise. A faithful quantization stays in the *same* sample; one that perturbs the trajectory enough lands in a different — and perfectly plausible — sample. Run 4 is a coherent video of a different scene, not a broken one, and its large mean-RGB delta is reporting the scene change, not damage. `corr` (frame-aligned against BF16) is the column that separates "same scene, degraded" from "different scene". Mean RGB *was* the right instrument earlier, when the failure being chased was a genuinely washed-out output from a misapplied `smooth_factor` permutation. It stopped being the right instrument the moment every candidate started producing a real video. ## The metric that picks the winner also hides the cost The calibration's own error column scores each candidate `lambda` on rows drawn uniformly from the packed sequence — which is what `deepcompressor`'s `OutputsError` objective specifies. Uniform means *proportional*, and video is 98.6% of the rows. So that column ranks the runs like this (relative L2 vs bf16, median over 312 layers, rank-32 branch included): | | lambda=1 | lambda=1 +LoRA | lambda=all | lambda=video | lambda=text | | --- | --- | --- | --- | --- | --- | | calibration's own column | 0.1019 | 0.0951 | 0.0942 | **0.0903** | 0.0953 | **`lambda=video` wins that metric and loses end to end.** The metric is not wrong; it is answering a question about the 98.6%, and the damage is in the other 1.4%. Re-scoring with the same quantizer on rows sampled *per modality* (`scripts/calib_sample_modal.py` + `scripts/score_modal_lambda.py`, 4096 rows per modality per layer, 312 layers) shows what the proportional column could not: | modality | lambda=1 | lambda=all | lambda=video | lambda=text | | --- | --- | --- | --- | --- | | video | 0.0980 | 0.0976 | **0.0937** | 0.0981 | | text | 0.0891 | **0.0871** | 0.1140 | **0.0871** | | audio | 0.0875 | **0.0857** | 0.0975 | 0.0874 | Change against no smoothing (negative = helps): | modality | lambda=all | lambda=video | lambda=text | | --- | --- | --- | --- | | video | −0.4 % | **−4.3 %** | +0.1 % | | text | −2.2 % | **+27.9 %** | −2.2 % | | audio | −2.1 % | **+11.5 %** | −0.2 % | `lambda=all` is the only column negative in all three rows. `lambda=video` buys video 4.3 % and charges text 27.9 % and audio 11.5 % for it. `lambda=text` is free but pointless — it does nothing for video. The worst layers make the trade obvious. `lambda=video` vs `lambda=all`, text error: | layer | text | video | | --- | --- | --- | | `blocks.13.ff.net.2` | 0.0267 → **0.3396** (12.7x) | 0.0886 → 0.0876 | | `blocks.12.ff.net.2` | 0.0702 → **0.1429** | 0.0898 → 0.0897 | | `blocks.8.ff.net.2` | 0.0807 → **0.1105** | 0.0951 → 0.0875 | | `blocks.33.attn.to_q` | 0.0702 → **0.0955** | 0.0669 → 0.0657 | A 12.7x text error for a 1 % video gain, in one layer. ### Why NVFP4 puts 16 consecutive input channels under one FP8 scale set by that group's absmax, so a channel whose own absmax is far below its group's loses `log2(group_absmax / channel_absmax)` bits, and `lambda` reshapes exactly that profile because the kernel sees `X / lambda`. The per-modality channel profiles are uncorrelated to anti-correlated (`b0.to_q` video~text −0.066; `b25.to_v` video~audio −0.390), and **dividing by a vector uncorrelated with your own profile sharpens it rather than flattening it**. The statistics-only proxy (`scripts/diag_lambda_crossmodal.py`, bits lost, no GPU) agrees with the measurement above: video −0.437 bits, text +0.216, audio +0.243 under `lambda=video`. Text is ~1.4% of the rows, but under full self-attention every video row attends to it. Degrading the conditioning degrades the video conditioned on it — which is how the per-layer video error can fall while the generated video gets worse. ## Calibration Same as `minimax-h3-svdquant-calib`: 128 video prompts, 768x1344, 124 frames, 50 steps, rank 32, 39-candidate lambda grid, 100 iterations of low-rank refit against `OutputsError`. The only thing varied across runs 2–5 is which modality mask the activation statistics were accumulated under (`scripts/calib_stats_modal.py`), and hence which `absmax` feeds the lambda grid. Run 1 is `lambda = 1` with the low-rank branch zeroed — the floor, plain W4A4. Run 2 is `lambda = 1` with the rank-32 branch refit against the raw weight — it isolates the low-rank contribution from the smoothing contribution. Base model: `MiniMaxAI/MiniMax-H3` @ `bfc8ed0353f5a9733be73e6b2c98ec0948195b86`. ## One thing a consumer must not get wrong `SVDQW4A4Linear.smooth_factor` is addressed by the kernel in **MMA-interleaved** channel order, while every other tensor picks that interleave up from `NunchakuWeightPacker`. Writing `lambda` in natural order gives 12 of every 16 channels another channel's lambda. Details and the permutation are in the `minimax-h3-svdquant-calib` README. Everything here was produced with that fix applied. Note the interaction with this page: **a sharper lambda makes that bug worse**, so before the fix, `lambda=video` looked catastrophic and `lambda=text` looked fine — for a reason that had nothing to do with modality.