yitongl's picture
Upload folder using huggingface_hub
0692312 verified
|
Raw
History Blame Contribute Delete
6.9 kB
---
license: other
tags:
- minimax-h3
- svdquant
- nvfp4
- quantization
---
# MiniMax-H3 NVFP4 β€” what the smoothing scale should be calibrated on
Five end-to-end 50-step generations, one BF16 reference, same prompt and same seed (1101), all on
one RTX 5090. They differ in **one thing**: which rows of the packed `[text | video | audio]`
sequence were used to derive the per-input-channel smoothing scale `lambda`.
The question they answer: H3 runs full self-attention over a packed sequence of three modalities,
but a linear layer has one weight, so `W * lambda` is shared by all three. Per-channel activation
profiles differ per modality β€” measured correlations between them are near zero or negative β€” so a
single `lambda` cannot be right for all of them. **Is favouring video worth it?**
Answer: **no.** `lambda` calibrated on video alone is the worst of the four smoothed configs
end-to-end. Calibrating on all rows is the right default.
## The runs
| # | file | lambda from | LoRA | lambda spread (p50 max/min) |
| --- | --- | --- | --- | --- |
| 0 | `0_bf16.mp4` | β€” | β€” | β€” |
| 1 | `1_plain_w4a4.mp4` | none (`lambda = 1`) | none | 1.0 |
| 2 | `2_lambda_none.mp4` | none (`lambda = 1`) | rank 32 | 1.0 |
| 3 | `3_lambda_all.mp4` | all rows | rank 32 | 3.1 |
| 4 | `4_lambda_video.mp4` | video rows | rank 32 | 5.9 |
| 5 | `5_lambda_text.mp4` | text rows | rank 32 | 1.6 |
`grid.mp4` is all of them on one timeline, captioned.
## Read the numbers with care
`RESULTS.md` has the table. One column in it is easy to misread, so it is worth saying here:
**Mean RGB is not a quality metric between two working quantizations.** Every run shares BF16's
seed and therefore its initial noise. A faithful quantization stays in the *same* sample; one that
perturbs the trajectory enough lands in a different β€” and perfectly plausible β€” sample. Run 4 is a
coherent video of a different scene, not a broken one, and its large mean-RGB delta is reporting
the scene change, not damage. `corr` (frame-aligned against BF16) is the column that separates
"same scene, degraded" from "different scene".
Mean RGB *was* the right instrument earlier, when the failure being chased was a genuinely
washed-out output from a misapplied `smooth_factor` permutation. It stopped being the right
instrument the moment every candidate started producing a real video.
## The metric that picks the winner also hides the cost
The calibration's own error column scores each candidate `lambda` on rows drawn uniformly from the
packed sequence β€” which is what `deepcompressor`'s `OutputsError` objective specifies. Uniform
means *proportional*, and video is 98.6% of the rows. So that column ranks the runs like this
(relative L2 vs bf16, median over 312 layers, rank-32 branch included):
| | lambda=1 | lambda=1 +LoRA | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- | --- | --- |
| calibration's own column | 0.1019 | 0.0951 | 0.0942 | **0.0903** | 0.0953 |
**`lambda=video` wins that metric and loses end to end.** The metric is not wrong; it is answering
a question about the 98.6%, and the damage is in the other 1.4%.
Re-scoring with the same quantizer on rows sampled *per modality*
(`scripts/calib_sample_modal.py` + `scripts/score_modal_lambda.py`, 4096 rows per modality per
layer, 312 layers) shows what the proportional column could not:
| modality | lambda=1 | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- | --- |
| video | 0.0980 | 0.0976 | **0.0937** | 0.0981 |
| text | 0.0891 | **0.0871** | 0.1140 | **0.0871** |
| audio | 0.0875 | **0.0857** | 0.0975 | 0.0874 |
Change against no smoothing (negative = helps):
| modality | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- |
| video | βˆ’0.4 % | **βˆ’4.3 %** | +0.1 % |
| text | βˆ’2.2 % | **+27.9 %** | βˆ’2.2 % |
| audio | βˆ’2.1 % | **+11.5 %** | βˆ’0.2 % |
`lambda=all` is the only column negative in all three rows. `lambda=video` buys video 4.3 % and
charges text 27.9 % and audio 11.5 % for it. `lambda=text` is free but pointless β€” it does nothing
for video.
The worst layers make the trade obvious. `lambda=video` vs `lambda=all`, text error:
| layer | text | video |
| --- | --- | --- |
| `blocks.13.ff.net.2` | 0.0267 β†’ **0.3396** (12.7x) | 0.0886 β†’ 0.0876 |
| `blocks.12.ff.net.2` | 0.0702 β†’ **0.1429** | 0.0898 β†’ 0.0897 |
| `blocks.8.ff.net.2` | 0.0807 β†’ **0.1105** | 0.0951 β†’ 0.0875 |
| `blocks.33.attn.to_q` | 0.0702 β†’ **0.0955** | 0.0669 β†’ 0.0657 |
A 12.7x text error for a 1 % video gain, in one layer.
### Why
NVFP4 puts 16 consecutive input channels under one FP8 scale set by that group's absmax, so a
channel whose own absmax is far below its group's loses `log2(group_absmax / channel_absmax)`
bits, and `lambda` reshapes exactly that profile because the kernel sees `X / lambda`. The
per-modality channel profiles are uncorrelated to anti-correlated (`b0.to_q` video~text βˆ’0.066;
`b25.to_v` video~audio βˆ’0.390), and **dividing by a vector uncorrelated with your own profile
sharpens it rather than flattening it**. The statistics-only proxy
(`scripts/diag_lambda_crossmodal.py`, bits lost, no GPU) agrees with the measurement above: video
βˆ’0.437 bits, text +0.216, audio +0.243 under `lambda=video`.
Text is ~1.4% of the rows, but under full self-attention every video row attends to it. Degrading
the conditioning degrades the video conditioned on it β€” which is how the per-layer video error can
fall while the generated video gets worse.
## Calibration
Same as `minimax-h3-svdquant-calib`: 128 video prompts, 768x1344, 124 frames, 50 steps, rank 32,
39-candidate lambda grid, 100 iterations of low-rank refit against `OutputsError`. The only thing
varied across runs 2–5 is which modality mask the activation statistics were accumulated under
(`scripts/calib_stats_modal.py`), and hence which `absmax` feeds the lambda grid.
Run 1 is `lambda = 1` with the low-rank branch zeroed β€” the floor, plain W4A4.
Run 2 is `lambda = 1` with the rank-32 branch refit against the raw weight β€” it isolates the
low-rank contribution from the smoothing contribution.
Base model: `MiniMaxAI/MiniMax-H3` @ `bfc8ed0353f5a9733be73e6b2c98ec0948195b86`.
## One thing a consumer must not get wrong
`SVDQW4A4Linear.smooth_factor` is addressed by the kernel in **MMA-interleaved** channel order,
while every other tensor picks that interleave up from `NunchakuWeightPacker`. Writing `lambda` in
natural order gives 12 of every 16 channels another channel's lambda. Details and the permutation
are in the `minimax-h3-svdquant-calib` README. Everything here was produced with that fix applied.
Note the interaction with this page: **a sharper lambda makes that bug worse**, so before the fix,
`lambda=video` looked catastrophic and `lambda=text` looked fine β€” for a reason that had nothing to
do with modality.