File size: 6,897 Bytes
0692312
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
---
license: other
tags:
  - minimax-h3
  - svdquant
  - nvfp4
  - quantization
---

# MiniMax-H3 NVFP4 β€” what the smoothing scale should be calibrated on

Five end-to-end 50-step generations, one BF16 reference, same prompt and same seed (1101), all on
one RTX 5090. They differ in **one thing**: which rows of the packed `[text | video | audio]`
sequence were used to derive the per-input-channel smoothing scale `lambda`.

The question they answer: H3 runs full self-attention over a packed sequence of three modalities,
but a linear layer has one weight, so `W * lambda` is shared by all three. Per-channel activation
profiles differ per modality β€” measured correlations between them are near zero or negative β€” so a
single `lambda` cannot be right for all of them. **Is favouring video worth it?**

Answer: **no.** `lambda` calibrated on video alone is the worst of the four smoothed configs
end-to-end. Calibrating on all rows is the right default.

## The runs

| # | file | lambda from | LoRA | lambda spread (p50 max/min) |
| --- | --- | --- | --- | --- |
| 0 | `0_bf16.mp4` | β€” | β€” | β€” |
| 1 | `1_plain_w4a4.mp4` | none (`lambda = 1`) | none | 1.0 |
| 2 | `2_lambda_none.mp4` | none (`lambda = 1`) | rank 32 | 1.0 |
| 3 | `3_lambda_all.mp4` | all rows | rank 32 | 3.1 |
| 4 | `4_lambda_video.mp4` | video rows | rank 32 | 5.9 |
| 5 | `5_lambda_text.mp4` | text rows | rank 32 | 1.6 |

`grid.mp4` is all of them on one timeline, captioned.

## Read the numbers with care

`RESULTS.md` has the table. One column in it is easy to misread, so it is worth saying here:

**Mean RGB is not a quality metric between two working quantizations.** Every run shares BF16's
seed and therefore its initial noise. A faithful quantization stays in the *same* sample; one that
perturbs the trajectory enough lands in a different β€” and perfectly plausible β€” sample. Run 4 is a
coherent video of a different scene, not a broken one, and its large mean-RGB delta is reporting
the scene change, not damage. `corr` (frame-aligned against BF16) is the column that separates
"same scene, degraded" from "different scene".

Mean RGB *was* the right instrument earlier, when the failure being chased was a genuinely
washed-out output from a misapplied `smooth_factor` permutation. It stopped being the right
instrument the moment every candidate started producing a real video.

## The metric that picks the winner also hides the cost

The calibration's own error column scores each candidate `lambda` on rows drawn uniformly from the
packed sequence β€” which is what `deepcompressor`'s `OutputsError` objective specifies. Uniform
means *proportional*, and video is 98.6% of the rows. So that column ranks the runs like this
(relative L2 vs bf16, median over 312 layers, rank-32 branch included):

| | lambda=1 | lambda=1 +LoRA | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- | --- | --- |
| calibration's own column | 0.1019 | 0.0951 | 0.0942 | **0.0903** | 0.0953 |

**`lambda=video` wins that metric and loses end to end.** The metric is not wrong; it is answering
a question about the 98.6%, and the damage is in the other 1.4%.

Re-scoring with the same quantizer on rows sampled *per modality*
(`scripts/calib_sample_modal.py` + `scripts/score_modal_lambda.py`, 4096 rows per modality per
layer, 312 layers) shows what the proportional column could not:

| modality | lambda=1 | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- | --- |
| video | 0.0980 | 0.0976 | **0.0937** | 0.0981 |
| text | 0.0891 | **0.0871** | 0.1140 | **0.0871** |
| audio | 0.0875 | **0.0857** | 0.0975 | 0.0874 |

Change against no smoothing (negative = helps):

| modality | lambda=all | lambda=video | lambda=text |
| --- | --- | --- | --- |
| video | βˆ’0.4 % | **βˆ’4.3 %** | +0.1 % |
| text | βˆ’2.2 % | **+27.9 %** | βˆ’2.2 % |
| audio | βˆ’2.1 % | **+11.5 %** | βˆ’0.2 % |

`lambda=all` is the only column negative in all three rows. `lambda=video` buys video 4.3 % and
charges text 27.9 % and audio 11.5 % for it. `lambda=text` is free but pointless β€” it does nothing
for video.

The worst layers make the trade obvious. `lambda=video` vs `lambda=all`, text error:

| layer | text | video |
| --- | --- | --- |
| `blocks.13.ff.net.2` | 0.0267 β†’ **0.3396** (12.7x) | 0.0886 β†’ 0.0876 |
| `blocks.12.ff.net.2` | 0.0702 β†’ **0.1429** | 0.0898 β†’ 0.0897 |
| `blocks.8.ff.net.2` | 0.0807 β†’ **0.1105** | 0.0951 β†’ 0.0875 |
| `blocks.33.attn.to_q` | 0.0702 β†’ **0.0955** | 0.0669 β†’ 0.0657 |

A 12.7x text error for a 1 % video gain, in one layer.

### Why

NVFP4 puts 16 consecutive input channels under one FP8 scale set by that group's absmax, so a
channel whose own absmax is far below its group's loses `log2(group_absmax / channel_absmax)`
bits, and `lambda` reshapes exactly that profile because the kernel sees `X / lambda`. The
per-modality channel profiles are uncorrelated to anti-correlated (`b0.to_q` video~text βˆ’0.066;
`b25.to_v` video~audio βˆ’0.390), and **dividing by a vector uncorrelated with your own profile
sharpens it rather than flattening it**. The statistics-only proxy
(`scripts/diag_lambda_crossmodal.py`, bits lost, no GPU) agrees with the measurement above: video
βˆ’0.437 bits, text +0.216, audio +0.243 under `lambda=video`.

Text is ~1.4% of the rows, but under full self-attention every video row attends to it. Degrading
the conditioning degrades the video conditioned on it β€” which is how the per-layer video error can
fall while the generated video gets worse.

## Calibration

Same as `minimax-h3-svdquant-calib`: 128 video prompts, 768x1344, 124 frames, 50 steps, rank 32,
39-candidate lambda grid, 100 iterations of low-rank refit against `OutputsError`. The only thing
varied across runs 2–5 is which modality mask the activation statistics were accumulated under
(`scripts/calib_stats_modal.py`), and hence which `absmax` feeds the lambda grid.

Run 1 is `lambda = 1` with the low-rank branch zeroed β€” the floor, plain W4A4.
Run 2 is `lambda = 1` with the rank-32 branch refit against the raw weight β€” it isolates the
low-rank contribution from the smoothing contribution.

Base model: `MiniMaxAI/MiniMax-H3` @ `bfc8ed0353f5a9733be73e6b2c98ec0948195b86`.

## One thing a consumer must not get wrong

`SVDQW4A4Linear.smooth_factor` is addressed by the kernel in **MMA-interleaved** channel order,
while every other tensor picks that interleave up from `NunchakuWeightPacker`. Writing `lambda` in
natural order gives 12 of every 16 channels another channel's lambda. Details and the permutation
are in the `minimax-h3-svdquant-calib` README. Everything here was produced with that fix applied.

Note the interaction with this page: **a sharper lambda makes that bug worse**, so before the fix,
`lambda=video` looked catastrophic and `lambda=text` looked fine β€” for a reason that had nothing to
do with modality.