Instructions to use yitongl/minimax-h3-nvfp4-lambda-modality with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MiniMax H3
How to use yitongl/minimax-h3-nvfp4-lambda-modality with MiniMax H3:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
| # Results | |
| All six clips: 768x1344, 124 frames, 50 steps, seed 1101, same prompt, one RTX 5090. | |
| Runs 1β5 are NVFP4 W4A4 through the Nunchaku runtime with the `smooth_factor` interleave fix. | |
| ## Offline, per modality β the measurement that decides it | |
| Relative L2 of the quantized layer output against bf16, median over 312 layers, 4096 sampled rows | |
| per modality per layer, rank-32 branch included, same NVFP4 fake-quantizer the calibration used. | |
| | modality | lambda=1 | lambda=all | lambda=video | lambda=text | | |
| | --- | --- | --- | --- | --- | | |
| | video | 0.0980 | 0.0976 | **0.0937** | 0.0981 | | |
| | text | 0.0891 | **0.0871** | 0.1140 | **0.0871** | | |
| | audio | 0.0875 | **0.0857** | 0.0975 | 0.0874 | | |
| vs no smoothing: `lambda=all` β0.4 % / β2.2 % / β2.1 %; `lambda=video` **β4.3 % / +27.9 % / | |
| +11.5 %**; `lambda=text` +0.1 % / β2.2 % / β0.2 %. | |
| `lambda=all` is the only candidate that is not negative-sum. This is 312 layers x 3 modalities of | |
| independent measurements and it is the basis for the recommendation. | |
| ## End to end, one prompt | |
| | run | mean RGB | per-frame loβhi | span | std | Ξmean | PSNR vs BF16 | corr vs BF16 | | |
| | --- | --- | --- | --- | --- | --- | --- | --- | | |
| | 0 BF16 | 36.27 | 27.78β45.96 | 18.18 | 43.90 | β | β | β | | |
| | 1 plain W4A4 | 28.42 | 27.40β29.11 | 1.71 | 41.60 | β7.85 | 15.02 | 0.438 | | |
| | 2 lambda=1 +LoRA | 30.43 | 25.50β33.72 | 8.21 | 42.76 | β5.84 | 15.12 | 0.436 | | |
| | 3 lambda=all | **36.67** | 31.64β43.28 | 11.64 | 46.14 | **+0.40** | 14.84 | **0.449** | | |
| | 4 lambda=video | 24.59 | 21.62β26.69 | 5.07 | 27.07 | β11.67 | 14.50 | 0.139 | | |
| | 5 lambda=text | 24.58 | 21.73β26.99 | 5.26 | 29.73 | β11.69 | 14.56 | 0.196 | | |
| **Read this table as a coarse split, not a ranking.** All five quantized runs produce coherent | |
| video β none is degraded or noisy; see `contact_sheet.png`. What differs is which sample the run | |
| lands in. Runs 1, 2 and 3 stay in BF16's bedroom framing (`corr` 0.44β0.45); runs 4 and 5 both go | |
| to a different floor-level scene (`corr` 0.14β0.20, and `corr` 0.645 to *each other* β they are | |
| closer to one another than either is to BF16). | |
| Two limits worth stating rather than papering over: | |
| * **n = 1 prompt.** Which basin a run falls into is close to a coin flip at this sample size. The | |
| table cannot separate run 1 from run 3, and it should not be read as doing so. | |
| * **`lambda=text` diverges end to end while being benign offline.** The offline table says it is | |
| harmless; this clip says it moved. One clip is not enough to call that a real effect. | |
| What the end-to-end runs *do* establish: run 3 is the only one whose mean RGB, per-frame spread and | |
| std all land on BF16's, and the two runs that diverged are not the two that the calibration's own | |
| error column ranked best. That is consistent with the offline per-modality table and inconsistent | |
| with the proportional one. | |
| ## The trap this repeats | |
| The calibration's own error column β the one in `svdquant_calib_768p*.pt` β ranks the runs: | |
| | | lambda=1 | lambda=1 +LoRA | lambda=all | lambda=video | lambda=text | | |
| | --- | --- | --- | --- | --- | --- | | |
| | median `err_smooth_lowrank` | 0.1019 | 0.0951 | 0.0942 | **0.0903** | 0.0953 | | |
| `lambda=video` wins it. That column is measured on rows drawn uniformly from the packed sequence, | |
| and video is 98.6 % of them, so it is a video error β it cannot see the 27.9 % it charged text. | |
| Picking a calibration by that number alone selects the configuration this study rejects. | |
| ## Files | |
| `0_bf16.mp4` β¦ `5_lambda_text.mp4`, `grid.mp4` (all six on one timeline, captioned), | |
| `contact_sheet.png` (5 frames each), `README.md` (what each run is and why). | |
| Scripts, on the H100 side: `calib_stats_modal.py`, `calib_sample_modal.py`, | |
| `score_modal_lambda.py`, `score_modal_report.py`, `diag_lambda_crossmodal.py`, | |
| `lambda_exps_report.py`, `mk_lambda_grid.py`. | |