aryan5v commited on
Commit
9edcb65
路
verified 路
1 Parent(s): dc4d366

Release model card for FastH3 Trim

Browse files
Files changed (1) hide show
  1. README.md +20 -9
README.md CHANGED
@@ -1,12 +1,23 @@
1
- # FastH3 Pruned 8-Step (bf16, checkpoint 300)
 
 
 
 
 
2
 
3
- Reference bf16 transformer; source for the FP8/NVFP4 exports.
 
 
 
 
4
 
5
- ## Model
6
- - MiniMax-H3 student pruned from 50 to **42 transformer blocks** (removed blocks 6, 7, 9, 13, 15, 16, 22, 23), rank-16 AdaLN,
7
- trained with 8-step DMD (timesteps 999, 874, 749, 624, 500, 375, 250, 125) and VSA sparse attention (sparsity 0.8, 64-token tiles).
8
- - Training checkpoint 300 of run `s42-r16-dmd8-vsa80` (selected over 800 for prompt adherence on a multi-seed evaluation).
9
- - Sampling contract: video/audio scheduler shift 10/3, guidance 1.0, VSA_sparsity 0.8, VSA_tile_size 64.
10
- - `text_encoder` is not included; it is identical to the one in `FastVideo/FastVideo-FastH3-8-Step-V2`.
11
 
12
- **Status: internal, private evaluation build.** License: MiniMax H3 Community License (inherited).
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ base_model: FastVideo/FastVideo-FastH3-8-Step-V2
4
+ tags: [video-generation, text-to-video, audio, fastvideo, bf16, pruned]
5
+ ---
6
+ # FastH3 Trim, 8-step, BF16 source weights
7
 
8
+ FastH3 Trim is an experimental, smaller version of [FastH3 V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2):
9
+ eight-step text-to-video with synchronized audio from 42 of the 50 MiniMax H3 transformer blocks. We removed the eight
10
+ blocks whose removal changed video and audio predictions the least, replaced each block's AdaLN timestep projection with
11
+ a shared rank-16 basis, and trained the result with eight-step DMD2 (checkpoint 300). Removing blocks makes the model
12
+ faster and smaller but costs some quality; use FastH3 V2 when quality matters most.
13
 
14
+ - **Sampling:** 8 DMD steps (999, 874, 749, 624, 500, 375, 250, 125), sparse attention keeping 20% of tiles.
15
+ `fastvideo_inference.json` holds the schedule, which FastVideo reads automatically.
16
+ - **Text encoder:** NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads.
17
+ - **VAE:** LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.
 
 
18
 
19
+ Blog post: [FastH3 on Consumer Hardware](https://haoailab.com/blogs/fasth3-rtx/) 路 Code: [FastVideo](https://github.com/hao-ai-lab/FastVideo)
20
+
21
+ - **Transformer:** 34.9 GiB in BF16 (FP16 AdaLN factors). Source for the quantized releases and for local MLX conversion on Apple Silicon.
22
+
23
+ Full-quality counterpart in the same format: [FastVideo/FastVideo-FastH3-8-Step-V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2).