Add diffusers format weights

#5
by multimodalart HF Staff - opened
README.md CHANGED
@@ -192,27 +192,19 @@ Each checkpoint is distributed as a self\-contained Hugging Face\-style reposito
192
  └── audio_vae/
193
  ```
194
 
195
- Download the model. The repository hosts the original checkpoint (`FL2VA/`, `Ref2VA/`) and the diffusers format side by side, so scope the download to what your framework needs:
196
-
197
- `model_index.json` is the repository-level modular index. The task-family-specific diffusers indexes remain under `FL2VA/model_index.json` and `Ref2VA/model_index.json`.
198
 
199
  ```bash
200
- # Original checkpoint, both task families (SGLang, vLLM):
201
- hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "modular_model_index.json" "FL2VA/*" "Ref2VA/*" --local-dir MiniMax-H3
202
-
203
- # Or a single task family:
204
- hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "modular_model_index.json" "FL2VA/*" --local-dir MiniMax-H3
205
  ```
206
 
207
- diffusers users do not need a manual download: `ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3")` fetches exactly the components it needs. See the [diffusers documentation](https://github.com/huggingface/diffusers/blob/minimax-h3/docs/source/en/api/pipelines/minimax_h3.md) for loading recipes.
208
-
209
  We recommend the following inference frameworks to serve the model:
210
 
211
  - [SGLang](https://docs.sglang.io/) \- see [cookbook](https://docs.sglang.io/cookbook/diffusion/MiniMax/MiniMax-H3)
212
 
213
  - [vLLM](https://github.com/vllm-project/vllm) \- see [vllm recipes](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3)
214
 
215
- - [diffusers](https://github.com/huggingface/diffusers) \- see [diffusers docs](https://github.com/huggingface/diffusers/blob/minimax-h3/docs/source/en/api/pipelines/minimax_h3.md)
216
 
217
  - [ComfyUI](https://github.com/Comfy-Org/ComfyUI) \- see [Comfy tutorial](https://docs.comfy.org/tutorials/video/minimax/minimax-h3); use [R2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json) / [T2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json)
218
 
@@ -277,11 +269,8 @@ TOKEN="<token>"
277
 
278
  MiniMax platform:
279
 
280
- API docs:
281
- - Create H3-2K: use /video-generation-v2-create [EN-docs](https://platform.minimax.io/docs/api-reference/video-generation-v2-create), [CN-docs](https://platform.minimaxi.com/docs/api-reference/video-generation-v2-create)
282
- - H3-Context-IR:use /video-generation-v2-h3-context-ir [EN-docs](https://platform.minimax.io/docs/api-reference/video-generation-v2-h3-context-ir), [CN-docs](https://platform.minimaxi.com/docs/api-reference/video-generation-v2-h3-context-ir)
283
- - H3-Regenerate-2K:use /video-generation-v2-regeneration [EN-docs](https://platform.minimax.io/docs/api-reference/video-generation-v2-regeneration), [CN-docs](https://platform.minimaxi.com/docs/api-reference/video-generation-v2-regeneration)
284
-
285
 
286
  The examples below encode local H3\-Base output files as Base64 Data URLs\. For production use, uploading the video to a publicly accessible URL and passing that URL as `base_video` is recommended\.
287
 
@@ -414,7 +403,7 @@ For each case below, we provide reference outputs at both 2K and 768p generated
414
 
415
  ## License
416
 
417
- MiniMax H3 is released under the [MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE). [Q&A about the License](docs/QA-about-License.md)
418
 
419
  ## Contact Us
420
 
 
192
  └── audio_vae/
193
  ```
194
 
195
+ Download the model:
 
 
196
 
197
  ```bash
198
+ hf download MiniMaxAI/MiniMax-H3 --local-dir MiniMax-H3
 
 
 
 
199
  ```
200
 
 
 
201
  We recommend the following inference frameworks to serve the model:
202
 
203
  - [SGLang](https://docs.sglang.io/) \- see [cookbook](https://docs.sglang.io/cookbook/diffusion/MiniMax/MiniMax-H3)
204
 
205
  - [vLLM](https://github.com/vllm-project/vllm) \- see [vllm recipes](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3)
206
 
207
+ - [diffusers](https://github.com/huggingface/diffusers) \- see [diffusers docs](https://huggingface.co/docs/diffusers/en/api/models/autoencoderkl_minimax_h3)
208
 
209
  - [ComfyUI](https://github.com/Comfy-Org/ComfyUI) \- see [Comfy tutorial](https://docs.comfy.org/tutorials/video/minimax/minimax-h3); use [R2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json) / [T2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json)
210
 
 
269
 
270
  MiniMax platform:
271
 
272
+ - CN: https://platform.minimaxi.com
273
+ - Global: https://platform.minimax.io
 
 
 
274
 
275
  The examples below encode local H3\-Base output files as Base64 Data URLs\. For production use, uploading the video to a publicly accessible URL and passing that URL as `base_video` is recommended\.
276
 
 
403
 
404
  ## License
405
 
406
+ MiniMax H3 is released under the [MiniMax H3 Community License Agreement](LICENSE). [Q&A about License](docs/QA-about-License.md)
407
 
408
  ## Contact Us
409
 
assets/action-reference.mov ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:116e0f08a399834e7ffc3472d036659b33250f4ba4f0b7e63f17ec07cb58e4dc
3
+ size 31446151
assets/character-action-reference.png ADDED

Git LFS Details

  • SHA256: 01a6c6f29e1ce9d414276e6e2eca06b171e5a68dd54a86074a4ab77bb9cb8883
  • Pointer size: 132 Bytes
  • Size of remote file: 5.24 MB
assets/character-replacement-action-reference.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:796c264126110d87992dcb54213ac0697920cb4ddf3d5a06aa36069386fc4fa1
3
+ size 8223683
assets/fashion-glasses-ad.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a11e6b0095dd7766b724bb886edf4d7d9930af44da6839ee195269ac4fc60ba4
3
+ size 10938085
assets/fashion-glasses-reference-1.png ADDED

Git LFS Details

  • SHA256: 5665f6529dd73339c0ac9397f0bd52df0cd44bd5b63d68af88f95855879d3d8b
  • Pointer size: 132 Bytes
  • Size of remote file: 1.01 MB
assets/fashion-glasses-reference-2.png ADDED

Git LFS Details

  • SHA256: 3a74bdab0bff0e126aae3757833fd2ec8cfba9afcb2ca36f63ee86977f7f7c00
  • Pointer size: 131 Bytes
  • Size of remote file: 889 kB
assets/fashion-glasses-reference-3.png ADDED

Git LFS Details

  • SHA256: 7b1491e5a227f226fa1395e71ac89a84358a435d974acb9f9065b9ded0982cbe
  • Pointer size: 132 Bytes
  • Size of remote file: 1.06 MB
assets/fashion-glasses-reference-4.png ADDED

Git LFS Details

  • SHA256: c9db6d6788855029ef9ac4109b05e17693ddb3bdf6ae3f24cec44de478c00f13
  • Pointer size: 131 Bytes
  • Size of remote file: 368 kB
assets/fl2va-clay-fox-reference.png ADDED

Git LFS Details

  • SHA256: 55eff74469b79a100e63bdd358596f6cff671e18fca9b276d6faef16faf22d27
  • Pointer size: 132 Bytes
  • Size of remote file: 1.73 MB
assets/fl2va-clay-fox.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4604b0838b736ecaa209c978f880bfacbb3a0fc82a8a0517c1f5aa16454b50ef
3
+ size 10931936
assets/h3-architecture.png ADDED

Git LFS Details

  • SHA256: 9765612d331f5fd31068a9283ffe28f11963f6a418c9f4093f81b408ea756630
  • Pointer size: 131 Bytes
  • Size of remote file: 336 kB
assets/h3-cinematic-shot.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:436defc81cfa7d53aef423f368be82ead20056088087e0be308b6e3865e8fb81
3
+ size 4361547
assets/h3-suspense-title.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:90ebbd7edc71c9a0151c3064126acd2dbb5b7cba58ce459e223e29f2c04f9186
3
+ size 16809220
assets/logo.svg ADDED
assets/reference-image-1.png ADDED

Git LFS Details

  • SHA256: 49b89115da4e253228969830ca491e4ca1b79331eb241a905d1ebca3dc6de7be
  • Pointer size: 131 Bytes
  • Size of remote file: 721 kB
assets/reference-image-2.png ADDED

Git LFS Details

  • SHA256: c9db6d6788855029ef9ac4109b05e17693ddb3bdf6ae3f24cec44de478c00f13
  • Pointer size: 131 Bytes
  • Size of remote file: 368 kB
assets/robot-arm-red-cube.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1a751e6100dbf6502a99a2adc0b12171304ca7b17eecea70e2ed035a73ea692c
3
+ size 2259732
assets/t2va-768p-demo.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d66903241362e224085cc93f7a5e70fba6ab378d0ac6ff6af83acf2559849a42
3
+ size 1637373
model_index.json DELETED
@@ -1,139 +0,0 @@
1
- {
2
- "_class_name": "MiniMaxH3ModularPipeline",
3
- "_diffusers_version": "0.36.0.dev0",
4
- "_blocks_class_name": "MiniMaxH3Blocks",
5
- "_minimax_h3": {
6
- "schema_version": 1,
7
- "index_scope": "repository",
8
- "task_family_indexes": {
9
- "fl2va": "FL2VA/model_index.json",
10
- "ref2va": "Ref2VA/model_index.json"
11
- }
12
- },
13
- "text_encoder": [
14
- "transformers",
15
- "Qwen3VLForConditionalGeneration",
16
- {
17
- "type_hint": [
18
- "transformers",
19
- "Qwen3VLForConditionalGeneration"
20
- ],
21
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
22
- "subfolder": "text_encoder",
23
- "variant": null,
24
- "revision": null
25
- }
26
- ],
27
- "tokenizer": [
28
- "transformers",
29
- "Qwen2TokenizerFast",
30
- {
31
- "type_hint": [
32
- "transformers",
33
- "Qwen2TokenizerFast"
34
- ],
35
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
36
- "subfolder": "tokenizer",
37
- "variant": null,
38
- "revision": null
39
- }
40
- ],
41
- "processor": [
42
- "transformers",
43
- "Qwen3VLProcessor",
44
- {
45
- "type_hint": [
46
- "transformers",
47
- "Qwen3VLProcessor"
48
- ],
49
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
50
- "subfolder": "processor",
51
- "variant": null,
52
- "revision": null
53
- }
54
- ],
55
- "vae": [
56
- "diffusers",
57
- "AutoencoderKLMiniMaxH3",
58
- {
59
- "type_hint": [
60
- "diffusers",
61
- "AutoencoderKLMiniMaxH3"
62
- ],
63
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
64
- "subfolder": "vae",
65
- "variant": null,
66
- "revision": null
67
- }
68
- ],
69
- "audio_vae": [
70
- "diffusers",
71
- "AutoencoderKLMiniMaxH3Audio",
72
- {
73
- "type_hint": [
74
- "diffusers",
75
- "AutoencoderKLMiniMaxH3Audio"
76
- ],
77
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
78
- "subfolder": "audio_vae",
79
- "variant": null,
80
- "revision": null
81
- }
82
- ],
83
- "transformer": [
84
- "diffusers",
85
- "MiniMaxH3Transformer3DModel",
86
- {
87
- "type_hint": [
88
- "diffusers",
89
- "MiniMaxH3Transformer3DModel"
90
- ],
91
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
92
- "subfolder": "transformer",
93
- "variant": null,
94
- "revision": null
95
- }
96
- ],
97
- "transformer_ref": [
98
- "diffusers",
99
- "MiniMaxH3Transformer3DModel",
100
- {
101
- "type_hint": [
102
- "diffusers",
103
- "MiniMaxH3Transformer3DModel"
104
- ],
105
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
106
- "subfolder": "transformer_ref",
107
- "variant": null,
108
- "revision": null
109
- }
110
- ],
111
- "scheduler": [
112
- "diffusers",
113
- "MiniMaxH3Scheduler",
114
- {
115
- "type_hint": [
116
- "diffusers",
117
- "MiniMaxH3Scheduler"
118
- ],
119
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
120
- "subfolder": "scheduler",
121
- "variant": null,
122
- "revision": null
123
- }
124
- ],
125
- "audio_scheduler": [
126
- "diffusers",
127
- "MiniMaxH3Scheduler",
128
- {
129
- "type_hint": [
130
- "diffusers",
131
- "MiniMaxH3Scheduler"
132
- ],
133
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
134
- "subfolder": "audio_scheduler",
135
- "variant": null,
136
- "revision": null
137
- }
138
- ]
139
- }