Instructions to use MATLOWAI/MiniMax-H3-Motion-Adapter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use MATLOWAI/MiniMax-H3-Motion-Adapter with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("MATLOWAI/MiniMax-H3-Motion-Adapter") prompt = "A man with short gray hair plays a red electric guitar." input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png") output = pipe(image=input_image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
Download examples/fight_quad/README.md from MATLOWAI/MiniMax-H3-Motion-Adapter: direct link, hf CLI and curl.
- Browser
- Download file 6.51 kB
-
https://huggingface.co/MATLOWAI/MiniMax-H3-Motion-Adapter/resolve/main/examples/fight_quad/README.md
- Command line
-
hf download hf://MATLOWAI/MiniMax-H3-Motion-Adapter/examples/fight_quad/README.md
-
curl -L -o README.md https://huggingface.co/MATLOWAI/MiniMax-H3-Motion-Adapter/resolve/main/examples/fight_quad/README.md
Reproduce the historical fight quad
The top-left tile is the supplied source video, not another processing pass.
Copy seedhunt_20260963_00001_.mp4 into your ComfyUI input directory. Keep its
filename unchanged. It is the original 124-frame H3 generation (1152 × 640,
24 fps, 5.167 s), recovered from the local experiment outputs.
Open it on the ComfyUI canvas
- Winner only — easy starting point: the verified adapter recipe with setup notes and editable controls.
- All four tiles — complete comparison: shared source/model loaders, color-coded processing lanes, four individual video outputs, and an automatically assembled comparison with source audio.
Download the JSON, then drag it onto ComfyUI or use Workflow → Open. Put the
source MP4 in ComfyUI/input, select your model files, and click Run. The
amber lane is the full-clip pass, rose is the windowed control, and green is the
adapter winner. Advanced nodes are collapsed; double-click a title to expand.
The full comparison takes longer and uses more memory than the winner alone.
The normal Save Video nodes embed both the prompt and frontend workflow when
run through the canvas with metadata enabled.
Requirements
Use a ComfyUI build with MiniMax-H3 support, ComfyUI-MAINodes, and ComfyUI-KJNodes with SageAttention available. Python 3 and FFmpeg are needed for the commands below. These API JSONs can also be imported through a frontend that supports API-format import; they are not canvas-layout JSONs.
Model names in the supplied graphs:
- diffusion model:
minimax_h3/minimax_h3_fl2va_pruned_int8_convrot.safetensors - text encoder:
minimax_h3/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors - VAE:
minimax_h3/minimax_h3_video_vae_fp16.safetensors - LoRA:
minimax_h3/minimax_h3_motion_adapter_pilot_r16.safetensors
Place this repository's adapter in models/loras/minimax_h3/. The historical
training filename p4_pilot_k100.safetensors refers to this published adapter.
Use the exact base/encoder/VAE variants for historical comparison; adjust only
folder paths if your installation uses different folders.
Render the three processed tiles
Download this directory, then run these commands inside it with ComfyUI running.
Use --server http://127.0.0.1:8189 on each command if that is your server port.
The runner uses the public /prompt, /history, and /view endpoints, assigns
unique output prefixes, waits for success, and downloads the resulting video.
python3 run.py full_clip.api.json --output full_clip.mp4
python3 run.py windowed.api.json --output windowed.mp4
python3 run.py adapter.api.json --output adapter.mp4
| Tile | Input / workflow | Settings |
|---|---|---|
| Top left | seedhunt_20260963_00001_.mp4 |
Original source, unchanged |
| Top right | full_clip.api.json |
T2C arm A, full-clip pixel smear and regeneration, inject 0.60, no adapter |
| Bottom left | windowed.api.json |
Windowed v3.1, inject 0.45, no adapter |
| Bottom right | adapter.api.json |
Same window, adapter 0.75, inject 0.30 |
All processing arms use seed 20260817, beta schedule, 25 total steps, and
gradient_estimation. The windowed arms crop world frames 68–123, expand the
56-frame window to 107 frames, retain the historical tail guide, recover to
56 frames, and splice after the untouched 68-frame head. These are the old
recipes, not the later two-boundary pinned graph.
expand_to_end=false is explicit so current node defaults cannot rewrite the
historical hold maps. The windowed map already expands through its final frame.
Descriptive node titles from the archive were removed because some still said
inject 0.45 even when the actual input was 0.30.
To queue the combined graph and download its assembled comparison directly:
python3 run.py fight_quad.api.json --output-node 9005 --output fight_quad.mp4
Assemble the four tiles
This rebuilds the panel arrangement with the source audio. It deliberately has no historical timing overlays: your new runs are not the old benchmark.
ffmpeg -n -i seedhunt_20260963_00001_.mp4 -i full_clip.mp4 -i windowed.mp4 -i adapter.mp4 \
-filter_complex '[0:v][1:v]hstack=inputs=2[top];[2:v][3:v]hstack=inputs=2[bottom];[top][bottom]vstack=inputs=2[v]' \
-map '[v]' -map '0:a?' -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac -movflags +faststart fight_quad.mp4
Metadata, provenance, and verification
winner_verified_20260921.mp4 is the September 21 functional rerun of the
adapter arm: success, 65.385 s server execution, all 124 frames decoded without
errors. Only the text-encoder and VAE loader nodes were reported cached; the
sampler executed. This does not replace the historical matched-cache 49.9 s
measurement. The full-clip and no-adapter workflows are recovered historical graphs.
See verification.json for the separate full-canvas verification run.
The verification used MAINodes checkout commit
4e957e267795c23c4d5544cf7d62bb4fa736c4c3. Hardware, software, kernels, and cache
state affect timing and pixels; identical results across environments are not
guaranteed. Historical benchmark labels remain in the original
comparison video.
Both supplied MP4s embed API JSON under the container's prompt tag. To inspect:
ffprobe -v error -show_entries format_tags=prompt -of json winner_verified_20260921.mp4
The embedded winner graph preserves the exact local run, including the original
LoRA training filename. The separate adapter.api.json uses the published
filename and a fresh output prefix. A workflow canvas-layout tag is not present.
The assembled FFmpeg quad does not automatically inherit all four graphs.
source_generation.api.json is extracted from the original source MP4 for
provenance and optional regeneration. It additionally requires the LightX2V
LoRA named in that graph. Reusing the supplied source is the way to reproduce
this comparison without changing the plate. Base-model licensing remains as
specified in the parent model card.
The complete browser-exported comparison graph also passed a fresh render: all three samplers executed, all five saved videos decode, and the assembled output is 2304 × 1280 at 24 fps with 124 frames. Total server execution was 279.09 s for this combined run. Watch the newly assembled comparison.