213 GB
6,291 files
Updated about 6 hours ago
Name
Size
README.md6.65 kB
xet
README.md

Plaguekind/Minimax-H3 — the V1.5 workflow, as a Space

Plaguekind/Minimax-H3 ships no weights. It is a ComfyUI graph (PlagueKind-MinimaxH3-V1.5.json) over Comfy-Org/MiniMax-H3, and everything it contributes is in the sampling and the post chain. So what this Space reproduces is the graph, on the MiniMaxAI/MiniMax-H3 diffusers checkpoint.

MiniMax-H3 is a 33B-parameter single-stream omni DiT that denoises video and a synchronized stereo soundtrack — ambience, foley, speech — as one packed sequence, in one pass. Text-to-video, first frame, last frame, or both.

What the workflow changes

ComfyUI node widget here
KSamplerSelect euler MiniMax-H3's only sampler; the checkpoint is CFG-distilled, so one forward per step and no negative prompt
BasicScheduler linear_quadratic, 15 steps, denoise 1.0 Sigma schedule / Stepspk_workflow.linear_quadratic_sigmas
MiniMaxH3ImageToVideo prompt, first/last frame Prompt / First frame / Last frame
UnifiedResizeImageMask ("Target Dimension") longer side 1344 Target dimension
ImageSharpenKJ rcas, 0.3 RCAS sharpeningpk_workflow.rcas
FrameInterpolate + FrameInterpolationModelLoader film_net_fp16.safetensors, multiplier 2 FILM frame interpolationpk_workflow.interpolate
CreateVideo fps 24 * 2 48 fps output
ComfyMathExpression max(5, round(a*24)) + (5 - (… % 17)) % 17 Duration snapped to 17n + 5 frames
RTXVideoSuperResolution 2x, ULTRA not reproduced
PathchSageAttentionKJ sageattn_qk_int8_pv_fp8_cuda++ cuDNN fused attention

The sigma schedule is the part that changes the pixels most, and the part that is easy to get subtly wrong. linear_quadratic is Mochi's schedule, ported from comfy/samplers.py: half the steps crawl through the first 2.5 % of the trajectory and the rest sprint the remaining 97.5 %, which is why PlagueKind's 15 steps hold up against ~28 of MiniMax-H3's native grid.

Transplanting it into diffusers exactly needs one observation. MiniMax-H3 carries two rectified-flow schedules per request, shift = 12 for the video rows and shift = 3 for the audio rows. diffusers builds both from one linspace(1, 0, steps) base grid; ComfyUI instead samples the video schedule and derives the audio one in closed form (comfy/ldm/minimax/model.py::time_shift_sigma). The two agree, because the exponential shift is a bijection of the base grid that fixes both 0 and 1 — so handing MiniMaxH3Scheduler.set_timesteps the linear_quadratic grid for the video stream and time_shift_sigma(grid, 12, 3) for the audio stream is the ComfyUI path, not an approximation of it.

Two deliberate deviations

  • RTXVideoSuperResolution is NVIDIA's NGX super-resolution, shipped as a driver-level Windows/RTX component with no Linux Python path. The workflow's 2x upscale is therefore missing; pick a larger Target dimension instead of upscaling a small one.
  • SageAttention's qk_int8_pv_fp8_cuda++ kernel is not built for this pool's sm120 cards. Attention runs cuDNN's fused kernel, which is both the fastest available here and the numerically faithful choice — SageAttention is a quantized approximation of it.

One upgrade

The workflow loads minimax_h3_fl2va_pruned_int8_convrot.safetensors and a qwen3vl_32b_…_int8_convrot text encoder because that is what fits a consumer card. This Space runs both unquantized bfloat16.

Why the deployment is split

MiniMax-H3 at bfloat16 is 195.9 GiB, and a Space is evicted above 150 GB of storage. So the halves live apart:

  • this Space — the 61.73 GiB transformer and the two autoencoders (10.43 GiB, float32: a bfloat16 audio VAE decodes the soundtrack about 20 dB too quiet). 77.3 GB downloaded.
  • multimodalart/qwen3vl-conditioner — the 62.14 GiB Qwen3-VL text encoder, called over the gradio API once per request. prompt_embeds + text_token_tags in a safetensors file is the whole wire format. gradio_client forwards the caller's own ZeroGPU token, so that booking is billed to whoever asked for the video.

H3_PLACEMENT=pack moves only the transformer to CUDA at startup: spaces packs every startup-resident CUDA tensor into a second on-disk copy, and packing all 77.3 GB busts the storage quota while the 61.7 GB transformer alone fits. The VAEs move on the first GPU call. H3_GPU_SIZE=xlarge is required — large does not fit. H3_AOTI=1 loads multimodalart/minimax-h3-aoti, one ahead-of-time-compiled transformer block serving all 50, which removes roughly 0.5 s/step.

Measured

The default request — 960x544, 124 frames (5.167 s), 15 steps of linear_quadratic, RCAS 0.3, FILM 2x to 48 fps:

startup 97–153 s (77.3 GB over Xet, 50 AoTI blocks patched)
denoise + both decoders 61 s, 4.1 s/step
RCAS + FILM + h264 mux 9 s
booked ZeroGPU 89 s for a 70 s call

First-and-last-frame at the same canvas costs 79 s — the two keyframes add 510 conditioning rows each — and its output's first and last frames reproduce the supplied keyframes to a mean absolute error of 2–3/255. The estimator is quadratic in the canvas, so 1344x768 is roughly 2.7x the denoise of 960x544; every run prints its measured GPU time against what it booked.

Files

file
app.py the demo
pk_workflow.py the workflow's own parts: linear_quadratic sigmas, RCAS, FILM
film_net.py FILM, vendored from ComfyUI (GPL-3.0)
h3_split_blocks.py the modular blocks that skip the text encoder
h3_aoti.py the AoTI package loader

License

The demo code is GPL-3.0, because film_net.py is vendored from ComfyUI and pk_workflow.py ports kernels from ComfyUI and ComfyUI-KJNodes, both GPL-3.0. The workflow itself is MIT; the MiniMaxAI/MiniMax-H3 weights carry their own license.

Total size
213 GB
Files
6,291
Last updated
Aug 10
Pre-warmed CDN
US EU US EU

Contributors