DuoMatching
Joint–Marginal Distribution Matching for Few-Step Video Generation
Code and installation · Project page
DuoMatching combines joint video distribution matching with image-teacher supervision on frame-level marginals. LatentBridge adapts temporally compressed video latents for the image teacher. This repository provides the video student and the matching Qwen LatentBridge weights. Inference uses only the student, Wan text encoder, and Wan VAE.
Checkpoint
- Source experiment:
cf2__q-adapter-qfake-f4-rand-g4-w04-s1800. - Released step: 1020, EMA (
generator_ema), FP32. - Backbone: Wan2.1 T2V 1.3B, with a causal framewise student.
- Resolution: 832 × 480; 21 latent frames decode to 81 video frames.
- Denoising: four steps for the first latent, two for subsequent latents.
- Image supervision: Qwen-Image, four random sampled latent frames, CFG 4.0, marginal loss weight 0.4, and a dedicated Wan 1.3B frame-marginal fake critic.
The s1800 suffix describes the source experiment's configured training horizon,
not the released checkpoint step. This checkpoint used random sampling;
it does not claim the Latent Variation Sampling configuration described on the
project page. The released weights are intended for generation and fresh
fine-tuning, not exact training resumption. Critic and optimizer states are not
included. The original training prompt corpus is not included.
Qwen-Image LatentBridge
The matching pretrained Qwen-Image bridge is included in
latent_bridge/. Download it independently with:
hf download JohnZhan/DuoMatching --include 'latent_bridge/*' --local-dir checkpoints/DuoMatching
The released LatentBridge implementation is scoped to Qwen-Image. See the folder's README for tensor shapes, loading, first-frame bypass, and gradient behavior.
Files
| File | Description |
|---|---|
model.pt |
Original EMA student checkpoint |
latent_bridge/adapter.pt |
Frozen Qwen LatentBridge |
latent_bridge/adapter_config.json |
Bridge architecture |
inference_config.yaml |
Inference settings |
training_config.yaml |
Portable source training settings |
release.json |
Source commit, settings, file sizes and SHA-256 hashes |
Use
Install the code, then run these commands from its root:
hf download Wan-AI/Wan2.1-T2V-1.3B --local-dir wan_models/Wan2.1-T2V-1.3B
hf download JohnZhan/DuoMatching --local-dir checkpoints/DuoMatching
python inference.py \
--config_path checkpoints/DuoMatching/inference_config.yaml \
--checkpoint_path checkpoints/DuoMatching/model.pt \
--use_ema \
--data_path prompts/demos.txt \
--output_folder outputs/demo \
--end_index 1 \
--num_output_frames 21 \
--seed 0
A CUDA GPU is required. The checkpoint is not a Diffusers pipeline; use the DuoMatching inference script, which normalizes FSDP key prefixes and loads weights strictly. Training instructions, teacher downloads, and LatentBridge training are provided in the code repository.
Validation and limitations
The release process checks student tensor names/shapes against the causal model and checks bridge loading, numerical compatibility and input gradients. The release environment has no available CUDA device, so end-to-end video generation has not been rerun there. There is no checkpoint-specific benchmark table in this release. Do not infer numerical performance from the project page's comparisons. Longer rollouts and resolutions beyond the trained 480p setting are not validated. Generated video may exhibit visual, motion, and semantic errors.
Attribution
Built on Causal Forcing, Self Forcing, Wan2.1 and Qwen-Image. The code release retains upstream notices. Apache-2.0 applies to this release; separately downloaded upstream models and datasets retain their applicable terms.
Model tree for JohnZhan/DuoMatching
Base model
Wan-AI/Wan2.1-T2V-1.3B