DuoMatching

Joint–Marginal Distribution Matching for Few-Step Video Generation

Code and installation · Project page

DuoMatching combines joint video distribution matching with image-teacher supervision on frame-level marginals. LatentBridge adapts temporally compressed video latents for the image teacher. This repository provides the video student and the matching Qwen LatentBridge weights. Inference uses only the student, Wan text encoder, and Wan VAE.

Checkpoint

  • Source experiment: cf2__q-adapter-qfake-f4-rand-g4-w04-s1800.
  • Released step: 1020, EMA (generator_ema), FP32.
  • Backbone: Wan2.1 T2V 1.3B, with a causal framewise student.
  • Resolution: 832 × 480; 21 latent frames decode to 81 video frames.
  • Denoising: four steps for the first latent, two for subsequent latents.
  • Image supervision: Qwen-Image, four random sampled latent frames, CFG 4.0, marginal loss weight 0.4, and a dedicated Wan 1.3B frame-marginal fake critic.

The s1800 suffix describes the source experiment's configured training horizon, not the released checkpoint step. This checkpoint used random sampling; it does not claim the Latent Variation Sampling configuration described on the project page. The released weights are intended for generation and fresh fine-tuning, not exact training resumption. Critic and optimizer states are not included. The original training prompt corpus is not included.

Qwen-Image LatentBridge

The matching pretrained Qwen-Image bridge is included in latent_bridge/. Download it independently with:

hf download JohnZhan/DuoMatching --include 'latent_bridge/*' --local-dir checkpoints/DuoMatching

The released LatentBridge implementation is scoped to Qwen-Image. See the folder's README for tensor shapes, loading, first-frame bypass, and gradient behavior.

Files

File Description
model.pt Original EMA student checkpoint
latent_bridge/adapter.pt Frozen Qwen LatentBridge
latent_bridge/adapter_config.json Bridge architecture
inference_config.yaml Inference settings
training_config.yaml Portable source training settings
release.json Source commit, settings, file sizes and SHA-256 hashes

Use

Install the code, then run these commands from its root:

hf download Wan-AI/Wan2.1-T2V-1.3B --local-dir wan_models/Wan2.1-T2V-1.3B
hf download JohnZhan/DuoMatching --local-dir checkpoints/DuoMatching
python inference.py \
  --config_path checkpoints/DuoMatching/inference_config.yaml \
  --checkpoint_path checkpoints/DuoMatching/model.pt \
  --use_ema \
  --data_path prompts/demos.txt \
  --output_folder outputs/demo \
  --end_index 1 \
  --num_output_frames 21 \
  --seed 0

A CUDA GPU is required. The checkpoint is not a Diffusers pipeline; use the DuoMatching inference script, which normalizes FSDP key prefixes and loads weights strictly. Training instructions, teacher downloads, and LatentBridge training are provided in the code repository.

Validation and limitations

The release process checks student tensor names/shapes against the causal model and checks bridge loading, numerical compatibility and input gradients. The release environment has no available CUDA device, so end-to-end video generation has not been rerun there. There is no checkpoint-specific benchmark table in this release. Do not infer numerical performance from the project page's comparisons. Longer rollouts and resolutions beyond the trained 480p setting are not validated. Generated video may exhibit visual, motion, and semantic errors.

Attribution

Built on Causal Forcing, Self Forcing, Wan2.1 and Qwen-Image. The code release retains upstream notices. Apache-2.0 applies to this release; separately downloaded upstream models and datasets retain their applicable terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JohnZhan/DuoMatching

Finetuned
(97)
this model