DrewJin0827's picture
Upload folder using huggingface_hub
908c4e0 verified
|
Raw
History Blame Contribute Delete
1.83 kB

WaveForcing-T2V-1.3B-5step-Preview

Preview release of the WaveForcing 1.3B text-to-video model (5 denoising steps). Trained from Wan2.1-T2V-1.3B with Rolling-Forcing + DMD distillation. This is the raw generator_ema checkpoint (FP32, FSDP layout) converted to a flat model.safetensors (BF16, 825 tensors) for preview purposes.

VBench scores

Metric Rolling Forcing Rolling Forcing (causal inference) Wave Forcing
Temporal Flickering 0.976686 0.970226 0.976579
Subject Consistency 0.971557 0.966446 0.980122
Background Consistency 0.961879 0.958950 0.966289
Motion Smoothness 0.987435 0.983091 0.984803
Aesthetic Quality 0.532904 0.528120 0.530068
Imaging Quality 0.707718 0.673502 0.711773
Quality Drift 0.01525 0.02915 0.00016

Wave Forcing improves subject/background consistency and quality drift over the Rolling-Forcing baselines.

Generation throughput (8 GPUs, 399 latent frames / 100 s video)

Topology Single-GPU baseline Fastest config Steady fps (p50) E2E fps Speedup
4+3 (4-step) 15.0 fps (106.5 s) 1.3B sage causal paged 125.7 117.7 7.9×
5+2 (5-step) 11.9 fps (134.4 s) 1.3B bf16 causal paged 86.5 84.2 7.1×

Notes:

  • E2E fps = 1596 pixel frames / wall time (diffusion + VAE decode, video saved). Steady fps = 12×1000 / p50 tick, steady-state window.
  • Single-GPU baseline: full Rolling-Forcing generation (diffusion + VAE) on one GPU.
  • On 5+2, sage/sagefp8 attention reaches ~127 fps steady state, but the 2-stage VAE is a structural bottleneck, capping E2E at ~54 fps; bf16 causal paged is the best E2E config there.
  • Real generation runs (BF16 causal overlap, 399 frames): 83.0 fps E2E on 4+3, 81.9 fps E2E on 5+2.