Dancing Stick Figures — reference models

Reference checkpoints for Dancing Stick Figures, the dataset introduced in Dancing Stick Figures: An Introductory Dataset for Training Video Generation Models. The current paper release is under paper-v6/. Historical image, short-video, and autoregressive baselines remain at the repository root.

Start here: Dataset · Colab · Code

Paper v6 checkpoints

All learned video rows below generate text-conditioned 64-frame, 64×64 clips at 20 fps. The Pixel DiTs operate directly on RGBA pixels. Mini-Wan generates normalized video latents and must be used with the released f4t4d8 codec and latent statistics.

file model parameters initialization training
paper-v6/image_dit_30k.pt single-frame Pixel DiT 39.9M random 30k updates
paper-v6/pixel_dit_factorised_random_30k.pt factorized Pixel DiT 39.9M random 30k updates
paper-v6/pixel_dit_factorised_image_30k.pt factorized Pixel DiT 39.9M image checkpoint 30k updates
paper-v6/pixel_dit_local_mixer_image_30k.pt factorized Pixel DiT + local 3×3×3 mixer 41.8M image checkpoint 30k updates
paper-v6/mini_wan_40m_decode_30k.pt compact Wan-style latent DiT 39.4M random 30k updates
paper-v6/vae_f4t4d8_10k.pt frozen video codec for Mini-Wan random 10k updates

The Pixel video runs use batch 8 × accumulation 2, velocity-prediction flow matching, foreground-weighted pixel loss, and an RGBA clean-prediction auxiliary term of weight 1. Mini-Wan uses a latent flow objective plus its decoded-RGBA auxiliary. These auxiliary constructions are analogous but not mathematically identical.

Exact hashes, architecture fields, provenance, and evaluation-file mappings are in paper-v6/release_manifest.json. Canonical evaluation JSONs are under paper-v6/evaluations/, and fixed-seed 64-frame samples are under paper-v6/samples/.

Table 3 reference results

The video protocol uses 128 held-out source motions, sampling seeds 0/1/2, 64 frames at stride 1, 50 Euler steps, and CFG 3. Lower is better for TVR, LIE, CPE, jerk, and FVD. Speed and motion fraction are two-sided diagnostic signals to compare with the real-reference row.

model TVR↓ LIE↓ CPE↓ speed motion fraction jerk FVD↓
real reference windows .116 .093 .037 .373 .501 .073 114.7 / ref.
codec reconstruction floor .153 .118 .053 .383 .507 .106 127.1
factorized Pixel DiT, random .259 .042 .032 .366 .431 .172 483.7
factorized Pixel DiT, image init .161 .033 .029 .350 .447 .151 319.0
Pixel DiT + local mixer, image init .149 .050 .035 .355 .470 .146 282.9
Mini-Wan 40M, decode loss .199 .101 .058 .423 .534 .148 182.6

Download the current release

hf download sprited/dancing-stick-figures-baselines \
  --include "paper-v6/*" --local-dir checkpoints

Checkpoint dictionaries contain EMA model weights, the optimizer step, and saved architecture arguments. The codec checkpoint additionally contains its model state and training metadata. Use the matching trainer/evaluator revision from the linked code repository; these are research checkpoints rather than a packaged inference API.

Historical baselines

The repository root retains the completed v0.1 image, 8-frame UNet, and long autoregressive UNet references for reproducing earlier dataset-card results. Files explicitly labelled interim, superseded 10k Mini-Wan checkpoints, and the no-decode-loss Mini-Wan checkpoint are not part of the current release.

Data are CC0-1.0 and code/checkpoint wrappers are MIT. See the dataset card and paper for scope and limitations.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train sprited/dancing-stick-figures-baselines