Visual Generative Lab (VGL) checkpoints
Checkpoints behind Table 2 of Do Diffusion Models Learn to Generalize Basic Visual Skills? Code, datasets and evaluation: https://github.com/AmishSethi/visual-generative-lab
Layout: table2/<skill>/<variant>/seed_<n>/final.pt with the run's run_config.json beside it.
Skills: size, position, rotation, count. Variants: baseline, sinusoidal, rotary, adaln, vae, flow, unet, dit_large
(rotation also has baseline_nz, the ten-seed zero-null baseline). Each .pt holds model, ema, opt and the
training state; load with the training scripts in the repository. License CC BY-NC 4.0 (inherited from DiT).
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support