metadata
license: mit
Semantic VAE
Purpose: Extremely fast convergence on illustraton data for a downstream DiT model
Trained from DINOv2 in 12 hours as follows:
Frozen DINOv2 -> LayerNorm -> 32 channels
Frozen DINOv2 with patch embed unfrozen -> LayerNorm -> 32 channels
64 channel bottleneck with Sigreg loss (1e-3 weight) from LeJEPA
VA-VAE style convolution decoder
Loss is all layers of DINOv2 and DINOv3 equally weighted MSE on input vs reconstruction.
There are no pixel losses/LPIPS/GAN used, this is DINO space loss only.