File size: 538 Bytes
1cb6cea 10f73ab 1c9bf60 b6fdbf0 1c9bf60 4de4887 1c9bf60 4de4887 1c9bf60 4de4887 1c9bf60 4de4887 396fad9 4de4887 1c9bf60 4de4887 1cb6cea | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 | ---
license: mit
---
Semantic VAE
Purpose: Extremely fast convergence on illustraton data for a downstream DiT model
Trained from DINOv2 in 12 hours as follows:
Frozen DINOv2 -> LayerNorm -> 32 channels
Frozen DINOv2 with patch embed unfrozen -> LayerNorm -> 32 channels
64 channel bottleneck with Sigreg loss (1e-3 weight) from LeJEPA
VA-VAE style convolution decoder
Loss is all layers of DINOv2 and DINOv3 equally weighted MSE on input vs reconstruction.
There are no pixel losses/LPIPS/GAN used, this is DINO space loss only. |