File size: 538 Bytes
1cb6cea
 
 
10f73ab
1c9bf60
b6fdbf0
 
1c9bf60
4de4887
1c9bf60
4de4887
1c9bf60
4de4887
1c9bf60
4de4887
396fad9
4de4887
1c9bf60
4de4887
1cb6cea
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
---
license: mit
---
Semantic VAE

Purpose: Extremely fast convergence on illustraton data for a downstream DiT model

Trained from DINOv2 in 12 hours as follows:

Frozen DINOv2 -> LayerNorm -> 32 channels

Frozen DINOv2 with patch embed unfrozen -> LayerNorm ->  32 channels

64 channel bottleneck with Sigreg loss (1e-3 weight) from LeJEPA

VA-VAE style convolution decoder

Loss is all layers of DINOv2 and DINOv3 equally weighted MSE on input vs reconstruction.

There are no pixel losses/LPIPS/GAN used, this is DINO space loss only.