Buckets:
| # step 5000 — root cause: wrong Phi bootstrap latent | |
| The rollout seeded its autoregressive Phi history with `zero_flow_latent.pt`, which is the | |
| VAE encoding of a WHITE image (the legacy Middlebury colour wheel renders zero flow white). | |
| causal_mode's Phi is raw (u,v,dz) clipped to [-1,1] — the VAE's own input range — so its | |
| "no motion" point is the ZEROS image, i.e. mid-GREY. | |
| | | norm | L2 to quietest real GT Phi | | |
| |---|---|---| | |
| | old `zero_flow_latent` (white) | 88.4 | **90.5** | | |
| | new `zero_phi_latent` (grey) | 46.1 | **33.0** | | |
| | quietest real GT Phi frame | 44.2 | 0 | | |
| Effect (same checkpoint, same seed, only the bootstrap changed): | |
| | | PSNR mean | min | | |
| |---|---|---| | |
| | white bootstrap | 14.14 dB | 10.55 | | |
| | grey bootstrap | **16.08 dB** | **14.62** | | |
| Video drift is gone. TRAINING WAS NEVER AFFECTED — it takes cond_f from the dataset's GT | |
| Phi latent, never from the bootstrap — so no training compute was wasted. | |
| Still open: decoded Phi is high-frequency noise rather than the sparse grey+patches of GT, | |
| at 8% of training. Watch; do not conclude yet. | |
Xet Storage Details
- Size:
- 1.09 kB
- Xet hash:
- 36ebb042bfa68d69098e0a0c38f1efb1f97e73c36e1649c613265a1ad584eede
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.