| V0_1: 2,800 more training steps compared to V0. |
| |
| Current V0 went through : |
| - 100 initial steps at LR 5e-5 (batch size 8) |
| - 1100 more steps at LR 1e-4 (batch size 32) |
| - 600 more steps at LR 1e-4 (batch size 8) |
| - 600 more steps at LR 7.5e-5 (batch size 8) |
| |
| Dataset has 30,000 examples covering: |
| |
| - Text-to-video with no reference: 6,500 |
| - Text-to-image with no reference: 3,500 |
| |
| - Still-image-to-video using: |
| - First frame from the target video: 750 |
| - Middle frame from the target video: 750 |
| - Last frame from the target video: 750 |
| - Frame from the same source but outside the target interval: 750 |
| |
| - Face-reference-to-video using a close-up crop of the primary face: 3,000 |
| |
| - Image-to-image using a different image from the same gallery: 3,000 |
| |
| - Image restoration and inpainting: |
| - Masked image reference: 800 |
| - Degraded image reference: 800 |
| |
| - Image outpainting using an aggressively cropped reference: 2,400 |
| |
| - Video outpainting: |
| - Static spatial crop: 501 |
| - Tracked spatial crop: 499 |
| |
| - Temporal video completion: |
| - Reference video immediately before the target: 467 |
| - Reference segment inside the complete target: 467 |
| - Reference video immediately after the target: 466 |
| |
| - Audio-to-video using the target audio as conditioning: 2,000 |
| |
| - Video-to-audio / sound generation using a muted, |
| low-resolution video reference: 1,400 |
| |
| - Structural conditioning: |
| - Single-frame depth: 160 |
| - Single-frame optical flow: 147 |
| - Single-frame pose: 145 |
| - Single-frame segmentation: 148 |
| - Video depth: 146 |
| - Video optical flow: 153 |
| - Video pose: 149 |
| - Video segmentation: 152 |