607 MB
306 files
Updated 27 days ago
Name
Size
README.md1.51 kB
xet
SIDEBYSIDE_val199_towel.mp42.78 MB
xet
SIDEBYSIDE_val299_best.mp42.23 MB
xet
SIDEBYSIDE_val399_napkin.mp44 MB
xet
SIDEBYSIDE_val99_handle.mp43 MB
xet
time_20260727_071232_traj_99_8_5_psnr17_Pull_the_wooden_handle_towards_ck20k.mp4484 kB
xet
time_20260727_071232_traj_99_8_5_psnr17_Pull_the_wooden_handle_towards_ck20k_flow.mp4557 kB
xet
time_20260727_071609_traj_199_8_5_psnr17_Move_the_green_towel_to_the_ri_ck20k.mp4461 kB
xet
time_20260727_071609_traj_199_8_5_psnr17_Move_the_green_towel_to_the_ri_ck20k_flow.mp4535 kB
xet
time_20260727_072238_traj_299_8_5_psnr18__ck20k.mp4324 kB
xet
time_20260727_072238_traj_299_8_5_psnr18__ck20k_flow.mp4372 kB
xet
time_20260727_072615_traj_399_8_5_psnr15_Use_one_napkin_to_wipe_the_lap_ck20k.mp4718 kB
xet
time_20260727_072615_traj_399_8_5_psnr15_Use_one_napkin_to_wipe_the_lap_ck20k_flow.mp4861 kB
xet
time_20260727_074650_traj_99_8_5_psnr19_Pull_the_wooden_handle_towards_base.mp4373 kB
xet
time_20260727_075030_traj_199_8_5_psnr21_Move_the_green_towel_to_the_ri_base.mp4326 kB
xet
time_20260727_075354_traj_299_8_5_psnr23__base.mp4196 kB
xet
time_20260727_075714_traj_399_8_5_psnr17_Use_one_napkin_to_wipe_the_lap_base.mp4641 kB
xet
README.md

BASE Ctrl-World vs CAUSAL checkpoint-20000 — same 4 val episodes, same seed

SIDEBYSIDE_*.mp4: TOP half = base, BOTTOM half = causal. Within each half: GT over prediction.

val task BASE CAUSAL-20k delta
99 pull the wooden handle 19.28 16.92 -2.36
199 move the green towel 20.68 16.86 -3.82
299 (unlabelled) 23.05 18.42 -4.63
399 wipe with a napkin 16.92 15.22 -1.70
mean 20.0 16.9 -3.1

Base wins on 4/4, by 1.7-4.6 dB. That is a systematic per-episode gap, not the 3.2 dB episode-to-episode spread, so it is a real effect.

Two competing explanations -- NOT separable from this data

(a) The Phi branch is costing appearance quality (the original concern: a shared UNet asked to emit both a sharp frame and a smooth flow field). (b) Plain fine-tuning loss: base was trained on the FULL DROID corpus; the causal run fine-tunes on 387 trajectories, ~200x narrower. Any fine-tune on a set that much smaller degrades PSNR even with no Phi at all (catastrophic forgetting).

Deciding between them needs the ablation-table baseline row that has never been run: Ctrl-World fine-tuned on these same 387 trajectories with lambda_flow=0. Until that exists, this table does not show that the causal design is harmful -- only that this checkpoint is worse than the corpus-trained base.

Also note the causal run is at 33% of its schedule (20k/60k).

Total size
607 MB
Files
306
Last updated
Jul 28
Pre-warmed CDN
US EU US EU

Contributors