twanghcmut/geotoken / RESULTS.md
twanghcmut's picture
|
download
raw
3.88 kB

Final table — 4 systems on the same val episodes (GT-action replay)

system psnr ssim lpips fid fvd_s3d psnr_arm
base_10k (60eps) 19.345 0.706 0.262 35.478 32.973 13.883
vanilla_40k
stageB_geo_40k (60eps) 19.133 0.705 0.265 42.878 41.132 13.657
stageC_aux_40k (60eps) 19.010 0.703 0.266 40.189 36.620 13.563

↑ psnr/ssim/psnr_arm better · ↓ lpips/fid/fvd better

Verdict (2026-07-04)

Geometry grounding (tokens = Stage B; tokens+aux = Stage C) does NOT improve the world model on in-distribution GT-action replay — both are 0.2-0.3 dB worse than base-10k on every metric. Mechanisms verified active (adapter fk_to_out‖W‖0.9, aux q-decode 0.17 rad). FID/FVD drop is consistent with finetune-on-subset drift (aux partially recovers it), not geometry per se — a vanilla control would confirm. arm-PSNR flat because GT actions already determine the arm; the zero-shot-camera thesis (where geometry should help) was NOT tested here.


Follow-up: WHY does geometry not help? (redundancy vs conflict)

Three experiments isolate the cause:

system psnr fid fvd psnr_arm geometry Δpix (toggle)
base_10k 19.35 35.5 33.0 13.88
geo AUGMENT (Stage B) 19.13 42.9 41.1 13.66 0.0075 (inert)
geo REPLACE (drop action) 15.26 61.6 61.3 11.03 0.115 (controls)
  1. AUGMENT (geometry + action): toggling geometry on/off changes <1% of pixels (Δpix 0.0075) and arm-PSNR by +0.03 dB — geometry is inert.
  2. REPLACE (geometry is the sole conditioner, action dropped): a counterfactual swap (episode A's video context + episode B's geometry) changes 11.5% of pixels — 15× the augment toggle. Geometry is causally load-bearing: the generated arm follows the (wrong) geometry.
  3. Therefore the augment failure is pure redundancy with the action token, not an inability to read geometry: on GT-action replay the 7-D EE action already determines the arm, so gradient descent never uses the (redundant) geometry. Geometry can be read, but it is a weaker standalone conditioner than the dense EE action (REPLACE reconstructs worse: psnr 15.3 vs 19.4).

Bottom line: not destructive conflict — redundancy + a weaker signal. The model uses geometry when forced (REPLACE), ignores it when action makes it redundant (AUGMENT). Geometry's real edge — viewpoint invariance / zero-shot new camera — is untested here (DROID has no extrinsics) and is where it would beat the robot-frame EE action.

Convergence / undertraining check (REPLACE)

Is REPLACE worse because it is undertrained? Eval of replace checkpoints (n=20):

step psnr psnr_arm
8k 17.90 11.48
24k 17.67 11.62
40k 17.19 11.50

arm-PSNR is flat 8k→40k (already converged by 8-24k); whole-frame psnr even declines slightly (mild overfit). So more steps at this finetune scale do not close the gap to base. BUT the deeper confound stands: geometry conditioning was only ever finetuned for ~8 GPU-h at unet_lr 5e-6, whereas the action conditioning is baked into the base **pretrain (8×H100 × 3 days ≈ 570 GPU-h)**. A modest finetune cannot re-teach the cross-attn to read geometry as well as 3 days of pretraining taught it to read action. Fairly answering "can geometry match action" requires pretraining a model with geometry as the native conditioner at comparable scale — out of scope here.

Robust regardless of compute: (1) geometry is readable/controllable (counterfactual Δpix 0.115), (2) the augment failure is redundancy with action. Only the claim "geometry is inherently weaker than action" is unresolved (never trained at a fair budget).

Xet Storage Details

Size:
3.88 kB
·
Xet hash:
dd5e8eb78ac3e440db2d0287be1ae556880b70b1dc4f21c4dcd37d0480ce1ec5

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.