twanghcmut's picture
|
download
raw
3.94 kB

Where does object motion come from? — three layers, ground truth only, no masks

Run on cached WAFT flow with the robot silhouette removed. Exterior views only (wrist Phi is camera ego-motion). No SAM, no semantic labels, no training.

Layer 1 — global camera motion: prediction REFUTED

A per-frame homography was fitted by RANSAC on the flow outside the robot and subtracted.

distance raw residual delta
0-10 px +0.781 +0.785 +0.004
10-25 +0.780 +0.784 +0.003
25-50 +0.745 +0.748 +0.004
50-100 +0.698 +0.702 +0.004
>100 +0.658 +0.656 -0.002

Only 6.3% of the flow magnitude was global motion, and the far-field +0.66 is untouched. The hypothesis that the far-field direction agreement is camera drift is wrong — the exterior cameras really are static (non-robot flow p90 = 0.19 px).

Layer 2 — Tomasi-Kanade rank: at most one object moves

Trajectories built by chaining flow from a seed grid; W is (2T x P) over points that moved more than 2 px. Rank <= 4 per rigid object.

median rank 3 (p25 2, p75 3, max 5)
rank <= 4 (one object) 98.9%
rank 5-8 (two objects) 1.1%
rank > 8 0.0%

Normalised spectrum: 1.00, 0.43, 0.21, 0.12, 0.08, 0.06, 0.05, 0.04, … — the knee is soft, with no rank-4 cliff. That is what cloth looks like, not a rigid body, and it is consistent with the manipulated objects in these trajectories being towels and cloths.

Layer 3 — motion mass by source, and a trap worth recording

Common-fate clustering on full-clip trajectories, then a directed graph by lagged cross-correlation with the gripper's image-velocity series as root.

k forced to 4 k from rank
m0 gripper -> object (hop 1) 43.9% 75.0%
m1 object -> object (hop >=2) 48.1% 0.0%
m_perp lag-0, attaches to nothing 0.1% 4.3%
m_none 7.9% 20.7%
hop-1 lag 0.02 frames 0.00 frames
hop-2 lag 1.21 frames no edges

The first column looked like a textbook result — hop-1 lag 0.02, hop-2 lag 1.21, exactly the predicted monotonic increase. It was an artifact. Fixing k at 4 shreds one deformable object into four clusters whose parts lead and lag each other, manufacturing 48% of fake hop-2 edges. Deriving k from the trajectory rank instead makes m1 collapse to exactly zero, which is what Layer 2 independently said. The lag-by-hop test is sound; it was the clustering feeding it that was wrong.

What this means for the project

  1. There are no interaction chains in this data. m1 = 0.0%, and 98.9% of scenes have rank <= 4. Nothing here teaches robot -> A -> B, so do not expect a model trained on it to generalise to tool use. That has to come from different data, not a better loss.
  2. m_perp = 4.3% is the artifact budget measured without any mask — motion that locks to lag 0 but attaches to no cause (shadow, bleed, estimator noise). Small, and now quantified.
  3. m_none = 20.7% is unattributable motion. Worth a look before trusting any Phi-based metric to be complete.
  4. The right target metric is the mass vector matched against GT, not a single scalar to maximise: a hallucinating model dumps everything into m0 and produces m1 = 0 — which on this dataset is indistinguishable from correct, so m1 cannot discriminate here at all.

Caveat on the earlier EDA

Re-running the direction/magnitude EDA over all 391 trajectories, panel C's pearson r fell from 0.348 (60-trajectory sample) to 0.026. The cause is outliers: gripper image speed reaches 5.06e7 px/frame when the EE projects near z=0. The quintile relationship survives (0.0027 -> 0.0242, monotonic across all five), so the coupling is real, but that pearson value is meaningless without clipping. Layers 1 and 3 here clip gripper speed at 60 px/frame.

Xet Storage Details

Size:
3.94 kB
·
Xet hash:
e5c47f09494080e3e664a6b60f63700d94f227c9d5adf41856d3beb2b27d8ef8

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.