Buckets:
Where does object motion come from? — three layers, ground truth only, no masks
Run on cached WAFT flow with the robot silhouette removed. Exterior views only (wrist Phi is camera ego-motion). No SAM, no semantic labels, no training.
Layer 1 — global camera motion: prediction REFUTED
A per-frame homography was fitted by RANSAC on the flow outside the robot and subtracted.
| distance | raw | residual | delta |
|---|---|---|---|
| 0-10 px | +0.781 | +0.785 | +0.004 |
| 10-25 | +0.780 | +0.784 | +0.003 |
| 25-50 | +0.745 | +0.748 | +0.004 |
| 50-100 | +0.698 | +0.702 | +0.004 |
| >100 | +0.658 | +0.656 | -0.002 |
Only 6.3% of the flow magnitude was global motion, and the far-field +0.66 is untouched. The hypothesis that the far-field direction agreement is camera drift is wrong — the exterior cameras really are static (non-robot flow p90 = 0.19 px).
Layer 2 — Tomasi-Kanade rank: at most one object moves
Trajectories built by chaining flow from a seed grid; W is (2T x P) over points that moved more than 2 px. Rank <= 4 per rigid object.
| median rank | 3 (p25 2, p75 3, max 5) |
| rank <= 4 (one object) | 98.9% |
| rank 5-8 (two objects) | 1.1% |
| rank > 8 | 0.0% |
Normalised spectrum: 1.00, 0.43, 0.21, 0.12, 0.08, 0.06, 0.05, 0.04, … — the knee is soft, with no rank-4 cliff. That is what cloth looks like, not a rigid body, and it is consistent with the manipulated objects in these trajectories being towels and cloths.
Layer 3 — motion mass by source, and a trap worth recording
Common-fate clustering on full-clip trajectories, then a directed graph by lagged cross-correlation with the gripper's image-velocity series as root.
| k forced to 4 | k from rank | |
|---|---|---|
| m0 gripper -> object (hop 1) | 43.9% | 75.0% |
| m1 object -> object (hop >=2) | 48.1% | 0.0% |
| m_perp lag-0, attaches to nothing | 0.1% | 4.3% |
| m_none | 7.9% | 20.7% |
| hop-1 lag | 0.02 frames | 0.00 frames |
| hop-2 lag | 1.21 frames | no edges |
The first column looked like a textbook result — hop-1 lag 0.02, hop-2 lag 1.21, exactly the predicted monotonic increase. It was an artifact. Fixing k at 4 shreds one deformable object into four clusters whose parts lead and lag each other, manufacturing 48% of fake hop-2 edges. Deriving k from the trajectory rank instead makes m1 collapse to exactly zero, which is what Layer 2 independently said. The lag-by-hop test is sound; it was the clustering feeding it that was wrong.
What this means for the project
- There are no interaction chains in this data. m1 = 0.0%, and 98.9% of scenes have rank <= 4. Nothing here teaches robot -> A -> B, so do not expect a model trained on it to generalise to tool use. That has to come from different data, not a better loss.
- m_perp = 4.3% is the artifact budget measured without any mask — motion that locks to lag 0 but attaches to no cause (shadow, bleed, estimator noise). Small, and now quantified.
- m_none = 20.7% is unattributable motion. Worth a look before trusting any Phi-based metric to be complete.
- The right target metric is the mass vector matched against GT, not a single scalar to maximise: a hallucinating model dumps everything into m0 and produces m1 = 0 — which on this dataset is indistinguishable from correct, so m1 cannot discriminate here at all.
Caveat on the earlier EDA
Re-running the direction/magnitude EDA over all 391 trajectories, panel C's pearson r fell from 0.348 (60-trajectory sample) to 0.026. The cause is outliers: gripper image speed reaches 5.06e7 px/frame when the EE projects near z=0. The quintile relationship survives (0.0027 -> 0.0242, monotonic across all five), so the coupling is real, but that pearson value is meaningless without clipping. Layers 1 and 3 here clip gripper speed at 60 px/frame.
Xet Storage Details
- Size:
- 3.94 kB
- Xet hash:
- e5c47f09494080e3e664a6b60f63700d94f227c9d5adf41856d3beb2b27d8ef8
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.