Buckets:
| # Visualizations: query-graph planner and subgoal-conditioned controller on RoboMME | |
| These figures and videos show what each part of the system sees and predicts on real RoboMME episodes. The architecture diagram is drawn separately. | |
| The system has two parts. A planner (7.7M parameters) watches the episode as a growing graph of nodes and, at every controller query, says which step the robot is on (pick, place, press, ...), which repetition it is ("for the second time"), and where on the image the target is. A controller (GR00T N1.6, about 1.47B parameters) receives that subgoal and produces the actions. | |
| All planner figures use the final 16-task planner (`planner16`) on validation episodes it was not trained on. The rollout videos use the final controller (`ctrl5wngml`, 72k steps) driven by the planner, without oracle subgoals, on the official test split. | |
| Image coordinates: the front camera image is 256 x 256 px. Points are stored as (row, col) normalised over the central 243.2 px crop, so pixel = 6.4 + 243.2 x value. | |
| ## 01_graph_node | |
| `node_anatomy.png` shows one graph node. A node is added every 4 frames and holds three groups of tokens: | |
| - 81 frame tokens: Eagle features of the front image on a fixed 9 x 9 grid. Slot 42 is always the same patch of the table, so comparing slot 42 across nodes tells the model what changed there. | |
| - 1 wrist token: the mean of the 81 wrist-camera patches. | |
| - 192 entity tokens: points tracked by CoTracker, seeded on foreground pixels (objects, targets, the robot). Each entity token sees a 12 x 12 colour patch around the point now and at the first node, its motion, and whether it is visible. The right panel shows four entities. | |
| Every token also reads the instruction through query edges, so the same video can produce different nodes for different instructions (see 03). | |
| `edges.png` shows what the edges connect on episode 214. Row 1 is a temporal edge of a frame token: one fixed slot of the table across nodes, where the cube arrives and leaves. Row 2 is a temporal edge of an entity token: the crop follows one tracked point (here the green cube) across the demo and execution. The bottom panel is a spatial edge between entities in one node, where the red star reads the other entities with a Gaussian weight on distance. Each attention head learns its own width; the picture uses 0.08 of the image. | |
| ## 02_planner_timeline | |
| One figure per validation episode (VideoPlaceOrder 214, PickXtimes 100, VideoUnmaskSwap 5, StopCube 1206, ButtonUnmaskSwap 501, BinFill 901, VideoRepick 608, InsertPeg 1305). The x axis is the node index; the black vertical line is where the demo video ends and the robot starts acting. | |
| - Filmstrip: frames at selected nodes, green circle = label target, red x = planner target. | |
| - Step kind: colour band per node, label on top, planner below. | |
| - Ordinal: the "n-th time" counter for tasks that repeat a step (PickXtimes, SwingXtimes). It stays at "-" for tasks without repetition. | |
| - Event counter: the planner keeps a learned running count of how many steps of each kind have happened. Solid lines are the model's counts, dotted lines count the steps in the labels. In VideoPlaceOrder the counter reaches 4 places during the demo, which is how the model can tell the third drop from the second. | |
| - Point error in pixels; the dashed line is 8 px. Gaps are nodes without a target. | |
| ## 03_query_dependence | |
| `VideoPlaceOrder_214_ordinal_swap.png` runs the planner four times on the same video and the same tracks. Only the ordinal word in the instruction changes ("first", "second", "third", "fourth" target). The numbered circles are where the cube was dropped during the demo, in order. The dots are the planner's pointer probabilities over the 192 entities, and the cyan x is its target. The target moves to drop 1, 2, 3 and 4 as the word changes, although the video does not change. The instruction text for each variant is the encoded instruction of another episode whose instruction differs only in that word. | |
| ## 04_controller_conditioning | |
| `mark_fovea_tokens.png` shows how one subgoal reaches the controller, for five tasks: | |
| 1. The frame with the label target. | |
| 2. The mark: a Gaussian weight (sigma 0.06 of the image) over the 81 front patches, centred on the target, which tells the action model which image tokens matter. | |
| 3. The fovea: a 64 x 64 crop at the target at full resolution. A small CNN turns it into 81 extra image tokens. | |
| 4. The memory tokens: a step token (kind, ordinal and the words of the subgoal), a point token (Fourier features of the target coordinates) and a goal token computed from the fovea crop and the scene features. Every DiT block cross-attends to these tokens and uses them to scale and shift its feed-forward input. | |
| ## 05_planner_videos | |
| One mp4 per episode in 02, at 8 nodes per second. Left: the front frame with the pointer probabilities over entities (bright and large = likely target), label target (green circle) and planner target (cyan x). Middle: wrist camera. Right: the event counter. The text below gives the label step and the planner step; the planner line turns red when they differ. | |
| ## 06_results | |
| `robomme16_per_task.png`: success rate per task on the 16 RoboMME tasks. Our numbers are the planner-driven run (50 episodes per task, mean 63.1) and the run with oracle subgoals (20 episodes per task, mean 72.5). Baseline numbers are from Table 8 of the SimpleMemVLA paper. Our weak tasks are MoveCube, InsertPeg, PatternLock and RouteStick; the oracle-subgoal run is also low on them, which points at the controller. | |
| `training_compute.png`: training GPU-hours on a log scale. Ours is about 31 A100 GPU-hours in total (feature caches, controller and planner). SimpleMemVLA reports 20 hours on 128 H100. HAMLET's number is from its own paper and benchmarks, not RoboMME. | |
| ## 07_rollouts | |
| 32 closed-loop rollouts, two per task, on the official test split, with the planner in the loop (no oracle). The file name ends with the outcome: success, fail, or timeout. Each frame shows the front and wrist cameras at 2x. Below the frames: | |
| - the label subgoal at that step (green; the policy does not see it), | |
| - the planner's latest step prediction, red if it differs from the label, | |
| - the number of graph nodes and the step of the last planner call. | |
| Green circle = label target, cyan x = planner target from the latest call. | |
| Outcomes in this sample: | |
| | Task | Episodes | | |
| |---|---| | |
| | BinFill | success, timeout | | |
| | ButtonUnmask | success, success | | |
| | ButtonUnmaskSwap | fail, fail | | |
| | InsertPeg | timeout, timeout | | |
| | MoveCube | success, success | | |
| | PatternLock | fail, success | | |
| | PickHighlight | success, success | | |
| | PickXtimes | success, timeout | | |
| | RouteStick | fail, success | | |
| | StopCube | fail, success | | |
| | SwingXtimes | success, success | | |
| | VideoPlaceButton | success, success | | |
| | VideoPlaceOrder | success, success | | |
| | VideoRepick | success, timeout | | |
| | VideoUnmask | success, success | | |
| | VideoUnmaskSwap | success, success | | |
| Two episodes per task is a small sample for showing behaviour. Use the 50-episode numbers in 06 for success rates. | |
| ## Scripts | |
| `tools/run_planner.py` loads the planner and runs it on cached episodes. `tools/make_figs.py` makes 01 to 06. `tools/annotate_rollouts.py` makes 07 from the rollout videos and the planner log. | |
Xet Storage Details
- Size:
- 7.3 kB
- Xet hash:
- 51c5bf8e1bd854fb4b045b521d51ff7cacf0c164ca200ca4f053bb8a8c1efe3c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.