twanghcmut's picture
|
download
raw
3.84 kB
# Foundation Physics Graph Model — demo outputs
Perception outputs from DROID episodes (lab `AUTOLab`, task *"Put brick in drawer
shelf and close drawer"*). Everything here is generated by the repo's pipelines;
nothing is hand-edited.
## ⚠️ Read this before trusting the numbers
**A time-base bug was found after these were rendered and is still being fixed.**
PointWorld's scene-flow clip frame `t` maps to video/trajectory frame
`2 * (start + t)`, **not** `start + t` — the annotations run at half the video
rate. Verified across all 42 clips of all 5 episodes with joint-angle error
`0.00e+00`.
Consequences for the files here:
- `01_robot_mesh_overlay/`**unaffected.** Driven purely by `trajectory.h5`
indexed by video frame, which is genuinely 1:1.
- `02_object_velocity/`**affected.** The depth seed came from an annotation
frame paired with the wrong video frame. Treat the m/s figures as provisional.
- `03_object_pipeline/`**affected** (same reason), and the clip shown happens
to contain no robot–object contact, so `state=free` throughout.
---
## 01_robot_mesh_overlay
Franka Panda + Robotiq 2F-85 posed from recorded joint angles by URDF forward
kinematics, rendered with pyrender/OSMesa and composited on the real RGB, plus a
per-frame robot mask.
Uses **NVlabs/PointWorld**'s `franka_panda_robotiq_2f85.urdf`. An earlier attempt
with fairo/polymetis's `panda_robotiq_85.urdf` was wrong: it puts the flange at
`z=0.045` where the real `panda_link8` is at `0.107`. That 62 mm shortfall was
measured, not guessed — FK vs DROID's own recorded end-effector gave a residual of
exactly `[0, 0, 0.062] m` with **zero variance over 128 frames**, and
`0.107 − 0.045 = 0.062`.
| URDF | FK residual vs recorded EE |
| --- | --- |
| fairo/polymetis | 0.062 m (std 0) |
| NVlabs/PointWorld | **1.6e-8 m** (std 7e-9) |
## 02_object_velocity
SAM 3.1 (text prompt derived from the episode's task instruction) → TAPNext++
point tracking → rigid-body PnP lifting → metric 3D velocity in the robot-base
world frame.
Depth is needed at **one seed frame only**: a rigid body has 6 DOF, so the 2D
tracks determine every later pose. This replaced a per-frame depth lookup that
was both weaker and dependent on the dataset annotating the object at every
instant — which it does not (the nearest annotated point to the brick is 40.07 px
away, just outside the search radius).
## 03_object_pipeline
Five independently re-runnable stages, each leaving its artefacts on disk:
| dir | stage |
| --- | --- |
| `1_mask` | SAM 3.1 mask → tight RGBA crop |
| `2_mesh` | crop → mesh (proxy hull here; TRELLIS.2 not yet installed) |
| `3_align` | mask × depth → world point cloud → scale + 6-DOF fit |
| `4_act` | robot FK → per-frame object pose (free / pushed / grasped) |
| `5_render` | robot + object + point cloud, recorded and free-camera views |
Alignment quality on this clip: **RMSE 3.25 mm, inlier fraction 1.00, scale 1.009**.
`scene_freecam.mp4` orbits the world independently of the recording viewpoint —
the recorded camera cannot show whether the object sits at the right *depth*, so a
second viewpoint is the only honest check of the 3D placement.
The HUD's `push_gain=4.0x` is labelled **VISUALISATION ONLY, not physical**: it
amplifies pushed motion (5 cm of arm travel → 20 cm of object travel) so the
effect is visible. It is not a simulation result.
The object mesh here is a **convex hull fitted to the observed points**, not a
reconstruction — `MeshSource` records this so a proxy is never mistaken for real
geometry.
## 04_verification_images
Alignment checks drawn as mask outlines over the *raw* frames. A filled overlay
hides misalignment by covering the thing it should match, so these outline views
are what the robot alignment was actually judged on.

Xet Storage Details

Size:
3.84 kB
·
Xet hash:
492ea699cf19f31c790cf01cc98cb281b5aaf890d6b51e6543091ab828534eda

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.