Buckets:
| # Method | |
| One method, two regimes. A **demonstration graph** built from the training demos is the only | |
| model of "where the arm should be"; a small trained network turns a window of joint states into a | |
| distribution over its nodes. At **t=0** that distribution is used once, to rewind an out-of- | |
| distribution start pose. At **t>0** it is filtered over time and used to decide when the rollout has | |
| left the demonstrated manifold and where to put it back. | |
| Nothing here is policy-specific: the graph and its head are trained once per suite from demo data | |
| alone, and the same artifacts drive both a frozen StableVLA and a frozen GR00T-N1.7 with no | |
| retraining of either policy. | |
| --- | |
| ## 1. The demonstration graph | |
| **Nodes.** Raw demo frames are coarsened `COARSEN = 5` frames to a node, with a forced split at every | |
| gripper-state change (`onf/graph/nodes.py`). Each node carries its joint configuration `q`, finite- | |
| differenced velocity, gripper flag, and three labels used everywhere downstream: | |
| | label | meaning | | |
| |---|---| | |
| | `task_id` | which base task the demo belongs to — the **task lane** | | |
| | `owner` | which demo (strand) it came from | | |
| | `phase` | position along the episode, `arange(T)/(T-1)`, so `0.0` = first frame, `1.0` = last | | |
| The `long` suite's graph has **28,476 nodes** over 500 demos / 10 tasks. | |
| **Edges** (`onf/graph/edges.py`, `schema.RELATIONS`): temporal successors at dilations | |
| `next1, next2, next4, next8, next16` (plus their reverses), `sibling` (kNN across demos of the same | |
| task, degree `K_SIBLING = 8`), and `align` (kNN within a phase bin, degree `K_ALIGN = 4`). | |
| **Head** (`onf/graph/gnn.py`, trained by `onf/graph/train.py` → `g_head.npz`). Input is a window of | |
| `HIST_H = 8` consecutive joint states. A query encoder seeds node states with a joint-space kernel | |
| (bandwidth `KERNEL_BW = 0.15`, sparsified to `SEED_TOPK = 256` nodes), `LAYERS = 3` rounds of | |
| relation-typed message passing follow, and a logit head scores every node. Nodes never reached by | |
| message passing get `-inf`, which is the head's *reachability* mask. | |
| The objective is **phase-primary**: cross-entropy on the phase-bin marginal `p(T|Q)` | |
| (`PHASE_CE_W = 1.0`) plus a phase-expectation regression (`PHASE_EXPECT_W = 0.5`), with node identity | |
| only auxiliary (`NODE_W = 0.1`). This ordering is deliberate — which *demo* a window came from is | |
| close to unlearnable and barely matters (same-task, same-phase demos sit ~0.05–0.6 rad apart), while | |
| *how far along* is both learnable and decisive (a phase error of 0.5 costs ~0.5 rad of target error). | |
| Training queries also include an `ENTRY_STATIC` class: windows of the form `repeat(q0, 8)`, i.e. | |
| identically zero velocity. Ordinary training windows are consecutive demo frames and always carry | |
| some velocity, so without this class the t=0 query shape never appears in training at all. | |
| --- | |
| ## 2. t = 0 — `entry` | |
| **WHEN.** Purely geometric, no learning: at episode start compute `qerr0 = ||q_0 − q*||`, and act only | |
| if `qerr0 > JQ_GATE = 0.08` rad. An unperturbed start is left completely untouched. | |
| **WHERE.** The head's distribution is masked down to the **entry stratum** | |
| ``` | |
| task lane ∩ phase ≤ JQ_GRAPH_ENTRY_BAND (0.05) ∩ reachable (logit > −inf) | |
| ``` | |
| and the survivors are averaged **uniformly** in joint space (`EucRawReadout`): | |
| `q* = Σ_v w_v · q_v` with `w_v = 1/n`. On the `long` graph this is ~229 of 28,476 nodes. | |
| So at t=0 the trained head contributes exactly one thing — reachability — and the mask, not the | |
| weighting, determines the answer. This was measured directly: reading the head's posterior over the | |
| same masked pool instead of averaging uniformly moves `q*` by 0.124 rad (larger than the gate itself) | |
| and yet does not improve success rate. The fixed rule is kept because it is simpler and does not | |
| depend on which head is loaded. | |
| **How it moves.** A joint-space PD servo (`onf/recovery/controller.py`, `JQ_MODE=pd`, `JQ_KP = 80`) | |
| drives the arm to `q*` with tolerance `JQ_EPS = 0.03` rad and a hard budget of `JQ_MAX = 60` control | |
| steps, then hands control to the frozen policy. The target is a **segment** (`SEG_K = 8` frames), not | |
| a point, so the policy receives a state that is in-distribution in both `q` and `qdot`. | |
| **`JQ_POSTURE=target` is load-bearing.** The operational-space controller keeps a null-space posture | |
| reference. If it is left at its default the servo lands correctly and then the null-space torque drags | |
| the arm back toward the perturbed pose within ~2 s of handoff, undoing the recovery. Pinning the | |
| reference to the pose the servo reached is what makes the whole regime work. | |
| --- | |
| ## 3. t > 0 — `sentinel` | |
| A per-worker `GraphTracker` (`onf/graph/track.py`) runs a belief filter over the same graph, one | |
| check every `SN_GRAPH_STRIDE = 1` window (matching the 8-raw-frame training window): | |
| ``` | |
| b_t ∝ [ (1 − ε) · P b_{t−1} + ε · L_t ] · L_t | |
| ``` | |
| - `L_t = softmax(logits)` — the **same trained head** as t=0, no separate model. | |
| - `P` is built from the graph's own relations, restricted to | |
| `TRACK_KERNEL_RELATIONS = (next1, next2, next4, next8, next16, sibling)`. | |
| Advances are mixed over `TRACK_ADVANCE_SET = (0,1,2,4,6,8,12,16)` raw frames because a check | |
| advances ~1.65 nodes, which integer node hops cannot express, and rollouts run slower than demos. | |
| `align` and backward relations are deliberately excluded: the arm does not change task mid-episode, | |
| and a backward jump must pay transition cost. | |
| - The restart term `ε · L_t` restarts toward the **current observation**, not toward uniform — a | |
| uniform restart puts mass on impossible phases and discards the aliasing suppression being bought. | |
| Why filter at all: the retriever is memoryless, and `long` tasks are two-object ("put both X and Y"), | |
| so reaching for object 2 looks kinematically like reaching for object 1. Over 1,372 measured | |
| consecutive checks only 67% of the memoryless phase transitions are physically plausible | |
| (`|Δphase| ≤ 0.05`); 8.8% teleport by >0.3 and 8.0% jump *backwards* by >0.1. | |
| **WHEN** (`FilterRule`, `onf/sentinel/when.py`). Two e-processes accumulated by the tracker — a | |
| **basin** test (is the arm inside the certified radius of the demo manifold?) and a **progress** test | |
| (is phase still advancing?) — each fire when `log E ≥ −log(α_half)` (Ville's inequality). The rule | |
| fires on their union, subject to three gates: | |
| - `α_half = α / 2 / max_interventions` (α = `TRACK_ALPHA_EVIDENCE` = 0.05). Resetting `E_t` after a | |
| fire refunds spent α, so the per-episode false-alarm guarantee only holds if the budget is split | |
| across the allowed fires. Without this split the certified claim is simply false. | |
| - The progress test is suppressed once `phase_hat ≥ 1 − 1/NBINS_ALIGN`: a task that legitimately | |
| finished has stopped advancing, and that is not a stall. | |
| - No fire while the gripper is closed — interrupting mid-grasp drops the object. The evidence | |
| bookkeeping keeps running; only the fire is suppressed. | |
| **WHERE** (`BasinReadout`). The belief `b_t` is projected onto the basin geometry fitted per | |
| `(task, phase-bin)` from held-out clean demo replays: certified radius `r` = p95 and KDE bandwidth | |
| `h` = p50 of leave-one-demo-out nearest-neighbour distances. The same PD servo then runs to that | |
| target. | |
| **Artifacts** are fit offline by `scripts/build_sentinel_artifacts.py` into `g_track.npz` | |
| (`π`, `β`, `leak`, the three null banks, `basin_r`, `basin_h`). The file carries a `graph_hash` stamp | |
| and is **refused on mismatch**: the advance operator is node-id positional, so a foreign file would | |
| advance the belief along the wrong strands with no shape error and no exception. | |
| --- | |
| ## 4. Constants: chosen vs derived | |
| | constant | value | how | | |
| |---|---|---| | |
| | `JQ_GATE` | 0.08 rad | CHOSEN | | |
| | `JQ_GRAPH_ENTRY_BAND` | 0.05 phase | CHOSEN (one `NBINS_ALIGN` bin) | | |
| | `JQ_KP` / `JQ_EPS` / `JQ_MAX` | 80 / 0.03 / 60 | CHOSEN | | |
| | `HIST_H`, `SEG_K` | 8, 8 | CHOSEN — one policy action chunk | | |
| | `COARSEN` | 5 | CHOSEN; splits forced at gripper changes | | |
| | `TRACK_ALPHA_EVIDENCE/PROGRESS` | 0.05 | CHOSEN (false-alarm budget) | | |
| | `TRACK_LAM` | 0.8 | CHOSEN | | |
| | `α_half` | `α/2/max_interventions` | DERIVED (union bound over fires) | | |
| | `BASIN_RADIUS_QUANTILE` / `BANDWIDTH` | p95 / p50 | DERIVED from LODO demo replays | | |
| | `π`, `β`, `leak` | per suite | DERIVED — fit by `build_sentinel_artifacts.py` | | |
| | `TRACK_ADVANCE_SET` upper end | 16 raw frames | DERIVED (~2× nominal advance rate) | | |
| --- | |
| ## 5. Results | |
| Every number below is on LIBERO-Plus. Nothing else in this repo carries measured numbers. | |
| | cell | StableVLA | GR00T-N1.7 | | |
| |---|---|---| | |
| | **t=0 `entry`**, `long` × Robot_Initial_States (n=393) | 378/393 = **96.2%** | 353/393 = **89.8%** | | |
| | **t>0 `sentinel`**, `long` × Objects_Layout (n=312) | not measured | 208/312 = **66.7%** | | |
| Read with the following caveats, all of which matter: | |
| - **The 66.7% cell has no same-session `base` baseline.** The nearest comparator (204/312) was run | |
| four days earlier; the difference is +1.3 pp. GR00T-N1.7 is flow-matching and **stochastic** — | |
| ~11.2% episode churn, ±2.4 pp on n≈393 — so +1.3 pp is inside noise, and cross-session GR00T | |
| comparisons are not valid at all. Only same-session arms may be compared. | |
| - StableVLA is deterministic (0 episode churn), so its arms can be compared paired and exactly, even | |
| across sessions. | |
| - **On GR00T at Robot-Init, a fixed-home baseline (94.7%) still beats this method (89.8%).** The | |
| graph-based entry target wins on StableVLA but not on GR00T. | |
| ### Not measured | |
| - `sentinel` has exactly **one** measured cell (GR00T, `long` × Objects_Layout) and **no** full-suite | |
| number on StableVLA. `configs/sr_ladder.yaml` nevertheless makes it the default for all six t>0 | |
| axes — that is an extrapolation beyond what has been measured, stated here rather than hidden. | |
| - Only the `long` suite currently has a trained `g_head.npz` and a fitted `g_track.npz`. `object`, | |
| `spatial` and `goal` have `g_nodes.npz`/`g_edges.npz` only, so the 4-suite benchmark is complete in | |
| code but cannot run on three of four suites until those artifacts are trained | |
| (see `docs/usage.md`). | |
| - `ENTRY_STATIC`'s effect on the t>0 regime has never been ablated. It is retained because the | |
| shipped `g_head.npz` was trained with it. | |
| - `src/onf/field/` (`ONFField`) is a **training-time dependency only** — it supplies the cleanliness | |
| weighting in `onf/graph/train.py`. It is on no inference path in either regime. | |
Xet Storage Details
- Size:
- 10.5 kB
- Xet hash:
- 689c50d651eb808a7959cdd0597166cd0194cdb5c97b14c5c0548a3abe9d062c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.