twanghcmut's picture
|
download
raw
5.76 kB

Usage

Copy-pasteable path from a fresh checkout to a scored benchmark. Every command assumes the repo root as cwd. Always set OMP_NUM_THREADS=4 — without it the training loop oversubscribes threads and slows down by an order of magnitude.


0. Environments

Four separate runtimes; they cannot be merged (conflicting pins). See envs/README.md.

env file used for
onf-core envs/onf-core.yml the onf package: graph build/train, artifacts, tests
stablevla envs/stablevla.yml the StableVLA policy server
libero-plus via scripts/setup_external.sh the simulator client
gr00t separate conda env, path given by GR00T_PY the GR00T-N1.7 policy server
conda env create -f envs/onf-core.yml
pip install -e .                       # inside onf-core
bash scripts/setup_external.sh         # links LIBERO-plus and the vendored deps into external/

Data and checkpoints are not downloaded by the setup script. Point the repo at them with:

export ONF_DATA=/path/to/data          # demo hdf5 + per-suite fwm/ artifacts
export ONF_OUTPUTS=/path/to/outputs    # graph artifacts land here
export ONF_RESULTS=/path/to/results    # rollout mp4s land here

onf.config.Paths resolves everything from these three; nothing else is hardcoded.


1. Build the demonstration graph

OMP_NUM_THREADS=4 python -m onf.graph build --suite long

Writes g_nodes.npz + g_edges.npz into $ONF_OUTPUTS/long/<run>/artifacts/, and points $ONF_OUTPUTS/long/latest at it. Use --suite all for all four suites, or --limit-demos 8 --limit-tasks 2 for a fast smoke build.

2. Train the retrieval head

OMP_NUM_THREADS=4 python -m onf.graph train --suite long

Writes g_head.npz next to the graph it trained on. Roughly one hour per suite on one GPU. Add --epochs 1 for a smoke run.

3. Fit the sentinel artifacts

OMP_NUM_THREADS=4 python scripts/build_sentinel_artifacts.py \
    "$ONF_OUTPUTS/long/latest/artifacts" --suite long

Fits the transition kernel (π, β, leak), the three null banks and the basin geometry into g_track.npz, plus a sentinel_artifacts.json provenance sidecar.

t=0 needs steps 1–2; t>0 needs 1–3. A missing g_track.npz is a hard error at launch, not a silent fallback. Today only long has a trained head and a fitted tracker; object, spatial and goal need steps 2–3 run once each before they can be benchmarked.


4. Benchmark

Three modes exist, and only three: base (frozen policy, no recovery), entry (t=0), sentinel (t>0). They are defined in one place, evals/common/modes.sh.

Result and log names share one grammar:

results/plus_{bench}/{axis}/{policy}_{mode}_{suite}_{axis}/     # rollout mp4s + run.json
logs/{utc_timestamp}_{policy}_{mode}_{suite}_{axis}/            # server.log + client_shard*.log

StableVLA

# smoke first (12 episodes per regime)
OMP_NUM_THREADS=4 python scripts/run_sr.py --rung R1 --gpus 2,3

# the full 4 suites x 7 axes grid
OMP_NUM_THREADS=4 python scripts/run_sr.py --rung R4 --gpus 2,3

Useful flags: --dry-run (print the schedule and exit), --check-env (validate paths/ports only), --resume (skip cells already complete in the ledger), --mode <mode> (override the per-axis default from configs/sr_ladder.yaml), and --cells suite:axis[,...] — on its own it builds ad hoc full-axis cells; combined with --rung it runs a subset of that rung, keeping each cell's own episode filter, expected count, comparators and gate:

# just the two `long` smoke cells (12 episodes each), one per regime
python scripts/run_sr.py --rung R1 --cells long:Robot_Initial_States,long:Sensor_Noise --gpus 2,3

GR00T-N1.7

GR00T_PY=/path/to/gr00t/env/bin/python \
OMP_NUM_THREADS=4 python evals/gr00t/run_gr00t.py \
    --suites object,spatial,goal,long --mode sentinel --gpus 2,3

Refuses to start if g_track.npz is missing, and refuses to write into a result directory that already holds rollouts unless --resume/--force — re-running a cell in place would otherwise leave both the success and the failure mp4 of a flipped episode and inflate n.

Score

# success-rate grid; --methods takes name:tag_prefix pairs
python evals/libero_plus/score.py --suites long \
    --methods base:stablevla_base entry:stablevla_entry sentinel:stablevla_sentinel

# paired McNemar between two arms of one suite x axis
python evals/libero_plus/mcnemar.py --suite 10 --axis Sensor_Noise \
    --tags stablevla_base_long_Sensor_Noise stablevla_sentinel_long_Sensor_Noise

McNemar is only meaningful for a deterministic policy (StableVLA). GR00T-N1.7 is stochastic (~11.2% episode churn); compare its arms only within one session, and read the sign, not the digits.


5. Reproducibility

OMP_NUM_THREADS=4 PYTHONPATH=src pytest tests/ -q
OMP_NUM_THREADS=4 PYTHONPATH=src pytest tests/test_parity.py -q     # must be 2 passed

tests/test_parity.py is the golden bit-parity net: it replays 64 recorded t=0 retrievals and 64 t>0 tracker steps against the real trained artifacts and asserts every float is exactly equal (via float.hex(), not approx) to tests/golden/parity.json. It skips cleanly when the artifacts are absent, so a fresh clone never fails it.

Every result directory also carries a run.json recording policy, mode, suite, axis, expected episode count, graph dir, graph_hash, head mtime, git commit and the full mode env — so any number can be traced back to the exact code and artifacts that produced it.

GPU etiquette on this box

GPUs 0 and 1 belong to another user. Use --gpus 2,3 only.

Xet Storage Details

Size:
5.76 kB
·
Xet hash:
38e30ac795062bca59b12f12a766e67792a01e55b8e1ccbe7295649e02b97fc5

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.