Buckets:
Setup
Copy-paste, top to bottom.
onf/ ← this repo
├── stablevla/ 1. clone the policy
│ └── hf_weights/{pretrained_models,object,spatial,goal,long}/ 3. weights
├── external/
│ ├── LIBERO-plus/ 1. clone the benchmark
│ └── Isaac-GR00T/ 1. clone optional, cross-policy arm
│ └── ckpts_n17_libero/{libero_object,libero_spatial,libero_goal,libero_10}/ 3. weights
├── data/ or $ONF_DATA if set
│ └── libero_datasets/{libero_object,libero_spatial,libero_goal,libero_10}/ 2. demos
├── outputs/ or $ONF_OUTPUTS — built in step 5
└── results/ or $ONF_RESULTS — rollout mp4s
long is called libero_10 everywhere outside this repo. Set these once:
export OMP_NUM_THREADS=4 PYTHONPATH=src
OMP_NUM_THREADS=4 is not optional — without it the training loop oversubscribes threads and runs an
order of magnitude slower.
ONF_DATA / ONF_OUTPUTS / ONF_RESULTS are optional. onf.config.Paths reads them if set and
otherwise falls back to the repo-local data/, outputs/, results/ — so leave them unset unless
those trees live somewhere else:
export ONF_DATA=/mnt/big/data # only if not $REPO/data
Ask the package where a tree resolved rather than assuming; this is the one command that always tells the truth:
python -c "from onf.config import default_paths as d; p=d(); print(p.data); print(p.graph('long'))"
1. Clone
git clone <this-repo> onf && cd onf
# the policy — only the StableVLA subtree of the HumanNet monorepo
git clone --filter=blob:none --no-checkout https://github.com/DAGroup-PKU/HumanNet.git /tmp/hn
git -C /tmp/hn sparse-checkout init --cone
git -C /tmp/hn sparse-checkout set src/model/StableVLA
git -C /tmp/hn checkout 68eb86d555fc4d3b09003319cff7644cd4c8b571
mkdir -p stablevla && cp -a /tmp/hn/src/model/StableVLA/. stablevla/ && rm -rf /tmp/hn
# the benchmark — the LIBERO fork with the perturbation suites
git clone https://github.com/sylvestf/LIBERO-plus.git external/LIBERO-plus
git -C external/LIBERO-plus checkout 4976dc30028e805ff8094b55501d532c48fec182
# optional: the cross-policy arm
git clone https://github.com/NVIDIA/Isaac-GR00T.git external/Isaac-GR00T
Commits are pinned in UPSTREAM.md.
2. Environments
Three, per each project's own README. They cannot be merged — the pins conflict.
# this repo's package. CPU is enough.
conda env create -f envs/onf-core.yml && conda activate onf-core && pip install -e .
# the policy AND the simulator client. Per stablevla/README.md.
conda create -n stablevla python=3.10.16 -y && conda activate stablevla
pip install torch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0
pip install -e stablevla
pip install packaging ninja && pip install "flash-attn==2.5.5" --no-build-isolation
pip install -e external/LIBERO-plus # ← the fork, NOT upstream LIBERO
# optional, GR00T only. Per external/Isaac-GR00T/README.md.
cd external/Isaac-GR00T && python3 -m venv .venv && .venv/bin/pip install -e . && cd -
export GR00T_PY=$PWD/external/Isaac-GR00T/.venv/bin/python
$GR00T_PY -m pip install websockets
Three things that bite if you deviate:
- The conda env must be named
stablevla.scripts/run_sr.pyresolves the interpreter by env name, not by$CONDA_PREFIX. StableVLA's README already names it that. external/LIBERO-plus, not upstream LIBERO. StableVLA's README points atLifelong-Robot-Learning/LIBERO. The perturbation axes exist only in the fork; upstream imports fine and then fails atget_ids_by_categorywhen the benchmark is constructed.- Headless rendering needs
MUJOCO_GL=osmesa PYOPENGL_PLATFORM=osmesa. The drivers set it themselves; export it only if you invoke the harness by hand.
envs/stablevla.yml is a frozen capture of a working env — use it if flash-attn fights you.
3. Weights
In the stablevla env (that is where huggingface_hub lives). On newer hub versions
huggingface-cli is spelled hf.
huggingface-cli download DAGroup-PKU/StableVLA \
--local-dir stablevla/hf_weights \
--include "pretrained_models/*" "long/*"
--include takes as many suites as you want. Omit it and you also pull calvin/, ~5 GB nothing here
loads. StableVLA's README says to put these under outputs/ — don't. configs/suites.yaml reads
stablevla/hf_weights/<suite>.
huggingface-cli download nvidia/GR00T-N1.7-LIBERO \
--local-dir external/Isaac-GR00T/ckpts_n17_libero \
--include "libero_10/*" \
--exclude "*/global_step*" "*/rng_state_*" "*/trainer_state.json" \
"*/training_args.bin" "*/scheduler.pt" "*/zero_to_fp32.py"
The --exclude matters. That repo is 143 GB because every suite ships its full DeepSpeed training
state, which inference never reads: 35.7 GB → 6.5 GB a suite. Measured against the live repo,
these patterns select 13 of a suite's 203 files.
Checkpoints elsewhere? export GR00T_CKPT_ROOT=/path/to/parent, nothing else moves.
4. Demonstrations
Raw LIBERO demo HDF5 — the *_demo.hdf5 files, not the RLDS conversion StableVLA trains on. Still in
the stablevla env, which is where the libero package lives; PYTHONPATH=src from the top is what
makes onf.config importable there. One call per suite (libero_object, libero_spatial,
libero_goal, libero_10):
python -c "
from onf.config import default_paths
from libero.libero.utils.download_utils import download_from_huggingface
dest = default_paths().data / 'libero_datasets'
download_from_huggingface('libero_10', str(dest), check_overwrite=False)
print('->', dest)"
The destination comes from onf.config.Paths, the same resolver every other command uses. Do not
interpolate $ONF_DATA into the path yourself: it is optional, so on a shell where it is unset the
string expands to empty and the download lands in /libero_datasets — the filesystem root.
Each suite lands 10 *_demo.hdf5; fewer means the download was interrupted, so re-run it. Do not
use external/LIBERO-plus/benchmark_scripts/download_libero_datasets.py — it prompts on stdin, and
its libero_100 choice does not match the libero_10 folder the HuggingFace repo publishes.
5. Build artifacts
Four commands per suite, ~1.5 h on one GPU. Each consumes only what the one above wrote.
conda activate onf-core
export SUITE=long DEVICE=cuda:0
python -m onf.field.build suite --suite $SUITE # onf_head.npz ~10 s CPU
python -m onf.graph build --suite $SUITE # g_nodes, g_edges min CPU
python -m onf.graph train --suite $SUITE --device $DEVICE # g_head.npz ~1 h GPU
ART=$(python -c "from onf.config import default_paths; print(default_paths().graph('$SUITE'))")
python scripts/build_sentinel_artifacts.py "$ART" --suite $SUITE --device $DEVICE # g_track.npz
Run all four whenever any one artifact is missing. The field is the query encoder the head is trained against, and the graph's node ids are what the tracker's advance operator indexes, so rebuilding an earlier link invalidates every later one. A partially rebuilt set loads without error and answers with the wrong strands.
6. Run
python scripts/run_sr.py --check-env # preflight — catches a bad step 1-5
python scripts/run_sr.py --rung R1 --gpus 0 # smoke: three small instance-filtered cells
python evals/gr00t/run_gr00t.py --suites long --gpu 0 --num-clients 6 # the GR00T arm
--check-env resolves the stablevla interpreter, imports libero.libero.benchmark out of
external/LIBERO-plus, and asserts the per-suite episode counts match configs/sr_ladder.yaml. If it
passes, the setup is correct.
Scoring, the full 4 × 7 grid and the McNemar caveats: RUNBOOK.md.
Easy to get wrong
- StableVLA weights under
outputs/instead ofhf_weights/. §3. - Upstream LIBERO instead of the LIBERO-Plus fork. §2.
- Downloading the whole GR00T repo — 143 GB instead of 26 GB. §3.
- Renaming the conda env. §2.
- Omitting
OMP_NUM_THREADS=4. Not a correctness bug, just an order of magnitude slower. - Mixing artifacts across graph builds. A
g_track.npzfrom a different build is node-id positional and advances the belief along the wrong strands with no shape error. Thegraph_hashstamp is the only defence, and it is checked on load. Seetechnical/05-artifacts.md.
Xet Storage Details
- Size:
- 8.84 kB
- Xet hash:
- 61ebaade18daeb3c0ac808eed27f16ae8c82c5040501535de45e5ba8484ce506
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.