--- license: mit tags: - robotics - reinforcement-learning - safe-reinforcement-learning - reach-avoid - unitree-go2 - mujoco --- # ODD-conditioned safety filters — trained policies Trained policies for **ODD-conditioned safety filters** on a Unitree Go2 in MuJoCo (mjlab): safety filters whose guarantee survives runtime changes of the operating design domain — a payload loaded mid-mission, a motor that derates. The filter is an automaton over specification modes (stand/walk ↔ rest) with certified transition funnels between them. Code, documentation and the evaluation that reproduces every result: **[SafeRoboticsLab/odd-conditioned](https://github.com/SafeRoboticsLab/odd-conditioned/tree/cleanup)** (branch `cleanup`). These weights are meant to be used with it. ## Use ```bash # in a clone of the code repository, after `source activate.sh` bash scripts/fetch_weights.sh https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/c923466bd1c68132b643f6806b9e8a2c532e4f1f/odd-conditioned-weights-v1.tar.gz pytest -q tests/ && bash scripts/reproduce.sh payload figures ``` `fetch_weights.sh` installs the policies under `checkpoints/` and checks every file against the manifest committed in the code repository (`weights/MANIFEST.sha256`, also included here). ## Files | path | contents | |---|---| | `odd-conditioned-weights-v1.tar.gz` | everything below, as one archive (what `fetch_weights.sh` installs) | | `checkpoints//model.zip` | a two-player reach-avoid PPO twin (`ReachAvoidPPO2P`, safety-stable-baselines 0.4.0): control policy + state-value net, whose sign is the mode's certificate | | `checkpoints//tensornormalize.pt` | its frozen observation-normalization statistics (48-d actor observation) | | `checkpoints//config.yaml` | the exact training configuration of the run | | `checkpoints/{rest,getup_stage1}/final/` | final models used as warm starts when retraining `descend` and `getup` | | `MANIFEST.sha256` | sha256 of every file | | policy | role | |---|---| | `stand` | STAND expert under a carried load W ∈ [0, 120] N with a raised centre of mass; its value is `V_stand` | | `rest` | REST expert: lie down and settle under any load — the anchor safe set | | `getup` | certified get-up funnel REST → STAND; its value `V_up` gates the return and the abort | | `descend` | certified descent funnel STAND → REST | | `stand_wide` | STAND expert trained on W ∈ [0, 150] N (a noisier certificate; the single-spec comparison) | | `unified`, `unified_discounted` | single-specification baselines (one policy for stand-or-rest) | | `leg_stand` | STAND expert for a derated front-right leg (the negative control) | | `compound_stand`, `compound_rest` | STAND / REST experts for a leg that derates while carrying 80 N | The nominal walking policy is not here: it ships with [go2_atomic_skills](https://github.com/SafeRoboticsLab/go2_atomic_skills). ## Training Every policy was trained with the two-player safety-PPO recipe of safety-stable-baselines on tasks of [robot-safety-sandbox](https://github.com/SafeRoboticsLab/robot-safety-sandbox/tree/project/odd-conditioned) (branch `project/odd-conditioned`): 1024 environments, 50M steps, seed 0, a learned adversarial push of up to 25 N. The released checkpoint of each run is the one at 49,999,872 steps. `scripts/train.sh` in the code repository retrains any of them. **Training variance.** Each policy is a single training run, and the stance-type policies (`stand`, `compound_stand`) vary a lot from run to run: retrained with other seeds, most stand experts fail the acceptance test that the released one passes. The code repository documents an acceptance test (`scripts/check_policies.py`) and a threshold-calibration script for retrained policies. ## License MIT, as the code.