odd-conditioned-dev / README.md
buzinguyen's picture
Model card
c9a41ea verified
|
Raw History Blame Contribute Delete
3.82 kB
---
license: mit
tags:
- robotics
- reinforcement-learning
- safe-reinforcement-learning
- reach-avoid
- unitree-go2
- mujoco
---
# ODD-conditioned safety filters β€” trained policies
Trained policies for **ODD-conditioned safety filters** on a Unitree Go2 in MuJoCo (mjlab): safety filters whose
guarantee survives runtime changes of the operating design domain β€” a payload loaded mid-mission, a motor that
derates. The filter is an automaton over specification modes (stand/walk ↔ rest) with certified transition
funnels between them.
Code, documentation and the evaluation that reproduces every result:
**[SafeRoboticsLab/odd-conditioned](https://github.com/SafeRoboticsLab/odd-conditioned/tree/cleanup)** (branch
`cleanup`). These weights are meant to be used with it.
## Use
```bash
# in a clone of the code repository, after `source activate.sh`
bash scripts/fetch_weights.sh https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/c923466bd1c68132b643f6806b9e8a2c532e4f1f/odd-conditioned-weights-v1.tar.gz
pytest -q tests/ && bash scripts/reproduce.sh payload figures
```
`fetch_weights.sh` installs the policies under `checkpoints/` and checks every file against the manifest
committed in the code repository (`weights/MANIFEST.sha256`, also included here).
## Files
| path | contents |
|---|---|
| `odd-conditioned-weights-v1.tar.gz` | everything below, as one archive (what `fetch_weights.sh` installs) |
| `checkpoints/<name>/model.zip` | a two-player reach-avoid PPO twin (`ReachAvoidPPO2P`, safety-stable-baselines 0.4.0): control policy + state-value net, whose sign is the mode's certificate |
| `checkpoints/<name>/tensornormalize.pt` | its frozen observation-normalization statistics (48-d actor observation) |
| `checkpoints/<name>/config.yaml` | the exact training configuration of the run |
| `checkpoints/{rest,getup_stage1}/final/` | final models used as warm starts when retraining `descend` and `getup` |
| `MANIFEST.sha256` | sha256 of every file |
| policy | role |
|---|---|
| `stand` | STAND expert under a carried load W ∈ [0, 120] N with a raised centre of mass; its value is `V_stand` |
| `rest` | REST expert: lie down and settle under any load β€” the anchor safe set |
| `getup` | certified get-up funnel REST β†’ STAND; its value `V_up` gates the return and the abort |
| `descend` | certified descent funnel STAND β†’ REST |
| `stand_wide` | STAND expert trained on W ∈ [0, 150] N (a noisier certificate; the single-spec comparison) |
| `unified`, `unified_discounted` | single-specification baselines (one policy for stand-or-rest) |
| `leg_stand` | STAND expert for a derated front-right leg (the negative control) |
| `compound_stand`, `compound_rest` | STAND / REST experts for a leg that derates while carrying 80 N |
The nominal walking policy is not here: it ships with
[go2_atomic_skills](https://github.com/SafeRoboticsLab/go2_atomic_skills).
## Training
Every policy was trained with the two-player safety-PPO recipe of safety-stable-baselines on tasks of
[robot-safety-sandbox](https://github.com/SafeRoboticsLab/robot-safety-sandbox/tree/project/odd-conditioned)
(branch `project/odd-conditioned`): 1024 environments, 50M steps, seed 0, a learned adversarial push of up to
25 N. The released checkpoint of each run is the one at 49,999,872 steps. `scripts/train.sh` in the code
repository retrains any of them.
**Training variance.** Each policy is a single training run, and the stance-type policies (`stand`,
`compound_stand`) vary a lot from run to run: retrained with other seeds, most stand experts fail the acceptance
test that the released one passes. The code repository documents an acceptance test (`scripts/check_policies.py`)
and a threshold-calibration script for retrained policies.
## License
MIT, as the code.