|
Download README.md from buzinguyen/odd-conditioned-dev: direct link, hf CLI and curl.
- Browser
- Download file 3.82 kB
-
https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/main/README.md
- Command line
-
hf download hf://buzinguyen/odd-conditioned-dev/README.md
-
curl -L -o README.md https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/main/README.md
3.82 kB
| license: mit | |
| tags: | |
| - robotics | |
| - reinforcement-learning | |
| - safe-reinforcement-learning | |
| - reach-avoid | |
| - unitree-go2 | |
| - mujoco | |
| # ODD-conditioned safety filters β trained policies | |
| Trained policies for **ODD-conditioned safety filters** on a Unitree Go2 in MuJoCo (mjlab): safety filters whose | |
| guarantee survives runtime changes of the operating design domain β a payload loaded mid-mission, a motor that | |
| derates. The filter is an automaton over specification modes (stand/walk β rest) with certified transition | |
| funnels between them. | |
| Code, documentation and the evaluation that reproduces every result: | |
| **[SafeRoboticsLab/odd-conditioned](https://github.com/SafeRoboticsLab/odd-conditioned/tree/cleanup)** (branch | |
| `cleanup`). These weights are meant to be used with it. | |
| ## Use | |
| ```bash | |
| # in a clone of the code repository, after `source activate.sh` | |
| bash scripts/fetch_weights.sh https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/c923466bd1c68132b643f6806b9e8a2c532e4f1f/odd-conditioned-weights-v1.tar.gz | |
| pytest -q tests/ && bash scripts/reproduce.sh payload figures | |
| ``` | |
| `fetch_weights.sh` installs the policies under `checkpoints/` and checks every file against the manifest | |
| committed in the code repository (`weights/MANIFEST.sha256`, also included here). | |
| ## Files | |
| | path | contents | | |
| |---|---| | |
| | `odd-conditioned-weights-v1.tar.gz` | everything below, as one archive (what `fetch_weights.sh` installs) | | |
| | `checkpoints/<name>/model.zip` | a two-player reach-avoid PPO twin (`ReachAvoidPPO2P`, safety-stable-baselines 0.4.0): control policy + state-value net, whose sign is the mode's certificate | | |
| | `checkpoints/<name>/tensornormalize.pt` | its frozen observation-normalization statistics (48-d actor observation) | | |
| | `checkpoints/<name>/config.yaml` | the exact training configuration of the run | | |
| | `checkpoints/{rest,getup_stage1}/final/` | final models used as warm starts when retraining `descend` and `getup` | | |
| | `MANIFEST.sha256` | sha256 of every file | | |
| | policy | role | | |
| |---|---| | |
| | `stand` | STAND expert under a carried load W β [0, 120] N with a raised centre of mass; its value is `V_stand` | | |
| | `rest` | REST expert: lie down and settle under any load β the anchor safe set | | |
| | `getup` | certified get-up funnel REST β STAND; its value `V_up` gates the return and the abort | | |
| | `descend` | certified descent funnel STAND β REST | | |
| | `stand_wide` | STAND expert trained on W β [0, 150] N (a noisier certificate; the single-spec comparison) | | |
| | `unified`, `unified_discounted` | single-specification baselines (one policy for stand-or-rest) | | |
| | `leg_stand` | STAND expert for a derated front-right leg (the negative control) | | |
| | `compound_stand`, `compound_rest` | STAND / REST experts for a leg that derates while carrying 80 N | | |
| The nominal walking policy is not here: it ships with | |
| [go2_atomic_skills](https://github.com/SafeRoboticsLab/go2_atomic_skills). | |
| ## Training | |
| Every policy was trained with the two-player safety-PPO recipe of safety-stable-baselines on tasks of | |
| [robot-safety-sandbox](https://github.com/SafeRoboticsLab/robot-safety-sandbox/tree/project/odd-conditioned) | |
| (branch `project/odd-conditioned`): 1024 environments, 50M steps, seed 0, a learned adversarial push of up to | |
| 25 N. The released checkpoint of each run is the one at 49,999,872 steps. `scripts/train.sh` in the code | |
| repository retrains any of them. | |
| **Training variance.** Each policy is a single training run, and the stance-type policies (`stand`, | |
| `compound_stand`) vary a lot from run to run: retrained with other seeds, most stand experts fail the acceptance | |
| test that the released one passes. The code repository documents an acceptance test (`scripts/check_policies.py`) | |
| and a threshold-calibration script for retrained policies. | |
| ## License | |
| MIT, as the code. | |