File size: 8,520 Bytes
6bdeb28 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 | ---
license: mit
pipeline_tag: robotics
tags:
- reinforcement-learning
- robot-manipulation
- ppo
- maniskill
- action-interface
- domain-adaptation
- system-identification
---
# ActionShift baselines
Reference checkpoints for **ActionShift**, a benchmark for manipulation policies adapting to a hidden
action-interface contract. Code: [github.com/Archerkattri/actionshift](https://github.com/Archerkattri/actionshift).
Splits, eval configs, and the rollout logs these checkpoints produced:
[kattri15/actionshift](https://huggingface.co/datasets/kattri15/actionshift).
These are **reference implementations from one benchmark run, not claims of best-in-class performance.**
Exact numbers, intervals, and the honesty caveats for every method are in `reports/*.md` in the GitHub
repo, in particular `reports/gate1.md` and `reports/adaptation_tournament.md`.
## What's here
```
ppo/*_final_ckpt.pt frozen Gate-1 competent PPO backbones, one per task
learned/*.pt unprivileged learned contract-identifier regressors
delay_aware/*.pt delay-aware augmented-state PPO backbones (long_lag split)
```
### `ppo/` β the frozen base policies
The task backbone every adaptation method sits on top of. Official ManiSkill v3.0.1 PPO baseline,
control mode `pd_ee_delta_pose`, trained per task, then frozen for the whole tournament. This is
**a fixed reference policy competent at the clean task, not a SOTA manipulation policy** β it is
deliberately unmodified so that the only variable across experiments is the action-interface contract,
not the backbone.
| File | Task | sha256 |
|---|---|---|
| `ppo/pick_cube_final_ckpt.pt` | PickCube-v1 | `3e6c95d63a2132843323e24cf7ba962b8cf2610f04b2a5a43f3efb6fef8497a8` |
| `ppo/push_cube_final_ckpt.pt` | PushCube-v1 | `a4a02198b309e73cb877959079023d967d5f63ec78380de9703a10c9efafc0cf` |
| `ppo/pull_cube_final_ckpt.pt` | PullCube-v1 | `74ae6a09b9af5e9e50dc71944f2e99316a8b67b02f3a96ca45df4a6d53dc1bd7` |
| `ppo/stack_cube_final_ckpt.pt` | StackCube-v1 | `e63cc8d8ffdca3d03553a21ea615c759b2b224493a7e7e12bee7efc29d5bad9c` |
Clean-task success (no hidden contract, i.e. identity): Pick and Push both 1.000 on `seen`. Under a
hidden contract with no adaptation, all four collapse to the 0.000-0.007 floor β that collapse is the
premise the rest of the benchmark measures adaptation against. PegInsertionSide-v1 has **no** frozen
checkpoint here: it never cleared the Gate 0 competence floor at the official training budget and is
excluded from the benchmark (see `reports/peg_retry.md` on GitHub).
### `learned/` β unprivileged learned contract identifiers
No pool, no grammar enumeration: these regressors learn to identify the active contract from calibrated
response alone. Trained on 96 contracts / 34,560 env steps / 80 epochs (`osi`, `osi_random_excitation`,
`recurrent`), or 96 contracts with a probe phase (`probe_osi`); held out on 12 contracts.
| File | Method | Held-out field accuracy (lag / permutation / sign / target) |
|---|---|---|
| `learned/osi_regressor.pt` | passive UP-OSI-style, fixed excitation | 0.415 / 0.291 / 0.551 / 0.941 |
| `learned/osi_random_excitation_regressor.pt` | passive UP-OSI-style, random excitation | 0.773 / 0.519 / 0.558 / 0.599 |
| `learned/recurrent_regressor.pt` | recurrent, episode-length accumulation | see `result.json` in the source run (lag 0.917 at 15 steps; permutation plateaus at ~0.37) |
| `learned/probe_osi_regressor.pt` | probe-augmented (budget 6, amplitude 0.5), pick_cube + push_cube | permutation caps at 0.39 |
These are reported in the paper as a **load-bearing negative result**: on PickCube/seen, the passive
learned identifier scores 0.000 end-to-end, the same floor as no adaptation at all β active probing
against a declared hypothesis space (the belief-family methods, not included here as checkpoints, see
below) is not something these unprivileged learners recover on their own. Do not read these as "the
adaptation methods" β read them as the specific negative result the paper reports.
### `delay_aware/` β delay-aware augmented-state PPO (the `long_lag` split)
A **new task backbone**, not the frozen Gate-1 PPO above. Same official PPO hyperparameters, plus two
changes: the observation is augmented with the last 4 canonical actions (K=4, restores the Markov
property under delay), and training uses a per-episode action lag resampled from `{0, 1, 2, 4}`. This
is the standard state-augmentation reduction for delayed MDPs (Katsikopoulos & Engelbrecht 2003), not a
reproduction of DCAC or D-TRPO.
| File | Task | Variant | Long-lag success | Seen-split competence |
|---|---|---|---:|---:|
| `delay_aware/pick_cube_randomized_final_ckpt.pt` | pick_cube | randomized lag (headline) | 0.528 [0.488, 0.568] | 0.715 |
| `delay_aware/pick_cube_curriculum_final_ckpt.pt` | pick_cube | lag-curriculum (documented alt) | 0.455 (1 seed) | 0.950 |
| `delay_aware/push_cube_randomized_final_ckpt.pt` | push_cube | randomized lag (headline) | 0.415 [0.376, 0.455] | 0.990 |
| `delay_aware/push_cube_curriculum_final_ckpt.pt` | push_cube | lag-curriculum (documented alt) | 0.355 (1 seed) | 0.885 |
The randomized-lag checkpoint is the headline long-lag method (frozen oracle scores 0.027 Pick / 0.153
Push on the same split, so this is roughly 20x / 2.7x). The curriculum variant trades long-lag success
for near-perfect instantaneous competence and is included because the paper documents both, not because
it is the recommended checkpoint for `long_lag` evaluation.
### What's not here
**Belief-family methods (exact belief, entropy probing, fixed/random probes, DualABI and its
ablations) have no checkpoint in this repo because they are analytic, not learned** β they maintain a
Bayesian posterior over the declared contract pool (`configs/contracts/core.yaml` in the dataset repo)
and are fully specified by their config files (`configs/method/dualabi*.yaml`, `configs/method/*probes*.yaml`
in the dataset repo), with no trained weights to publish. This is also true of the scale-corrector and
grasp-channel adaptation variants. Diffusion Policy imitation backbones used for the brittleness study
(Round 12 in the paper) exist locally but are not uploaded here; ask if you need them.
## Intended use
Reproducing or extending the ActionShift tournament: load a `ppo/` checkpoint as the frozen backbone,
evaluate it through the hidden-contract wrapper on the splits in
[kattri15/actionshift](https://huggingface.co/datasets/kattri15/actionshift), and compare an adaptation
method's success against the oracle ceiling and no-adapt floor those splits report. The `learned/` and
`delay_aware/` checkpoints are reference points on that ladder, not standalone manipulation policies to
deploy.
## Honest limits
- **Sim-only, no hardware.** All checkpoints were trained and evaluated in ManiSkill (SAPIEN/PhysX).
No real-robot claim is made or implied.
- **Reference baselines, not SOTA.** The PPO backbones are deliberately plain official-baseline policies
so the interface is the only variable; they are not tuned for maximum task performance.
The `learned/` regressors are reported specifically as a **negative result** in the paper.
- **Method coverage is partial.** Only oracle, no-adapt, the four learned identifiers, and delay-aware
control have trained checkpoints wired to the frozen backbones as of this release. Domain-randomized,
RMA-style, and the belief-family methods have config files but not trained-weight checkpoints (see
above) β that gap is disclosed in `reports/gate1.md` on GitHub as a structured exclusion, not silently
omitted.
- **`delay_aware/` checkpoints are separate backbones**, not drop-in replacements for `ppo/` on splits
other than `long_lag` β do not benchmark them on `seen` and compare directly to the `ppo/` row without
reading the caveats in `reports/adaptation_delay_aware.md`.
## License
MIT, same as the code and paper. See `LICENSE` in the [GitHub repo](https://github.com/Archerkattri/actionshift).
## Citation
```bibtex
@software{attri2026actionshift,
author = {Attri, Krishi},
title = {ActionShift: Hidden compositional action-interface adaptation benchmark},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21500713},
url = {https://github.com/Archerkattri/actionshift}
}
```
Zenodo DOI: [10.5281/zenodo.21500713](https://doi.org/10.5281/zenodo.21500713). If you use the ManiSkill
baselines these checkpoints derive from, cite ManiSkill separately.
|