File size: 5,445 Bytes
78ed463 f7cf181 78ed463 49cd73c 78ed463 49cd73c 78ed463 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | ---
language:
- en
license: mit
pipeline_tag: robotics
tags:
- OpenRAL
- rskill
- openvla
- vision-language-action
- nf4
- 4-bit
- widowx
- openral
- openvla-oft
- vla
- simpler
- maniskill3
- manipulation
inference: false
---
# rskill-openvla-oft-simpler-widowx-nf4
> OpenVLA-OFT bridge policy (RLinf, PPO-tuned on ManiSkill3 PutOnPlateInScene25),
> packaged for OpenRAL and locally verified on SimplerEnv WidowX carrot-on-plate.
## What this skill does
Wraps [`RLinf/RLinf-OpenVLAOFT-PPO-ManiSkill3-25ood`](https://huggingface.co/RLinf/RLinf-OpenVLAOFT-PPO-ManiSkill3-25ood)
— an [OpenVLA-OFT](https://openvla-oft.github.io/) (arXiv:2502.19645) policy,
RL-tuned with PPO on the ManiSkill3 `PutOnPlateInScene25` task using a WidowX
250 S — and runs it on the [SimplerEnv](https://github.com/simpler-env/SimplerEnv)
WidowX (Bridge V2) carrot-on-plate task that shares its embodiment, EE-delta
control, and `bridge_orig` normalization. The sibling Bridge tasks are not
declared in `evaluated_tasks` until locally reproduced.
**Why WidowX and not Panda/PickCube:** this checkpoint is a *bridge* policy. The
ManiSkill3 Panda `PickCube-v1` scenes are a different embodiment and task; the
ADR-0060 task-data gate correctly refuses that pairing (it would produce a
plausible-but-unsolvable rollout). See [ADR-0063](../../docs/adr/0063-openvla-oft-policy-family.md)
for the full rationale.
## How it works
Loaded in-process by OpenRAL's `openvla` policy adapter
(`python/sim/src/openral_sim/policies/openvla.py`) as a transformers
*custom-code* model (`AutoModelForVision2Seq` + `trust_remote_code`, gated by
`OPENRAL_ALLOW_REMOTE_CODE=1`). NF4 (4-bit) quantization plus the CUDA
expandable-segments allocator bring the 7.5 B backbone within an 8 GB GPU.
The RLinf checkpoint currently needs a 4.40-era transformers runtime; the
default OpenRAL workspace pins transformers 5.3 for lerobot families, so keep
OpenVLA validation in a dedicated environment rather than syncing it together
with the default VLA groups.
### Observation → action contract
- **Input:** one 224×224 RGB frame (the SimplerEnv 3rd-view, surfaced as
`camera1`) and the prompt
`In: What action should the robot take to {instruction.lower()}?\nOut: `. No
proprioception (`use_proprio=False`).
- **Output:** 256-bin discrete action tokens decoded to `[-1, 1]`, then
de-normalized with the embedded `unnorm_key=bridge_orig` stats (BOUNDS_Q99):
6 end-effector deltas (3 position + 3 rotation) rescaled, gripper passed
through. The manifest drives RLinf's `generate_action_verl` path with
right-padded prompts (`max_length=30`), temperature sampling (`0.6`), torch
seed `0`, `action_scale=2.0` on the first six dimensions, and binary gripper
threshold `0.5`. Action chunk = 8 × 7-D, replayed open-loop.
## Upstream model / training
Upstream base `Haozhan72/Openvla-oft-SFT-libero-goal-trajall`, ManiSkill LoRA
SFT, then PPO on `PutOnPlateInScene25Main-v3` (WidowX 250 S). RLinf model-index
success: train 0.977; OOD vision 0.921 / semantic 0.648 / position 0.736. See
the upstream card for the full protocol.
## Supported robots
- `widowx` (WidowX 250 S, Bridge V2 flat-table setup).
## Sensors required
- One RGB stream, ≥224×224, mapped to `observation.images.camera1`.
## Manifest summary
See [`rskill.yaml`](./rskill.yaml). Key fields: `model_family: openvla`,
`license: mit`, `quantization.dtype: int4`, `chunk_size: 8`,
`evaluated_tasks` = `simpler_env/widowx_carrot_on_plate`,
`benchmarks.simpler_env_widowx: 0.4`, `policy_extras` = the RLinf generation
and action-transform knobs, `action_contract` = `delta_ee_6d_plus_gripper`
(dim 7).
## Quick start
```bash
just sync --all-packages --group simpler-env
hf download RLinf/RLinf-OpenVLAOFT-PPO-ManiSkill3-25ood
OPENRAL_ALLOW_REMOTE_CODE=1 openral benchmark run \
--suite simpler_env_widowx --task simpler_env/widowx_carrot_on_plate \
--rskill openvla-oft-simpler-widowx-nf4
```
## Reproduction
```bash
# Single SimplerEnv WidowX scene (carrot-on-plate):
OPENRAL_ALLOW_REMOTE_CODE=1 openral benchmark run \
--suite simpler_env_widowx --task simpler_env/widowx_carrot_on_plate \
--rskill openvla-oft-simpler-widowx-nf4
```
## Evaluation
Local seeded validation on an RTX 4070 Laptop GPU (8 GB), NF4, SimplerEnv
ManiSkill3 `PutCarrotOnPlateInScene-v1`, 5 episodes, seeds 0..4, 60-step
horizon, `generate_action_verl`, torch seed 0 reapplied on each policy reset,
`action_scale=2.0`:
- `simpler_env/widowx_carrot_on_plate`: **2/5 success (40%)**.
Public `widowx_carrot_on_plate` without the RLinf action transform scored 0/5,
and the exact upstream `PutOnPlateInScene25Main-v3` registration needs RLinf
assets that were not present in the public source checkout. Those numbers are
not claimed here.
## License
MIT (upstream `RLinf/RLinf-OpenVLAOFT-PPO-ManiSkill3-25ood`). The OpenRAL
packaging is Apache-2.0. The checkpoint is a `trust_remote_code` custom-code
model; loading executes repo-shipped Python and requires
`OPENRAL_ALLOW_REMOTE_CODE=1` (provenance: rSkill signature verification is not
yet implemented — ADR-0006).
## See also
- [ADR-0063 — OpenVLA / OpenVLA-OFT policy family](../../docs/adr/0063-openvla-oft-policy-family.md)
- [ADR-0060 — benchmark task-data compatibility gate](../../docs/adr/0060-benchmark-task-data-compatibility-gate.md)
- [`rldx1-ft-simpler-widowx-nf4`](../rldx1-ft-simpler-widowx-nf4) — the sibling WidowX bridge rSkill.
|