lift_act_sim_v2 β the no-foam cell
An action chunking transformer trained by behaviour cloning to lift a 2 Γ 6 Γ 2 cm red block 10 cm
off the table of a simulated Dobot X-Trainer dual-arm cell β omn4i/Lift-v0 in
Physical41/omn4i_envs.
81.0% success on held-out seeds, against the scripted expert's 99.0%.
This is published as a baseline to beat, not a good model. It has a 23-point gap between the arms and its cause is not yet known. See "The open problem" below.
Which cell this is for β read this before using either version
lift_act_sim_v1 |
lift_act_sim_v2 (this one) |
|
|---|---|---|
| Cell | 50 mm foam slab on the table | no foam β the cell as it is now |
| Trained at | pre-foam-removal tree | commit ebb10c8 (main HEAD) |
| Corpus | 2847 episodes | 1987 episodes |
| Overall | 93.8% | 81.0% |
| Left / right arm | 93.8% / 93.8% | 93.0% / 69.8% |
v1 is for a cell that no longer exists. The foam was removed from main in 326dc18, which
changed every object height by 50 mm and gave the arms back 50 mm of reach. v1's 93.8% is a real
measurement, but of a configuration you cannot reproduce from main today. Use v2 with current
main.
The weights here were verified to load and run under main's own omn4i_envs/baselines/act.py at
ebb10c8. Every file defining the lift task β lift.py, env.py, scene.xml, lift.xml, both
arm XMLs, contract.py, classical.py, benchmark.py, transforms.py β is byte-identical
between main and the tree this was trained on.
Results
500 held-out episodes (benchmark.TEST_SEEDS, range(10000, 10500)), by
python tools/benchmark.py act. None of these seeds was trained on.
act (this model) |
classical (its teacher) |
|
|---|---|---|
| Success | 81.0% (405/500) | 99.0% (495/500) |
| 95% Wilson interval | 77.3% β 84.2% | 97.7% β 99.6% |
| Per-step latency, median β p99 | 7.52 β 7.61 ms | 0.01 β 0.01 ms |
Against a 33 ms environment step the policy uses 7.5 ms. It replans every step.
By arm and band β where it goes wrong
| Band | left | right |
|---|---|---|
| under 0.40 m | 95% (52/55) | 95% (42/44) |
| 0.40 to 0.50 m | 92% (78/85) | 74% (62/84) |
| 0.50 to 0.62 m | 94% (80/85) | 54% (62/115) |
| 0.62 m and out | 88% (15/17) | 93% (14/15) |
| overall | 93.0% | 69.8% |
The expert is symmetric on the same seeds (98.8% left, 99.2% right), so nothing about the task favours one side.
The open problem
The right arm is 23 points behind the left and we do not know why.
What it is not:
- Not data scarcity. The corpus is balanced by arm β 978 left-arm episodes against 1009 right. The worst cell, right 0.50β0.62 m at 54%, has more training episodes (461) than its left-hand counterpart (407) which scores 94%.
- Not the two demonstration defects fixed in
mainPR #13 (the idle arm swinging 90Β° every episode, and the wrist roll scattering across its equivalents). Both are verified absent from this corpus β idle drift 0.3Β°, wrist roll within Β±90Β° on both arms. Fixing them coincided with the gap closing on the foam cell, but the gap returned when only the cell changed, so they were not the whole cause. - Not the expert.
classicalscores 99.2% on right-arm seeds.
The open lead is the right arm's joint 6 sign, which is unverified upstream. It turns about the tool axis, which the end-effector site sits on, so no trajectory data can settle it β but it does set which way the jaws straddle the block. That is consistent with a policy that reaches correctly and closes on nothing, and with the expert being unaffected, since its inverse kinematics is self-consistent inside the same model.
Architecture and training
18.7M parameters. One shared ResNet18 across the three cameras, feeding a 4-layer encoder/decoder at width 256, 8 heads. Predicts 32 actions from one observation, L1 loss, temporal ensembling at inference. The backbone starts from ImageNet weights and is fine-tuned at 1e-5 against 1e-4 for the rest β 11.2M of the 18.7M parameters, not frozen.
No CVAE (the demonstrations come from a deterministic planner) and no language conditioning (one task, so the prompt is a constant).
1900 demonstrations / 117,874 frames from Physical41/lift_classical_sim_v1, 87 episodes held out
stratified by reach band. 40,000 steps at batch 64, 59.4 minutes on one RTX 4090. Final held-out
chunk error 0.01123 rad.
Note that held-out chunk error has been a poor predictor of rollout success on this task: a run at 0.01138 scored 93.4% and one at 0.00928 scored 73.2%. Judge by the benchmark, not by the loss.
Loading
from omn4i_envs.baselines import act # import before mujoco
policy = act.ACTPolicy(checkpoint="act.pt")
observation, info = env.reset(seed=10000)
policy.reset(observation, info)
action = policy(observation) # (14,) float32, clamped into the action space
Actions are absolute joint position targets in the controller's convention, which differs per arm β the right is read at the negation of the left, and neither neutral is the zero vector. Gripper entries are normalized 1.0 open, 0.0 closed.
This is simulation only
Recorded real joint angles do not reproduce the real robot's kinematics in this simulation, and the joint sign and zero-offset convention between the vendor stack and the MJCF is only partly resolved β the right arm's joint 6 sign remains open, which may well be the same thing causing the asymmetry above. Do not cite these numbers as evidence about the physical robot.