|
Download README.md from harrywang01/CAMP-RMBench-Checkpoints: direct link, hf CLI and curl.
- Browser
- Download file 12.2 kB
-
https://huggingface.co/harrywang01/CAMP-RMBench-Checkpoints/resolve/main/README.md
- Command line
-
hf download hf://harrywang01/CAMP-RMBench-Checkpoints/README.md
-
curl -L -o README.md https://huggingface.co/harrywang01/CAMP-RMBench-Checkpoints/resolve/main/README.md
12.2 kB
| license: mit | |
| pipeline_tag: robotics | |
| tags: | |
| - robotics | |
| - imitation-learning | |
| - diffusion-policy | |
| - memory | |
| - bimanual-manipulation | |
| - rmbench | |
| - robotwin | |
| # π Overview | |
| **CAMP** (Compressed Action Memory Policy) gives a visuomotor policy a *behavioral memory*. A recurrent memory | |
| is pretrained to reconstruct a compressed (DCT) summary of the robot's own past actions, so its hidden state | |
| has to encode what the robot already did; a Diffusion Policy is then conditioned on that state. This lets the | |
| policy track task progress and learn from its own failed attempts, which a memoryless policy cannot do when the | |
| current image does not determine the next action. | |
| This repository holds the CAMP policies for all nine tasks of **[RMBench](https://github.com/RoboTwin-Platform/RMBench)**, | |
| a memory-dependent bimanual manipulation benchmark built on RoboTwin 2.0 (Aloha-AgileX dual-arm robot). Every | |
| policy is trained from the official RMBench demonstrations | |
| ([`TianxingChen/RMBench`](https://huggingface.co/datasets/TianxingChen/RMBench), `demo_clean`, 50 per task) | |
| and evaluated with RMBench's own protocol. | |
| The files contain inference weights only. | |
| # π Results | |
| Success rate over **100 test episodes** per task. We follow RMBench's evaluation protocol: the `demo_clean` | |
| configuration, test seeds from 100000 that the scripted expert can solve, and the per-task step limit. | |
| <table align="center"> | |
| <tr><th align="left">Task</th><th align="left">Memory needed for</th><th>Step limit</th><th>Success</th></tr> | |
| <tr><td><a href="#rearrange_blocks"><code>rearrange_blocks</code></a></td><td>task progress</td><td align="center">700</td><td align="center"><b>100%</b></td></tr> | |
| <tr><td><a href="#blocks_ranking_try"><code>blocks_ranking_try</code></a></td><td>learning from failure</td><td align="center">3500</td><td align="center"><b>100%</b></td></tr> | |
| <tr><td><a href="#put_back_block"><code>put_back_block</code></a></td><td>task progress</td><td align="center">500</td><td align="center"><b>100%</b></td></tr> | |
| <tr><td><a href="#battery_try"><code>battery_try</code></a></td><td>learning from failure</td><td align="center">1000</td><td align="center"><b>97%</b></td></tr> | |
| <tr><td><a href="#swap_t"><code>swap_T</code></a></td><td>task progress</td><td align="center">600</td><td align="center">24%</td></tr> | |
| <tr><td><a href="#swap_blocks"><code>swap_blocks</code></a></td><td>task progress</td><td align="center">1000</td><td align="center">19%</td></tr> | |
| <tr><td><a href="#cover_blocks"><code>cover_blocks</code></a></td><td>task progress</td><td align="center">1500</td><td align="center">17%</td></tr> | |
| <tr><td><a href="#observe_and_pickup"><code>observe_and_pickup</code></a></td><td>a past observation</td><td align="center">250</td><td align="center">9%</td></tr> | |
| <tr><td><a href="#press_button"><code>press_button</code></a></td><td>counting</td><td align="center">1500</td><td align="center">5%</td></tr> | |
| </table> | |
| # π¬ Tasks | |
| Each preview is a **demonstration by RMBench's scripted expert**: the first of the 50 released demonstrations of | |
| the task, re-rendered so each task and its success condition can be seen clearly. It is filmed from the head | |
| camera the policy is trained on, with the same pose and field of view, at 1920Γ1440 instead of 320Γ240. Every | |
| preview passes the task's success check. Long demos are sped up to at most about 24 s. The `observe_and_pickup` | |
| preview holds the first frame and plays at half speed, because the target is visible for only that one frame. | |
| The full-resolution MP4 files are in [`previews/`](previews). Success rates in the table above come only from the | |
| 100-episode policy evaluations. | |
| <a id="rearrange_blocks"></a> | |
| ## rearrange_blocks: 100% | |
| <p align="center"><img src="previews/rearrange_blocks.gif" width="70%"></p> | |
| Two blocks sit on mats next to a button, and one mat is empty. The robot moves the first block onto the empty | |
| mat and presses the button. It then moves the second block off its mat to the spot between the mats. | |
| - **Memory:** whether the button has already been pressed. The scene looks the same before and after the press. | |
| - **Success:** the first block is within 3 cm of the target mat and the second block within 3 cm of the spot | |
| between the mats. The button has been pressed exactly once and the gripper is open. | |
| <a id="blocks_ranking_try"></a> | |
| ## blocks_ranking_try: 100% | |
| <p align="center"><img src="previews/blocks_ranking_try.gif" width="70%"></p> | |
| Three colored cubes stand in a row in a random order, next to a check button. The robot does not know the target | |
| order. It presses the button to test the current arrangement. If the arrangement is rejected, it swaps two cubes | |
| and tests again, working through the orders until the button accepts one. | |
| - **Memory:** which arrangements have already been tried and rejected. Repeating a rejected arrangement never | |
| succeeds, and the scene does not show the history. | |
| - **Success:** the three cubes stand next to each other in the correct left-to-right order and the button has | |
| been pressed. | |
| <a id="put_back_block"></a> | |
| ## put_back_block: 100% | |
| <p align="center"><img src="previews/put_back_block.gif" width="70%"></p> | |
| A block starts on a mat. The robot moves the block to the center of the table and presses the button. It then | |
| puts the block back on the mat it came from. | |
| - **Memory:** the mat the block started on. Once the block is in the center, the image no longer shows where | |
| it came from. | |
| - **Success:** the button has been pressed once with the block in the center. The block then rests within | |
| 3 cm of its original mat and the gripper is open. | |
| <a id="battery_try"></a> | |
| ## battery_try: 97% | |
| <p align="center"><img src="previews/battery_try.gif" width="70%"></p> | |
| Two batteries must go into a slot whose correct polarity is hidden. The dashboard needle shows whether the current | |
| combination is correct. If it is not, the robot takes a battery out and re-inserts it the other way round. | |
| - **Memory:** which orientations have already been tried. | |
| - **Success:** both batteries are seated in the slot in the correct orientation and the dashboard turns on. | |
| <a id="swap_t"></a> | |
| ## swap_T: 24% | |
| <p align="center"><img src="previews/swap_T.gif" width="70%"></p> | |
| Two T-shaped blocks lie on the table. The robot picks them up and places each one at the other's initial position | |
| and orientation. | |
| - **Memory:** both initial poses. Once the robot moves a block, its original pose is no longer visible. | |
| - **Success:** each block is within 2.5 cm and 15Β° of the other block's initial pose, both are resting on the | |
| table, and both grippers are open. | |
| <a id="swap_blocks"></a> | |
| ## swap_blocks: 19% | |
| <p align="center"><img src="previews/swap_blocks.gif" width="70%"></p> | |
| Two blocks are in two of three trays. The robot may move one block at a time and each tray holds at most one | |
| block. It swaps the two blocks using the spare tray as a buffer, then presses the button. | |
| - **Memory:** where each block started and which step of the three-move swap comes next. | |
| - **Success:** each block is inside the tray the other block started in, the button has been pressed once | |
| and the gripper is open. | |
| <a id="cover_blocks"></a> | |
| ## cover_blocks: 17% | |
| <p align="center"><img src="previews/cover_blocks.gif" width="70%"></p> | |
| A red, a green and a blue block are arranged randomly together with three identical lids. The robot covers the | |
| blocks from left to right. Then it lifts the lids again in the order red, green, blue. | |
| - **Memory:** which block is under which lid. The lids are identical, so the colors are hidden once covered. | |
| - **Success:** the covering and uncovering sequence matches the required order exactly, with no wrong lid | |
| lifted. | |
| <a id="observe_and_pickup"></a> | |
| ## observe_and_pickup: 9% | |
| <p align="center"><img src="previews/observe_and_pickup.gif" width="70%"></p> | |
| A target object is shown on a shelf. A wall then drops in front of the shelf and hides it. The robot must pick | |
| up the matching object from several distractors on the table. | |
| - **Memory:** the target object's identity. During the evaluation it is visible only in the first frame, | |
| before the robot moves. | |
| - **Success:** the arms stay still while the target is shown, and the correct object is then lifted off the | |
| table. | |
| - **Note:** this is the only task whose memory uses every frame (action subsampling 1 instead of 4). Otherwise | |
| the single frame that shows the target would be skipped. | |
| <a id="press_button"></a> | |
| ## press_button: 5% | |
| <p align="center"><img src="previews/press_button.gif" width="70%"></p> | |
| Two number cards lie on the table. The robot presses the left button as many times as the left card shows and | |
| the middle button as many times as the right card shows, then presses the right button to confirm. | |
| - **Memory:** how many presses each button has received so far. A button looks the same after every press. | |
| - **Success:** both press counts match the cards exactly and the confirm button has been pressed. | |
| # π¦ Files | |
| ``` | |
| <task>/ | |
| βββ policy.ckpt # CAMP policy: EMA weights of the memory-conditioned Diffusion Policy + resolved config | |
| βββ memory/ | |
| βββ best_model.pt # Stage-1 behavioral-memory LSTM (weights + architecture args) | |
| βββ normalizer.pt # its input normalizer | |
| previews/<task>.gif | .mp4 # expert demonstrations shown above (1920Γ1440 MP4) | |
| ``` | |
| The checkpoints hold only what inference needs. There is no optimizer or scheduler state and no training | |
| bookkeeping. | |
| # π§ Training recipe | |
| | | | | |
| |:--|:--| | |
| | Data | the 50 RMBench demonstrations per task (`demo_clean`) | | |
| | Observation | head camera 240Γ320 (random crop 216Γ288) and the 14-D joint state; `n_obs_steps = 1` | | |
| | Action | 14-D absolute joint targets in 8-step chunks, normalized to a per-joint range | | |
| | Stage 1: memory | LSTM (hidden 128) pretrained to reconstruct the DCT coefficients of its past actions; action subsampling 4 (1 for `observe_and_pickup`) | | |
| | Stage 2: policy | Diffusion Policy conditioned on the memory through a 32-D projection. The memory is frozen for 400 epochs, then memory and policy are finetuned jointly (200 epochs; 600 for `put_back_block`) | | |
| | Augmentation | joint noise 0.01, image noise 0.02, brightness and contrast jitter 0.15 | | |
| | Checkpoint | the best of the evaluated epochs per task (every 100 epochs) | | |
| # π Usage | |
| Evaluation code, training launcher and full instructions are in the | |
| [RMBench section of the CAMP repository](https://github.com/KuanchengWang/CAMP#-rmbench). To evaluate | |
| one task on the 100 standard test seeds: | |
| ```bash | |
| git clone --recurse-submodules https://github.com/KuanchengWang/CAMP.git && cd CAMP | |
| ./docker_run.sh build && docker build -f docker/Dockerfile.rmbench -t camp-rmbench:latest . | |
| git clone https://github.com/RoboTwin-Platform/RMBench ~/rmbench && (cd ~/rmbench && bash script/_download_assets.sh) | |
| huggingface-cli download harrywang01/CAMP-RMBench-Checkpoints --local-dir ckpts/hf | |
| T=swap_T | |
| mkdir -p ckpts/stage2/$T/checkpoints ckpts/stage1/$T | |
| cp ckpts/hf/$T/policy.ckpt ckpts/stage2/$T/checkpoints/policy.ckpt | |
| cp ckpts/hf/$T/memory/* ckpts/stage1/$T/ | |
| export RMBENCH=~/rmbench STAGE1=$PWD/ckpts/stage1 STAGE2=$PWD/ckpts/stage2 EVAL_OUT=$PWD/data/rmbench_eval | |
| E=scripts/rmbench/eval | |
| $E/docker_rmbench.sh python /workspace/$E/rmbench_eval.py seeds --task $T # the 100 expert-solvable test seeds | |
| $E/docker_rmbench.sh python /workspace/$E/rmbench_eval.py eval --task $T --ckpt policy # -> data/rmbench_eval/results/$T/policy/results.json | |
| ``` | |
| To retrain a task from the official demonstrations with the same settings, run | |
| `bash scripts/rmbench/train_rmbench.sh <task>` inside the `camp` container. | |
| # π Citation | |
| If you find CAMP useful, please cite: | |
| ```bibtex | |
| @misc{wang2026rememberdidlearningbehavioral, | |
| title={Remember what you did?: Learning Behavioral Memories for Partially Observable Object Manipulation}, | |
| author={Kuancheng Wang and Seungho Yeom and Jinglin Cao and Yuheng Zhi and Nikhil Shinde and Michael Yip}, | |
| year={2026}, | |
| eprint={2606.21188}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.RO}, | |
| url={https://arxiv.org/abs/2606.21188}, | |
| } | |
| ``` | |