Robotics
LeRobot
Safetensors
so-100
imitation-learning
act

ACT: pick R2-D2 and put it in the box (SO-100)

ACT (Action Chunking with Transformers), trained from scratch.

Trained on van-i/r2d2_to_box_bg_20261003_210444: 50 teleoperated episodes on an SO-100 arm with three cameras (left and right overhead, grip on the wrist, 640×480 @ 30 fps). Task: "pick r2d2 and put to box": pick up a small R2-D2 figure from one of five taped start positions and drop it into a cardboard tray. Part of a school robotics test stand built with LeRobot and LeLab. Full write-up, comparison of all four policies and scripts: van-i/so100-imitation-learning-stand.

Training

15,000 steps, batch 8 (~8.5 epochs), mixed precision, LeLab web UI, 1 h 06 min on an RTX 3090 (250 W). Final training loss 0.166.

Results on the real arm

8/10 on the trained start positions, 2/4 on held-out positions (H1 2/2, H2 0/2), ~10 s per success. All failures were reaches to the wrong spot.

All four policies trained on the same dataset, 14 tries each (5 trained positions ×2, held-out H1 between two marks ×2, H2 just outside the marked area ×2):

Policy Trained spots (P1–P5) H1 (between marks) H2 (outside the marks) Avg. time to finish
ACT (15k) 8/10 2/2 0/2 ~10 s
Diffusion Policy (36k) 9/10 2/2 0/2 ~16.5 s
SmolVLA (25k) 10/10 2/2 0/2 ~8.4 s
GR00T N1.7 (18k) 10/10 2/2 2/2 ~9.8 s

Run it

LeRobot 0.6.0, SO-100 / SO-101 follower. Adjust the serial port, calibration id and camera indices to your setup:

lerobot-rollout \
  --strategy.type=base \
  --policy.path=van-i/act_r2d2_to_box_bg \
  --policy.device=cuda \
  --robot.type=so101_follower --robot.port=/dev/ttyACM1 --robot.id=<your-follower-calibration-id> \
  --robot.cameras="{left: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}, grip: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}, right: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --task="pick r2d2 and put to box" \
  --duration=60

The policy only works in a scene like the training one: dark matte table, the tray at its spot, similar lighting and camera placement.

Downloads last month
32
Safetensors
Model size
51.7M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train van-i/act_r2d2_to_box_bg