Robotics
LeRobot
Safetensors
so-100
imitation-learning
diffusion-policy

Diffusion Policy: pick R2-D2 and put it in the box (SO-100)

Diffusion Policy, trained from scratch.

Trained on van-i/r2d2_to_box_bg_20261003_210444: 50 teleoperated episodes on an SO-100 arm with three cameras (left and right overhead, grip on the wrist, 640×480 @ 30 fps). Task: "pick r2d2 and put to box": pick up a small R2-D2 figure from one of five taped start positions and drop it into a cardboard tray. Part of a school robotics test stand built with LeRobot and LeLab. Full write-up, comparison of all four policies and scripts: van-i/so100-imitation-learning-stand.

Training

36,000 steps, batch 8 (~20 epochs), mixed precision, LeLab web UI, 4 h 04 min on an RTX 3090 (250 W). Final training loss 0.006 (not comparable with other policies).

Results on the real arm

9/10 on the trained start positions, 2/4 on held-out positions (H1 2/2, H2 0/2), ~16.5 s per success.

All four policies trained on the same dataset, 14 tries each (5 trained positions ×2, held-out H1 between two marks ×2, H2 just outside the marked area ×2):

Policy Trained spots (P1–P5) H1 (between marks) H2 (outside the marks) Avg. time to finish
ACT (15k) 8/10 2/2 0/2 ~10 s
Diffusion Policy (36k) 9/10 2/2 0/2 ~16.5 s
SmolVLA (25k) 10/10 2/2 0/2 ~8.4 s
GR00T N1.7 (18k) 10/10 2/2 2/2 ~9.8 s

Run it

LeRobot 0.6.0, SO-100 / SO-101 follower. Adjust the serial port, calibration id and camera indices to your setup:

lerobot-rollout \
  --strategy.type=base \
  --policy.path=van-i/diffusion_r2d2_to_box_bg \
  --policy.device=cuda \
  --robot.type=so101_follower --robot.port=/dev/ttyACM1 --robot.id=<your-follower-calibration-id> \
  --robot.cameras="{left: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}, grip: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}, right: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --task="pick r2d2 and put to box" \
  --duration=60

The policy only works in a scene like the training one: dark matte table, the tray at its spot, similar lighting and camera placement.

Downloads last month
32
Safetensors
Model size
0.3B params
Tensor type
F32
·
Video Preview
loading

Dataset used to train van-i/diffusion_r2d2_to_box_bg