Bimanual puzzle pushing: Diffusion Policy

This is the policy I trained for my puzzleBimanual project. Two Franka Panda arms with stick end-effectors stand on either side of a table in MuJoCo. Each arm pushes one half of a randomly cut puzzle, and the goal is to push the two halves together until they fit into one rectangle. There is no target marker, so the policy has to work out from the shapes alone how the pieces go together.

Two rollouts on unseen puzzles
Two rollouts on test puzzles. Each row shows the top camera and both wrist cameras.

Results

I tested the policy in the simulator on puzzles that were never used in training. An episode counts as a success when the pieces are within 5 mm and 4° of the mated pose and stay there for half a second, which is the same rule I used when recording the demonstrations.

Split Puzzles Success 95% CI
dev (used to pick this checkpoint) 50 88% 76–94%
test (evaluated once) 200 82% 76–87%

When the policy fails, it's mostly because the pieces never got close: in 30 of the 36 failed test episodes they never came within 2 cm of the mated pose. Getting stuck on the final alignment is rare (3 episodes). Once the pieces are close, it usually finishes: the median final error in successful episodes is about 2 mm.

What the policy sees and does

  • Inputs: three 240×320 RGB cameras (a top view and one camera on each wrist) and an 18-number robot state (tip x, y and seven joint angles per arm), with the last 2 frames as context. The piece poses are not given to the policy; it has to see the pieces.
  • Output: absolute x, y targets for the two stick tips, 4 numbers per step at 20 Hz. The policy predicts 16 steps and executes 8 (0.4 s) before planning again.
  • Model: LeRobot's Diffusion Policy with a separate ResNet18 encoder per camera and a 1D U-Net with channels [256, 512, 1024]. 103M parameters in total. DDIM with 10 denoising steps at inference.

Training data and setup

I recorded 217 demonstrations myself with a DualShock 4 (one stick per arm), about 50 minutes in total. A processing step removed the pauses in each arm's motion and added mirrored and rotated copies of every demonstration, which gave 797 episodes. Every episode has a different random puzzle.

Training ran for 100k steps with batch size 32 and mixed precision on a laptop RTX 4050 (6 GB), which took about 12 hours. I compared the 50k, 70k and 100k checkpoints on the same 50 dev puzzles (64%, 70% and 88%) and kept the last one.

How to use it

The simulator and evaluation code live in the GitHub repository. You don't need the dataset; the files here include everything the evaluation needs (eval_profile.json holds the camera setup and the training seeds, so evaluation puzzles never overlap with training).

git clone https://github.com/HarunSMetin/puzzleBimanual.git
cd puzzleBimanual
uv sync
uv run python scripts/3_evaluate/eval_policy.py --checkpoint HSM1/puzzle-bimanual-diffusion --episodes 20

It runs on a GPU if there is one and on the CPU otherwise (much slower). Results and videos go to outputs/eval/.

Limitations

This policy only knows this simulated scene: the table, the arm placement, the camera positions and the puzzle generator's shapes. It has not been tried on a real robot. Rollouts are also not exactly repeatable, because tiny rendering differences between runs grow into different pushes, so compare success rates over many puzzles rather than single episodes.

Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
·
Video Preview
loading