Robotics
LeRobot
Safetensors
so-100
imitation-learning
gr00t
vla

GR00T N1.7: pick R2-D2 and put it in the box (SO-100)

NVIDIA GR00T N1.7 (3.1B parameters, built on NVIDIA Cosmos: nvidia/Cosmos-Reason2-2B backbone), fine-tuned with LeRobot (--policy.type=groot, embodiment new_embodiment). The vision-language backbone stayed frozen; 1.6B parameters were trained.

Trained on van-i/r2d2_to_box_bg_20261003_210444: 50 teleoperated episodes on an SO-100 arm with three cameras (left and right overhead, grip on the wrist, 640×480 @ 30 fps). Task: "pick r2d2 and put to box": pick up a small R2-D2 figure from one of five taped start positions and drop it into a cardboard tray. Part of a school robotics test stand built with LeRobot and LeLab. Full write-up, comparison of all four policies and scripts: van-i/so100-imitation-learning-stand.

Training

18,000 steps, batch 4 (~5 epochs), all weights in bf16 (--policy.model_params_fp32=false, needed to fit in 24 GB; NVIDIA's own recipe uses fp32 weights), 2 h 57 min on an RTX 3090 (250 W), 23.2 GB peak VRAM. Final training loss 0.035.

Results on the real arm

10/10 on the trained start positions and 4/4 on held-out positions, including H2 outside the trained area, where all other policies failed. ~9.8 s per success. It dropped R2-D2 a few times but went back and finished. Loading takes ~80 s before the arm moves.

All four policies trained on the same dataset, 14 tries each (5 trained positions ×2, held-out H1 between two marks ×2, H2 just outside the marked area ×2):

Policy Trained spots (P1–P5) H1 (between marks) H2 (outside the marks) Avg. time to finish
ACT (15k) 8/10 2/2 0/2 ~10 s
Diffusion Policy (36k) 9/10 2/2 0/2 ~16.5 s
SmolVLA (25k) 10/10 2/2 0/2 ~8.4 s
GR00T N1.7 (18k) 10/10 2/2 2/2 ~9.8 s

Instruction following: fine-tuned on a single instruction. With a new object (a roll of tape) and the instruction "pick tape roll and put into box", it almost completed the task. With both the roll and R2-D2 on the table, it went for R2-D2 first.

Inference: use real-time chunking (--inference.type=rtc --inference.rtc.execution_horizon=40 --inference.queue_threshold=0), which is what LeLab uses for GR00T.

License

This model is a derivative of NVIDIA GR00T N1.7 and is distributed under the NVIDIA Open Model License. Licensed by NVIDIA Corporation under the NVIDIA Open Model License. Built on NVIDIA Cosmos.

Run it

LeRobot 0.6.0, SO-100 / SO-101 follower. Adjust the serial port, calibration id and camera indices to your setup:

lerobot-rollout \
  --strategy.type=base \
  --policy.path=van-i/groot_r2d2_to_box_bg \
  --policy.device=cuda \
  --inference.type=rtc --inference.rtc.execution_horizon=40 --inference.queue_threshold=0 \
  --robot.type=so101_follower --robot.port=/dev/ttyACM1 --robot.id=<your-follower-calibration-id> \
  --robot.cameras="{left: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}, grip: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}, right: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --task="pick r2d2 and put to box" \
  --duration=60

The policy only works in a scene like the training one: dark matte table, the tray at its spot, similar lighting and camera placement.

Downloads last month
28
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for van-i/groot_r2d2_to_box_bg

Finetuned
(223)
this model

Dataset used to train van-i/groot_r2d2_to_box_bg