Robotics
LeRobot
English
lelab
so-100
imitation-learning
education
act
diffusion-policy
smolvla
gr00t

Building an Imitation Learning Test Stand for a School Robotics Class

How I built a robot-learning stand for a school robotics class, using LeRobot, Hugging Face's free toolkit for teaching real robots new skills.

Why a test stand, and why LeRobot

The main goal is learning. I wanted to understand how the LeRobot platform works, and to give students a first look at physical intelligence: AI that doesn't just write text or make pictures, but sees, decides and moves things in the real world. Robot learning is changing fast. It has gone from small programs that each learn one task to large "vision-language-action" models that take in camera images and a written instruction. A hands-on stand is the most direct way to see what's really behind the impressive demo videos.

In practice, that means a test stand: a fixed setup that works the same way every time, so students can go through the whole imitation learning loop on real hardware. Imitation learning means the robot learns by copying examples. Students show the robot a task, record those examples, train a model on them, and then watch the robot try the task on its own. I chose a real robot over a computer simulation on purpose. A robot that knocks the object off the table teaches you more than any graph.

LeRobot fits because it covers that whole loop for cheap, open-source robot arms: controlling the arm by hand, recording examples, training, and running the result. The recordings and trained models can also be shared on the Hugging Face Hub, a website for sharing AI models and data.

On top of LeRobot, the stand runs LeLab, LeRobot's web interface. It makes getting started much easier. Instead of beginning with Python setup, command-line options and config files, students start in a web browser with buttons for each step: calibrate, control, record, train and run. They can get a robot doing something in the very first lesson, and look at the commands and code underneath later, once they know what they're looking at.

Words you'll see in this article

Term What it means here
Policy The trained model that controls the robot. It looks at the camera images and the arm's position and decides the next move.
Leader / follower arm Two identical arms. You move the leader by hand, and the follower copies it. Later the trained policy drives the follower alone.
Teleoperation Controlling the robot remotely, here by moving the leader arm.
Servo The small motor in each joint of the arm. Each arm has six.
Episode One recorded example of the task, from start to finish (about 8–10 seconds here).
Dataset A collection of episodes: camera videos plus the arm's positions, saved together.
Frame One moment in a recording. At 30 frames per second (fps), a 10-second episode has 300 frames.
Training step / batch Training happens in small steps. In each step, the model looks at a small group of frames (a batch, here 4 or 8) and adjusts itself a little.
Epoch One full pass through the whole dataset during training.
Loss A score for how wrong the model's guesses are during training. It should go down as the model learns.
Checkpoint A saved copy of the model at some point during training, so you can test it later.
GPU / VRAM The graphics card that does the heavy math, and its own memory. The model and its training data must fit in that memory.
Pretrained / fine-tuning A pretrained model has already learned from a large amount of other data. Fine-tuning teaches it your specific task on top of that.
Rollout Letting the trained policy control the real robot.
Hz "Times per second." A 30 Hz control loop sends a new command to the arm 30 times per second.

The stand at a glance

Part What's used
Arms SO-100 leader (moved by hand) + SO-100 follower (does the task), built from parts
Cameras 3 USB cameras at 640×480 resolution, 30 fps: two overhead webcams (left, right) on the official SO-ARM100 overhead mount, plus a small InnoMaker 32×32 mm camera on the wrist (grip)
Computer HP Z440 workstation with an RTX 3090 graphics card (24 GB of memory), running Proxmox with an LXC container (explained below)
Software LeRobot (version 0.6.0), used through LeLab, its web interface
Task "pick r2d2 and put to box": pick up a small R2-D2 figure and drop it in a box

The build: SO-100, from parts

I decided to build the SO-100 arms myself instead of buying them already assembled. Ordering the parts and putting them together felt like the right way to really understand the hardware before trying to teach it anything. It's also something students can see, repair, and later build themselves.

For the build itself I followed the LeRobot SO-100 guide. It covers the whole process, from setting up the motors to assembly and calibration. For buying the parts, the SO-ARM100 repo from The Robot Studio is the main reference: it has the parts list with links to suppliers, and the STL files for 3D printing.

On top of the basic arm, I added two of the optional camera add-ons from the SO-ARM100 repo:

  • Overhead camera mount, used twice. It's a 3D-printed tower that holds a webcam above the work area. It's held together with the small M2 screws that come with the Feetech servos. I use two of them, one on each side, for the left and right views. Because the tower is attached to the arm's base, the camera always looks at the arm from the same spot, session after session. That matters for a classroom stand that gets bumped, moved, and put back together.
  • Wrist camera plug mount. A single printed adapter plugs into the gripper and holds a 32×32 mm USB camera (at least 720p, 30 fps). I used the InnoMaker module that the repo recommends. I picked this design over the main wrist-camera design because it has fewer parts and you don't have to replace any of the arm's existing pieces. Printing tips from the repo: use tree supports and 40% infill. The camera's lens is focused by hand: you twist it until the picture is sharp at gripping distance.

Both mounts are designed for 640×480 at 30 fps, which is also what all three cameras are set to in LeLab.

Three versions of the stand

Version 1: a desk extender. The first stand was built on part of a desk extender: a board with two upright posts, with a webcam clamped to one of them on an adjustable arm. It worked for controlling the arm, but keeping the cameras steady was hard. The clamp could shift, and the posts weren't attached to the arm's base. Every bump moved the camera views a little. For imitation learning that's a real problem. The policy learns from what the camera sees, pixel by pixel, so a camera that has moved a few centimetres shows a different scene than the one the policy trained on.

First version: a desk-extender board with two posts and a clamped webcam, next to the arm

Version 2: camera towers on a small table. The cameras sit on two red 3D-printed towers from the overhead camera mount. Both towers attach to the same base plate as the arm, so the follower arm and both overhead cameras form one solid unit. If you move the stand, the camera views stay the same relative to the arm. The work area was a small white board with a green basket on one side. (The Pororo plush toy in the corner was the first test object, before R2-D2 took over.) Dataset 1 was recorded on this version.

Version 2: SO-100 follower with two overhead webcams on red 3D-printed towers, wrist camera, the R2-D2 figure and the green basket on a small white board

Version 2 fixed the shaky cameras, but it caused problems for the trained models. The table was small, so the cameras saw past its edges, and right next to the window, so the street outside was in the picture too. The policies trained on that scene struggled as soon as anything in the room changed (more on that in "Reworking the scene" below).

Version 3: a large picnic table. The same arm and camera towers now stand on a large, dark picnic table, away from the window. The table is big enough to fill almost the whole picture from both overhead cameras, so there's no floor, furniture or street in view anymore. The green basket was replaced by a low cardboard tray, and both the tray and R2-D2's five starting spots are marked with tape on the table. Dataset 2 and all the comparison results below were recorded on this version.

Version 3: the large dark picnic table with the arm and its red camera towers at the edge, the cardboard tray, and the taped start positions

Version 3 from above, behind the arm: the tray with its tape corner mark, the small red tape marks for R2-D2's start positions, and the roll of tape used later in the "new instruction" test

Lessons for anyone building one: attach the cameras firmly to the arm's base before recording anything, because steady cameras matter more than good cameras. And use a table big enough to fill the cameras' view, away from windows, so the background never changes.

Running it in an LXC container

The arms and cameras are connected to an LXC container: a separate, walled-off space on the computer that I use for experiments, so nothing I install there can mess up the rest of the machine. It keeps LeRobot, its drivers and the USB connections to the arms together in one place that's easy to rebuild.

The computer is borrowed lab hardware, not mine: an HP Z440 workstation with an RTX 3090 graphics card (24 GB of memory). Training can only use what fits on that card. So choosing a policy (see below) isn't only about the task. It's also about what this hardware can actually train.

That workstation already runs Proxmox, a program that lets one computer run several separate "virtual" computers and containers. So LeRobot had to go inside a container on top of Proxmox instead of being installed directly on the machine. That caused its own problems, mainly getting the arms' USB connections and the USB cameras from the main system through to the container ("passthrough"). I worked through it with help from an AI, but the details are enough for their own article, so that's coming separately.

For a class, I'd skip the container. If you install LeLab directly on the computer (Linux, Windows or macOS), the arms and cameras simply show up as normal devices. There's no passthrough to fight with. The container only made sense here because the borrowed workstation was already set up with Proxmox.

The first check inside the container is just making sure the devices actually show up:

ls -l /dev/video* /dev/ttyACM* /dev/ttyUSB* 2>/dev/null

On this stand that shows the two arm controller boards, /dev/ttyACM0 (leader) and /dev/ttyACM1 (follower), and /dev/video0 to /dev/video5 for the three cameras. Each USB camera shows up twice, and the even-numbered ones (0, 2, 4) are the ones that actually send video.

Once the cameras are visible, the next step is checking what they see, which helps when you position them. ustreamer shows a live camera picture in your web browser, without starting LeRobot:

ustreamer --device=/dev/video2 --host=0.0.0.0 --port=8081 --allow-origin=*

Open http://<container-ip>:8081 and adjust the camera's angle and distance until the whole work area and the arm are in the picture.

Installing LeLab

LeLab is the only thing installed by hand. It's a web app that puts every step (calibrate → control → record → train → replay) in one browser page, so nobody has to remember commands or answer typed prompts in the middle of a recording. It installs with uv, a tool for installing Python programs:

uv tool install git+https://github.com/huggingface/leLab.git
lelab

LeLab brings its own copy of LeRobot (version 0.6.0), so there's no separate LeRobot install. That matters later: when a policy needs an extra Python package, it has to be added to LeLab's copy, for example with uv tool install --force --with scipy --with num2words git+https://github.com/huggingface/leLab.git (--force reinstalls LeLab with the extra packages).

Gotcha: skip the USB hub. At first everything (two arm controllers and three cameras) went through one USB hub, and the hub caused trouble. Three video streams plus two arm connections is a lot for one shared cable, and every unreliable connection showed up as a camera dropping out or an arm losing contact. The fix was simple: remove the hub and plug every device straight into the Z440, which has plenty of USB ports. Since then the connections have been stable.

Calibration and cameras

In LeLab the robot is set up once, under the name so-100: leader on /dev/ttyACM0, follower on /dev/ttyACM1, and three cameras named left, grip and right. Both arms are calibrated through the web page. Calibration teaches the software where each joint's limits and middle position are, so "this motor position" means the same physical pose on both arms. The calibration files are saved in ~/.cache/huggingface/lerobot/calibration/.

Using three cameras instead of one is on purpose. The two overhead cameras see the whole work area from different angles, which helps the policy judge depth and where the object is. The wrist camera moves with the gripper and shows the last few centimetres before a grab, which is where most failures happen.

These snapshots were taken on version 2 of the stand, the small white table (see above).

left (overhead) grip (wrist) right (overhead)
Left overhead camera, with the box in view Wrist camera looking past the gripper jaw Right overhead camera

Gotcha: camera numbers can change. LeLab refers to the cameras by number (0, 2, 4). Linux hands out those numbers in the order it finds the devices, not by which physical camera is which. After one camera was unplugged and plugged back in, the numbers shifted. The dataset recorded after that saved the wrist view under the name right and the left overhead view under grip (fixed since). A policy trained on mislabeled data gets the wrong picture in each slot as soon as the cameras are labeled correctly again. The real fix is to identify each camera by the USB port it's plugged into, not by its number. On a classroom stand where cables get pulled out and plugged back in, this is a must.

Gotcha: loose screws. During calibration, some of the arm's screws turned out to be loose. A loose joint wobbles in a way the servo can't detect: the motor reports one position while the arm part actually sits at another. That ruins the calibration, and later the recorded examples. Tightening the screws and adding thread-locker (a glue that stops screws from loosening) fixed it for good. On a stand that students will handle every week, I'd use thread-locker on every screw from the start, and check for wobble before each calibration.

Once the screws were locked, the calibration stayed good: there's no need to redo it between sessions on this computer. It does matter when the stand moves to a different computer. The calibration is saved in files on the computer (~/.cache/huggingface/lerobot/calibration/), not inside the arm, so a new computer starts without it. Either calibrate again there, or copy those files over. Keeping them with the stand (for example on a USB stick) saves a step when the class sets it up on their own computers.

The task: R2-D2 into the box

The first task is a simple pick-and-place: pick up a small R2-D2 figure and put it in a box. It's easy for a student to demonstrate, and it's easy to tell whether it worked: the figure is either in the box or it isn't.

The first object was actually a Pororo plush toy. Pororo works, but its shape is more complicated: a soft body, a big head, and a different way to grab it depending on how it's lying. A policy needs more examples to learn all those cases. The small R2-D2 figure looks about the same from every side, so a first dataset of 50 episodes is enough for it. That's a useful rule when picking a first task for a class: the simpler the object's shape, the fewer examples you need. Pororo stays on the list as a harder follow-up task.

Recording demonstrations

Examples are recorded by teleoperation: you move the leader arm and the follower copies it, while LeLab records the arm's joint positions, the movement commands, and all three camera videos at 30 frames per second.

The leader arm sits on its own desk right next to the stand, clamped to the edge with a red 3D-printed base. Instead of a gripper, it has a handle with a trigger, so you hold it like a tool and squeeze to close the follower's gripper. Keeping it on a separate desk means bumping the leader during a recording doesn't shake the follower's table or cameras, and it keeps the operator's hands out of the overhead views.

Teleoperation setup: the SO-100 leader arm with handle and trigger, clamped to a desk next to the stand

Dataset Episodes Frames Total length
1: r2d2_to_box_2 (white table, by the window) 50 11,433 6.4 min (7.6 s per episode)
2: r2d2_to_box_bg (reworked scene, see below) 50 14,199 7.9 min (9.5 s per episode)

Dataset 1 was recorded to test the whole stand from start to finish: arms, all three cameras, recording, uploading and training. It's public on the Hub:

(The Hub labels it so-101. The SO-101 is a newer version of the same arm with different motors, and for this stand the two work the same. The arms here are SO-100s.)

It was recorded right after the camera numbers got mixed up (described above), so it was first uploaded with the camera videos under the wrong names. That's fixed on the Hub now. Renaming the video folders and the camera labels was enough, with no need to record again. It's a good reminder to check every new dataset in the visualizer before training on it.

How does this compare with what LeRobot recommends? The LeRobot agent guide (a guide to choosing and training policies, written for AI assistants that help people with LeRobot, but just as useful for humans) suggests these starting settings for a first task:

Setting Recommended This stand
Episodes 50 to start, 100–300 later 50 ✅
Episode length 20–45 s ~8–9 s (the task is short)
Frames per second 30 30 ✅
Cameras 2: one fixed front camera + one wrist camera 3: two overhead + wrist
Task description short, specific, describes an action "pick r2d2 and put to box" ✅

Its rules for good data are worth printing out and taping to the stand, because they're easy to forget in the middle of a recording session:

  • "Good data beats clever models."
  • "Deliberate, high-quality execution beats fast sloppy runs."
  • Grab, approach and time it the same way every time. Consistent examples are much easier to learn from than very different ones.
  • Start with a simple version of the task: one object, a fixed position, fixed cameras, one operator.

The LeRobot community datasets blog post has a "what makes a good dataset" checklist, written with sharing in mind. Checking the stand against it was a useful exercise, and one I'd repeat with students:

Recommendation This stand
Preferably two camera views, steady, sharp focus ✅ three fixed views; the wrist lens is focused by hand
Even, steady lighting ⚠️ dataset 1: next to a window, so daylight changes through the day · ✅ dataset 2: curtains + a lamp that doesn't flicker
Plain background that doesn't distract ⚠️ dataset 1: window and street visible · ✅ dataset 2: a dark, non-shiny table fills the picture
Leader arm not in the picture; only the follower and the object move ✅ the leader sits on its own desk, outside every camera's view
At least 640×480, about 30 fps ✅ 640×480 at 30 fps
Correct robot type in the dataset's information ⚠️ dataset 1 is labeled so-101 on the Hub, but the arm is an SO-100
Camera names that say where the camera is (images.left, images.wrist, …), not what brand it is ✅ left / right; grip would be clearer as wrist
Clear, short task description, 25–50 characters, not task1 "pick r2d2 and put to box" is 24 characters; "Pick the R2-D2 figure and put it in the box" would fit the format better

The lighting and background points matter more than they look. A policy only sees pixels, so to the policy, a sunny afternoon or a passing car is a change just as big as moving the object.

Reworking the scene: dataset 2

The checklist's two warnings turned out to be the real story. An ACT policy (one of the policy types, explained below) trained on dataset 1 struggled as soon as the scene around it changed. It's a textbook case of a policy learning the room instead of the task. So before recording again, I changed the scene one thing at a time, checking a still picture from every camera after each change:

  1. A bigger table, away from the window. The stand moved to a large picnic table (version 3), and curtains went up, so no camera sees daylight or the street anymore.
  2. A lamp that doesn't flicker. The first new lamp put horizontal stripes across the wrist camera's picture. Changing the cameras' anti-flicker setting from 50 Hz to 60 Hz (to match the power grid here) didn't help, and neither did setting the exposure by hand. The lamp itself flickered too fast. A different lamp fixed it.
  3. A dark, non-shiny table. The white arm, the white gripper jaws and the grey R2-D2 almost disappeared against the white tabletop, and every shadow stood out strongly. A dark, matte (non-shiny) surface that fills almost the whole picture fixed both problems, and the picnic table's top is exactly that. Matte matters: a dark sheet I tried before the picnic table was slightly shiny and reflected the lamp into the wrist camera.
  4. Wrist camera brightness. On the dark table, the wrist camera's automatic brightness overcorrected and turned the white gripper jaws into flat white blobs (7% of the picture was pure white, with no detail). Lowering its brightness setting to -16 and turning off "backlight compensation" brought that down to about 0%.
  5. A low tray instead of a basket. A shallow cardboard tray keeps R2-D2 visible to both overhead cameras even when it's inside, it's easy to drop into, and it stands out well against the dark table. Its position is marked with tape so it always goes back to the same spot.
  6. Fixed start positions. Five marked starting spots for R2-D2, about ten episodes each, recorded in mixed order.

The left overhead camera before and after: dataset 1 on the white table next to the window (outside view blurred), and dataset 2 on the dark table with the cardboard tray and R2-D2 at a marked start position

Two practical lessons from that session:

  • Camera settings are lost when you unplug a camera. The anti-flicker and brightness settings are stored in the camera itself, not in LeLab, and they reset when the camera is unplugged. A small script (scripts/camera-setup.sh) finds each camera by the USB port it's plugged into and applies the settings again before every session.
  • A loose motor cable. Halfway through setup, one arm stopped responding past the elbow: servos 4 to 6 didn't answer at all. The servos are connected in a chain, one after another, so one loose plug between servo 3 and servo 4 cut off everything after it. Plugging it back in firmly fixed it. That cable bends every time the elbow moves, so it's worth securing at the plug.

Dataset 2 is public too:

Recording a 50-episode dataset takes about 40 minutes. Only about 8 of those minutes are the recorded episodes themselves (about 9.5 seconds each). The rest is resetting between episodes: putting R2-D2 back on its next starting mark, moving the leader arm back to its rest position, and the short pause LeLab leaves between recordings. Plan for that when a whole class wants to record: roughly one minute per episode.

Choosing a policy: what fits on a 3090

LeRobot 0.6.0 comes with about twenty kinds of policies. Some are small and learn one task from scratch, like ACT and Diffusion Policy. Others are large vision-language-action models (VLAs): models that were pretrained on lots of data, take in camera images plus a written instruction, and output arm movements, like SmolVLA, π0 and GR00T. The quickest way to see what a graphics card can handle is LeRobot's compute hardware guide. It groups the policies by how much graphics card memory (VRAM) they need when training on 8 frames at a time. The table below puts the guide's numbers next to what actually happened on this stand's RTX 3090 (24 GB, limited to 250 watts of power):

Policy Guide group Guide VRAM (batch of 8) On this 3090 Measured here
ACT (act) Light BC ~2–6 GB ✅ trained + tested 5.8 GB, 3.8 steps/s, 8/10 on the arm
Diffusion Policy (diffusion) Diffusion ~8–14 GB ✅ trained + tested 10.8 GB, 2.5 steps/s, 9/10
SmolVLA (smolvla, fine-tuned from smolvla_base) Small VLA ~10–16 GB ✅ trained + tested 3.6 GB (vision part frozen), 2.4 steps/s, 10/10
GR00T N1.7 (groot) Multimodal ~24–40 GB ✅ trained, only with 16-bit numbers + batch of 4 23.2 GB peak, 1.7 steps/s, 2 h 57 min for 18k steps, 15 GB per checkpoint, 14/14 on the arm
VQ-BeT (vqbet) Light BC ~2–6 GB 🟡 should fit, not tried
Multi-task DiT (multi_task_dit) Diffusion ~8–14 GB 🟡 should fit, not tried (built for several tasks with written instructions)
SmolVLA, vision part unfrozen Small VLA ~10–16 GB 🟡 should fit, not tried (guide: "usually substantially better" on specialized tasks)
π0, π0.5, π0-FAST (pi0, pi05, pi0_fast) Large VLA ~24–40 GB 🟠 ran out of memory with default settings π0: out of memory at a batch of 8; the guide calls 24 GB "tight at batch 1"
X-VLA, Wall-X (xvla, wall_x) Large VLA ~24–40 GB 🟠 maybe, with the same tricks as GR00T
EO-1 (eo1) Multimodal ~24–40 GB 🟠 maybe, with the same tricks as GR00T
TD-MPC (tdmpc) Light BC ~2–6 GB ⚪ fits, but not a match it needs a reward score for each try, not just recorded examples
evo1, fastwam, gaussian_actor, lingbot_va, molmoact2, vla_jepa not in the guide unknown ⚪ not tested a short test run would tell

So on this stand the easy choices are ACT, Diffusion Policy and SmolVLA, with GR00T possible if you squeeze. That's still a good mix for a class: a small policy trained from scratch, one based on diffusion, a small pretrained VLA, and a large one. The large VLAs (the π0 family, X-VLA, Wall-X) are better trained on a rented, bigger graphics card such as an A100, for example through Hugging Face Jobs. Only the finished model then comes back to the 3090 to run on the arm.

The four policies, briefly

ACT (Action Chunking with Transformers) comes from the paper Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (Zhao et al.), which was designed for exactly this kind of cheap hardware. A small image network (ResNet-18) turns each camera picture into a list of features. A transformer (the same kind of AI building block used in chatbots) combines those features with the arm's joint positions. It then predicts a whole chunk of future moves at once, instead of one move at a time, which keeps the motion smooth. ACT is fairly small (about 80 million adjustable numbers, called parameters), needs nothing extra installed, often works well with about 50 examples, and is the policy LeRobot recommends starting with. The LeRobot team's ACT tutorial video is a good introduction to show in class.

Diffusion Policy (Chi et al.) creates movements the same way AI image generators create pictures. It starts with random noise and cleans it up step by step until it becomes a sequence of arm movements. That makes it good at tasks where there's more than one right way to move. The cost is longer training (the agent guide suggests 80,000–150,000 steps, against 30,000–80,000 for ACT) and more computing time for each decision.

SmolVLA (paper) is Hugging Face's small vision-language-action model, with 450 million parameters. This is the stand's window into physical intelligence. It takes in the camera views, the arm's position and a written instruction, and a part called the "action expert" turns them into a chunk of moves. Unlike ACT, it doesn't start from scratch: you fine-tune the pretrained smolvla_base model on your own ~50 episodes. Its documentation makes a point worth repeating to students. Their example dataset used 50 episodes spread over 5 positions of a cube (10 each), and a 25-episode version "was not enough." Having enough examples for each variation matters more than the total number.

GR00T N1.7 is NVIDIA's open robot foundation model, originally built for humanoid robots but trained on data from many kinds of robots. It's the biggest model here: 3.1 billion parameters, almost seven times the size of SmolVLA. Like SmolVLA, it has two parts. A vision-language model (nvidia/Cosmos-Reason2-2B) looks at the camera pictures and reads the written instruction, and an action part turns that understanding into a chunk of arm movements, using the same "start from noise and clean it up" idea as Diffusion Policy. It's pretrained, so you fine-tune it on your own episodes, and LeRobot tells it that this arm is a "new embodiment" (a robot body it hasn't seen before). The catch is size: training it barely fits on a 24 GB graphics card, and only with some memory-saving tricks (see "Squeezing GR00T N1.7 onto a 3090" below).

Before finding the guide, my first round was simply to try everything on dataset 1. Here's what happened then, and where each policy stands now:

Policy First try Now
ACT ✅ trained right away (~5.8 GB of VRAM at a batch of 8, ~3.8 steps per second) ✅ trained on dataset 2, 8/10 on the arm
Diffusion Policy not part of the first round ✅ trained on dataset 2, 9/10 on the arm
SmolVLA ❌ a missing Python package (num2words) ✅ package installed, fine-tuned from smolvla_base, 10/10 on the arm
GR00T ❌ its base model (nvidia/Cosmos-Reason2-2B) needed permission to download, then it ran out of graphics memory (24 GB) ✅ access granted; trained with 16-bit numbers and a batch of 4 (see below)
π0 ❌ its base model (google/paligemma-3b-pt-224) needed permission to download, then it ran out of graphics memory ⛔ not retried; too big for this card with default settings
π0-FAST ❌ a missing Python package (scipy, part of lerobot[pi]) ⛔ package installed, but not retried; same size group as π0

The memory failures match the guide exactly. The other failures were setup problems, not hardware limits:

  • Some models need permission to download. You ask for access on the model's web page and log in with hf auth login. Sort this out before a lesson, not during one.
  • Missing packages are easy to fix. They just have to be installed into LeLab's copy of LeRobot (see the install section).
  • I also hit an early error where a video-reading library (torchcodec) wouldn't load (on my first, Pororo dataset), which stopped training before it started.

ACT is the reliable workhorse: small, fast to train, and made for short tasks learned from a few dozen examples. Diffusion Policy and SmolVLA are compared with it below, and GR00T gets its own section after that.

The LeRobot agent guide also says which policy to use when. ACT is for "first-time users, laptops, single-task. Fast and reliable." Diffusion Policy is for tasks with more than one right way to move. SmolVLA is for setups with several tasks and written instructions. For a 24 GB card, it suggests ACT for a single task and SmolVLA (with its vision part unfrozen) for several tasks. Its memory numbers are lower than the hardware guide's because they were measured with a simpler training method (SGD). LeRobot normally uses a different one (AdamW) that needs more memory, so plan with the hardware guide's numbers. One tip for later: letting SmolVLA's vision part keep learning (--policy.freeze_vision_encoder=false --policy.train_expert_only=false) "usually improves performance substantially on specialized tasks," but it needs more graphics memory.

Training ACT

Run Dataset Steps Batch ≈ Epochs Time Notes
1 dataset 1 (11,433 frames) 50,000 8 ~35 3 h 39 min (~3.8 steps/s) trained before the camera-name fix, so it doesn't match the corrected setup
2 dataset 1, names fixed 15,000 8 ~10.5 1 h 06 min length chosen with the epoch rule below

LeLab's job page for run 1, trained on the RTX 3090: progress at 8,530 / 50,000 steps, loss down to 0.278, a flat learning rate of 1e-5, a "Run on robot" button for any saved checkpoint, and the live training log

The hardware guide also changed how I decide how long to train. Imitation learning usually finishes learning after 5–10 epochs (full passes through the dataset), not after some fixed big number of steps. The math is simple:

steps_per_epoch = total_frames / batch_size
total_steps     = epochs × steps_per_epoch

For dataset 1 with a batch of 8, that's about 1,430 steps per epoch. Run 1's 50,000 steps were about 35 epochs, far more than needed, while run 2's 15,000 steps come to about 10.5. The speed matches the guide's estimate for a 3090, too: ACT with a batch of 8 takes about 30–60 minutes for 5 epochs on a ~50-episode dataset. One more tip from the guide: some policies slowly lower their learning rate (how big each adjustment is) as training goes on. If you shorten training, shorten that schedule too (scheduler_decay_steps ≈ steps), or the learning rate never gets lowered. The agent guide's suggested step counts (for example 30,000–80,000 steps for ACT on 50 episodes) assume 30-second episodes, about 45,000 frames. This stand's episodes are only about 8 seconds long, so the same 5–10 epochs is just ~7,000–14,000 steps with a batch of 8. Always work it out from epochs, not from someone else's step count.

Comparing ACT, Diffusion Policy, SmolVLA and GR00T

The plan: train the three policies that fit comfortably on the 3090 on the same data, dataset 2, and compare them on the real arm. GR00T, which only fits with some tricks, was trained on the same data afterwards. It's included in the tables below, and the tricks are explained in its own section further down. That's the most useful experiment the stand can offer a class: same data, same robot, three very different ideas about how to turn camera pictures into movement.

ACT Diffusion Policy SmolVLA GR00T N1.7
Main idea A transformer predicts a chunk of moves directly Starts from random noise and cleans it up into a sequence of moves A pretrained vision-language model plus an "action expert" A pretrained vision-language model plus an action part that cleans up noise into moves
Starting point Trained from scratch Trained from scratch Fine-tuned from smolvla_base Fine-tuned from nvidia/GR00T-N1.7-3B
Size ~80 million parameters an image network per camera + a noise-cleaning network ~450 million parameters, of which 100 million are trained (the action expert) ~3.1 billion parameters, of which 1.6 billion are trained
Uses the written task No No Yes ("pick r2d2 and put to box") Yes ("pick r2d2 and put to box")
Moves per prediction (defaults) chunk of 100, all used plans 64, uses 32, looks at the last 2 frames chunk of 50, all used chunk of 40, all used
Default learning rate 1e-5, stays the same 1e-4, slowly lowered 1e-4 1e-4, short warm-up, then slowly lowered
VRAM (hardware guide, batch of 8) ~2–6 GB (measured: 5.8 GB) ~8–14 GB (measured: 10.8 GB) ~10–16 GB (measured: 3.6 GB with the vision part frozen) ~24–40 GB (measured: 23.2 GB peak at a batch of 4, with 16-bit numbers)
Steps on dataset 2 15,000 (~8.5 epochs) 36,000 (~20 epochs) 25,000 (~14 epochs) 18,000 at a batch of 4 (~5 epochs)
Training time on the 3090 (limited to 250 W) 1 h 06 min (~3.8 steps/s) 4 h 04 min (~2.5 steps/s) 2 h 57 min (~2.4 steps/s) 2 h 57 min (~1.7 steps/s)
Final training loss 0.166 0.006 0.062 0.035
Checkpoint size ~0.6 GB ~3.3 GB ~1.3 GB ~15 GB
How it's trained here LeLab web page LeLab web page Command line (see below) Command line (see GR00T section)
Success, trained positions 8/10 9/10 10/10 10/10
Success, new positions 2/4 2/4 2/4 4/4
Average time to finish the task ~10 s ~16.5 s ~8.4 s ~9.8 s

Why the different step counts? Diffusion Policy is known to need longer training than ACT, while SmolVLA and GR00T start from a pretrained model and learn quickly. All four were planned in epochs first, then converted to steps for dataset 2's 14,199 frames: about 1,775 steps per epoch with a batch of 8, or about 3,550 with GR00T's batch of 4.

Training settings

All four were trained on dataset 2 (van-i/r2d2_to_box_bg_20261003_210444, 50 episodes, 14,199 frames) on the RTX 3090. The setting names match LeLab's training form. Anything not listed was left at its default.

Setting ACT Diffusion Policy SmolVLA GR00T N1.7
Started from LeLab web page LeLab web page command line command line
Policy act diffusion --policy.path=lerobot/smolvla_base --policy.type=groot (loads nvidia/GR00T-N1.7-3B)
Steps 15,000 36,000 25,000 18,000
≈ Epochs ~8.5 ~20 ~14 ~5
Batch size 8 8 8 4 (8 doesn't fit)
Learning rate and related settings policy default (1e-5, stays the same) policy default (1e-4, slowly lowered) policy default (1e-4) policy default (1e-4, warm-up, then slowly lowered)
Use policy training preset on on n/a (comes with smolvla_base) on (default)
Image transforms (random changes to the pictures) off off off (default) off (default)
Number precision mixed precision on mixed precision on default 16-bit (bf16) everywhere: --policy.model_params_fp32=false
Vision part trained (ResNet-18) trained (one ResNet-18 per camera) frozen (freeze_vision_encoder=true, train_expert_only=true) frozen (default: vision-language part not trained)
Camera inputs left, grip, right left, grip, right renamed to camera1/2/3 with --rename_map left, grip, right
Save a checkpoint every 1,000 steps 4,000 steps 5,000 steps 6,000 steps
Workers / seed / log frequency 4 / 1000 / 250 4 / 1000 / 250 4 / 1000 / 100 4 / 1000 / 100

A few notes on these choices:

  • No random picture changes this round. All four trained without image transforms, so the comparison is fair. Turning them on (random changes to brightness, contrast and color during training) is the obvious next experiment for ACT and Diffusion. Those two learn to see from scratch, so they're the most sensitive to lighting changes.
  • Mixed precision uses faster, slightly less precise math on the 3090. If the loss suddenly shows nan ("not a number") or jumps around, train again with it off.
  • How often to save is mostly about disk space. A SmolVLA checkpoint is about 1.3 GB and a GR00T checkpoint about 15 GB, so those two save less often.
  • Loss numbers can't be compared between policies. Each policy measures its errors in a different way, so 0.006 for Diffusion doesn't mean it's better than ACT at 0.166, or than GR00T at 0.035. Only the real arm can tell. For one policy on its own, though, the loss shows whether training worked. SmolVLA's loss dropped from 0.72 at step 100 to 0.28 within the first epoch, and settled around 0.06 after about 15,000 steps.
  • The 3090 is limited to 250 watts on this workstation (normally it can use 350 W), so these times are a bit slower than on a standard card.

A SmolVLA gotcha in LeLab. LeLab's training form offers SmolVLA, but it doesn't start from the pretrained model. In LeRobot 0.6.0, that means the part of SmolVLA that understands pictures and words starts out empty (random) instead of pretrained. By default that part is also frozen, so it never learns, and only the small action expert trains, on top of meaningless features. LeRobot itself warns that this "is unlikely to yield good results." The pretrained knowledge is the whole point of a VLA, so SmolVLA was trained from the command line instead, starting from smolvla_base. Its three expected camera inputs are mapped to this stand's cameras.

Training. This uses LeLab's own copy of LeRobot (~/.local/share/uv/tools/lelab/bin/lerobot-train), which needs the num2words package installed. Saving the results in LeLab's outputs/train folder keeps all training runs in one place:

OUT=$HOME/.cache/huggingface/lerobot/outputs/train/smolvla_van-i_r2d2_to_box_bg_$(date +%Y-%m-%d_%H-%M-%S)

~/.local/share/uv/tools/lelab/bin/lerobot-train \
  --policy.path=lerobot/smolvla_base \
  --dataset.repo_id=van-i/r2d2_to_box_bg_20261003_210444 \
  --rename_map='{"observation.images.left":"observation.images.camera1","observation.images.grip":"observation.images.camera2","observation.images.right":"observation.images.camera3"}' \
  --batch_size=8 --steps=25000 --save_freq=5000 --log_freq=100 \
  --policy.device=cuda --policy.push_to_hub=false \
  --output_dir=$OUT --job_name=smolvla_r2d2_bg

A run started this way doesn't appear in LeLab's list of jobs, because LeLab only lists jobs it started itself. For test runs, scripts/run-smolvla.sh wraps the command below: ./scripts/run-smolvla.sh [checkpoint] [seconds] ["task text"].

Running it on the robot. The trained model expects cameras named camera1/2/3, so running it needs the same --rename_map as training, plus the exact task text the dataset was recorded with. This is the same command LeLab would build, with the rename added: the follower arm's port, its calibration name, and the three cameras by number (left 2, grip 4, right 0 on this stand; check again after any unplugging). echo | answers the "press ENTER to use the calibration file" question automatically:

echo | ~/.local/share/uv/tools/lelab/bin/python -m lerobot.scripts.lerobot_rollout \
  --strategy.type=base \
  --policy.path=$HOME/.cache/huggingface/lerobot/outputs/train/smolvla_van-i_r2d2_to_box_bg_2026-10-03_22-54-34/checkpoints/025000/pretrained_model \
  --policy.device=cuda \
  --robot.type=so101_follower --robot.port=/dev/ttyACM1 --robot.id=so-100 \
  --robot.cameras="{left: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}, grip: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}, right: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --rename_map='{"observation.images.left":"observation.images.camera1","observation.images.grip":"observation.images.camera2","observation.images.right":"observation.images.camera3"}' \
  --task="pick r2d2 and put to box" \
  --duration=60

(so101_follower is what LeLab uses for both arm versions; the SO-100 works fine with it.) The log from the first run shows an uneven control loop, the cycle of "look, decide, move" that should repeat 30 times per second. It ran at a typical 25.9 times per second, with short drops below 1. The likely reason: SmolVLA plans 50 moves at once, and by default the arm waits while the next 50 are being calculated. LeRobot has a mode called real-time chunking (--inference.type=rtc) that calculates the next batch of moves while the current one is still playing. That's the next thing to try.

What to look for in the comparison, and to discuss with students:

  • Success rate over a fixed number of tries (for example 10) with the object in its trained positions, then a few tries with it moved a little, to see which policy can handle new situations.
  • Smoothness: ACT's chunks of moves versus Diffusion's cleaned-up movement paths.
  • Thinking time: Diffusion and SmolVLA do more calculating for each decision, which matters given the control-loop warnings below.
  • Training cost: about an hour for ACT versus several hours for the others, for whatever improvement in success they bring.

First rollouts

ACT and Diffusion rollouts are started from LeLab: pick a saved checkpoint, connect the follower arm and the cameras, and let the policy control the arm for 30–120 seconds. SmolVLA and GR00T are started with the scripts in scripts/, which run the same LeRobot command.

One thing the logs show clearly: the control loop often ran slower than its target of 30 times per second, from short dips down to a few times per second, to a steady ~26. LeRobot warns that this can cause missed camera frames and shaky control. It lists the usual causes: cameras that can't keep up, slow policy calculations, or a busy processor. With three cameras and the policy all running on one computer, this needs a closer look.

The agent guide also gives a clear goal to measure against: for a single pick-and-place task with 50 clean examples, "ACT should reach >70% success on the training configuration. Less → data problem, not model problem." Its quick troubleshooting list is a good lesson in itself:

  • Fails at one stage (for example, grabs the object but misses the box): record 10–20 more examples focused on that stage.
  • Wobbles back and forth: the examples weren't consistent, or training was too short. Record the worst examples again.
  • Ignores the object: a problem with camera position or lighting, not with the model.

That last one is a reminder that the camera mix-up described earlier would look exactly like a "bad policy."

First result: a scene change breaks a policy trained from scratch. The ACT run 2 model (15,000 steps, trained on dataset 1) was tested after the scene had been reworked: dark table, lamp light, a different box. It failed: it couldn't do the task in the new scene. That's expected. ACT learns to see only from your 50 examples, so to ACT, a scene that looks different is a different task. That's also why dataset 2 exists, and it makes a good experiment for a class: train in scene A, test in scene B.

Results: ACT vs Diffusion Policy vs SmolVLA vs GR00T

The full test protocol and every single try are in docs/test-plan.md. Each policy's final checkpoint (ACT at 15,000 steps, Diffusion at 36,000, SmolVLA at 25,000, GR00T at 18,000) got the same 14 tries: each of the five marked starting spots twice, plus two new spots twice each. H1 is between two of the marks. H2 is slightly outside the marked area. Each try was started separately: place R2-D2, the operator steps back, the policy starts. ACT, Diffusion and SmolVLA took turns at each spot, so any slow change in light or temperature affected them equally. GR00T was tested afterwards, once its training had finished, in the same setup.

ACT (15k) Diffusion (36k) SmolVLA (25k) GR00T N1.7 (18k)
Trained spots (P1–P5) 8/10 9/10 10/10 10/10
H1 (between marks) 2/2 2/2 2/2 2/2
H2 (outside the marks) 0/2 0/2 0/2 2/2
Average time to finish ~10 s ~16.5 s (13–18 s) ~8.4 s (7–10 s) ~9.8 s (5–20 s)
How it failed went to the wrong spot (all 4 failures) grab failed (2), didn't really move (1) grab failed (both at H2) never; at P3 and once at P4 it dropped R2-D2, went back, picked it up and finished
Control loop (target: 30 per second) typically 28.6 typically 28.8, one ~2 s pause per run typically 26.1, ~13 short pauses per 30 s run smooth: only 1–3 slow moments per run, all at the start
Time from start to the arm moving ~6 s ~16 s ~21 s ~82 s

For comparison: a human example in dataset 2 takes about 9.5 seconds.

All four trained policies are on the Hugging Face Hub, each with a model card that has its training settings, results and a command to run it:

Policy Model repo
ACT van-i/act_r2d2_to_box_bg
Diffusion Policy van-i/diffusion_r2d2_to_box_bg
SmolVLA van-i/smolvla_r2d2_to_box_bg
GR00T N1.7 van-i/groot_r2d2_to_box_bg (NVIDIA Open Model License)

To try one on your own SO-100 or SO-101, pass the repo name as --policy.path, for example --policy.path=van-i/act_r2d2_to_box_bg. Expect it to work only in a scene that looks like this one: dark matte table, the tray in the same place, similar light and camera positions.

What the numbers say:

  • All four learned the task. Even ACT, the smallest, clears the agent guide's ">70% on the training setup" goal, so the data is good enough.
  • The two pretrained VLAs were the best. SmolVLA and GR00T both got 10 out of 10 on the trained spots, about as fast as a human example. That's pretraining paying off: they already know what objects look like, and only had to learn this arm and this task.
  • Only GR00T handled the new spot outside the training area. All four handled H1, between two trained spots. At H2, just outside them, ACT, Diffusion and SmolVLA failed every time, and GR00T succeeded both times. Fifty examples from five spots mostly teach a policy those five spots and the space between them, not "find R2-D2 anywhere on the table." The biggest model, with the most pretraining, was the only one that stretched beyond that. It's the clearest lesson of the whole test for a class: here, the bigger pretrained model handled a new situation better, but spreading out the starting spots in the data is still the reliable fix.
  • GR00T can recover from mistakes. At P3, and once at P4, it dropped R2-D2 on the way to the box, turned back, picked it up again and finished. Those tries still count as successes, just slower ones.
  • The policies failed in different ways. ACT's misses were all reaches to the wrong place. Grabbing was easy for it, but it sometimes misjudged where R2-D2 was. Diffusion and SmolVLA almost always went to the right place and occasionally fumbled the grab itself. Diffusion's carefulness cost time: it was about 60% slower than ACT.
  • Smoothness was different. SmolVLA calculates 50 moves at a time, and with the default settings the arm waits briefly each time it calculates the next batch, about once every 2 seconds. GR00T runs in real-time chunking mode (--inference.type=rtc), which calculates the next moves while the current ones are still playing, and its control loop barely slowed down at all. The same mode should remove SmolVLA's pauses too.
  • Size has a price. GR00T needs about 80 seconds to get ready before every run, takes almost all of the graphics card's memory, and each saved checkpoint is 15 GB, against about 0.6 GB (ACT), 1.3 GB (SmolVLA) and 3.3 GB (Diffusion) for the others.

A note on fairness: Diffusion and SmolVLA had a 30-second time limit, ACT and GR00T had 60 seconds, so a slow Diffusion success at H2 might have been cut off (GR00T's slowest success took 20 s, so it would have passed with 30 s too). GR00T was also tested a few hours after the others, not mixed in with them. And with only 14 tries per policy, a difference of one try (8 vs 9 vs 10) is a hint, not proof. The big differences are the reliable findings: wrong-spot failures versus grab failures, and that only GR00T got past the wall at H2.

Do the VLAs follow a new instruction? Not yet

SmolVLA reads a written instruction, so the obvious next question was whether it can do something it wasn't trained on. I gave it a new object, a roll of tape, and a new instruction: "grab roll and move it to left spot".

It failed, but in an interesting way. It reached for the roll and closed the gripper on it, so its pretrained vision did recognize "an object to pick up," even though it had never seen a roll on this table. But as soon as it thought it was holding the object, it carried it to the box, exactly like in the R2-D2 task, and ignored the "left spot" part of the instruction.

That's what you'd expect from how it was fine-tuned. All 50 episodes in dataset 2 had the same instruction: "pick r2d2 and put to box." If the text never changes during training, it gives the model no useful information, so the model learns to act on the pictures and the arm's position alone. After 25,000 training steps on one task, the instruction is basically decoration, and the only thing the model knows to do after grabbing something is "go to the box." The pretrained smolvla_base had seen many different tasks, but fine-tuning on just one narrows it down to that one.

GR00T, the bigger model, got further, but not all the way. It got the same kind of test after its training finished, with longer, 6-minute runs (./scripts/run-groot.sh 018000 360 "..."):

  1. The roll alone. Instruction: "pick tape roll and put into box". GR00T picked up the tape roll and almost got it into the box. (SmolVLA got a different instruction, so the two tests aren't directly comparable, but both models were willing to grab an object they'd never trained on.) The roll has a very different shape from R2-D2, so carrying it almost all the way is GR00T's larger pretraining showing again, just like at the H2 spot.
  2. The roll and R2-D2 on the table together, with step-by-step coaching. Instruction: "pick tape roll and put into box, to pick move grip lower, close grip and hold until roll above box". GR00T ignored the roll at first and went for R2-D2, the object from its training, and almost put it in the box. Only after that did it turn to the roll and try to grab it several times.

The second test is the more telling one. The instruction named the roll clearly and even explained how to grab it, but as soon as the familiar object was on the table, the training habit ("grab R2-D2, take it to the box") won. The coaching didn't help either. The model never saw instructions like "move grip lower" connected to arm movements during training, so to it, they're just more words. A larger model is better at handling a new object, but it still hasn't learned to choose what to do based on the text, because every training example had the same instruction.

The lesson for a class: a vision-language-action model only follows instructions if its training data forces it to. That means several tasks in the same scene that can only be told apart by their instructions. The follow-up experiment is simple: record a second task in this scene with its own instruction (for example "move r2d2 to the left spot," with a taped target), train SmolVLA or GR00T on both datasets together, and check whether the same starting position leads to the box or the spot, depending on the text.

Squeezing GR00T N1.7 onto a 3090

NVIDIA's GR00T is the biggest model in LeRobot's lineup that's still worth trying on this hardware. In LeRobot 0.6.0, --policy.type=groot means GR00T N1.7: a model with 3.1 billion parameters, built on a vision-language model called nvidia/Cosmos-Reason2-2B. You need permission to download both, so first request access on their Hub pages and log in with hf auth login. That was the first thing that failed in the first-round table above.

The second failure was memory. By default the big vision-language part is frozen and "only" 1.6 billion parameters are trained. But by default, all the numbers in the model are stored in 32-bit precision (4 bytes each), even the frozen ones. Together with the extra memory training needs, that doesn't fit in 24 GB. Two changes make it fit:

  • --policy.model_params_fp32=false stores the model's numbers in a 16-bit format (bf16, 2 bytes each), which roughly halves the memory the model takes up;
  • a batch of 4 frames instead of 8.

A 50-step test and a 100-step timing run gave the numbers to plan with: 23.2 GB at the peak on a 24 GB card, about 1.64 steps per second, and 15 GB per checkpoint (8.7 GB for the model, 6.1 GB of extra training data). The full run is about 5 epochs, takes about 3 hours, and saves three checkpoints:

OUT=$HOME/.cache/huggingface/lerobot/outputs/train/groot_van-i_r2d2_to_box_bg_$(date +%Y-%m-%d_%H-%M-%S)

~/.local/share/uv/tools/lelab/bin/lerobot-train \
  --policy.type=groot --policy.model_params_fp32=false \
  --dataset.repo_id=van-i/r2d2_to_box_bg_20261003_210444 \
  --batch_size=4 --steps=18000 --save_freq=6000 --log_freq=100 \
  --policy.device=cuda --policy.push_to_hub=false \
  --output_dir=$OUT --job_name=groot_r2d2_bg

Two warnings. First, this is a compromise. NVIDIA's own training recipe keeps 32-bit precision because with 16-bit numbers, very small adjustments can get rounded away to nothing, so the model might stop learning. The loss over the first few thousand steps is the thing to watch. Second, with less than 1 GB of memory to spare, nothing else can use the graphics card while it trains, not even a LeLab camera preview. LeLab's training form can't change the 32-bit setting, so this is a command-line run like SmolVLA. Unlike SmolVLA, GR00T doesn't need its cameras renamed: it uses the dataset's camera names as they are.

Result. The full run finished in 2 h 57 min (~1.69 steps per second) with no errors. The 16-bit worry turned out fine: the loss dropped from 1.24 at step 100 to about 0.13 by step 1,500, then kept falling slowly to 0.035 at 18,000 steps. As with the other policies, that number says training worked, not how well the arm will do. The three checkpoints (6k, 12k, 18k) take 15 GB each.

To run it on the arm, GR00T needs LeRobot's real-time chunking mode, which LeLab switches on automatically for GR00T. scripts/run-groot.sh uses the same settings: ./scripts/run-groot.sh [checkpoint] [seconds] ["task text"].

On the arm, GR00T was the most capable of all four policies: 14 out of 14, including the spot outside the training area. See the results above.

What's next

  • Get past the H2 wall: record more examples from a wider spread of starting spots, and test the new spots again.
  • Make SmolVLA smoother with real-time chunking (--inference.type=rtc).
  • Make SmolVLA and GR00T actually follow instructions: record a second task in the same scene with a different instruction, and train on both.
  • Fix the slow control-loop warnings during rollouts.
  • Try VR control. LeRobot works with NVIDIA's Isaac Teleop (LeRobot guide), which lets a VR controller move the arm instead of the leader arm. You squeeze the grip button to take control (it works like a clutch, so the arm doesn't jump), move and twist the controller to position the gripper, and press the trigger to close the gripper gradually. The headset connects over NVIDIA CloudXR, and the same setup can record LeRobot datasets. For a class, it's an exciting alternative to the leader arm. It's also a good lesson on its own: suddenly you need inverse kinematics (working out the joint angles from where you want the gripper to be), limits on where the arm may go, and safety checks. One catch: the example is written for the SO-101 and comes with the LeRobot source code, not with LeLab, so it needs a separate install and some changes for the SO-100.

Bigger ideas: giving the arm wheels. An arm bolted to a table can only reach so far. The natural next step is putting it on wheels, which turns the stand into a small mobile robot, closer to what physical intelligence looks like outside the lab:

  • LeKiwi: a wheeled base that carries an SO-100/SO-101 arm and works with LeRobot, so the same record → train → run steps still apply.
  • XLeRobot: a step further, a two-armed household robot. It has two SO-101 arms on a wheeled base, is mostly 3D-printed, and the basic version starts at around $660. It's built on LeRobot, so policies trained on a single SO-101 arm can carry over.

What's in this repository

Path What it is
README.md This write-up
images/ Photos and camera views used above (resized; full-size originals are not included)
scripts/camera-setup.sh Re-applies the camera settings (60 Hz anti-flicker, wrist camera brightness), finding each camera by its USB port
scripts/run-smolvla.sh Runs one SmolVLA test on the follower arm and prints control-loop stats
scripts/run-groot.sh Same for GR00T N1.7 (with the real-time chunking settings it needs)
scripts/rollout-common.sh Shared helper for both run scripts: says when the model is loaded and when the arm starts moving
docs/test-plan.md Test protocol and the full results of every try
tools/preview.html Renders this README in a browser, with images (serve the folder, then open tools/preview.html)

The scripts are written for this stand (USB ports, serial port, checkpoint folder), so change those values for your own setup.

References

Software and docs

Hardware

Data

Trained models

Next-step projects

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Datasets used to train van-i/so100-imitation-learning-stand

Papers for van-i/so100-imitation-learning-stand