Robotics
LeRobot
Safetensors
smolvla
so101
imitation-learning
multitask

SmolVLA — SO-101 left arm, 11-task multitask

SmolVLA finetuned on the 11 datasets of the hungho77/SO101_LeftArmOnly_Task collection, merged into one multitask set. One checkpoint covers all 11 tasks; the language instruction selects the behaviour.

Tasks

Pass one of these strings verbatim as the task; the model has not been trained on paraphrases.

task_index instruction episodes
0 Pick up the red cube and place it into the steel pot. 50
1 Pick up the blue cube and place it into the steel pot 50
2 Take the banana and place it into the steel pot then close the lid. 50
3 Pick up the blue cube and place it next to the blue cube. 50
4 Take the blue cube out of the steel pot and place it on the blue cube. 51
5 Take the red cube out of the steel pot and place it on the red cube 50
6 Stack the yellow cube on top of the orange cube. 50
7 Pick up one orange block and stack it neatly on top of the other orange block 50
8 Pick up the cube that makes the number of red and blue cubes on the wooden board equal, and place it in the empty corner of the wooden board. 50
9 Pick up the red cube from the wooden board, put it into the orange hollow cube, then pick up the orange hollow cube and place it on the wooden board 50
10 Take out the blue cubes from shortest to tallest and place it on the table 50

Note the inconsistent trailing periods: they are part of the strings as recorded and must be reproduced exactly. These are the exact strings from meta/tasks.jsonl.

Robot

SO-101 follower arm, left arm only, 5 joints plus gripper, 30 fps.

State / action dim 6 — shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll + gripper
Cameras top (overhead, 360×640) → camera1, wrist (480×640) → camera2
Action chunk 50
Absolute action std 28.67 (mean over dims)

Camera renaming is required. smolvla_base declares camera1/camera2/camera3 while the dataset names its cameras top and wrist. Training used:

--rename_map='{"observation.images.top":"observation.images.camera1",
               "observation.images.wrist":"observation.images.camera2"}'

Without it, lerobot stops with Feature mismatch between dataset/environment and policy config. Inference must apply the same mapping.

Training

Base lerobot/smolvla_base
Data 551 episodes / 253,745 frames / 11 tasks
Steps × batch 20,000 × 128 = 2.56M samples = 10.1 epochs
Learnable params 100M of 450M — freeze_vision_encoder and train_expert_only are on
LR 1e-4, lerobot default schedule
Hardware one H100 MIG 3g.40gb slice (39.5 GiB), peak 27.3 GB
Runtime 4.6 h at 0.835 s/step, 153 samples/s
Final loss 0.050

Freezing the VLM is what makes this fit a 40 GB slice at batch 128 where 3B models cannot: no activations are stored for the frozen tower, so the cost is 0.19 GB per sample against 0.82 GB for X-VLA and ~0.36 GB for π₀.₅.

Open-loop evaluation

One trajectory per task, walked in chunk-length strides, scored over the first 16 of the 50 predicted steps. Errors are in the dataset's own action units.

traj task MAE
0 red cube → pot 1.592
50 blue cube → pot 1.369
100 banana → pot, close lid 1.905
150 blue cube next to blue cube 2.430
200 blue cube out of pot 1.854
251 red cube out of pot 1.784
301 stack yellow on orange 1.950
351 stack orange blocks 2.580
401 balance red/blue on board 1.881
451 red cube → hollow cube 1.834
501 blue cubes shortest→tallest 2.029
average 1.928

Dividing by the absolute action std (28.67) gives 6.72% of action scale.

Error grows along a chunk, so a horizon-50 number must never be compared against a horizon-16 one from another model. Use the absolute action std, not the checkpoint's normalizer statistics, which are a different quantity.

Limitations

No held-out split. All 551 episodes were used for training, so the numbers above are measured on training data. They show the policy fits its data and is not degenerate; they do not measure generalization and are not a success rate. Real-robot performance is unmeasured.

Scoring only three trajectories gives 5.17% — the first three tasks happen to be the easiest. The 11-trajectory figure of 6.72% is the one to quote.

Deployment note

The data is 30 fps. A client interpolating between model steps must derive its sub-step count from that: smooth_step = control_hz / 30. Reusing a value from a 15 fps robot stretches every trajectory by 2× and the arm creeps without finishing the task.

Downloads last month
13
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for twanghcmut/SmolVLA-SO101-LeftArm-Multitask

Dataset used to train twanghcmut/SmolVLA-SO101-LeftArm-Multitask