Instructions to use twanghcmut/SmolVLA-SO101-LeftArm-Multitask with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use twanghcmut/SmolVLA-SO101-LeftArm-Multitask with LeRobot:
# See https://github.com/huggingface/lerobot?tab=readme-ov-file#installation for more details git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e .[smolvla]
# Launch finetuning on your dataset python lerobot/scripts/train.py \ --policy.path=twanghcmut/SmolVLA-SO101-LeftArm-Multitask \ --dataset.repo_id=lerobot/svla_so101_pickplace \ --batch_size=64 \ --steps=20000 \ --output_dir=outputs/train/my_smolvla \ --job_name=my_smolvla_training \ --policy.device=cuda \ --wandb.enable=true
# Run the policy using the record function python -m lerobot.record \ --robot.type=so101_follower \ --robot.port=/dev/ttyACM0 \ # <- Use your port --robot.id=my_blue_follower_arm \ # <- Use your robot id --robot.cameras="{ front: {type: opencv, index_or_path: 8, width: 640, height: 480, fps: 30}}" \ # <- Use your cameras --dataset.single_task="Grasp a lego block and put it in the bin." \ # <- Use the same task description you used in your dataset recording --dataset.repo_id=HF_USER/dataset_name \ # <- This will be the dataset name on HF Hub --dataset.episode_time_s=50 \ --dataset.num_episodes=10 \ --policy.path=twanghcmut/SmolVLA-SO101-LeftArm-Multitask - Notebooks
- Google Colab
- Kaggle
SmolVLA — SO-101 left arm, 11-task multitask
SmolVLA finetuned on the 11 datasets of the
hungho77/SO101_LeftArmOnly_Task
collection, merged into one multitask set. One checkpoint covers all 11 tasks; the
language instruction selects the behaviour.
Tasks
Pass one of these strings verbatim as the task; the model has not been trained on paraphrases.
task_index |
instruction | episodes |
|---|---|---|
| 0 | Pick up the red cube and place it into the steel pot. |
50 |
| 1 | Pick up the blue cube and place it into the steel pot |
50 |
| 2 | Take the banana and place it into the steel pot then close the lid. |
50 |
| 3 | Pick up the blue cube and place it next to the blue cube. |
50 |
| 4 | Take the blue cube out of the steel pot and place it on the blue cube. |
51 |
| 5 | Take the red cube out of the steel pot and place it on the red cube |
50 |
| 6 | Stack the yellow cube on top of the orange cube. |
50 |
| 7 | Pick up one orange block and stack it neatly on top of the other orange block |
50 |
| 8 | Pick up the cube that makes the number of red and blue cubes on the wooden board equal, and place it in the empty corner of the wooden board. |
50 |
| 9 | Pick up the red cube from the wooden board, put it into the orange hollow cube, then pick up the orange hollow cube and place it on the wooden board |
50 |
| 10 | Take out the blue cubes from shortest to tallest and place it on the table |
50 |
Note the inconsistent trailing periods: they are part of the strings as recorded and must
be reproduced exactly. These are the exact strings from meta/tasks.jsonl.
Robot
SO-101 follower arm, left arm only, 5 joints plus gripper, 30 fps.
| State / action dim | 6 — shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll + gripper |
| Cameras | top (overhead, 360×640) → camera1, wrist (480×640) → camera2 |
| Action chunk | 50 |
| Absolute action std | 28.67 (mean over dims) |
Camera renaming is required. smolvla_base declares camera1/camera2/camera3
while the dataset names its cameras top and wrist. Training used:
--rename_map='{"observation.images.top":"observation.images.camera1",
"observation.images.wrist":"observation.images.camera2"}'
Without it, lerobot stops with Feature mismatch between dataset/environment and policy config. Inference must apply the same mapping.
Training
| Base | lerobot/smolvla_base |
| Data | 551 episodes / 253,745 frames / 11 tasks |
| Steps × batch | 20,000 × 128 = 2.56M samples = 10.1 epochs |
| Learnable params | 100M of 450M — freeze_vision_encoder and train_expert_only are on |
| LR | 1e-4, lerobot default schedule |
| Hardware | one H100 MIG 3g.40gb slice (39.5 GiB), peak 27.3 GB |
| Runtime | 4.6 h at 0.835 s/step, 153 samples/s |
| Final loss | 0.050 |
Freezing the VLM is what makes this fit a 40 GB slice at batch 128 where 3B models cannot: no activations are stored for the frozen tower, so the cost is 0.19 GB per sample against 0.82 GB for X-VLA and ~0.36 GB for π₀.₅.
Open-loop evaluation
One trajectory per task, walked in chunk-length strides, scored over the first 16 of the 50 predicted steps. Errors are in the dataset's own action units.
| traj | task | MAE |
|---|---|---|
| 0 | red cube → pot | 1.592 |
| 50 | blue cube → pot | 1.369 |
| 100 | banana → pot, close lid | 1.905 |
| 150 | blue cube next to blue cube | 2.430 |
| 200 | blue cube out of pot | 1.854 |
| 251 | red cube out of pot | 1.784 |
| 301 | stack yellow on orange | 1.950 |
| 351 | stack orange blocks | 2.580 |
| 401 | balance red/blue on board | 1.881 |
| 451 | red cube → hollow cube | 1.834 |
| 501 | blue cubes shortest→tallest | 2.029 |
| average | 1.928 |
Dividing by the absolute action std (28.67) gives 6.72% of action scale.
Error grows along a chunk, so a horizon-50 number must never be compared against a horizon-16 one from another model. Use the absolute action std, not the checkpoint's normalizer statistics, which are a different quantity.
Limitations
No held-out split. All 551 episodes were used for training, so the numbers above are measured on training data. They show the policy fits its data and is not degenerate; they do not measure generalization and are not a success rate. Real-robot performance is unmeasured.
Scoring only three trajectories gives 5.17% — the first three tasks happen to be the easiest. The 11-trajectory figure of 6.72% is the one to quote.
Deployment note
The data is 30 fps. A client interpolating between model steps must derive its
sub-step count from that: smooth_step = control_hz / 30. Reusing a value from a 15 fps
robot stretches every trajectory by 2× and the arm creeps without finishing the task.
- Downloads last month
- 13
Model tree for twanghcmut/SmolVLA-SO101-LeftArm-Multitask
Base model
HuggingFaceTB/SmolLM2-360M