TurboVLA β SO-101 left arm, 11-task multitask
TurboVLA trained on the 11 datasets of the
hungho77/SO101_LeftArmOnly_Task
collection, merged into one multitask set. One checkpoint covers all 11 tasks; the
language instruction selects the behaviour.
Trained with the official TurboVLA repository, not a port. The turbovla_so101 policy
that exists as a LeRobot plugin is a different model β it uses dinov2-base + BERT, while
official TurboVLA uses DINOv3 ViT-B with a GroundingDINO-initialised fusion stack.
Results
Open-loop, one trajectory per task, scored over the first 16 predicted steps, errors in the dataset's own action units.
| model | MAE | % of action scale |
|---|---|---|
| TurboVLA (this, EMA) | 1.326 | 4.62% |
| TurboVLA (non-EMA) | 1.609 | 5.61% |
| SmolVLA, same data/samples | 1.928 | 6.72% |
Per trajectory (EMA): 1.534, 1.379, 1.371, 1.298, 1.147, 1.026, 1.742, 1.692, 0.911, 1.050, 1.434. Normalized by the absolute action std of 28.67.
The EMA weights are clearly better and are what this repo ships.
Robot
SO-101 follower arm, left arm only, 5 joints plus gripper, 30 fps.
| State / action dim | 6 β shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll + gripper |
| Cameras | top (overhead, 360Γ640), wrist (480Γ640), both resized to 224Γ224 |
| Action horizon | 16 |
| Backbones | DINOv3 ViT-B (trained), BERT base (frozen), GroundingDINO init |
All six dimensions are continuous
so101_data_config.py normalizes every dimension with min_max, including the gripper.
This is the one place where copying the upstream RoboTwin recipe is wrong: RoboTwin
binarizes its gripper, and a first version of this model did too. Measured over 15,466
frames the SO-101 gripper spans 0.70 to 52.70 with std 17.41, spread across the whole
range rather than piled at two ends β it is a continuous aperture, not a switch.
Training it as binary made the model emit 0/1 against ground truth running to 52, which on its own took the open-loop error from roughly 2.5% to 14.5%. That run is not published.
eval_so101_open_loop.py carries the same assumption in CONTINUOUS_MASK, and it must
match. The mask is applied to the state going in as well as the actions coming back, so a
stale value corrupts the model's input, not just its output β a gripper-only mismatch
degraded every joint and made a correct checkpoint read as 24.87%.
Training
| Data | 551 episodes / 253,745 frames / 11 tasks |
| Steps Γ batch | 20,000 Γ 128 (64 per device Γ 2) = 2.56M samples = 10.1 epochs |
| Hardware | 2 Γ H100 MIG 3g.40gb slices, DeepSpeed ZeRO-2 |
| Runtime | 4h33m at 1.22 it/s |
| LR | 5e-5, warmup 1,000, warmup_constant_with_factor |
| EMA | 0.999 |
| Attention | sdpa β flash-attn 2.7.4 does not build in this environment |
The sample budget matches the SmolVLA run exactly, which is what makes the two comparable.
Files
| file | purpose |
|---|---|
checkpoints/steps_20000_ema_pytorch_model.pt |
EMA weights |
config.yaml, config.full.yaml |
the training configuration as run |
dataset_statistics.json |
required β min/max used to un-normalize predictions |
so101_data_config.py |
drop into experiments/so101/data_registry/ |
so101_leftarm.yaml |
drop into experiments/so101/configs/ |
modality.json |
goes in the dataset's meta/ |
eval_so101_open_loop.py |
the evaluation used above |
predict_action returns normalized actions; un-normalize with dataset_statistics.json
before sending anything to a robot.
Limitations
No held-out split. All 551 episodes were used for training, so the numbers above are measured on training data. They show the policy fits its data and is not degenerate; they are not a success rate, and real-robot performance is unmeasured. A sibling SmolVLA checkpoint with a comparable open-loop figure scored 0% on the real robot, which is exactly the gap this caveat is about.
Deployment note
The data is 30 fps: smooth_step = control_hz / 30. The action horizon is 16, about
0.53 s; executing a full chunk open-loop before re-planning is already a long time without
feedback, and executing more is worse.
- Downloads last month
- -
Model tree for twanghcmut/TurboVLA-SO101-LeftArm-Multitask
Base model
H-EmbodVis/TurboVLA