TurboVLA β€” SO-101 left arm, 11-task multitask

TurboVLA trained on the 11 datasets of the hungho77/SO101_LeftArmOnly_Task collection, merged into one multitask set. One checkpoint covers all 11 tasks; the language instruction selects the behaviour.

Trained with the official TurboVLA repository, not a port. The turbovla_so101 policy that exists as a LeRobot plugin is a different model β€” it uses dinov2-base + BERT, while official TurboVLA uses DINOv3 ViT-B with a GroundingDINO-initialised fusion stack.

Results

Open-loop, one trajectory per task, scored over the first 16 predicted steps, errors in the dataset's own action units.

model MAE % of action scale
TurboVLA (this, EMA) 1.326 4.62%
TurboVLA (non-EMA) 1.609 5.61%
SmolVLA, same data/samples 1.928 6.72%

Per trajectory (EMA): 1.534, 1.379, 1.371, 1.298, 1.147, 1.026, 1.742, 1.692, 0.911, 1.050, 1.434. Normalized by the absolute action std of 28.67.

The EMA weights are clearly better and are what this repo ships.

Robot

SO-101 follower arm, left arm only, 5 joints plus gripper, 30 fps.

State / action dim 6 β€” shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll + gripper
Cameras top (overhead, 360Γ—640), wrist (480Γ—640), both resized to 224Γ—224
Action horizon 16
Backbones DINOv3 ViT-B (trained), BERT base (frozen), GroundingDINO init

All six dimensions are continuous

so101_data_config.py normalizes every dimension with min_max, including the gripper. This is the one place where copying the upstream RoboTwin recipe is wrong: RoboTwin binarizes its gripper, and a first version of this model did too. Measured over 15,466 frames the SO-101 gripper spans 0.70 to 52.70 with std 17.41, spread across the whole range rather than piled at two ends β€” it is a continuous aperture, not a switch.

Training it as binary made the model emit 0/1 against ground truth running to 52, which on its own took the open-loop error from roughly 2.5% to 14.5%. That run is not published.

eval_so101_open_loop.py carries the same assumption in CONTINUOUS_MASK, and it must match. The mask is applied to the state going in as well as the actions coming back, so a stale value corrupts the model's input, not just its output β€” a gripper-only mismatch degraded every joint and made a correct checkpoint read as 24.87%.

Training

Data 551 episodes / 253,745 frames / 11 tasks
Steps Γ— batch 20,000 Γ— 128 (64 per device Γ— 2) = 2.56M samples = 10.1 epochs
Hardware 2 Γ— H100 MIG 3g.40gb slices, DeepSpeed ZeRO-2
Runtime 4h33m at 1.22 it/s
LR 5e-5, warmup 1,000, warmup_constant_with_factor
EMA 0.999
Attention sdpa β€” flash-attn 2.7.4 does not build in this environment

The sample budget matches the SmolVLA run exactly, which is what makes the two comparable.

Files

file purpose
checkpoints/steps_20000_ema_pytorch_model.pt EMA weights
config.yaml, config.full.yaml the training configuration as run
dataset_statistics.json required β€” min/max used to un-normalize predictions
so101_data_config.py drop into experiments/so101/data_registry/
so101_leftarm.yaml drop into experiments/so101/configs/
modality.json goes in the dataset's meta/
eval_so101_open_loop.py the evaluation used above

predict_action returns normalized actions; un-normalize with dataset_statistics.json before sending anything to a robot.

Limitations

No held-out split. All 551 episodes were used for training, so the numbers above are measured on training data. They show the policy fits its data and is not degenerate; they are not a success rate, and real-robot performance is unmeasured. A sibling SmolVLA checkpoint with a comparable open-loop figure scored 0% on the real robot, which is exactly the gap this caveat is about.

Deployment note

The data is 30 fps: smooth_step = control_hz / 30. The action horizon is 16, about 0.53 s; executing a full chunk open-loop before re-planning is already a long time without feedback, and executing more is worse.

Downloads last month
-
Video Preview
loading

Model tree for twanghcmut/TurboVLA-SO101-LeftArm-Multitask

Finetuned
(3)
this model