Instructions to use ServoVLA/ServoVLA-LIBERO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use ServoVLA/ServoVLA-LIBERO with LeRobot:
- Notebooks
- Google Colab
- Kaggle
ServoVLA LIBERO
ServoVLA policy for LIBERO-40 simulation. This release contains the EMA policy head from training step 1,190,000 and the action normalization statistics saved with that run.
Model Contract
- Benchmark: LIBERO-40 (
libero_10,libero_spatial,libero_object,libero_goal) - Action space: 7-dimensional direct environment action
- State space: 8-dimensional LIBERO proprioception
- Cameras:
observation.images.image,observation.images.image2 - Action chunk: 16 steps
- Inference steps: 3
- Vision encoder:
facebook/dinov3-vitb16-pretrain-lvd1689m - Language encoder:
Qwen/Qwen3.5-0.8B - Weights: EMA (
used_ema=true) - Action normalization: enabled, per-horizon dimension, 16 x 7 statistics
Evaluation
The selected checkpoint was evaluated on the fixed LIBERO-40 four-suite protocol with 30 episodes per task, 1,200 episodes total:
- Overall balanced success rate: 78.0%
libero_10: 79.0%libero_spatial: 75.0%libero_object: 84.33%libero_goal: 73.67%
Three additional 50-episode-per-task evaluations with different seeds produced overall balanced success rates of 75.6%, 79.15%, and 69.25% (mean 75.95%). Rollouts used relative LIBERO environment control, open-loop 16-step chunks, and observed maximum frame delay 0 in the reported runs.
Files
model.safetensors: LeRobot-compatible EMA policy weights.config.json: ServoVLA policy configuration.policy_preprocessor.jsonandpolicy_postprocessor.json: LeRobot processor configuration.checkpoint.pt: cleaned ServoVLA checkpoint containingmodel_state,ema_state,step, andaction_normalizationonly.
The checkpoint is weights-only for inference and fresh initialization. Optimizer, scheduler, W&B, and resume-training state are intentionally absent.
Usage
Use this directory with the ServoVLA LIBERO evaluation entry points. The LIBERO camera keys, state ordering, action dimensions, and direct-action semantics must match the model contract above. This policy is not compatible with SO101 robot observations or the SO101 delta-action contract.
License
ServoVLA is released under the MIT License. The vision and language encoders and the LIBERO benchmark assets are subject to their respective upstream licenses and terms.
- Downloads last month
- 13