SmolVLA Fine-tuned on LIBERO-Plus (Object)

This model is a fine-tuned version of lerobot/smolvla_libero specifically trained on the LIBERO-Plus task suite (libero_object). It is designed to perform robotic manipulation tasks by taking visual observations and language instructions as input to predict 7-DoF actions.

πŸ€– Model Details

  • Model Type: Vision-Language-Action (VLA)
  • Base Model: lerobot/smolvla_libero
  • Task Suite: LIBERO-Plus (libero_object)
  • Framework: LeRobot
  • Action Space: 7-DoF (dx, dy, dz, droll, dpitch, dyaw, gripper)
  • Language: English

πŸ“Š Training Information

The model was trained by freezing the vision and language encoders and fine-tuning only the action expert modules and state projection layers.

  • Dataset: lerobot/libero_plus (libero_object)
  • Training Steps: 10,000 steps
  • Batch Size: 8
  • Learning Rate: 1e-4 (Cosine decay with warmup)
  • Hardware: Single NVIDIA L4 GPU

πŸš€ How to Use

You can load and use this model directly with the Hugging Face lerobot library.

Installation

pip install lerobot
Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
F32
Β·
BF16
Β·
Video Preview
loading