SmolVLA Fine-tuned on LIBERO-Object

This model is a fine-tuned version of lerobot/smolvla_libero specifically trained on the LIBERO-Object task suite. It is designed to perform robotic manipulation tasks by taking visual observations and language instructions as input to predict 7-DoF actions.

πŸ€– Model Details

  • Model Type: Vision-Language-Action (VLA)
  • Base Model: lerobot/smolvla_libero (which is based on SmolVLM2)
  • Task Suite: LIBERO-Object
  • Framework: LeRobot
  • Action Space: 7-DoF (dx, dy, dz, droll, dpitch, dyaw, gripper)
  • Language: English (Task Instructions)

πŸ“Š Training Information

The model was trained by freezing the vision and language encoders and fine-tuning only the action expert modules and state projection layers.

  • Dataset: lerobot/libero (filtered for LIBERO-Object tasks)
  • Training Steps: 10,000 steps
  • Batch Size: 8 (per device)
  • Learning Rate: 1e-4 (Cosine decay with warmup)
  • Hardware: Single NVIDIA T4 GPU
  • Training Time: ~4 hours 50 minutes

πŸš€ How to Use

You can load and use this model directly with the Hugging Face lerobot library.

Installation

pip install lerobot
Downloads last month
20
Safetensors
Model size
0.5B params
Tensor type
F32
Β·
BF16
Β·
Video Preview
loading