Instructions to use TeleEmbodied/SMART-VLA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TeleEmbodied/SMART-VLA with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TeleEmbodied/SMART-VLA", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from TeleEmbodied/SMART-VLA: direct link, hf CLI and curl.
- Browser
- Download file 2.93 kB
-
https://huggingface.co/TeleEmbodied/SMART-VLA/resolve/main/README.md
- Command line
-
hf download hf://TeleEmbodied/SMART-VLA/README.md
-
curl -L -o README.md https://huggingface.co/TeleEmbodied/SMART-VLA/resolve/main/README.md
2.93 kB
metadata
license: mit
base_model: Qwen/Qwen3-VL-4B-Instruct
library_name: transformers
datasets:
- TeleEmbodied/SMART-Data
tags:
- robotics
- vision-language-action
- embodied-ai
- manipulation
- simulation
- arxiv:2610.07652
SMART-VLA
SMART-VLA is a PRTS-architecture vision-language-action model pretrained on SMART-Data for articulated-object manipulation and sim-to-real transfer.
Model Details
- Base model: Qwen/Qwen3-VL-4B-Instruct
- Architecture:
PRTS_Qwen3VL - Training data: SMART-Data
- Parameters: approximately 4.44B
- Precision:
bfloat16 - Action chunk size: 50
- Maximum action dimension: 32
- Released checkpoint:
checkpoint-final-278972
Links
- Paper: SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
- Project page: https://teamillusion-smart.github.io/
- Dataset: TeleEmbodied/SMART-Data
- Code and loading instructions: TeleHuman/PRTS
Usage
Please follow the installation, loading, and inference instructions in the PRTS repository.
Intended Use
SMART-VLA is intended for research on vision-language-action pretraining, articulated-object manipulation, and simulation-to-real robot learning.
License
This model is released under the MIT License.
Citation
@article{SMART2026ao,
title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining},
author={Jicong Ao and Shuhan Jiang and Yuling Zhong and Yanwen Liu and Yuhan Gao and Jiangyuan Zhao and Yang Zhang and Shiqiang Zhu and Chenjia Bai and Xuelong Li},
year={2026},
eprint={2610.07652},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2610.07652},
}
