TeleEmbodied/SMART-Data
Viewer • Updated • 16.1k • 19
How to use TeleEmbodied/SMART-VLA with Transformers:
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("TeleEmbodied/SMART-VLA", device_map="auto")SMART-VLA is a PRTS-architecture vision-language-action model pretrained on SMART-Data for articulated-object manipulation and sim-to-real transfer.
PRTS_Qwen3VLbfloat16checkpoint-final-278972Please follow the installation, loading, and inference instructions in the PRTS repository.
SMART-VLA is intended for research on vision-language-action pretraining, articulated-object manipulation, and simulation-to-real robot learning.
This model is released under the MIT License.
@article{SMART2026ao,
title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining},
author={Jicong Ao and Shuhan Jiang and Yuling Zhong and Yanwen Liu and Yuhan Gao and Jiangyuan Zhao and Yang Zhang and Shiqiang Zhu and Chenjia Bai and Xuelong Li},
year={2026},
eprint={2610.07652},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2610.07652},
}
Base model
Qwen/Qwen3-VL-4B-Instruct