Instructions to use TeleEmbodied/SMART-VLA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TeleEmbodied/SMART-VLA with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TeleEmbodied/SMART-VLA", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from TeleEmbodied/SMART-VLA: direct link, hf CLI and curl.
- Browser
- Download file 2.93 kB
-
https://huggingface.co/TeleEmbodied/SMART-VLA/resolve/main/README.md
- Command line
-
hf download hf://TeleEmbodied/SMART-VLA/README.md
-
curl -L -o README.md https://huggingface.co/TeleEmbodied/SMART-VLA/resolve/main/README.md
2.93 kB
| license: mit | |
| base_model: Qwen/Qwen3-VL-4B-Instruct | |
| library_name: transformers | |
| datasets: | |
| - TeleEmbodied/SMART-Data | |
| tags: | |
| - robotics | |
| - vision-language-action | |
| - embodied-ai | |
| - manipulation | |
| - simulation | |
| - arxiv:2610.07652 | |
| # SMART-VLA | |
|  | |
| <p align="center"> | |
| <a href="https://teamillusion-smart.github.io/"><img alt="Project: teamillusion-smart.github.io" src="https://img.shields.io/badge/Project-teamillusion--smart.github.io-2563eb?logo=github&logoColor=white&style=flat-square"></a> <a href="https://arxiv.org/abs/2610.07652"><img alt="arXiv: 2610.07652" src="https://img.shields.io/badge/arXiv-2610.07652-b31b1b?logo=arxiv&logoColor=white&style=flat-square"></a> | |
| </p> | |
| <p align="center"> | |
| <a href="https://huggingface.co/datasets/TeleEmbodied/SMART-Data"><img alt="Dataset: TeleEmbodied/SMART-Data" src="https://img.shields.io/static/v1?label=Dataset&message=TeleEmbodied%2FSMART-Data&color=ffcc4d&logo=huggingface&logoColor=black&style=flat-square"></a> | |
| </p> | |
| SMART-VLA is a PRTS-architecture vision-language-action model pretrained on [SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data) for articulated-object manipulation and sim-to-real transfer. | |
| ## Model Details | |
| - **Base model:** [Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) | |
| - **Architecture:** `PRTS_Qwen3VL` | |
| - **Training data:** [SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data) | |
| - **Parameters:** approximately 4.44B | |
| - **Precision:** `bfloat16` | |
| - **Action chunk size:** 50 | |
| - **Maximum action dimension:** 32 | |
| - **Released checkpoint:** `checkpoint-final-278972` | |
| ## Links | |
| - **Paper:** [SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining](https://arxiv.org/abs/2610.07652) | |
| - **Project page:** [https://teamillusion-smart.github.io/](https://teamillusion-smart.github.io/) | |
| - **Dataset:** [TeleEmbodied/SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data) | |
| - **Code and loading instructions:** [TeleHuman/PRTS](https://github.com/TeleHuman/PRTS) | |
| ## Usage | |
| Please follow the installation, loading, and inference instructions in the [PRTS repository](https://github.com/TeleHuman/PRTS). | |
| ## Intended Use | |
| SMART-VLA is intended for research on vision-language-action pretraining, articulated-object manipulation, and simulation-to-real robot learning. | |
| ## License | |
| This model is released under the MIT License. | |
| ## Citation | |
| ```bibtex | |
| @article{SMART2026ao, | |
| title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining}, | |
| author={Jicong Ao and Shuhan Jiang and Yuling Zhong and Yanwen Liu and Yuhan Gao and Jiangyuan Zhao and Yang Zhang and Shiqiang Zhu and Chenjia Bai and Xuelong Li}, | |
| year={2026}, | |
| eprint={2610.07652}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.RO}, | |
| url={https://arxiv.org/abs/2610.07652}, | |
| } | |
| ``` | |