File size: 2,925 Bytes
3cae1a1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22f92c6
 
f1e544f
 
 
 
800c13f
 
 
 
3cae1a1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
license: mit
base_model: Qwen/Qwen3-VL-4B-Instruct
library_name: transformers
datasets:
  - TeleEmbodied/SMART-Data
tags:
  - robotics
  - vision-language-action
  - embodied-ai
  - manipulation
  - simulation
  - arxiv:2610.07652
---

# SMART-VLA

![SMART-Data overview](assets/slide4-4k.png)

<p align="center">
  <a href="https://teamillusion-smart.github.io/"><img alt="Project: teamillusion-smart.github.io" src="https://img.shields.io/badge/Project-teamillusion--smart.github.io-2563eb?logo=github&logoColor=white&style=flat-square"></a> <a href="https://arxiv.org/abs/2610.07652"><img alt="arXiv: 2610.07652" src="https://img.shields.io/badge/arXiv-2610.07652-b31b1b?logo=arxiv&logoColor=white&style=flat-square"></a>
</p>

<p align="center">
  <a href="https://huggingface.co/datasets/TeleEmbodied/SMART-Data"><img alt="Dataset: TeleEmbodied/SMART-Data" src="https://img.shields.io/static/v1?label=Dataset&message=TeleEmbodied%2FSMART-Data&color=ffcc4d&logo=huggingface&logoColor=black&style=flat-square"></a>
</p>

SMART-VLA is a PRTS-architecture vision-language-action model pretrained on [SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data) for articulated-object manipulation and sim-to-real transfer.

## Model Details

- **Base model:** [Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct)
- **Architecture:** `PRTS_Qwen3VL`
- **Training data:** [SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data)
- **Parameters:** approximately 4.44B
- **Precision:** `bfloat16`
- **Action chunk size:** 50
- **Maximum action dimension:** 32
- **Released checkpoint:** `checkpoint-final-278972`

## Links

- **Paper:** [SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining](https://arxiv.org/abs/2610.07652)
- **Project page:** [https://teamillusion-smart.github.io/](https://teamillusion-smart.github.io/)
- **Dataset:** [TeleEmbodied/SMART-Data](https://huggingface.co/datasets/TeleEmbodied/SMART-Data)
- **Code and loading instructions:** [TeleHuman/PRTS](https://github.com/TeleHuman/PRTS)

## Usage

Please follow the installation, loading, and inference instructions in the [PRTS repository](https://github.com/TeleHuman/PRTS).

## Intended Use

SMART-VLA is intended for research on vision-language-action pretraining, articulated-object manipulation, and simulation-to-real robot learning.

## License

This model is released under the MIT License.

## Citation

```bibtex
@article{SMART2026ao,
  title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining},
  author={Jicong Ao and Shuhan Jiang and Yuling Zhong and Yanwen Liu and Yuhan Gao and Jiangyuan Zhao and Yang Zhang and Shiqiang Zhu and Chenjia Bai and Xuelong Li},
  year={2026},
  eprint={2610.07652},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.07652},
}
```