SMART-VLA / README.md
BraveBlackbird's picture
Add SMART-Data dataset badge
800c13f verified
|
Raw History Blame Contribute Delete
2.93 kB
metadata
license: mit
base_model: Qwen/Qwen3-VL-4B-Instruct
library_name: transformers
datasets:
  - TeleEmbodied/SMART-Data
tags:
  - robotics
  - vision-language-action
  - embodied-ai
  - manipulation
  - simulation
  - arxiv:2610.07652

SMART-VLA

SMART-Data overview

Project: teamillusion-smart.github.io arXiv: 2610.07652

Dataset: TeleEmbodied/SMART-Data

SMART-VLA is a PRTS-architecture vision-language-action model pretrained on SMART-Data for articulated-object manipulation and sim-to-real transfer.

Model Details

  • Base model: Qwen/Qwen3-VL-4B-Instruct
  • Architecture: PRTS_Qwen3VL
  • Training data: SMART-Data
  • Parameters: approximately 4.44B
  • Precision: bfloat16
  • Action chunk size: 50
  • Maximum action dimension: 32
  • Released checkpoint: checkpoint-final-278972

Links

Usage

Please follow the installation, loading, and inference instructions in the PRTS repository.

Intended Use

SMART-VLA is intended for research on vision-language-action pretraining, articulated-object manipulation, and simulation-to-real robot learning.

License

This model is released under the MIT License.

Citation

@article{SMART2026ao,
  title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining},
  author={Jicong Ao and Shuhan Jiang and Yuling Zhong and Yanwen Liu and Yuhan Gao and Jiangyuan Zhao and Yang Zhang and Shiqiang Zhu and Chenjia Bai and Xuelong Li},
  year={2026},
  eprint={2610.07652},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.07652},
}