robotwin_dual_system_joint_self_attention

An OpenWAM-Study checkpoint on RoboTwin 2.0, and the common reference the controlled comparisons vary against — so the same checkpoint takes part in several of them.

Its configuration is Dual-System / Joint Self-Attention — the hidden states of the video and action streams merge into a joint sequence at designated bridge layers, allowing bidirectional token-level interaction before the states return to their own streams — with a Wan2.2-TI2V-5B video backbone, its native Wan-VAE encoder, and the action-sees-video attention mask.

The same checkpoint therefore serves as the reference point in four comparisons at once: architecture (Q1), video backbone (Q2), visual encoder (Q3) and interaction mask (Q4). Every other robotwin_* checkpoint changes exactly one of these axes relative to it. Note that OpenWAM-α keeps this architecture but adopts the mutual mask instead.

Citation

@article{wang2026openwam,
  title   = {OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining},
  author  = {Yuran Wang and Siqiao Huang and Mingleyang Li and Chenhao Zhang and Jiaqi Liang and Weiyang Jin and Yue Chen and Xuemin Chi and Donghao Zhou and Qize Yu and Yu-Kai Wang and Yuhan Rui and Shenzhe Yao and Zhen Yuan and Zhenhao Shen and Kefei Zhu and Zijie Zhu and Ning Gao and Xiaowei Chi and Guanqi He and Shanghang Zhang and Hao Dong and Lin Shao and Hang Zhao},
  year    = {2026},
  journal = {arXiv preprint arXiv: 2609.07398}
}
Downloads last month
16
Video Preview
loading

Collections including OpenWAM/robotwin_dual_system_joint_self_attention

Paper for OpenWAM/robotwin_dual_system_joint_self_attention