RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Paper • 2607.26991 • Published • 4
For more information on usage, please refer to the RL2-VLA Github repository here.
@article{tan2026rl2,
title = {RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models},
author = {Derek Ming Siang Tan and Shailesh Shailesh and Srikrishna Iyer and William Wei Jie Teo and Yuanliang Ju and Qiao Gu and Guillaume Sartoretti},
year = {2026},
journal = {arXiv preprint arXiv:2607.26991}
}