OraRL/Video-ORA-9B
Image-Text-to-Text • 9B • Updated • 26 • 1
Video understanding, multimodal large language models, reinforcement learning, temporal and spatial grounding, tracking, segmentation, and spatial intelligence.
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Data] [💻 Code]
We build OraRL, a reinforcement-learning framework for unified video multimodal large language models.
Our goal is to make broad video RL both reliable and efficient by treating task annotations as positive rollouts while retaining policy samples for the on-policy baseline. The resulting Video-ORA family handles seven task families without chain-of-thought decoding.
▶ OraRL project overview (1:38)