AI & ML interests

Video understanding, multimodal large language models, reinforcement learning, temporal and spatial grounding, tracking, segmentation, and spatial intelligence.

Recent Activity

lyhisme  updated a model about 9 hours ago
OraRL/Video-ORA-4B
lyhisme  published a model about 9 hours ago
OraRL/Video-ORA-4B
lyhisme  updated a dataset about 15 hours ago
OraRL/OraRL-Data
View all activity

Organization Card

OraRL

[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Data] [💻 Code]

We build OraRL, a reinforcement-learning framework for unified video multimodal large language models.

Our goal is to make broad video RL both reliable and efficient by treating task annotations as positive rollouts while retaining policy samples for the on-policy baseline. The resulting Video-ORA family handles seven task families without chain-of-thought decoding.

▶ OraRL project overview (1:38)

What we release

  • Video-ORA-9B: a unified video MLLM for temporal and spatial grounding, segmentation, tracking, video QA, and spatial intelligence.
  • OraRL-Data: canonical evaluation annotations and raw media for reproducible comparison across the seven task families.
  • OraRL resources: project documentation, model cards, inference examples, and evaluation guidance.

Research interests

  • Unified video understanding
  • Multimodal reinforcement learning
  • Temporal & spatial grounding
  • Video segmentation, tracking & spatial intelligence