CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Abstract
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Community
We introduce CoEvoWhen, a policy-tool coevolution framework for ultra-long video temporal grounding. Our key idea is to jointly evolve high-level policies and executable media tools from agentic reasoning trajectories, forming a reusable skill without updating VLM parameters. At inference, the frozen VLM uses the evolved skill to coordinate long-range image-based search and fine-grained video-based observations, without relying on a separate, stronger planning model. Experiments across five benchmarks and three VLMs demonstrate improved grounding accuracy and reduced visual token costs, as well as transfer to general long-video QA without additional task-specific evolution.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics (2026)
- Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding (2026)
- AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research (2026)
- VideoResearcher: Self-Improving Tool Design for Long-Video Understanding (2026)
- SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning (2026)
- MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding (2026)
- Online Video Agent Harness for Long Video Understanding (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper