APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
Abstract
To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.
Community
Streaming video models increasingly support continuous perception, real-time interaction, and proactive assistance, showing their potential as personal assistants; memory is key to making such assistants truly personal by retaining and reusing user-specific experience over time. However, most existing benchmarks and methods study memory within a single continuous video, typically over a limited time span. In the real world, interactions are intermittent: users may turn off smart glasses and resume using the assistant hours or days later. The assistant must therefore retain and use relevant visual evidence from earlier interactions to answer later questions and provide proactive assistance.
This calls for persistent memory: a storable record of past experience that remains available after an interaction ends and can be reused in later interactions. For real-world assistants, such memory must support later tasks while keeping storage and response latency manageable.
Get this paper in your agent:
hf papers read 2609.37559 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper