MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Abstract
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
Community
We introduced MemLife, an agentic memory system for long-term egocentric video that constructs time- and entity-anchored, first-person episodes and accesses them through time-scoped retrieval and chronological evidence organization. We further proposed MemOpt, which applies reinforcement learning only to the memory writer through the FIRM objective for faithful, informative, and retrievable memories. MemLife outperforms state-of-the-art training-free systems, while MemOpt provides further gains that transfer across writer and reader backbones, memory systems, and out-of-domain video and question distributions. Together, the results demonstrate that learning what to remember is an effective and generalizable approach to long-term video question answering.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EM^2Mem: Event-Centric Multimodal Memory for Large Language Models (2026)
- REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering (2026)
- CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video (2026)
- Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning (2026)
- LeanMem: Simple and Efficient Long-Term Memory for LLM Agents (2026)
- EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory (2026)
- Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.40195 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper