VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Abstract
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Community
Can your voice assistant remember who told it something, not just what was said?
Mostly, no. We built VoxMem, a benchmark where the answer to a question hides in a multi-session spoken history. Sometimes it's in the words, but often it's in the speaker's voice, their tone, or a sound in the background. Across 15 LALMs, none exceeds 40% accuracy at 32K. Proprietary models reach 55.6% on what was said, but only 32.7% on who said it, 20.0% on how it was said, and 21.9% on what was audible.
The gap is not about transcription. When we replace the audio with exact transcripts, speaker accuracy falls from 69.8% to 10.3%, while speech-semantics accuracy barely moves (75.9% → 71.0%). Three of our four evidence types exist only in the audio.
Memory operations make it sharper. Models track how a stated fact changes across sessions reasonably well (44.5%). Tracking how someone's vocal state changes collapses to 3.4%, and a changing background sound to 1.2%.
Because every question and its evidence stay fixed from 8K to 64K, the decline with length comes from the growing history, not from harder questions. The failures also differ in kind: speaker errors are mostly binding failures (right fact, wrong person), while paralinguistic errors are mostly localization failures (the cue is never retrieved at all).
The punchline: longer context windows do not give you spoken memory. A "transcribe first, remember later" system is blind to most of what makes a conversation spoken.
📄 3,196 instances · 34,743 sessions · 177 h of audio · 4 evidence types × 4 memory operations · 8K–64K
🌐 https://swagshaw.github.io/voxmem/ · 💻 https://github.com/swagshaw/voxmem · 🤗 https://huggingface.co/datasets/AudioMemory/voxmembench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models (2026)
- MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing (2026)
- UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory (2026)
- When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue (2026)
- VoiceLongMemEval: Do Assistants Remember How You Sounded? (2026)
- Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark (2026)
- HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32607 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper