OmniMemBench: Towards Scalable Evaluation of Long-Term Omni-Modal Agent Memory
Abstract
Multimodal agents that interact with users over days or months must remember not only what was said, but what was shown, heard, and played. Current memory benchmarks are predominantly text-only or cover at most two modalities. Those that do incorporate multimodal content still evaluate with aggregate scores that conflate retrieval and reasoning errors, and provide no mechanism to test whether performance degrades as context scales, leaving our understanding of agent memory fundamentally incomplete. We introduce OmniMemBench, to our knowledge the first long-term memory benchmark that spans image, audio, and video in multi-session dialogue. Its design is guided by three requirements for reliable evaluation. First, an anti-leakage mechanism filters questions answerable from text alone, reducing text-only solvability to a negligible level. Second, two LLM-judged retrieval metrics—Clue Coverage and Entry Relevance—operate on content rather than entry IDs, enabling unified retrieval diagnosis across RAG and memory paradigms. The gap between retrieval coverage and downstream QA scores further localizes failures to the memorization stage. Third, content-isolated distractor injection scales context from 128K to 1M tokens while holding evidence constant, so scalability can be measured without confounding difficulty. The benchmark covers 103 characters and 6,655 QAs across 8 task types and 4 context tiers. Evaluating representative long-context, RAG, and structured-memory methods, we find that no single paradigm dominates: long-context models lead overall, yet RAG surpasses them on targeted fact extraction. Further analysis shows that memory methods retrieve clues at coverage rates comparable to RAG but score substantially lower on QA, revealing that the bottleneck is not what is retrieved but what is lost during memorization. Caption-based processing irreversibly discards perceptual details; raw embeddings degrade retrieval accuracy; and scaling the memorization model yields negligible gains.