According to Me: Long-Term Personalized Referential Memory QA
Abstract
Personalized AI assistants must recall and reason over long-term user memory, which naturally spans multiple modalities and sources such as images, videos, and emails. However, existing Long-term Memory benchmarks focus primarily on dialogue history, failing to capture realistic personalized references grounded in lived experience. We introduce ATM-Bench, the first benchmark for multimodal, multi-source personalized referential Memory QA. ATM-Bench contains four years of privacy-preserving personal memory data and human-annotated question–answer pairs with ground-truth memory evidence, including queries that require resolving personal references, multi-evidence reasoning from multiple sources, and handling conflicting evidence. To tackle the problem, we propose Schema-Guided Memory (SGM) to structurally represent memory items originating from different sources. In experiments, we implement nine state-of-the-art systems to evaluate different memory ingestion, retrieval, and answer generation techniques. We find that prior systems achieve limited performance on ATM-Bench-Hard, while our proposed SGM consistently improves performance across all evaluated systems.