Same Conversation, Different Languages: Localizing Multilingual Gaps in Agent Memory
Abstract
Most agent-memory benchmarks are in English, leaving open both whether equivalent conversations produce consistent memory behavior across languages and where any difference appears in the pipeline. We introduce a parallel benchmark of 20 synthetic user profiles, each with four memory-building sessions and one evaluation session, aligned across English, Spanish, Hindi, and Simplified Chinese: 400 sessions and the same 71 evaluation questions in every language. Across LangMem, Mem0, Graphiti, and Cognee over three trials, producing 3,408 answers, non-English accuracy is lower in 11 of 12 system–language comparisons by 9.8 percentage points on average. The diagnostics show that this difference is not uniform across languages or pipeline stages. We probe where the difference appears using three diagnostics. First, we bypass the providers, supply all correct evaluation-relevant memories directly to the LLM, and require only opaque memory identifiers. The model still selects the complete required set 7.0–12.2 points less often in matched non-English conditions, showing that provider behavior alone cannot explain the endpoint difference. Within this provider-free task, the statistically supported Spanish and Chinese contrasts occur when the supplied memories are localized, whereas the Hindi contrast occurs when the question is localized. Saved provider traces show a compatible pattern: conditioned on sufficient stored evidence in both language copies, retrieval retains that evidence 1.1 and 2.0 points less often for Spanish and Chinese, but 13.5 points less often for Hindi, pooled across three systems and driven by Mem0 and Graphiti. Finally, an isolated extraction test does not reproduce the endpoint disparity, with expected-fact recall remaining within 1.0–1.2 points of English. Together, the provider-bypass and retrieval diagnostics are consistent with question-side sensitivity for Hindi and evidence-language sensitivity for Spanish and Chinese. Because the retrieval analysis does not independently manipulate query and stored-entry language, this is an observational pattern rather than a causal localization. Nevertheless, it shows that a single multilingual endpoint score can obscure language- and stage-specific behavior.