MemoryVCD: Benchmarking Multimodal Self-Evolving Memory for Personalized Vision-Language Agents
Abstract
Personalized AI assistants must continually update the memory as multi-modal user behaviors accumulate over time and across domains. Yet existing memory benchmarks are largely text-only, rely on simulated or single-session interactions, and evaluate memory at a single static snapshot, making it difficult to study a central question for lifelong personalization: how should a memory system evolve as new interactions arrive? We introduce MemoryVCD, a benchmark that studies personalized memory along three separable design dimensions: what modality represents memory, what rule updates it, and along which axis it grows. MemoryVCD is grounded in real users with time-stamped histories of hundreds to thousands of interactions paired with product images. It combines five memory input formats, spanning plain text, rendered text, and product images, with a self-evolving evaluation protocol that expands memory along two orthogonal axes: a within-domain temporal axis and a cross-domain axis, while holding test items fixed so that performance changes can be attributed to memory evolution itself. We evaluate four canonical update methods across up to twelve vision-language backbones from three providers, twelve domains, and six user-facing tasks. Our results reveal four findings that challenge common practice: visual memory formats win 83% of storage cells and benefit smaller models most; longitudinal evaluation reverses the single-snapshot “recency wins” conclusion, as retrieval overtakes recency once user history grows; no update policy wins a majority of settings, yet the same policy-to-task mapping remains consistent across both growth axes; and cross-domain memory alone is sufficient for pattern- and style-oriented tasks but becomes a distractor for visual prediction, producing a U-shaped performance collapse that we further verify with an independent judge. We release all traces, memory representations, the evaluation harness, and every per-cell prediction.