CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
Abstract
We present CustomerSim, a framework and benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-driven customer behavior in conversational retail environments such as ChatGPT Shopping. While prior work treats user simulation as surface-level dialog generation, our focus is on model ability to seek information on and make decisions that adhere to customer specifications in interactive, agentic simulations. We create CustomerSim, an environment consisting of a human-curated set of 360 personas over five product categories alongside a suite of metrics centered around measuring consistency between a customer simulator's actions and its specifications, and conversational quality. We find several behavioral gaps after benchmarking five open and closed-source state-of-the-art models. First, while models produce fluent conversations, they display significantly lower lexical diversity than human shoppers, and open-source backbones overdisclose their criteria in the opening turn. Second, models tend to be persuaded by sales agent tone and drift from persona specifications. Claude Opus 4.8 achieves <74\% average alignment with its underlying persona specifications. To make progress on these limitations, we propose UserGRPO, a multi-turn, multi-objective reinforcement learning recipe to optimize both conversational fluency and decision alignment under persona specifications. Our experiments demonstrate that UserGRPO raises the decision alignment of the baseline model from 0.417 to 0.652, a gain of 23.5 points, without meaningful degradation in conversational quality, and that these gains transfer to product categories held out from training. We further find that stylistic prompting, the natural alternative, trades directly against decision alignment: it is the only intervention that makes surface form more human-like, yet it nearly halves persona adherence in certain models. By introducing CustomerSim, we provide a testbed for the community to investigate and improve the adherence of user simulators in goal-oriented settings.