One Candidate Changes the Rest: Non-Local Instability in LLM-based Web Shopping Ranking
Abstract
Large language models (LLMs) and LLM-based agents are increasingly used to recommend and rank products in e-commerce settings. In these LLM-as-a-Recommender scenarios, models are often asked to map an unordered set of candidate products to an ordered ranking. While prior works have shown that LLM-based rankers can be sensitive to the presentation order of candidates, we identify a distinct and previously underexplored feature of these rankers: \textbf{Non-Local Instability}. Ideally, with a non-local stable ranker, modifying a single candidate should only affect its own relative position. However, we show that current autoregressive LLMs systematically violate this property. Small, localized changes to the description of a single product can reorder products whose descriptions remain completely unchanged. We characterize this phenomenon across multiple perturbation types and evaluate its prevalence in both reasoning and non-reasoning models. Our experiments show that non-local instability is widespread, persisting even in recent state-of-the-art models. These findings reveal a fundamental robustness issue in LLM-based ranking systems: the relative ranking assigned to an item can be largely affected by seemingly irrelevant semantic changes to other candidates.