SituRecBench: A Benchmark for Situated Recommendation in 3D Interactive Environments
Abstract
Recommender systems are transitioning from static retrieval engines into interactive assistants grounded in physical environments. However, current paradigms largely treat recommendation as a decoupled ranking task, neglecting the situated complexities of spatial navigation, visual grounding, and multi-turn social interaction. To bridge this gap, we introduce SituRecBench, a novel interactive benchmark designed to evaluate situated recommendation agents within high-fidelity 3D environments. SituRecBench reformulates the recommendation process as a closed-loop interaction between a recommendation assistant and a user with evolving latent preferences and realistic behaviors. Our benchmark features diverse real-world shopping scenarios (e.g., supermarkets and home furnishings stores), requiring assistants to integrate preference elicitation, spatial reasoning, and long-horizon conversation. We introduce a simulator-in-the-loop evaluation protocol along with a multi-dimensional metric taxonomy that quantifies communication fluency, recommendation utility, and navigation effectiveness. Through extensive evaluation of state-of-the-art vision-language models, we reveal critical challenges in situated recommendation, such as inefficient preference elicitation, coherent conversation transition, and poor spatial grounding. SituRecBench provides a rigorous, reproducible foundation for advancing the frontier of situated recommendation in interactive, visually-grounded environments.