ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models
Abstract
Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial reliability maintenance and real-time fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to the role it serves: taking actions and providing information that respect the knowledge and capability boundary of the given role. As agent responses are oftentimes followed as procedure and enacted on physical equipment, an uncalibrated response is not a recoverable coding error but equipment damage, production loss, or personnel harm. However, existing benchmarks largely overlook the vital requirement on LLM agents' Perspective Awareness. To this end, we introduce ReFract, a benchmark of 150 expert-validated entries in which an agent must act differently on the same query depending on who is asking. Entries of ReFract is grounded in anonymized real user queries, against which we construct Text World Model that simulates the agent's operating environment and assemble perspective-aware action trajectories. State-of-the-art LLMs solve at most 69% of the tasks, yet almost 60% of their trajectories contain perspective-violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and motivating agents that calibrate not just how to act, but for whom.