Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve
Abstract
Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, retained 0.65 success, the highest observed under social access. It fell by 0.11 from centralized indirect access, with an interval that included zero. Exploratorily, social access descriptively separated five of six model pairs, while the two centralized settings and native anchor separated none at this sample size; post-hoc task-blocked tests found three multiplicity-adjusted interactions. A reference-relative reader associates the wider gaps with failures to obtain needed information. Because it cannot distinguish a poor request from a poor simulated reply, we treat that association as a diagnostic, not a mechanism.