Read and Ignore: What Self-Directed LLM Agents Do in Public Service Scenarios
Diana Kozachek ⋅ Kinjalk Srivastava
Abstract
As LLM agents enter the offices where public services are made: city budgeting teams, hospital wards, college administration, environmental monitoring; With novice users, they routinely receive ambiguous instructions that leave them discretion over what to do next. We ask what agents do with that discretion. We build four expert-validated public-service scenarios, each a frozen operational cycle seeded with verifiable decision levers, collect $2{,}400$ trajectories from ten LLMs under four underspecified prompt conditions, and compare them against $1{,}000$ runs of the same models in a context-free sandbox \cite{anon}. The context is read but rarely acted on: agents inspect the records in $96.9\%$ of open-ended runs, yet produce an artifact or self-chosen task in only $18.7\%$ ($86.0\%$ once a task is explicit). Against the baseline, domain context changes \emph{which} models act rather than \emph{how many} ($18.7\%$ vs.\ $16.0\%$, not significant, while individual models move by up to $21.5$ points in opposite directions), and under the shared verbatim prompts it suppresses constraint probing by $60\%$ ($p{<}10^{-14}$). The artifacts that do appear can be polished yet ungrounded, citing budgets and counts that contradict the records. The rule-based taxonomy reads only action-level evidence, which four human raters apply consistently in the source benchmark (Fleiss' $\kappa{=}0.88$), and needs no labeled data, so a public-service team can rebuild a scenario from its own records and read the same metrics on the same scale before it procures a model. All scenarios, code, and trajectories are released.
Chat is not available.
Successful Page Load