The Safety Decisions Behind LLM Agents: Options, Conventions, and Open Questions
Abstract
Builders of LLM-based agents face many safety decisions — how much autonomy to grant, where a safety check should sit, what enforces a policy while the agent acts, when to ask a person — and the literature offers several answers to each. This paper is a guide to those decisions: the available answers, which of them have been tested, which have not, and where a researcher can add a comparison the field still lacks. From a corpus of 12,060 admitted works, a decision map organizes the safety-facing literature into 212 decisions over 1,753 works, with 1,176 mapped option rows. A census then reads, end to end, every reachable open-access work placed at the options of 4 questions covering 16 of those decisions. The answers differ sharply in support. Some conventions are both widely held and tested: adversarial robustness is evaluated against the deployed agent system, and 4 of 5 full-text positions behind that convention run a control. The most debated choice — where safety checks should sit in the pipeline — is tested one side at a time: every full-text position on it runs a control, yet each tests its own placement, so no two leading options have been compared at matched tokens, dollars and latency. The largest open question, how constraints are enforced while the agent acts, has little controlled evidence behind its leading option: 2 of 6 positions there run a control, and those controls vary the attack or the detector, not the enforcement choice. Approval gates lead the human-oversight literature, and none of the 18 census works on that decision compares a gate against no gate on a safety outcome. 98 works carry safety decisions so thin that only one or two works address them at all. The comparisons that would settle the open and contested options are known, small, and largely unrun.