Bounded Mediation: Limiting LLM Execution Authority in Cyber-Physical Systems
Abstract
LLM agents are increasingly embedded in cyber-physical systems, where even schema-valid but semantically incorrect outputs can alter physical behavior. Existing safeguards validate outputs or shield individual actions, but do not bound how much an admissible LLM proposal may alter a coupled multi-agent controller. We introduce bounded mediation, an interface pattern that makes execution authority an explicit budget between LLM reasoning and trusted control. DAMA (Dual-Agent Mediation for Adaptive Incentive Design) uses role-separated LLM agents to propose coordination schedules; a deterministic mediator restricts them through capped action masks and bounded reward bonuses. Under local-sensitivity and feasible-fallback assumptions, we derive a degradation bound that scales with the masked-agent fraction and shaping budget, even for repeated schema-valid adversarial schedules. In mixed-autonomy traffic, DAMA achieves the highest mean throughput among LLM interfaces in five of six regimes with zero observed collisions. It matches or improves bounded reward shaping under schema-valid hallucinations, while channel stress reveals the robustness cost of its veto channel. These results point to authority budgeting as a principled way to retain useful LLM coordination while limiting its effect on physical control.