Beyond Runtime Guardrails: A Pre-Execution Agent Qualification Framework for Task-Criticality-Aware AI Deployment
Abstract
Autonomous AI agents increasingly plan multi-step tasks, call tools and execute actions with limited human supervision. Existing safety approaches largely focus on general capability benchmarks, output-level trust evaluation or runtime authorization of individual tool calls, each of which only takes action after an agent has already begun interpreting and started planning a task. This paper presents PEAQ (Pre-Execution Agent Qualification), a framework that evaluates whether an agent should be allowed to enter the planning and tool-use phase for a given task and criticality level, using a weighted Agent Trust Score (competence, policy adherence, adversarial robustness) combined with a set of non-negotiable hard constraints. A criticality-indexed decision engine returns one of four outcomes: approve, conditional approve, human review, or reject. We report a substantially expanded empirical evaluation relative to an earlier prototype. Instead of 80 handwritten cases evaluated with five prompt-conditioned personas of a single model, PEAQ is now evaluated against 2,079 cases drawn from five independently published agent-safety and tool-use benchmarks (AgentHarm, InjecAgent, ToolEmu, τ-bench, SWE-bench), normalized into a unified schema spanning five realistic agent capability domains (customer service, software engineering, financial operations, communications and data/systems access) and split 70/30 into training and held-out test sets. Two of the six evaluated agent profiles are literature-derived, with competence, policy-adherence and adversarial-robustness figures taken directly from published per-model tables in the τ-bench, AgentHarm and InjecAgent papers. The remaining four are disclosed synthetic archetypes retained for sensitivity analysis. Decision thresholds θ_κ are recalibrated against the literature-derived profiles rather than left as illustrative constants, with the recalibration's circularity explicitly discussed rather than hidden. On the held-out test split, PEAQ with runtime authorization blocks 98.3–100.0% of malicious or policy-violating actions (0% for the no-gate baseline, by construction), with the two internal layers shown by a dedicated three-way ablation to be complementary rather than redundant. This gain comes at a substantial and honestly reported cost: false-rejection rates of 36.4–95.2% depending on agent profile, traced to two concrete, disclosed root causes rather than left unexplained. We do not present PEAQ as a validated, deployable system but present a rebuilt empirical foundation with real benchmark data, literature-grounded agent profiles, a calibrated (not arbitrary) threshold policy and a completed component-wise ablation on top of which the original prototype's own stated future-work items are now either resolved or measured rather than merely proposed.