Halluci-action: Evaluating Deception in Claim–Action Alignment of Agents
Kaifang Mao ⋅ Chuke Liu ⋅ Wenlan Gu ⋅ Ryan Calo ⋅ Hua Shen
Abstract
Tool-using large language model (LLM) agents are increasingly evaluated by task success and factual accuracy, yet these criteria do not test whether an agent's account of its own behavior matches the actions it actually executed. We study this gap, which we call Halluci-action: an agent may claim an operation unsupported by its execution trace or execute an operation that it never discloses. We therefore introduce HalluciActionLens, a framework that evaluates action-related claims against recorded tool use in both directions, and derive the Halluci-Action Rate (HAR), together with $\mathrm{HAR}_{\mathrm{severity}}$, which restricts scoring to the gaps that carry consequences: required actions claimed but not performed, and forbidden actions performed but not disclosed. Across five frontier agents in two independent environments---an Odoo enterprise sandbox and AppWorld---per-model HAR reaches 45.9%, and with neither gap direction consistently dominating across environments, underscoring the need to evaluate both. Existing criteria miss the gap: 54.2% of factuality-clean Odoo trajectories and 52.1% of goal-complete AppWorld trajectories still contain an unsupported action claim or an undisclosed action. A failure analysis further locates these mismatches at four gates---execution, grounding, outcome verification, and disclosure.
Chat is not available.
Successful Page Load