Closing the Deal Is Not Enough: A Trajectory-Level Behavioral Rubric for Interactive Agents
Abstract
Interactive agents are increasingly judged by the outcome an interaction reaches rather than the trajectory that produced it. Agent-to-agent (A2A) marketplaces now let autonomous agents negotiate, transact, and settle payments on a principal's behalf with no human in the loop. In a recent real-world deployment, stronger and weaker models earned markedly different amounts per item yet were rated equally fair, a behavioral gap invisible to outcome metrics and legible only in the trajectory. Still, these systems are judged the way capability benchmarks judge everything: by what agents achieve, not how they behave to get there. In this position paper, we argue that A2A marketplaces should be evaluated as behavioral systems: the same model can look identical on outcomes while exploiting a weaker counterpart, ignoring reputation signals, leaking its principal's private information, or paying a scammer under pressure. Holding personas, prompts, model assignments, and seeds fixed enables paired comparisons of behavior across exchange rules. We propose a seven-dimension behavioral rubric and two scenarios, MarketDeal and SwapShop, spanning four stages from basic trading to an executed payment step under an adversary. Each dimension is scored by a grader combining deterministic checks with judged sub-metrics. Across 140 market trajectories across controlled conditions, the same model trades competently under one set of rules yet falters under another, mutual-win rates fall sharply, and a payment step under an adversary separates configurations on settlement safety. Behavior depends heavily on the rules of exchange, not just the raw capability of the model.