Evaluating Agent Fidelity Beyond Capability for Local Multi-Agent Decisions
Abstract
Multi-agent systems increasingly use independent judgments followed by critique, challenge, or debate to improve the reliability of decisions on challenging questions. Such interaction can correct an erroneous judgment, but it can also persuade a capable agent away from a correct one. An agent’s contribution therefore depends on two distinct properties: capability, how reliably its independent assessments are correct, and fidelity, how its later public responses relate to those assessments under conflicting input. We ask whether an agent’s fidelity provides useful information once the system has access to that agent’s capability history. We introduce a controlled evaluation of one transition within this setting rather than full debate: each agent’s independent assessment is recorded before a fixed public-reporting step driven by a standardized opposing cue, separating initial correctness from subsequent reporting behavior. A decision-making system then makes matched decisions using population-average histories, agent-specific capability history, agent-specific fidelity history, or both. All conditions reuse the same frozen tasks, independent assessments, and public responses, isolating the value of the historical information available to the decision-making system. Across multiple model–provider configurations and 80 constructed tasks, agent-specific fidelity predicted held-out maintain/reverse/abstain behavior. Under the conflict-reporting regime, Brier loss was 12.8% lower when fidelity history rather than capability history was agent-specific; adding fidelity history to capability history reduced Brier loss by 11.8%. Only the latter comparison is directly incremental, and the prespecified decision-value criteria were not met. Therefore, these relative differences do not establish incremental decision value beyond agent-specific capability history.