Do Reasoning-Quality Metrics Predict Better Decisions? Verifying Multi-Agent Debate Against External Market Outcomes
Abstract
Multi-agent LLM systems increasingly use structured debate to improve reasoning, and their outputs are commonly evaluated with internal reasoning-quality rubrics and consensus measures. Whether those proxies predict better decisions is rarely tested, because most benchmarks lack an external outcome to check them against. We use simulated portfolio allocation as a testbed where complete reasoning traces can be compared against an objective downstream outcome, evaluating specialist LLM agents under single-agent, independent-ensemble, and multi-round debate conditions, scored with a CRIT-style four-pillar rubric and Jensen-Shannon portfolio divergence against realized return, Sharpe, and drawdown. Across 210 runs, CRIT reasoning scores are essentially uncorrelated with financial outcomes (descriptive r=+0.07 with Sharpe, r=+0.03 with return), even though structured prompting moves those scores substantially (mean rho from 0.72 to 0.84, Cohen's d≈2.0) and debate portfolios beat an equal-weight baseline only 33% of the time. Improved outcomes align with interventions that limit consensus collapse rather than with interventions that raise reasoning scores: a divergence-collapse trigger improves Sharpe by +0.14 (p=0.028), while an intervention that verifiably repairs causal-reasoning scores slightly degrades performance. That trigger also constrained revision magnitude, so we cannot isolate disagreement preservation as the mechanism. Across our tested configurations the clearest benefits concern risk and risk-adjusted performance rather than consistent alpha, and consensus and rubric scores can both be maximized while decision quality stagnates. These results suggest internal reasoning proxies should be validated against external task outcomes before being treated as evidence of agent decision quality.