Routing, Not Redundancy: Capability Activation in LLM Agent Collectives
Abstract
Agent collectives are usually scaled by adding agents, debate rounds, or redundant judgments. We argue that a class of failures is invisible to all three, and that evaluating collectives requires measuring not what their members can do but what their operating roles elicit. We make the case with a controlled study. Language agents review manuscripts whose reported evidence is held fixed while the disclosed size of a configuration search varies, so that the statistically correct verdict flips at a known point. Generic-role agents detect the problem readily - they raise the search in 66% of reviews - but across 166 responses not one names a correction procedure or computes an adjusted threshold, and their judgments barely move across the decision boundary. Routed to a specialist role, the same models express the correction in 49 of 74 responses. Redundancy does not close the gap: panels of three drive detection to saturation while their discrimination across the boundary falls to zero. The failure is therefore one of calibration rather than detection, and it is a property of the elicitation a role provides rather than of the models. We propose activation coverage and routing recall as ecosystem-level measures, give a protocol for estimating them, and argue that benchmarks reporting only directly elicited competence will systematically overstate what a deployed collective does.