Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Abstract
Many recent math- and science-oriented agent systems adopt hierarchical designs with specialized reviewer roles, motivated by the idea that routing critique through a dedicated review stage should help turn wrong candidates into correct ones on hard problems. We study that expectation on 4,181 verifier-grounded Omni-MATH problems, a hard ten-tier benchmark with enough headroom to separate protocols, using matched gpt-oss-120b actors as the primary actor family. On tiers 1-2, multi-agent collaboration adds at most about 2 percentage points over the matched single-agent baseline. From tier 4 onward, the gains open sharply, reaching about 10 to 20 percentage points on tiers 6-9. In that regime, broadcast-style peer discussion attains higher final accuracy than a planner-executor-reviewer pipeline (PER). We use that divergence as a starting observation and ask when reviewer quality translates into effective solver updates. On this benchmark, the PER-broadcast accuracy gap is not accounted for by reviewer precision alone. Here reviewer precision asks how often a reviewer warning truly points to a real error. PER's reviewer has higher precision (0.861 vs. 0.644), yet correct critique is much less likely to change the next candidate the protocol carries forward, and reviewer-guided repair is correspondingly lower. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy instead of improving follow-through, while placing reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Taken together, these interventions point in the same direction: critique appears more likely to be acted upon when it is presented more directly in the solver's working context, although this evidence is directional rather than causal. Under reviewer-centric evaluation, a system can look strong at spotting errors yet still fail to solve more problems if the protocol does not act on those critiques.