PeerConf: Self-Calibrating Early Stopping for Efficient Parallel LLM Reasoning
Abstract
Sampling many reasoning traces in parallel and voting is known to improve ac curacy in Large Language Models (LLMs), but every trace runs to completion regardless of quality. Confidence-based filtering terminates weak traces early. Cur rent methods, like Deep Think with Confidence (DeepConf), set the threshold from a dedicated warm-up phase that must finish before any trace can be filtered. We introduce Peer Think with Confidence (PeerConf), which estimates the same per- centile threshold from traces finishing within the run itself. The threshold arms at the first finisher, any trace that finishes with an answer, and is recomputed at every finisher thereafter, so filtering and consensus checks begin immediately. PeerConf also probes running traces for their current answer; a trace whose probe clears a confidence threshold and closes the model’s own reasoning commits early, casts that answer as its vote, and frees its seat. We evaluate DeepConf, PeerConf, and Self-Consistency on DeepSeek-R1-8B across MATH500, AIME25, and HMMT25, and on GPT-OSS-20B on AIME25. PeerConf uses 60.7% to 79.8% fewer tokens than self-consistency at 32 traces and 23.2% to 61.3% fewer than DeepConf. The completion probes alone use 42–50% fewer tokens on MATH500, 10–12% fewer on AIME25, and 5–9% fewer on HMMT25. More broadly, calibration in confidence- filtered reasoning does not need to be paid for separately. Code and benchmarks are available at https://anonymous.4open.science/r/PeerConf.