When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Resources, and Task Difficulty
Nicolas Leins ⋅ Nico Pelleriti ⋅ Jana Gonnermann-Müller ⋅ Sebastian Pokutta
Abstract
LLM orchestration allocates additional inference-time computation to improve reasoning, but its gains may not justify its cost, and unequal optimization effort complicates comparisons. We compare Self-Refine, Best-of-$N$, and Debate with task-only and chain-of-thought (CoT) single-call baselines across five LLM backbones and three domains: programming, chess, and mathematics. To ensure comparability, we optimize each method with GEPA under the same budget and evaluate all methods on identical difficulty-stratified benchmark items. Orchestration yields moderate, benchmark-dependent gains. Averaged across backbones, the largest benchmark-level improvement is 6.2 percentage points over optimized CoT and 4.5 points over task-only inference, while requiring approximately 2 to 4 times the mean tokens of task-only inference. Item difficulty is associated with lower absolute accuracy in all benchmarks. We find no evidence that the relative benefit of orchestration increases with difficulty on Codeforces or Lichess; on ArXivMath, the relative benefit of Best-of-$N$ decreases with difficulty. Exploratory analyses reveal method-by-backbone interactions on all three benchmarks, indicating that orchestration effectiveness depends substantially on the underlying model. These results motivate model-specific orchestration decisions and evaluations that control optimization effort and report accuracy--cost trade-offs.
Chat is not available.
Successful Page Load