Cost-Aware Multi-LLM Aggregation under Hierarchical Pitman–Yor Dependence
Abstract
Repeated sampling can improve LLM performance, but self-consistency is tied to a single model: repeated responses may reinforce a model-specific mistake. Consulting several LLMs provides broader evidence, but raises two linked questions: which LLMs should be queried, and how should queries be allocated among them? These choices are difficult because outputs can be marginally correlated within and across models. We introduce the cross-consistency problem: identifying a population-level consensus answer at minimum expected query cost, subject to a prescribed confidence level. A hierarchical Pitman--Yor model describes answer generation across LLMs and formalizes the problem as a two-stage stochastic program. Stage 1 selects an ensemble and Stage 2 adaptively queries within it. We bound the overall consensus-recovery error, characterize the asymptotically optimal query cost for a fixed ensemble, and, since exact ensemble selection requires nested HPY inference, develop a tractable linear approximation separating the value of including an LLM from the value of querying it repeatedly.