The Monitor Is Not Free: When Equal-Call and Equal-Cost Evaluations Disagree in SLM-to-LLM Cascades
Abstract
A panel cascade runs several small language models (SLMs) before deciding whether to call a stronger model. The panel cost is therefore incurred on every query. We use a frozen three-SLM Pyramid MoA configuration to study how the choice of resource-matching criterion affects system comparisons. Equal-call matching holds the number of strong-model calls fixed; matched-cost evaluation instead holds fixed the mean token-priced cost of recorded model-generation calls. Over their respective supported sets of operating points (anchor supports), the median difference against RouteLLM favors Pyramid R3 at equal calls on GSM8K and MMLU but favors RouteLLM after cost matching; RouteLLM is favored under both views on a restricted MBPP cohort. At matched modeled cost, the median comparator-minus-R3 gap over each policy's supported R3 anchors is positive on every benchmark for four comparator methods and a seeded-random baseline. A strong-model price-sensitivity analysis tests whether and where these median gaps change sign. These descriptive results concern one frozen configuration and one dated set of token prices; they do not measure full deployment cost or rank intrinsic router quality.