Who Wrote the Final Decision? Auditing Decision Provenance in Small Language Model Agents
Abstract
Current benchmarks for small language models (SLMs) typically consider multiple factors, including task success, energy use, latency, and memory usage. Among these, task success is often measured from the final pipeline output, even when the outcome is influenced by multiple steps, such as validators, fallbacks, and ranking rules. This makes it difficult to determine whether the final outcome actually reflects the model's performance. To address this problem, we audit all 15 decision points in our drug-repurposing agent pipeline across four model sizes (27B, 9B, 4B, and 2B parameters). Our analysis reveals that end-to-end results can be misleading in assessing the performance of SLMs. While five of eight conditions yielded the same aggregated recommendation, run-to-run agreement drops from 0.82 to 0.52 and citation rejection increases drastically from 2.8% to 50.0%. Furthermore, deterministic code wrote the final value in 47-82% of match outcomes, while aggregation obscures much of the remaining variation. We also find that the aggregator can favor candidates that reach the final stage rather than candidates that win more tournaments. Changing the aggregation policy removes the apparent differences between the 9B and 27B models but not the difference for 2B. We identify and fix nine evaluation defects, including four that introduce model-size-dependent effects. These findings demonstrate that rigorous SLM evaluation requires auditing intermediate decision nodes rather than relying exclusively on final pipeline outputs.