Same Answer, Different Path: Measuring Prompt Sensitivity Beyond Task Outcomes in Small Language Models
Afroditi Papadaki ⋅ Kaibang Zhang
Abstract
The same request can be phrased in many ways and a model that answers well only for some of those phrasings is unreliable in practice. Current ways of quantifying this problem watch only the final score, asking whether an answer flips rather than how the generation itself shifts. However, the score is a proxy, as two generations it rates identically can still differ in what the user actually receives, or in whether the model reached the answer for a reason that will hold on the next input. We therefore propose three measures that look inside the decoding process, tracking where two generations diverge token by token, how sharply the model's confidence moves where they differ, and how much its probability trajectory shifts overall. We pair these with a taxonomy of five prompt-variation styles, and apply both to $14$ small language models between $0.27$B and $8$B parameters across three benchmarks. On multiple-choice benchmarks, our measures capture variation that the outcome view misses, separating instances a model answers reliably from those it does not, even where outcome scores rate them identically. We also show that prompting a small model to reason under a fixed generation budget helps only a minority of them and becomes less effective more as they grow larger.
Chat is not available.
Successful Page Load