Test-Time Augmentation for LLMs: Where to Inject Diversity at Inference Time
Abstract
Aggregating several LLM responses to the same question can improve accuracy, but every method that does so must decide where the responses should differ. Established test-time scaling methods such as self-consistency vary only the reasoning path, resampling from a fixed input. The input modification offers several further options: the wording of the question can be rephrased, the spelling can be perturbed, and the accompanying images can be transformed. We study Test-Time Augmentation (TTA), which aggregates predictions across transformed versions of the input, and use it to ask where diversity is best injected. Comparing these options across six benchmarks at a matched number of inference calls, we find that where the diversity comes from matters more than how much of it there is. Meaning-preserving paraphrases of the input yield the largest accuracy gains (+1.80 pp over a single call) and reach peak accuracy with the fewest samples. Yet, they are not the most diverse: character-level noise makes the responses disagree far more often (37.8% versus 27.4%) while gaining less accuracy, and perturbing two modalities at once falls below the single-call baseline. Not all diversity is equally useful: perturbations that keep the meaning intact produce disagreement that aggregation can use, while those that corrupt it produce disagreement that is wrong. Overall, our results suggest that smaller LLMs such as Claude Haiku, varying how a question is phrased is a better use of an inference call than varying the reasoning path. Our TTA implementation is available at (masked-for-review).