When Does Fine-Tuning Overtake Prompting? An Accuracy and Cost Analysis of Small Language Models on Narrow Tasks
Abstract
When small language models (SLMs) are deployed for narrow tasks, adapting them with a small number of labeled samples can substantially improve performance. Three adaptation families compete in this regime: in-context learning (ICL), parameter-efficient fine-tuning, and automatic prompt optimization (APO), yet no prior study compares all three at matched conditions. We compare three ICL selection strategies, fine-tuning using LoRA with per-task hyperparameter optimization, and the APO method GEPA across 17 narrow tasks, three Qwen3 model sizes, and data budgets up to 200 labeled samples. ICL sample efficiency largely saturates by 25-50 examples regardless of selection strategy, while fine-tuning keeps improving across the full budget range. Averaged over our task suite, approximately 50 labeled samples suffice for fine-tuning to overtake all ICL strategies, with the crossover occurring earlier on 0.6B and 1.7B than on 4B, although task category can shift or reverse this ordering. GEPA, in our configuration, is not competitive. A throughput-calibrated two-component cost analysis shows that, under single-task serving with merged adapters, data availability rather than deployment volume is the main driver of the cost-optimal method.