Smart Fine-Tune, Then Rectify
Abstract
Driven by recent advances in artificial intelligence (AI), a growing literature has demonstrated the potential for using large language models (LLMs) as scalable surrogates to generate human-like responses in business applications. Two common approaches improve LLM performance: fine-tuning, which updates the LLM toward human responses, and rectification, which corrects biases in LLM outputs. We develop a two-stage framework that combines both methods and optimally allocates limited labeled samples across them. Unlike the conventional objective of minimizing mean squared prediction errors, we minimize residual variance, which is optimal for downstream rectification. We then leverage the fine-tuning scaling law to determine the optimal allocation. On the Wine Reviews dataset, the rule allocates 10.3% of labels to fine-tuning, closely matching the 10% empirical optimum on the tested grid. Our framework saves 48%-54% of labeled samples relative to the sample mean estimator and reduces estimator variance by 22%-42% relative to employing the standard MSE objective within the same framework. We further extend the framework to general M-estimation problems and evaluate the extension empirically.