Off-Policy Evaluation of Large Language Models via Learned Semantic Bottleneck Embeddings
Abstract
Evaluating large language models (LLMs) via online human feedback is prohibitively expensive and time-consuming. Off-policy evaluation (OPE) addresses this by estimating LLM policy performance from offline data. However, traditional OPE estimators such as Inverse Propensity Score (IPS) suffer from high variance. Marginalized IPS (MIPS) addresses this variance issue by reweighting the IPS weights via condensed semantic embeddings. However, MIPS assumes embeddings contain sufficient reward-predictive information, an assumption that off-the-shelf embeddings often fail to satisfy. To address this, in this paper, we propose Semantic Bottleneck Embedding (SBE), which learns representations for marginalized reweighting via a conditional information bottleneck, compressing reward-irrelevant information while preserving reward-predictive content to achieve both bias and variance reduction. Our objective targets minimizing the derived mean squared error upper bound of the resulting MIPS estimator. Empirically, evaluation on the HelpSteer2 and UltraFeedback datasets across Qwen, Gemma, and Llama policies show SBE-MIPS reduces MSE over prior baselines.