Support Mismatch as a Benchmark Failure Mode for In-Context Prediction
Jie Liu ⋅ Lanlan Liang ⋅ Qinan Bao ⋅ ZiXi Yan ⋅ Linghao Meng ⋅ Wenbo Gong
Abstract
We analyze support mismatch in language-model in-context prediction as a frozen-prior Bayes benchmark with finite discrete support. The goal is diagnostic rather than literal: we do not identify production LMs with exact Bayesian updating, but characterize a failure mode whose signatures can be checked in learned in-context predictors. Our main theorem gives the finite-sample predictor-risk bound $R_k^\pi \le C e^{-k\rho_{\min}} + 2\varepsilon_{\mathrm{approx}}^2$, so risk decomposes into an exponentially decaying transient and a nonvanishing support-mismatch floor under exponential-moment control of log-likelihood ratios. In scalar and aligned Gaussian location families, exact two-component analyses match the upper exponent up to a tight factor 2 and show a regime change at $\mu_2=3\mu_1$. Controlled learned-model experiments track the benchmark near crossover. Decoder-only LM diagnostics then test the benchmark signatures: restricted support creates a persistent high-error regime; restoring the missing support sharply lowers error; and modern size-matched distractor controls on Qwen and Llama show that adding a wrong same-cardinality candidate does not reproduce this gain. Multiple-choice, instruction-tuned, cross-family, calibration, null-control, and near-tie checks preserve the same diagnostic picture.
Chat is not available.
Successful Page Load