Resilient Latent Readouts for Long-Context Question Answering
Abstract
Long-context question answering often asks a model to answer from extended documents, multi-turn dialogues, or persistent histories. Directly prefilling the full history preserves all observed tokens, but it is expensive and can expose generation to substantial irrelevant or spuriously salient context, which may hinder evidence localization and answer synthesis. This motivates compact memory interfaces that provide a short, query-conditioned readout of the history. Existing systems predominantly instantiate this readout as text, where the reader sees only what was selected and any evidence left out is unrecoverable, leaving the interface fragile to selection errors. We argue that memory readout has two coupled axes that should be optimized separately: evidential coverage, whether all right support is selected, and inferential resilience, how much answerability remains under imperfect selection. To jointly address these two axes, we propose LIRA, a latent-memory reader that exposes full-prefix contextualized states rather than re-encoded text. It amortizes one full-prefix pass into a reusable latent store; for each query, it localizes candidate positions with calibrated endogenous attention, repairs incomplete support with coverage-aware planning, and packs selected states into a position-consistent readout for autoregressive generation. Across three open-weight backbones on LoCoMo and Loong, LIRA is the strongest non-oracle compact-memory method and surpasses Full Context on 8 of 9 LoCoMo metrics. When localization misses all supporting evidence, LIRA improves F1 by up to 13.3 points over a text-native control with identical positions, suggesting that full-prefix states preserve answer-relevant influence discarded by text-native re-encoding.