Latent-Context Distillation: Recovering Chain-of-Thought Effect with Latent Thoughts
Abstract
Reasoning traces, such as from chain-of-thought (CoT), improve the performance of LLMs on tasks which require reasoning, but substantially increase the number of tokens used, driving up inference cost and context use. The trace is then discarded: its value lies not in itself, but in its effect on the distribution over final answers. A short latent context offers an appealing alternative. The existing latent-reasoning literature focuses on \emph{generating} latent reasoning, whether through fine-tuning or training-free methods; performance is mixed at best. None asks the fundamental question: \emph{does latent reasoning exist?} We run a feasibility study to answer this question: we use a frozen LLM and, per question, optimise a (short) latent context---the only trained parameters---to minimise the KL divergence between the student and teacher answer distributions. The results are positive: we establish the existence of latent thoughts which recover much of the CoT-induced performance. On MATH-Hard (Qwen3-8B), a single latent thought recovers 56\% of the CoT-induced accuracy gap; using 128 latent thoughts leaves the student accuracy indistinguishable from the teacher. On HumanEval+ (Qwen3-1.7B), recovery plateaus near 90\%.