Self-Correction as Transition Geometry: Internalizing Reasoning via Lifted State Policy Optimization
Abstract
Language models routinely revise initial answers through intermediate edits, yet current alignment objectives restrict credit assignment to complete outputs or entire rollouts. To formalize the local geometry of improvement, we introduce Lifted State Policy Optimization (LSPO), which augments each textual answer with a continuous auxiliary coordinate to form a lifted state that encodes refinement context beyond surface text. Training rewards transitions that descend an energy landscape, subject to edit and step penalties, and internalizes the resulting answers into the generative model. Across the Qwen3.5 family on mathematics, science, code, logic, and broad reasoning benchmarks, LSPO improves aggregate accuracy while bypassing the latency of explicit revision loops. Component ablations, together with evaluations of internalization and energy ordering on unseen tasks, confirm that this gain originates from the lifted representation and from credit assigned at the transition level.