Geometry of Validity: Analyzing Trajectories Beyond Linear Probes
Abstract
Large language models often produce reasoning that looks convincing but is invalid, and their written explanations are unreliable witnesses of what they actually compute. The truth of a single statement can be read from a model's internal state with a linear probe, but a reasoning chain is a sequence, and as a model reads one its internal state traces a path through the network's layers, so we measure the shape of that path with the tools of curve geometry. Valid chains trace smoother paths than broken ones in almost every model and dataset we tested, and the effect survives strict controls for text length, which otherwise imitates geometric signal. A probe on a single pooled internal state detects the validity of an entire chain, and trajectory geometry never significantly improves that probe, yet it beats a linear compression of the representation to the same number of dimensions. The geometric detector also responds to corrupted content while remaining at chance on corrupted order, so the shape of a failure carries information about its kind that a detection score does not. Reasoning validity is written into the curve geometry of a model's computation, but the model compresses that same information into the linear geometry of every single state.