The Fluency–Reasoning Dissociation in Depth-Pruned LLMs: A Representational Geometry Perspective
Saksham Kapoor
Abstract
Contiguous depth pruning reduces memory and latency in large language models by removing intermediate layers, but severs residual stream continuity at the incision interface. Standard recovery fine-tunes all remaining layers with LoRA, indiscriminately perturbing parameters in distant, unaffected blocks. We study post-pruning recovery through a controlled experimental vehicle, Incision Boundary Healing, which confines parameter adaptation to the $t$ layers directly flanking the cut. On standard global-attention architectures (Mistral-7B-v0.1, Meta-Llama-3-8B), boundary healing ($t=2$) improves C4 perplexity by 0.608–1.534 points over 6-layer global sensitivity LoRA while updating 33.3% fewer parameters; Gemma-2-2B needs $t=3$ (sliding-window attention), with no parameter savings. However, restoring language modeling fluency exposes a severe fluency–reasoning dissociation: while text perplexity returns to near-dense levels (9.12 vs. 7.86 on C4), multi-step reasoning accuracy on GSM8K collapses from 39.6% to 1.8% and recovers to only 6.2% ($N=500$, 8-shot CoT). On Big-Bench Hard Multi-Step Arithmetic, reasoning accuracy drops to $\le 1\%$ across all pruning depths. Representation geometry ($N=50$ sequences) reveals that depth pruning induces a $33.93^\circ$–$45.60^\circ$ canonical subspace rotation across model families, which boundary healing only partially realigns ($\Delta\theta = -1.90^\circ \pm 0.32^\circ$). In global-attention models, $\mathrm{PC}_1$ preserves its orientation ($\le 0.75^\circ$ deviation) but undergoes architecture-dependent variance collapse, compressing the primary reasoning subspace. In an intact model with zero weights modified, applying the pruning-derived inverse Procrustes rotation $(R^*)^T$ reproduces the reasoning collapse ($39.6\% \to 9.0\%$, $p = 5.6 \times 10^{-35}$), whereas magnitude-matched planar rotations leave accuracy intact ($40.3\%$) and random $\text{SO}(d)$ rotations destroy it indiscriminately ($0.0\%$), confirming selective vulnerability to the pruning-derived transformation. These findings establish that post-pruning reasoning failure is caused by coordinated subspace misalignment rather than scalar drift or capacity limits.
Chat is not available.
Successful Page Load