TOM-Pruning: Target-aware Output Manifold for LLM Pruning
Abstract
Post-training pruning compresses pretrained large language models (LLMs) by removing redundant weights without full retraining. However, existing pruning criteria primarily rely on parameter statistics or generic reconstruction errors, fundamentally ignoring the underlying geometric structure of how weight removals perturb the layer output. In this paper, we rethink LLM pruning from a local manifold perspective and reveal that representative methods (e.g., SparseGPT and Wanda) implicitly assume a standard Euclidean output metric. This geometry-agnostic assumption yields an isotropic manifold that assigns uniform costs to all perturbation directions, thereby failing to capture target-dependent sensitivity. To overcome this structural limitation, we propose Target-Aware Output Manifold Pruning (TOM-Pruning). TOM-Pruning equips the output perturbation space with a target-aware local Riemannian metric, assigning direction-dependent geometric costs based on the perturbation's alignment with the target response. By mathematically pulling this metric back to the parameter space, we derive a pruning sensitivity score that inherently captures a cosine-based input-output channel affinity. To resolve the poor discriminability of this raw affinity in high-dimensional spaces, we further introduce a Competitive-Hubness Channel Affinity Mechanism to reconstruct a highly distinguishable final pruning criterion. Extensive experiments across multiple LLM families demonstrate that our manifold-driven approach consistently outperforms existing no-weight-update pruning baselines under various sparsity settings, notably reducing WikiText-2 perplexity by 26.8\% on LLaMA-3.2-1B and achieving an 11.0\% relative zero-shot accuracy gain on LLaMA-3-8B at 70\% unstructured sparsity.