Faithfulness Is Not Actionability: Extrinsic Validation of Circuit Selectors Under Structured Pruning
Abstract
Circuit-discovery methods are usually assessed using patch faithfulness (patchF), which means that a discovered subnetwork retains the target behaviour when all the parts outside it are removed. In practice, though, downstream pipelines tend to use the same component scores when deciding to remove components, for example in the case of structured pruning. We investigate whether patchF actually predicts downstream pruning utility throughout a range of selectors, models, and tasks. It does not do so in a reliable way: on one task it strongly anti-predicts post-pruning accuracy, on another it predicts in the opposite direction, and in the remaining cases it gives no meaningful indication. This negative finding remains true even when post-pruning recovery is applied and a higher sparsity budget is used. A causal cross-swap test identifies the source of the failure as lying in the components where faithful and utility-maximising selectors come to different conclusions, demonstrating that components which are discarded by high-faithfulness selectors can be considerably more damaging to remove. Further permutation-based ablation analysis also shows three qualitatively different regimes of response to removal: harmless, presence-dependent, and bottleneck. PatchF fails to discern among these regimes. Our results therefore indicate that intrinsic circuit faithfulness is not a stable indicator of intervention utility. Even if a faithfulness score is perfectly stable, it still needs to be validated externally before it can be used as a justification for making a removal decision.