Cheaper Faithfulness: A Gradient-Free Behavioral Metric Substitutes for Costly Parametric Unlearning under Model Compression
Avinash Kumar Sharma
Abstract
Measuring whether a language model's chain-of-thought (CoT) is faithful, causally driving its answer rather than rationalizing it, is central to trustworthy reasoning, but the most rigorous method is prohibitively expensive for resource-constrained settings. Parametric faithfulness via step-unlearning (FUR) requires gradient-based optimization for every reasoning step of every problem. We show that a gradient-free behavioral metric (ASAND, $\mathcal{O}(|\theta|)$, no backpropagation) recovers the same load-bearingness ranking as FUR-based FF-HARD under L1 magnitude pruning, at roughly an order of magnitude less compute. On two small open-weight models (Qwen2.5-0.5B, Qwen2.5-Math-1.5B) across four pruning levels, the two metrics achieve perfect rank agreement (Spearman $\rho_S = +1.00$). Crucially, we pair this with a lightweight three-test diagnostic that determines when the cheap substitute is valid: models that fail the diagnostic (e.g. due to high baseline prediction leakage) are correctly flagged as unsafe for substitution. Our results offer a practical, low-compute path to faithfulness measurement for edge deployment and compute-limited research environments.
Chat is not available.
Successful Page Load