Interpretability-Conditioned Value Priors for Evaluation-Efficient Intervention Search
Abstract
Interpretability is usually treated as a diagnostic that describes what information a learned representation contains. We ask whether concept-level interpretability measurements can instead support decision making by predicting which representation interventions are worth evaluating before target-task feedback is available. We study this question in a controlled intervention-search benchmark built on frozen CLIP representations of the CelebA dataset, a large collection of celebrity faces. We train a gradient-boosted model to predict intervention value from concept-level interpretability features and compare it with two stringent controls. One is an equal-dimensional, capacity-matched, and data-matched generic value model. The other is a task-shuffled interpretability model that preserves the feature family while breaking the alignment between interpretability measurements and task value. On an untouched final split of eight task specifications spanning four held-out target attributes, the interpretability-conditioned prior identifies an intervention whose feedback incumbent reaches 95% of the attainable feedback improvement available over the no-intervention baseline while using about 64\% fewer evaluations than the matched generic model, 63% fewer than the shuffled control, and 73% fewer than random search. The benefit is specific to intervention ranking rather than a general advantage for interpretability in sequential optimization. An adaptive surrogate-UCB optimizer nearly matches the fixed interpretability ranking. Using the same learned values for potential-based reward shaping in tabular Q-learning does not produce an interpretability-specific advantage. Directly ranking interventions with a naive interpretability proxy also performs poorly. Through this proof-of-concept work, we conclude that interpretability-derived features can improve value estimation and intervention prioritization in a sequential decision problem, but the same signal does not automatically translate into interpretability-specific gains when used for potential-based reward shaping.