GraviLogic: geometric audit of decision paths and the limits of what an interpretability metric can certify
Abstract
Interpretability metrics are increasingly cited as evidence in accountability settings — regulatory reporting, clinical model approval, incident review. We ask what such a metric can certify, and prove that a broad family of them certifies less than is usually claimed. GraviLogic measures the geometry of the decision path: the trajectory a record traverses in latent space from its observed state to the nearest counterfactual that flips the model’s decision. Its two principal quantities are the total turning Θ of that path and its arc length L. We establish a negative result — representational underdetermination: two implementations of the same predictive function, agreeing to machine precision on outputs, gradients, integrated gradients, SHAP values and counterfactual locations (max |fA−fB |= 2e−16, max |SHAPA−SHAPB |= 5e−17), disagree on Θ by a median of 23.5% with rank correlation ρs = 0.853 and top-decile agreement of only 0.33. The disagreement is not numerical noise; it scales with the nonlinearity of the encoder and vanishes in a linear control (ρs = 1.0000). We then run a negative control on a clinical cohort (UCI Cleveland heart disease, n= 297, MLP, test AUC 0.906): transformations that provably cannot change the model — permuting the input columns, adding an inert record-number column — shift the audit as much as a genuine change in the model does. A column permutation alone moves the flagged-patient set from Jaccard 1.00 to 0.60. Two invariances do hold to machine precision, and we state them as the metric’s certified core. From this we derive a two-tier reporting protocol: what a path-geometry audit may claim, what it may not, and which invariance tests must accompany any such claim before it is admissible as evidence.