TabPFN Through The Looking Glass: An interpretability study of TabPFN and its internal representations
Abstract
Tabular foundation models achieve strong predictive performance across domains, but their internal computations remain poorly understood. We study TabPFN v2 representations through probing experiments on synthetic tasks with known functional structure. Specifically, we test whether hidden states encode (i) coefficients of linear relationships, (ii) intermediate terms in compositional arithmetic expressions, and (iii) the final answer before native output alignment. Across experiments, coefficients and intermediates are strongly decodable in middle layers, while answer signals become linearly recoverable earlier than they are aligned by the output head. We further compare cross-fit and within-fit probing, showing that some representations are fit-specific even when decodable within a shared context. Using a logit-lens analysis, we track when intermediate activations enter output-space geometry, complementing linear decodability with compatibility to TabPFN's native readout. Overall, our findings support the view that TabPFN performs structured multi-step computation in its residual stream.