The Tangled Web They Weave: Exploring Neural Representations with Directed Simplicial Complexes
Abstract
Understanding how artificial neural networks (ANNs) represent and combine features internally is a central challenge in mechanistic interpretability. Existing neural circuit-discovery methods represent neuron interactions as pairwise edges in a graph. But distributed neural representations consist of large groups of neurons that are inherently linked by higher-order causal dependencies: suppressing a subgroup of neurons in concert may shift the activation of other neurons in the feature representation. We propose a method to construct filtered directed simplicial complexes (DSCs) that give a topological representation of such dependencies, built from the activation profiles of an ANN on a given set of inputs. Our approach treats ANNs as structural causal models and uses counterfactual ablations to test for causal links between groups of internal neurons within specific input contexts. As a proof of concept, we apply the approach to a pair of small multilayer perceptron (MLP) architectures in a synthetic two-dimensional quadrant classification task. In these examples, DSCs recover the designed feature pathways and coordinated class transitions, and the complexes of trained networks are higher-dimensional than those of their handcrafted references. As validation, we compare our method with graph-based circuit-discovery baselines and find that their 1-skeletons agree within the relations that a weight graph can represent. These results support DSCs as context-dependent descriptions of multi-neuron computation and may thus provide a novel, orthogonal framework to study the phenomenon of distributed feature representations and polysemanticity.