Activation Traces as Neural Artifacts: Visualizing RNN Learning with MM-PHATE
Abstract
Neural network artifacts are often studied through weights, checkpoints, or scalar training records, but executing saved models on a shared set of inputs produces a richer artifact: a computational trace of how internal representations develop through learning. For recurrent neural networks, this trace is structured simultaneously by training epoch, sequence time, hidden-unit identity, and probe input. Treating these axes and their relations as part of the data makes it possible to study not only where a model ends, but how its computation emerges. We introduce Multiway Multislice PHATE (MM-PHATE), a graph-based method for representing and visualizing these activation traces. MM-PHATE relates units with similar responses across shared inputs within each epoch and time-step, while following corresponding units across sequence and training time. A joint diffusion geometry integrates these complementary relations into a low-dimensional view of how unit organization and temporal differentiation evolve during learning. In controlled dynamical systems, MM-PHATE organizes activations according to known regime transitions. In task-trained recurrent models, it preserves the intended relational neighborhoods and reveals changes in internal geometry that correspond to learning phases and task-relevant structure. These results support treating activation traces as a structured neural-artifact dataset: shared inputs define what is observed, tensor indices provide structure, and the graph defines how observations are compared. MM-PHATE thereby connects internal representations, optimization trajectories, and model behavior at unit and time-step resolution.