Opening Up a New Layer: A Deeper Look into "Interpreting CLIP with Hierarchical Sparse Autoencoders"
Abstract
Sparse Autoencoders (SAEs) have become essential for decomposing model activations into interpretable concepts. However, despite their effectiveness, SAEs lack a natural ordering of features, making it difficult to prioritize important concepts under compute constraints. The Matryoshka Sparse Autoencoder (MSAE) was introduced as a means of learning nested subspaces, theoretically forcing high-level features into earlier dimensions to enable adaptive granularity. In this paper, we reproduce and analyze the MSAE framework. The reproduction of the study suffers from certain challenges, but the main claims still hold. Furthermore, this study extends on the original paper by implementing feature absorption as a metric, reducing computational cost and emissions with a method inspired by the Sandwich Rule, and deeply analyzing MSAE's ability to learn hierarchical information.