Towards Compositional Binding in Sparse Autoencoder Features via Holographic Reduced Representations
Abstract
Sparse autoencoders (SAEs) provide learned dictionaries of directions that often admit human-interpretable descriptions, but current uses of these features are largely additive and do not explicitly represent relations between them. We investigate whether SAE dictionaries can instead serve as vocabularies for structured binding. Using holographic reduced representations (HRR), we bind, superpose, and retrieve features from Gemma Scope and Llama Scope dictionaries. Across both SAE suites, multiple layers, and varying memory loads, features used as both roles and fillers achieve retrieval comparable to matched random HRR codebooks. The clearest cleanup failures occur among a small number of near-duplicate Gemma Scope features; Llama Scope lacks this extreme tail and shows little corresponding ambiguity. Overall, our results suggest that learned SAE features can support structured binding beyond purely additive operations despite the non-random geometry of real SAE dictionaries.