Where SAE Interpretations Fail: Readout Fragility Under Object Scale
Abstract
Sparse autoencoders (SAEs) are increasingly used to interpret vision-model representations through small sets of latent features associated with human-interpretable concepts. In practice, these interpretations rely on a readout, a procedure that maps selected SAE features to a concept-level score. Yet the reliability of such readouts is rarely stress-tested under changes that preserve concept identity. We audit a released vision SAE on a frozen CLIP ViT-B/32, using COCO instance masks to measure object spatial extent, defined as the fraction of the image occupied by the object. On held-out images, frozen-readout reliability declines as objects become smaller. Controlled same-object rescaling, with identity and background fixed, establishes a graded causal effect. To test whether the frozen readout is especially fragile, we compare the frozen SAE readout with trained linear probes on the raw CLIP representation and full SAE latent. Spatial-extent-linked degradation persists under trained readout, indicating that part of the loss is inherent to the representation, but the frozen readout degrades 2.1× more than the trained-probe mean. We also test this on a second SAE from the same family and observe the same pattern. These results show that apparent concept absence at small extent partly reflects readout failure, rather than simply a loss of representation-level information, motivating controlled reliability audits of frozen SAE interpretations.