ConceptBasis
Abstract
The linear representation hypothesis proposes that semantic concepts are encoded as directions in a model's embedding space. Directions estimated with linear probes or differences between group means already support useful concept readout, steering, and debiasing. However, directions that work individually need not form a separable, additive coordinate system. Such a coordinate system would make output embeddings natively interpretable and enable more precise test-time interventions. Group-mean-derived directions can be confounded by co-occurring concepts, while directions estimated after controlling for those co-occurrences can still overlap geometrically. We introduce ConceptBasis, which constructs a fixed concept dictionary and dense label matrix, uses joint ridge regression to estimate partial concept effects, and trains lightweight image and text adapters to reduce the remaining overlap while retaining contrastive retrieval. We evaluate direction overlap and class-disjoint compositional retrieval on the THINGS object dataset. Compared with orthogonality training on group-mean-derived directions, ConceptBasis reduces RMS overlap from .122 to .056 and raises ten-attribute compositional Recall@5 from .624 to .730, while bidirectional image-text Recall@5 remains virtually unchanged.