Invariance Is Inherited: What a Contrastive Encoder Keeps From a Misspecified Simulator
Matteo Cartiglia ⋅ Wouter Botermans ⋅ Sandro Kuppel ⋅ Sanjin Marion
Abstract
A contrastive encoder learns invariance to factors that differ between its positive pairs. With image augmentations, the designer specifies how views differ through explicit transformations; with two simulated measurements of the same object, the designer specifies the generation procedure, while the simulator's nuisance model determines the distribution of their differences. The resulting invariances are therefore inherited from the simulator and are only as reliable as its representation of nuisance variation. We study this effect using a calibrated simulator of labeled DNA molecules in solid-state nanopore sensing, in which physical components can be replaced individually. Trained only to identify whether two traces share a barcode identity, the contrastive encoder suppresses nuisance factors that vary independently of identity while retaining factors that affect how identity is expressed in the trace. Replacing the measured noise spectrum with equal-power white noise increases nuisance decodability by $0.249\ R^2$, while replacing the calibrated velocity profiles with constant velocity reduces identity decodability by $0.026\ R^2$. Aggregate task metrics may reveal degradation but not its representational cause. Linear probes on the frozen embedding provide this distinction. We turn this observation into a simple diagnostic that requires no real data and can detect when simulator misspecification changes the representation learned by an encoder.
Chat is not available.
Successful Page Load