Microstructure Descriptor Fields as Supervision for Scientific Images
Abstract
Scientific images are often governed less by exact pixel fidelity than by preservation of microstructure: morphology, connectivity, anisotropy, interface complexity, phase organization, and characteristic length scales. Yet modern image-learning objectives, including masked image modeling and reconstruction losses, supervise models in pixel, token, or latent feature space rather than in the space of scientific observables. We introduce MICROFIELD, a supervision framework that represents a scientific image by a field of microstructure descriptors computed over spatial regions. A descriptor field can be induced on a uniform grid, a quadtree, an octree, a learned partition, or a hierarchy produced by a high-resolution image model. Rather than proposing a new masked autoencoder architecture, MICROFIELD contributes an objective-level mechanism: it trains models to preserve local descriptors, cross-scale descriptor transitions, and level-wise descriptor distributions. Because descriptor fields are defined over region systems rather than model architectures, MICROFIELD applies to uniform grids, adaptive partitions, and hierarchical models alike. The framework is compatible with self-supervised pretraining, supervised learning, and reconstruction; in this paper, we focus on masked self-supervised pretraining as the primary instantiation. MICROFIELD provides a practical path toward physics-inspired representation learning for scientific images without requiring explicit governing equations or domain labels.