Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
Enzhi Zhang ⋅ Du Wu ⋅ Rui Zhong ⋅ Cong Ma ⋅ Isaac Lyngaas ⋅ Amir K Ziabari ⋅ Xiao Wang ⋅ Peng Chen ⋅ Tao Luo ⋅ Toshio Endo ⋅ Fumiyoshi Shoji ⋅ Kento Sato ⋅ Kentaro Uesugi ⋅ Takayuki Nonoyama ⋅ Ryuji Kiyama ⋅ Masahiro Yoshida ⋅ Tezuka Masaru ⋅ Tetsuya Ishikawa ⋅ Satoshi Matsuoka ⋅ Masaharu Munetomo ⋅ Mohamed Wahib
Abstract
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose \method{}, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. \method{} couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, \method{} consistently outperforms MAE baselines. It achieves 95.68\% Dice on the 8K${\times}$8K${\times}$28K SpringXCT dataset, improving over the strongest MAE baseline by +9.70 points, and 83.21\% Dice on the $32\text{K}^2$ WSI PAIP dataset, improving by +16.84 points, while providing up to a $24.8\times$ inference speedup
Chat is not available.
Successful Page Load