Pathway-Aligned Regulator Tokens for Interpretable Spatial Gene Expression Prediction from Histologyatial Transcriptomics Prediction from Histology
Abstract
Predicting spatial gene expression from histopathology images enables transcriptomic analysis on archival H&E samples where matched spatial transcriptomics data is unavailable. Gene expression is organized through cyclic regulatory circuits: transcription factor complexes are themselves products of regulated genes, aggregating upstream gene activity and modulating co-expressed downstream programs (e.g., the ISGF3 complex coordinating hundreds of interferon-stimulated genes through shared cis-regulatory elements). Existing approaches to histology-to-expression prediction, both deterministic and generative, employ generic architectures that ignore this structure. We introduce regulator attention, a mechanism that encodes this cycle by pooling gene representations into a small set of content-dependent tokens via a learned selector and routing them back through cross-attention to modulate gene-level features, in contrast to learnable token arrays whose identity is independent of the input. Within a flow matching framework conditioned on pathology foundation model embeddings, regulator attention consistently outperforms deterministic and generative baselines on correlation-based metrics across four tissue types spanning human cancer, normal tissue, and cross-species samples, with substantial gains on highly variable genes. Beyond predictive accuracy, 87.5\% of regulator tokens align with known biological pathways without supervision. Replacing them with content-independent learned queries eliminates this specialization and degrades predictive performance, isolating content-dependent token generation as the causal mechanism.