Minimax-Optimal Transformer Classification for Functional Data with Dense-Sparse Phase Transition
Abstract
Despite the remarkable empirical success of transformer models in natural language processing, their theoretical foundations for functional data classification remain largely unexplored. This paper takes a first step toward closing this gap by developing a rigorous statistical framework for transformer-based functional classifiers. We show that a transformer architecture with an expanding attention window in the latent representation attains minimax-optimal excess-risk rates up to logarithmic factors in infinite-dimensional functional classification, without relying on conventional dimension-reduction procedures. Our analysis further reveals a dense-to-sparse phase transition: when the sampling frequency exceeds a critical threshold relative to the sample size, the proposed classifier achieves the optimal dense-observation rate; otherwise, we precisely characterize the degradation in convergence under sparse sampling. Extensive simulations and real-data experiments support the theory and demonstrate that transformers provide a competitive and theoretically justified approach to functional classification.