Rethinking ECG Tokenization: A Unified Taxonomy and Clinically-Grounded ECG Foundation Model
Abstract
Deep learning models for electrocardiograms (ECGs) commonly process raw waveforms directly or partition them into fixed-duration patches. Such representations are not explicitly aligned with cardiac physiology and do not naturally encode the irregular timing of cardiac events. Moreover, existing ECG tokenization strategies lack a common terminology and have received limited systematic comparison. We address these gaps by first introducing the \textbf{ECG Tokenization Taxonomy}, a unified framework for describing and evaluating ECG tokenizers. Within this, we propose \textbf{{ECGTokenizer}}, a physiologically grounded tokenization strategy that decomposes each lead into clinically meaningful waveform regions, including the P wave, QRS complex, T wave, and PR, ST, and TP segments. Building on ECGTokenizer, we introduce \textbf{ECGTFormer}, a compact causal ECG foundation model with continuous-time rotary position encoding and self-supervised V-JEPA pretraining. {ECGTFormer} supports both single- and multi-lead inputs, while its causal architecture enables streaming inference with prefix caching for low-latency deployment. Through a comprehensive empirical study, we systematically evaluate tokenization strategies, region coverage, position encoding, scalar-attribute representation, morphology--attribute fusion, and multi-lead configurations. {ECGTFormer} consistently achieves a strong performance compared to baselines and ablations. We release the code, pretrained checkpoints, and experimental configurations to support reproducible research.