TabBioMed: A Large-Scale Benchmark for Biomedical Tabular Learning
Abstract
Recent advances in machine learning have accelerated biomedical prediction, yet evaluation on structured biomedical data remains fragmented, especially for datasets with high dimensionality, missingness, class imbalance, and small-sample, many-feature regimes. We introduce TabBioMed, a large-scale benchmark of 95 curated public biomedical tabular datasets spanning electronic health records \cite{data2016secondary}, drug response \cite{wu2018moleculenet}, genomics, transcriptomics, proteomics, metabolomics \cite{yang2025mlomics}, single-cell omics \cite{heumos2023best}, and systems biology. TabBioMed unifies these datasets through a standardized, dataset-aware preprocessing and evaluation framework, enabling reproducible comparison across classification, regression, and multi-target tasks. We benchmark 23 models across classical baselines, gradient-boosted trees, neural tabular architectures, and emerging tabular foundation models. Foundation models deliver the strongest family-level predictive performance. However, their gains are regime-dependent: tree-based models and tuned MLPs remain highly competitive, often dominating under inference-time or training-cost constraints. We further show that foundation-model advantages are most meaningful in low-noise biomedical tasks, where small absolute gains translate into large reductions in residual error. Overall, TabBioMed reveals that biomedical tabular learning is strongly context-dependent, with optimal model choice shaped by data regime, task type, and computational budget. By releasing curated datasets, preprocessing code, baselines, and results, TabBioMed provides a reproducible foundation for practical model selection and future methods development in biomedical machine learning, available at https://anonymous.4open.science/r/tabbiomed.