Distilling Tabular Foundation Models to Simple Decision Trees
Abstract
With recent advances in tabular data modeling, specialized foundation models leveraging in-context-learning can regularly outperform tree-based models on tabular benchmarks, particularly for small i.i.d. datasets. These advances in predictive accuracy, however, can come at a cost to model size and explainability. On the other extreme of compactness and interpretability is a single sparse decision tree. Recent work has established that particularly well-optimized trees can achieve reasonable performance despite their compact nature. In high risk, interpretability-critical domains (or domains where especially fast inference time is needed), there is an opportunity to combine these methodologies by using a foundation model as a teacher for knowledge distillation. We benchmark decision trees with a foundation model teacher, demonstrating how the model size and performance tradeoff changes with access to foundation models at training time. We also introduce modifications to optimization approaches for decision trees to facilitate distillation.