Principled Federated Random Forests for Heterogeneous Data
Abstract
Random Forests (RF) are among the most powerful and widely used predictive models for centralized tabular data, yet few methods exist to adapt them to the federated learning setting. Unlike most federated learning approaches, the piecewise-constant nature of RF prevents exact gradient-based optimization. As a result, existing federated RF implementations rely on unprincipled heuristics: by aggregating decision trees trained independently on clients, they fail to optimize the centralized decision-tree impurity criterion, which constitutes the ground-truth in federated settings, even under simple covariate shifts. We propose FedForest, a new federated RF algorithm for horizontally partitioned data that naturally accommodates diverse forms of client data heterogeneity, from covariate shift to more complex concept shift mechanisms. We prove that our splitting procedure, based on aggregated client statistics, closely approximates the split selected by a centralized algorithm. Moreover, FedForest allows splits on client indicators when beneficial, enabling a non-parametric form of personalization that is absent from prior federated random forest methods. Empirically, we demonstrate that the resulting federated forests closely match centralized performance across heterogeneous benchmarks while remaining communication-efficient.