Random-Effects Centroids for Domain Generalization
Abstract
Domain generalization methods train on multiple source environments but usually deploy a single pooled predictor. We study a different deployment object: one head per source environment on a shared representation, combined by fixed weights when the test environment label is unavailable. Under a same-meta random-effects model, where environment-specific population minimizers are iid deviations around a shared meta-mean, the optimal fixed simplex rule is the uniform centroid. The same analysis separates this centroid from pooled ERM: under a common feature second moment, pooled ERM is sample-weighted, while the centroid is environment-weighted, giving a closed-form gap controlled by sample-weight imbalance. We introduce REC (Random-Effects Centroid), a training objective that fits source-specific experts while directly scoring their averaged deployment head. In the common-second-moment case, the centroid term controls the trained-centroid discrepancy and the dispersion of environment-specific optima. Synthetic experiments isolate the population estimator effect and learned-representation behavior; WILDS experiments show that REC is competitive with established DG baselines and that the centroid term is necessary for a useful averaged predictor.