PINNBench: A Benchmark and Evaluation Study of Training Policy Selection in Hybrid PINN-Operator Solvers
Abstract
We introduce PINNBench, a benchmark and evaluation study for training-policy selection in hybrid PINN-operator PDE solvers, contributed to the NeurIPS 2026 Evaluations and Datasets (ED) Track. PINNBench comprises 1,539 recorded runs and probes across 13 PDE configurations drawn from 8 equation families, 3–5 training policies per PDE, and 10 seeds per condition, with seven standardized metrics, Wilcoxon signed-rank tests, paired bootstrap 95% confidence intervals, Cohen's d, and Benjamini–Hochberg FDR correction. Baseline evaluation of 5 routing methods establishes three findings that bear directly on how PDE solvers should be evaluated. (i) No single policy dominates: hybridfull wins 7/13, hybridnogate 4/13, pinnonly 2/13. (ii) Accuracy and regret dissociate: under relative L2, a physics-loss heuristic and a learned Random Forest router reach identical 62% top-1 accuracy, but the heuristic suffers mean regret 0.0850 versus 0.0011 — a 77x gap. Evaluation by accuracy alone is misleading for PDE policy routing. (iii) Causal-loss effectiveness is stage-mediated: 7/9 PDEs improve with pretraining (+5.9% mean) while 5/9 degrade without it (-4.2%). A new leave-one-PDE-family-out evaluation shows that the strongest family-level generalizer uses no PDE-identity features, addressing the concern that PDE-aware routing memorizes PDE identity. All 1,539 result JSONs, PDE implementations, policy configurations, statistical pipeline, and routing infrastructure are released as open-source artifacts.