Paired Significance Testing for Active Experimental Design Comparisons in Wildfire Model Calibration
Rhea Senan
Abstract
Autonomous and adaptive experimentation systems are increasingly proposed for the physical sciences, often on the intuition that choosing which measurement to take next, rather than measuring at random or on a fixed schedule, should make better use of a limited observation budget. This project tests this intuition directly on a physical-parameter calibration problem: a stochastic wildfire simulator with three unknown physical parameters, calibrated from field observations under a fixed budget. This project compares an active, uncertainty-seeking sensing policy against random sampling and a fixed monitoring-grid sweep, evaluating all three on held-out calibration loss, burned-area agreement with the true fire, and parameter-recovery error. An initial run with 5 independent fire draws suggested active sensing had a real edge. A properly-powered rerun with 12 independent fire draws did not reproduce it: none of the three strategies are statistically distinguishable on any of the three metrics (paired Wilcoxon signed-rank test, $p \geq 0.456$ pooled over the full budget sweep). A second scenario with a different wind direction, ignition point, and terrain map does show a significant active-sensing advantage on calibration loss ($p=0.0009$), so the effect is not universally absent, it is scenario-dependent and easy to mistake for a general property of the method from a single underpowered pilot. This project argues that claims about adaptive experimental design improving physical-parameter calibration should be treated as empirical, checkable questions specific to a deployment scenario, not assumed a prior, and it reports the initial false positive as a worked example of why.
Chat is not available.
Successful Page Load