Beyond Low Error: Lessons from Tumor Growth Simulation
Abstract
A growing body of work uses a pharmacokinetic–pharmacodynamic tumor-growth simulator to train and evaluate models that predict tumor volume under different treatment plans. In this work, we consider whether low error when predicting simulated tumors provides evidence that a model has learned how tumor volume changes with treatment. We evaluate thirteen published models, two controls that ignore the proposed plan, and Masked Growth Network (MGN), a model matched to known properties of the simulated dynamics. We compare evaluation with widely used sliding plans, which contain one treatment dose on one forecast day, with random plans that vary treatments independently across all forecast days. We show that across data generation settings, proposed sliding plans account for only a limited share of variation in tumor outcomes. In the setting of our main results, proposed plans account for only 5.5% of outcome variation under sliding plans, compared with 21.5% under random plans. With sliding treatment plans, persistence, which repeats the last observed tumor volume, outperforms ten of thirteen published models. We also assess performance separately within deciles defined by tumor volume immediately before the proposed plan begins. With random treatment plans, the group with the largest tumors accounts for 86.8%–98.3% of squared prediction error across MGN and the thirteen published models. In the five smallest-tumor deciles, the published models' predictions reflect how simulated outcomes differ between plans no better than a reference returning the same tumor volume for every plan. In contrast, the model inspired by simulator dynamics improves on this reference across all deciles. It also follows the simulator's average local dose response. Together, these results show that low prediction error alone does not establish that a model has learned how treatment affects tumor volume. Simulator-based evaluations should be interpreted in light of the treatment assignments observed during training, the plans proposed at evaluation, and the errors emphasized by the metric.