When Should Resource-Constrained AI Abstain? Confidence Calibration for Reliable Multimodal Decisions
Abstract
Efficient multimodal inference often uses confidence to decide when computation can stop, but raw neural confidence can be poorly calibrated. In resource-constrained settings, an overconfident early exit may therefore save compute by returning an avoidable error. We study a complementary question: when should the system refuse to exit or abstain? We build a seeded controlled simulator with 4,000 benchmark-inspired multimodal task profiles, equally divided across ScienceQA-, VizWiz-, TextVQA-, and AI2D-like workloads. The simulator contains three progressively more capable evidence stages and an intentionally overconfident raw score. We compare raw early exit, global temperature scaling, stage-specific temperature scaling, calibrated abstention, and a fixed stage-2 baseline. On a held-out 3,200-profile evaluation set, global temperature scaling reduces ECE from 0.171 to 0.029 at essentially unchanged computation; stage-specific scaling further reduces ECE to 0.021 and lowers premature-exit harm from 0.123 to 0.115. Adding abstention increases selective accuracy from 0.637 to 0.649 at 79.4% coverage. The result is a reliability-coverage trade-off rather than a universal accuracy gain. We present the study explicitly as a reproducible simulation, not measured VLM accuracy or hardware performance, and use it to motivate calibrated control signals for low-resource multimodal AI.