HumblePRM: Improving VLM Abstention Due to Uncertainty in Ambiguous Safety Situations through Process Reward Modeling
Abstract
Trustworthy deployment of Vision Language Models (VLMs) requires that a model is aware when a situation is too ambiguous to respond confidently. Standard VLM post-training uses outcome reward models (ORMs) that supervise only the final decision, which tends to reward a confident answer regardless of possible ambiguity present in an image. This is an even more critical problem for safety-related images, as a model that cannot admit it does not know if an image is safe will often guess, leading to unpredictable safety behavior. Existing in-training safety reasoning methods assume that a confident verdict exists, while methods that target uncertainty more directly tend to detect it post-hoc without building new behavior into the policy. We instead split multimodal reasoning into two stages: perception, where the model must factually describe the scene, and calibration, where the model must judge whether the observed evidence is sufficient to commit to a conclusion. We introduce HumblePRM, a lightweight process reward model (PRM) that scores each stage independently, rewarding agreement between evidence and expressed confidence rather than outcome correctness alone. We find that training a model with HumblePRM substantially increases the rate at which the model abstains specifically due to ambiguity on genuinely unclear inputs, while leaving accuracy on clear images roughly the same as standard outcome reward supervision. We also find that the performance improvement from HumblePRM transfers to Gemma 4 E4B, a larger, more popular edge model. Overall, our findings show that process reward supervision is a promising route to training more trustworthy VLMs whose confidence is better grounded in visual evidence.