When Should Vision-Language Models Abstain? Evaluating Self-Reported Uncertainty Under Visual Degradation in Industrial Inspection
Abstract
Vision language models (VLMs) allow for quick zero-shot industrial visual inspection, but their reliability under degraded environments is less understood. We conduct a study on 102 source images from the MVTec Anomaly Detection (MVTec AD) screw category. Each source image is evaluated in five visual conditions: clean, mild blur, heavy blur, mild brightness reduction, and heavy brightness reduction. We test two models, Gemini 3.5 Flash-Lite and Gemma 4 26B, using the same prompt that requires a NORMAL, DEFECTIVE, or UNCERTAIN decision with a self reported confidence score. Across 1,020 model evaluations (102 images × 5 conditions × 2 models), degradation causes losses in classification accuracy and defect recall, but reported confidence does not reduce proportionally. Under heavy blur, Gemini 3.5 Flash-Lite accuracy falls from 81.4% to 61.8% but mean reported confidence still remains around 94.8%. Gemma 4 26B displays significantly more abstention under severe darkness, yet its calibration error reaches 0.390. This suggests a huge gap between identifying objects in degraded images and expressing that uncertainty quantitatively.