Do Vision-Language Models Understand Construction Hazards Beyond Object Recognition?
Abstract
Construction sites present complex and dynamic environments in which visual hazards arise not only from the presence of unsafe objects but also from spatial relationships, human–equipment interactions, and violations of safety rules. While recent vision-language models (VLMs) demonstrate strong capabilities in visual recognition and multimodal reasoning, their ability to understand safety-critical situations in construction environments remains underexplored. This study investigates whether VLMs can move beyond object recognition toward meaningful construction-hazard understanding. Using images and safety-related annotations from a construction-site dataset, we formulate a hierarchical evaluation framework spanning three levels of visual understanding: object perception, safety-violation recognition, and contextual hazard reasoning. We evaluate VLMs with different model sizes and computational requirements to examine the relationship between multimodal reasoning capability and deployment efficiency. Particular attention is given to scenarios involving workers, personal protective equipment, heavy equipment, and hazardous spatial configurations. Beyond overall predictive performance, we analyze failure cases in which models correctly identify relevant objects but fail to infer the associated safety risk. By considering both safety reasoning and computational accessibility, this work aims to provide insights into the feasibility of deploying multimodal AI for construction safety monitoring in resource-constrained environments. The study contributes toward accessible and socially beneficial AI systems capable of supporting safer construction practices across diverse operational settings.