Flood Damage Assessment Without the Cloud: What a 2B-Parameter Vision-Language Model Can Be Trusted With
Abstract
Floods are among the most frequent and damaging climate driven disasters, and a warming climate is making them more frequent, more severe, and harder to predict. After a flood, aerial drones can survey the damage in hours, but someone still has to turn thousands of images into decisions: which roads are passable, which buildings are inundated, and where help should go first. The vision-language models that answer these questions most accurately are also the hardest to deploy-they are large, and the response teams who most need them are the least likely to have the hardware or connectivity to run them. We ask whether a small, 2B-parameter model, cheap enough to run on a field laptop, can do this job well enough to be useful. Fine-tuned on FloodNet-VQA, a visual question answering benchmark, it reaches 0.825 accuracy at 0.43 seconds per question, outperforming zero-shot GPT-4.1 and a 30B model at a fraction of the latency. Accuracy is distributed unevenly across question types: it is reliable on the condition questions that dominate early triage and unreliable on counting, and when asked something outside its training it answers confidently rather than signalling that it cannot. For responders, knowing where a model can and cannot be trusted matters more than a point of average accuracy, and models this small put flood-damage assessment within reach of the agencies that actually respond.