VisText-VLM2Plastic: A Bilingual Benchmark for Visual Grounding and Environmental Reasoning in Vision-Language Models
Abstract
Vision-language models (VLMs) increasingly support open-ended visual reasoning, yet their behavior in low-resource languages and sustainability-oriented domains remains underexplored. We introduce a controlled English–Bangla plastic-waste VQA benchmark comprising 402 image records derived from 136 base images and 2,412 question–answer pairs across five waste categories and three task types: recognition, visual description, and environmental reasoning. We evaluate GPT-4.1 mini, Gemini 2.5 Flash, Qwen2.5-VL-7B-Instruct, and Gemma-3-4B zero-shot using task-specific answer-quality metrics, bilingual consistency, and image abla- tion. Gemini 2.5 Flash yields the highest recognition point estimate, with 93.5% accuracy and 91.3% macro-F1. Across models, English normalized performance exceeds Bangla by 3.78 points on average, with the largest observed gap for en- vironmental reasoning. Removing the image reduces normalized performance by 66.29 points for recognition, 52.89 for description, and 11.42 for environmental reasoning. The smaller environmental-reasoning ablation gap indicates that, in this benchmark, many environmental questions remain partly answerable from linguistic and domain priors. These results show why answer quality alone is insuf- ficient for multilingual multimodal evaluation and motivate reporting cross-lingual consistency and sensitivity to visual evidence alongside conventional metrics.