UrduGrad Schemas: A Benchmark for Urdu Commonsense Reasoning
Abstract
Commonsense reasoning is essential for natural language understanding, yet bench marks for low-resource languages remain limited. To address this gap, we in troduce UrduGrad, the first dedicated Winograd-style benchmark for Urdu, de veloped through a human-in-the-loop process to evaluate coreference resolution and commonsense reasoning. To establish baseline performance, we evaluate two instruction-tuned LLMs, Phi-3.5-mini and Qwen2.5-3B, under zero- and few-shot settings using accuracy. Both models perform near chance level in the zero-shot set ting (51.48% and 51.31%), while few-shot prompting yields modest improvements to 56.37% and 52.12%, respectively. These findings indicate that the evaluated models struggle with Urdu Winograd-style commonsense reasoning and establish UrduGrad as a foundation for future multilingual and low-resource NLP research.