SportD: How do VLMs physically strategize?
Jasin Cekinmez ⋅ Addison J. Wu ⋅ Haotian Xia ⋅ Kyumin A Shim ⋅ Anay Putty ⋅ Jinglin Xiao ⋅ Zhuohan Liu ⋅ Leo Liu ⋅ Weining Shen
Abstract
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action. Models only select the optimal action around 30% of the time, even less frequently than humans do. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 83-92% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed ($\rho=+0.30$ to $+0.52$), despite no such relationship in the ground truth ($\rho=-0.08$). The conservatism therefore reflects a mis-calibration of value. Replacing a single deliberation sentence with one that steers toward risk lifts the frontier models towards the real players' skills. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.
Chat is not available.
Successful Page Load