SportD: How do VLMs physically strategize?
Quick summary
arXiv:2607.14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action. Models only select the optimal action around 30% of the time, even less frequently than humans do. Furthermore, they
Key takeaways
- arXiv:2607.14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one?
- We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions.
- We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments