SportD: How do VLMs physically strategize?
Quick summary
arXiv:2607.14616v5 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next. Models on average select the optimal action around 27% of the time, less often than the professional players, and capture markedly less of the valu
Key takeaways
- arXiv:2607.14616v5 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one?
- We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions.
- We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments