arXiv Artificial Intelligence

SportD: How do VLMs physically strategize?

SportD: How do VLMs physically strategize?

Quick summary

arXiv:2607.14616v5 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next. Models on average select the optimal action around 27% of the time, less often than the professional players, and capture markedly less of the valu

Key takeaways

  • arXiv:2607.14616v5 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one?
  • We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions.
  • We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗