arXiv Artificial Intelligence

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Quick summary

arXiv:2610.07646v1 Announce Type: new Abstract: Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-s

Key takeaways

  • arXiv:2610.07646v1 Announce Type: new Abstract: Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning.
  • Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing.
  • We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals.

Why it matters

“Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗