arXiv Artificial Intelligence

Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

Quick summary

arXiv:2608.09654v1 Announce Type: new Abstract: GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates. This task faces a persistent dual obstacle: conventional grounding models lack the semantic richness to interpret abstract instructions, while end-to-end MLLMs suffer from coordinate hallucinations caused by deficient fine-grained perception. We propose a regression-free fr

Key takeaways

  • arXiv:2608.09654v1 Announce Type: new Abstract: GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots.
  • The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates.
  • This task faces a persistent dual obstacle: conventional grounding models lack the semantic richness to interpret abstract instructions, while end-to-end MLLMs suffer from coordinate hallucinations caused by deficient fine-grained perception.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗