arXiv Artificial Intelligence

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

Quick summary

arXiv:2610.08401v1 Announce Type: cross Abstract: While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis ac

Key takeaways

  • arXiv:2610.08401v1 Announce Type: cross Abstract: While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context.
  • In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective.
  • \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces.

Why it matters

“GeoPID: Decomposing and Steering Visual Information in Vision-Language Models” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗