arXiv Artificial Intelligence

(How) Do MLLMs Report Bistable Images Like Humans?

(How) Do MLLMs Report Bistable Images Like Humans?

Quick summary

arXiv:2609.13254v1 Announce Type: cross Abstract: Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both

Key takeaways

  • arXiv:2609.13254v1 Announce Type: cross Abstract: Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans.
  • We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it.
  • Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation.

Why it matters

“(How) Do MLLMs Report Bistable Images Like Humans?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗