Vision-Language Models are Fragile Multilingual Associators
Quick summary
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, wi
Key takeaways
- arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes.
- Whether these associations or concept bindings remain stable when the language of the input changes is unexplored.
- We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments