Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs
Quick summary
arXiv:2608.00076v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction? Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning. We formulate modality attribution as a complementary explainability objecti
Key takeaways
- arXiv:2608.00076v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text.
- While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction?
- Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments