Rethinking Multi-Image Re-Representation in Multi-Image Understanding
Quick summary
arXiv:2609.39363v1 Announce Type: cross Abstract: Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-re
Key takeaways
- arXiv:2609.39363v1 Announce Type: cross Abstract: Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them.
- We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning.
- We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations.
Why it matters
“Rethinking Multi-Image Re-Representation in Multi-Image Understanding” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments