MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Quick summary
arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However, existing benchmarks such as BrowseComp primarily focus on textual content, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities. Unlike prior work, MM-BrowseComp incorporates visual prompts and necessitates the extra
Key takeaways
- arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search.
- However, existing benchmarks such as BrowseComp primarily focus on textual content, overlooking the prevalence of multimodal content.
- To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities.
Why it matters
“MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments