MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
Quick summary
arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and val
Key takeaways
- arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion.
- We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers.
- Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and val
Why it matters
“MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments