arXiv Artificial Intelligence

An Empirical Study of VLM Pipelines for Long-Document QA

An Empirical Study of VLM Pipelines for Long-Document QA

Quick summary

arXiv:2609.29933v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model agentically or as a static pipeline. We study these choices on two long-document QA benchmarks with both frontier API and open-weight VLMs. First, on MMLongBench-Doc our six-tool agent with page, table, figure, and search calls p

Key takeaways

  • arXiv:2609.29933v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts.
  • Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model agentically or as a static pipeline.
  • We study these choices on two long-document QA benchmarks with both frontier API and open-weight VLMs.

Why it matters

“An Empirical Study of VLM Pipelines for Long-Document QA” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗