arXiv Artificial Intelligence

Query-Driven Multimodal Information Extraction from Long Documents

Query-Driven Multimodal Information Extraction from Long Documents

Quick summary

arXiv:2608.22214v1 Announce Type: new Abstract: In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing evidence regions, rather than outputting query-specific textual attribute values and corresponding images. To address this gap, we propose query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxe

Key takeaways

  • arXiv:2608.22214v1 Announce Type: new Abstract: In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone.
  • However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing evidence regions, rather than outputting query-specific textual attribute values and corresponding images.
  • To address this gap, we propose query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxe

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗