arXiv Artificial Intelligence

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

Quick summary

arXiv:2609.17043v1 Announce Type: cross Abstract: Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents. Whether this holds at the level of individual reasoning steps remains largely unexamined. We investigate this across three standard multi-hop QA benchmarks and find that failures decompose into two distinct modes: retrieval failures, where the needed passage was not retrieved, and extraction failure

Key takeaways

  • arXiv:2609.17043v1 Announce Type: cross Abstract: Multi-hop question answering requires combining information from multiple documents to answer complex questions.
  • These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents.
  • Whether this holds at the level of individual reasoning steps remains largely unexamined.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗