Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
Quick summary
arXiv:2609.17043v1 Announce Type: cross Abstract: Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents. Whether this holds at the level of individual reasoning steps remains largely unexamined. We investigate this across three standard multi-hop QA benchmarks and find that failures decompose into two distinct modes: retrieval failures, where the needed passage was not retrieved, and extraction failure
Key takeaways
- arXiv:2609.17043v1 Announce Type: cross Abstract: Multi-hop question answering requires combining information from multiple documents to answer complex questions.
- These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents.
- Whether this holds at the level of individual reasoning steps remains largely unexamined.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments