arXiv Artificial Intelligence

MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects

MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects

Quick summary

arXiv:2608.29477v1 Announce Type: cross Abstract: Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability. When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects. We present MUDDLE, a controlled benchmark that separates them. MUDDLE uses 270 human-annotated questions, each tied to a single source document, and instantiate

Key takeaways

  • arXiv:2608.29477v1 Announce Type: cross Abstract: Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability.
  • When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects.
  • We present MUDDLE, a controlled benchmark that separates them.

Why it matters

“MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗