MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
Quick summary
arXiv:2608.29477v1 Announce Type: cross Abstract: Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability. When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects. We present MUDDLE, a controlled benchmark that separates them. MUDDLE uses 270 human-annotated questions, each tied to a single source document, and instantiate
Key takeaways
- arXiv:2608.29477v1 Announce Type: cross Abstract: Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability.
- When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects.
- We present MUDDLE, a controlled benchmark that separates them.
Why it matters
“MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments