Redesigning and Auditing Deep Research Writing for Faithful Reports
Quick summary
arXiv:2608.28643v1 Announce Type: cross Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, we find that strong DR pipelines can omit key evidence and misattribute claims even when their rubric scores remain stable. We then propose CLAIMWRITER, a hierarchical claim-based writer that extracts source fac
Key takeaways
- arXiv:2608.28643v1 Announce Type: cross Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports.
- We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence.
- Using CLAIMPROBE, we find that strong DR pipelines can omit key evidence and misattribute claims even when their rubric scores remain stable.
Why it matters
“Redesigning and Auditing Deep Research Writing for Faithful Reports” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments