GANDR: Claim Auditing for Verifiable Legal Answer Generation
Quick summary
arXiv:2609.10293v1 Announce Type: cross Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-agent system in which a Drafter writes an answer in
Key takeaways
- arXiv:2609.10293v1 Announce Type: cross Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites.
- Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well.
- Closing this gap requires both a system built for per-claim verification and an evaluation that measures it.
Why it matters
“GANDR: Claim Auditing for Verifiable Legal Answer Generation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments