arXiv Artificial Intelligence

Building Legal Reward Models for Grounding and Abstention

Building Legal Reward Models for Grounding and Abstention

Quick summary

arXiv:2609.14739v1 Announce Type: cross Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (L

Key takeaways

  • arXiv:2609.14739v1 Announce Type: cross Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient.
  • However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings.
  • We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (L

Why it matters

“Building Legal Reward Models for Grounding and Abstention” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗