arXiv Artificial Intelligence

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

Quick summary

arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under

Key takeaways

  • arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
  • Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback.
  • We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges.

Why it matters

“Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗