arXiv Artificial Intelligence

Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence

Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence

Quick summary

arXiv:2608.12645v2 Announce Type: replace Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability un

Key takeaways

  • arXiv:2608.12645v2 Announce Type: replace Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
  • Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback.
  • We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges.

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗