Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Quick summary
arXiv:2608.12645v2 Announce Type: replace Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability un
Key takeaways
- arXiv:2608.12645v2 Announce Type: replace Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
- Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback.
- We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges.
Why it matters
AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Member comments