Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?
Quick summary
arXiv:2605.22148v3 Announce Type: replace Abstract: A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar per skill. The answer is exact: a judge scoring failures as passes at rate $(1-\tau)/2$ or above retires nothing, at any sample size, for eviction margin $\tau$. Audits find that machinery is rarely built: LLM-written skills are worth $+0.0$ percentage points (pp) against a no-skill control, human-written ones $+16.2$pp. Unmaintained, a library enters \emph{library drift}, growing until injecting a skill
Key takeaways
- arXiv:2605.22148v3 Announce Type: replace Abstract: A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar per skill.
- The answer is exact: a judge scoring failures as passes at rate $(1-\tau)/2$ or above retires nothing, at any sample size, for eviction margin $\tau$.
- Audits find that machinery is rarely built: LLM-written skills are worth $+0.0$ percentage points (pp) against a no-skill control, human-written ones $+16.2$pp.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments