Binarization Flattens the Score Space
Quick summary
arXiv:2609.35797v1 Announce Type: cross Abstract: Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch
Key takeaways
- arXiv:2609.35797v1 Announce Type: cross Abstract: Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture.
- However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion.
- We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Binarization Flattens the Score Space” may reshape data collection, model training, output accountability and market access.

Member comments