Guard Models Are Overconfident Where Base Models Are Uncertain
Quick summary
arXiv:2609.36477v1 Announce Type: cross Abstract: Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressi
Key takeaways
- arXiv:2609.36477v1 Announce Type: cross Abstract: Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions.
- We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections.
- Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressi
Why it matters
“Guard Models Are Overconfident Where Base Models Are Uncertain” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments