When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Quick summary
arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior. Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance. We call the gap between what static audits certify and what an intervention can reach the audit gap, and we show it is realizable: one can build
Key takeaways
- arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones.
- But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior.
- Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance.
Why it matters
“When Behavioral Safety Evaluation Fails: A Representation-Level Perspective” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments