arXiv Artificial Intelligence

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Quick summary

arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior. Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance. We call the gap between what static audits certify and what an intervention can reach the audit gap, and we show it is realizable: one can build

Key takeaways

  • arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones.
  • But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior.
  • Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance.

Why it matters

“When Behavioral Safety Evaluation Fails: A Representation-Level Perspective” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗