Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
Quick summary
arXiv:2610.11766v1 Announce Type: cross Abstract: Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding
Key takeaways
- arXiv:2610.11766v1 Announce Type: cross Abstract: Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR).
- Yet the same non-harmful outcome can arise for very different reasons.
- A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely.
Why it matters
This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Member comments