Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration
Quick summary
arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention ent
Key takeaways
- arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions.
- Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms.
- In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention.
Why it matters
This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Member comments