arXiv Artificial Intelligence

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration

Quick summary

arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention ent

Key takeaways

  • arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions.
  • Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms.
  • In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗