arXiv Artificial Intelligence

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

Quick summary

arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that stand

Key takeaways

  • arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse.
  • While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood.
  • In this work, we investigate the representation-level mechanisms underlying alignment degradation.

Why it matters

“RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗