arXiv Artificial Intelligence

Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks

Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks

Quick summary

arXiv:2604.18510v2 Announce Type: replace-cross Abstract: Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feature abliteration in Qwen2.5-7B and Llama-3.1-8B using harmfulness, capability, self-audit, safety reflection, representation similarity, and refusal-direction repair. All three routes reach near-ceiling harmfulness, but SF

Key takeaways

  • arXiv:2604.18510v2 Announce Type: replace-cross Abstract: Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes.
  • We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feature abliteration in Qwen2.5-7B and Llama-3.1-8B using harmfulness, capability, self-audit, safety reflection, representation similarity, and refusal-direction repair.
  • All three routes reach near-ceiling harmfulness, but SF

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗