arXiv Artificial Intelligence

Cat-DPO: Category-Adaptive Safety Alignment

Cat-DPO: Category-Adaptive Safety Alignment

Quick summary

arXiv:2604.17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair. The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories. We cast safety alignment as a per-category constrained optimization problem and derive Cat-DPO, a direct-preference-optimizatio

Key takeaways

  • arXiv:2604.17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones.
  • Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair.
  • The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗