arXiv Artificial Intelligence

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets

Quick summary

arXiv:2609.38909v1 Announce Type: cross Abstract: Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit buil

Key takeaways

  • arXiv:2609.38909v1 Announce Type: cross Abstract: Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it.
  • Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs.
  • We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit buil

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗