Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
Quick summary
arXiv:2609.06649v1 Announce Type: cross Abstract: Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. As an alternative, we propose studying emergent misalignment from iterative DPO, which preserves important properties of RLVR while reducing costs and enabling training on popular finetuning APIs. In practice, we find that training GPT-4.1
Key takeaways
- arXiv:2609.06649v1 Announce Type: cross Abstract: Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models.
- Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models.
- As an alternative, we propose studying emergent misalignment from iterative DPO, which preserves important properties of RLVR while reducing costs and enabling training on popular finetuning APIs.
Why it matters
“Inducing Emergent Misalignment from Reward Hacks with Iterative DPO” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments