arXiv Artificial Intelligence

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

Quick summary

arXiv:2606.16914v2 Announce Type: replace Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal.

Key takeaways

  • arXiv:2606.16914v2 Announce Type: replace Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best.
  • We measure what that omission hides.
  • In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality.

Why it matters

“Greed Is Learned: Visible Incentives as Reward-Hacking Triggers” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗