arXiv Artificial Intelligence

CheatBench: Measuring Reward Gaming in AI Agents

CheatBench: Measuring Reward Gaming in AI Agents

Quick summary

arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating

Key takeaways

  • arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended.
  • In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems.
  • As agents become more capable, this behavior could pose increasingly serious risks.

Why it matters

“CheatBench: Measuring Reward Gaming in AI Agents” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗