hacktrace: behavior-supervised detection of reward hacking during code generation
Quick summary
arXiv:2610.03055v1 Announce Type: new Abstract: A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection. We introduce HACKTRACE, a behavior-supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enable
Key takeaways
- arXiv:2610.03055v1 Announce Type: new Abstract: A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug.
- Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail.
- We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection.
Why it matters
“hacktrace: behavior-supervised detection of reward hacking during code generation” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments