CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization
Quick summary
arXiv:2610.02527v1 Announce Type: cross Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five train
Key takeaways
- arXiv:2610.02527v1 Announce Type: cross Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task.
- We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy.
- We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization” may reshape data collection, model training, output accountability and market access.

Member comments