arXiv Artificial Intelligence

Action Shaping: Policies Absorb What They Can Express

Action Shaping: Policies Absorb What They Can Express

Quick summary

arXiv:2609.32752v2 Announce Type: replace Abstract: Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its

Key takeaways

  • arXiv:2609.32752v2 Announce Type: replace Abstract: Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy.
  • The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem.
  • Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee.

Why it matters

“Action Shaping: Policies Absorb What They Can Express” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗