WAM-OPD: On-Policy Distillation for World Action Models
Quick summary
arXiv:2608.22364v1 Announce Type: new Abstract: World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. We study whether on-policy distillation (OPD) can repair such a student without requiring sparse-reward reinforcement learning. We introduce WAM-OPD, a deployment-consistent post-training recipe for a video-first WAM. The student acts in the environment and therefore determines the history distribution. A frozen teach
Key takeaways
- arXiv:2608.22364v1 Announce Type: new Abstract: World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data.
- We study whether on-policy distillation (OPD) can repair such a student without requiring sparse-reward reinforcement learning.
- We introduce WAM-OPD, a deployment-consistent post-training recipe for a video-first WAM.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “WAM-OPD: On-Policy Distillation for World Action Models” may reshape data collection, model training, output accountability and market access.

Member comments