arXiv Artificial Intelligence

AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation

AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation

Quick summary

arXiv:2609.22332v2 Announce Type: cross Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act. Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance. We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance He

Key takeaways

  • arXiv:2609.22332v2 Announce Type: cross Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act.
  • Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance.
  • We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance He

Why it matters

“AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗