WCM: World-Cognition Model for Generalizable Human-Robot Interaction
Quick summary
arXiv:2607.22999v2 Announce Type: replace-cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, are mainly optimized for instruction execution, leaving users with little visibility into why an action is chosen and few mechanisms to redirect, correct, or teach the robot through interaction. To solve this problem, we present the World-Cognition Model (WCM), a human-centered embodied agen
Key takeaways
- arXiv:2607.22999v2 Announce Type: replace-cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks.
- Current robot-control paradigms, including vision-language-action policies and world-model-based planners, are mainly optimized for instruction execution, leaving users with little visibility into why an action is chosen and few mechanisms to redirect, correct, or teach the robot through interaction.
- To solve this problem, we present the World-Cognition Model (WCM), a human-centered embodied agen
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments