Gondola: Grounded Vision Language Planning for Robotic Manipulation
Quick summary
arXiv:2506.11261v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object gr
Key takeaways
- arXiv:2506.11261v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation.
- However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks.
- To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments