arXiv Artificial Intelligence

Grounded World Model: Latent Planning with Language Goals

Grounded World Model: Latent Planning with Language Goals

Quick summary

arXiv:2604.11751v2 Announce Type: replace-cross Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding s

Key takeaways

  • arXiv:2604.11751v2 Announce Type: replace-cross Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks.
  • We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone.
  • Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model.

Why it matters

“Grounded World Model: Latent Planning with Language Goals” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗