arXiv Artificial Intelligence

Do-JEPA: From Masking to Intervention in Latent World Models

Do-JEPA: From Masking to Intervention in Latent World Models

Quick summary

arXiv:2609.37378v1 Announce Type: cross Abstract: Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $\Delta z=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, h

Key takeaways

  • arXiv:2609.37378v1 Announce Type: cross Abstract: Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it.
  • Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens.
  • From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $\Delta z=z^{a}-z^{a_{\varnothing}}$ between the two latent futures.

Why it matters

“Do-JEPA: From Masking to Intervention in Latent World Models” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗