Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks
Quick summary
arXiv:2609.24955v1 Announce Type: cross Abstract: Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findi
Key takeaways
- arXiv:2609.24955v1 Announce Type: cross Abstract: Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment.
- We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow.
- A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments