WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
Quick summary
arXiv:2606.18847v2 Announce Type: replace Abstract: To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions,
Key takeaways
- arXiv:2606.18847v2 Announce Type: replace Abstract: To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions.
- Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments.
- We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance.
Why it matters
“WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments