Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Quick summary
arXiv:2609.37938v1 Announce Type: cross Abstract: Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions der
Key takeaways
- arXiv:2609.37938v1 Announce Type: cross Abstract: Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination.
- Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers.
- We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments