MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances
Quick summary
arXiv:2610.12416v1 Announce Type: cross Abstract: Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factoriza
Key takeaways
- arXiv:2610.12416v1 Announce Type: cross Abstract: Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion.
- However, supervision for these capabilities is rarely available jointly at scale.
- Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce.
Why it matters
“MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments