arXiv Artificial Intelligence

Identity-Aware Human-Object Interaction Motion Captioning

Identity-Aware Human-Object Interaction Motion Captioning

Quick summary

arXiv:2608.20690v1 Announce Type: cross Abstract: Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI motion. For example, the model generates "Sub_ID lifts the chair" rather than "A person lifts the chair". F

Key takeaways

  • arXiv:2608.20690v1 Announce Type: cross Abstract: Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity.
  • To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task.
  • This task requires each generated caption to specify both the subject identity and the corresponding HOI motion.

Why it matters

“Identity-Aware Human-Object Interaction Motion Captioning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗