arXiv Artificial Intelligence

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Quick summary

arXiv:2609.39575v1 Announce Type: cross Abstract: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support tra

Key takeaways

  • arXiv:2609.39575v1 Announce Type: cross Abstract: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion.
  • To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts.
  • Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities.

Why it matters

“ECHO-G: Embodied Co-speech Humanoid mOtion Generation” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗