ECHO-G: Embodied Co-speech Humanoid mOtion Generation
Quick summary
arXiv:2609.39575v1 Announce Type: cross Abstract: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support tra
Key takeaways
- arXiv:2609.39575v1 Announce Type: cross Abstract: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion.
- To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts.
- Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities.
Why it matters
“ECHO-G: Embodied Co-speech Humanoid mOtion Generation” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments