Co-speech HumanoidIntermediate
ECHO-G: Embodied Co-speech Humanoid mOtion Generation
ECHO-G is a preprint framework that generates full-body speech-synchronized motion for humanoid robots from audio and timed transcripts. It models motion directly in robot space with a Speech-Grounded Diffusion Transformer and reports a Fréchet Gesture Distance of 2.278 versus 4.976 for EMAGE+GMR, with 5.96 ms/frame inference time. This matters for making expressive, deployable co-speech gestures on humanoids without human-motion retargeting.