ECHO-G: Embodied Co-speech Humanoid mOtion Generation
ECHO-G is a preprint framework that generates full-body speech-synchronized motion for humanoid robots from audio and timed transcripts. It models motion directly in robot space with a Speech-Grounded Diffusion Transformer and reports a Fréchet Gesture Distance of 2.278 versus 4.976 for EMAGE+GMR, with 5.96 ms/frame inference time. This matters for making expressive, deployable co-speech gestures on humanoids without human-motion retargeting.
Can jointly conditioning full-body humanoid motion generation on speech audio and timed transcripts, while generating directly in robot space, improve co-speech characteristics and physical robot deployment compared with human-motion generation plus retargeting or learned mapping?
Full-body co-speech motion for humanoids must align speech prosody and linguistic content while respecting robot embodiment, but existing methods mostly generate human skeletons and retarget them, focus on upper-body motion, omit utterance-specific audio-text conditioning, or lack paired audio-text-robot training data and a common benchmark.
Human co-speech generators such as EMAGE and GestureLSM synthesize human SMPL-X motion, which is then retargeted to robots via GMR or learned mappings. Robot-oriented methods like RoboGesture and PhysDrift address speech-driven robot gestures but are often upper-body or do not fully integrate frame-level audio with token-level transcript structure.
SGDiT is a rectified-flow diffusion transformer trained directly on 39-dimensional Unitree G1 motion features. It adds frame-aligned wav2vec2 acoustic features to motion tokens and retrieves token-level Qwen3.5-4B transcript embeddings through global-local cross-attention with a Gaussian temporal prior. The authors build a BEAT2-derived robot-space dataset with 14,987 training and 3,242 validation clips retargeted to a 29-DoF Unitree G1 and quality filtered; training uses 12 transformer blocks, AdamW, 63,000 steps on RTX 4090 hardware, and inference runs on Jetson AGX Orin for physical deployment through a fixed SONIC whole-body tracker.
On the BEAT2 speaker-held-out evaluation, ECHO-G achieved FGD 2.278, ΔDiv 0.320, MM 1.786, and ΔBA 0.063, outperforming EMAGE+GMR (FGD 4.976, ΔDiv 0.749, MM 0, ΔBA 0.172) and GestureLSM+GMR (FGD 5.008, ΔDiv 0.561, MM 1.016, ΔBA 0.158). It also recorded 5.96 ms/frame end-to-end time versus 21.3 ms/frame for EMAGE+GMR and 20.5 ms/frame for GestureLSM+GMR. Human-Retargeted had better robot-quality scores on jerk gap, foot-ground error, and contact sliding (8.001, 0.003 m, 0.050 m/s) than ECHO-G (9.039, 0.008 m, 0.052 m/s). In ablations, joint audio-text conditioning outperformed audio-only and text-only on FGD, ΔDiv, MM, and ΔBA. A 45-participant video study rated ECHO-G overall 3.49/5, above pooled unimodal 3.05/5, Human-Retargeted 2.65/5, and EMAGE+GMR 1.80/5. Physical deployment was demonstrated on a Unitree G1.
The authors state the co-speech gains do not consistently reduce jerk or improve foot-contact measures, the training data has few explicit deictic or instructional gestures, and the system needs complete audio plus timed transcripts rather than supporting causal streaming. Additional evident limits are a single robot embodiment (Unitree G1), evaluation mainly on one BEAT2-derived dataset, high peak RAM around 18.5 GB on Jetson AGX Orin, and no quantitative real-robot success metric beyond qualitative trials.
Humanoid robot developers and service-robot makers, including teams using the Unitree G1, could use this for speech-driven full-body gestures in reception, presentation, companion, or telepresence robots. The code and dataset support prototyping now, but product deployment in 1–3 years will require causal streaming inference, improved motion consistency and safety for untethered humanoids, broader semantic gesture data, and validation across multiple robot platforms.
Full paper
ECHO-G: Embodied Co-speech Humanoid mOtion Generation
Published under CC BY 4.0. Reproduced with attribution; original at arXiv:2609.39575 (PDF).
Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance–motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio–text–robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.
Index Terms— Human and Humanoid Motion Analysis and Synthesis, Gesture, Posture and Facial Expressions, Humanoid Robot Systems, co-speech gesture generation, flow matching.
I Introduction
In human communication, gestures complement spoken content, convey emphasis, and organize the temporal structure of speech [1, 2]. Inspired by this coordination, we aim to equip speaking humanoids with body motion that reflects both how an utterance is spoken and what it conveys, while respecting the robot’s embodiment. We study full-body co-speech motion generation from the audio and timed transcript of the robot’s own utterance.
Human co-speech research provides a foundation for learning speech–motion relationships. Representative methods explicitly combine acoustic features with transcript-derived linguistic features to model speech rhythm and content [3, 4, 5]. These methods primarily generate human-motion representations. Extending speech-conditioned generation to humanoids requires robot-specific motion references and an interface for their physical execution. Recent robot-oriented methods address the generation of such references from speech. RoboGesture [6] studies audio-driven streaming generation of upper-body and hand gestures, while PhysDrift [7] explores robot-native generation with speech and text encoders.
These advances suggest a full-body robot co-speech generator should combine densely sampled acoustic cues with token-level linguistic content while preserving their distinct granularities. The one-to-many relationship between utterances and gestures [8] further motivates a generative formulation, and real-robot deployment favors direct prediction in robot space. In addition, existing public releases do not consistently provide paired audio–text–robot training data together with a benchmark spanning co-speech characteristics, robot-motion quality, and runtime efficiency.
Motivated by these considerations, we present ECHO-G, a framework for full-body humanoid co-speech generation from speech audio and timed transcripts (Fig. ). Its Speech-Grounded Diffusion Transformer (SGDiT) adds frame-aligned acoustic features to motion tokens and retrieves token-level linguistic context through global–local cross-attention, preserving the distinct granularities of the two conditions. Trained with rectified flow matching [9], SGDiT models the one-to-many relationship between utterances and full-body robot motion, enabling different motions to be sampled for the same utterance. A fixed pretrained whole-body motion tracker executes the joint-position components of the generated references.
To support training and evaluation in humanoid co-speech generation, we construct a BEAT2-derived robot-space dataset through retargeting and embodiment-specific quality filtering [4, 10]. We publicly release the dataset and code for training, inference, and evaluation, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Pipeline comparisons and modality ablations assess the generation-space and conditioning choices, while video ratings and physical demonstrations provide complementary perceptual and deployment evidence.
Our contributions are threefold:
- We present ECHO-G, a full-body humanoid co-speech generation framework that jointly uses speech audio and timed transcripts, and demonstrate its deployment on a physical humanoid.
- We develop SGDiT, a rectified-flow model combining frame-aligned acoustic conditioning with global–local transcript cross-attention for one-to-many full-body robot-motion generation.
- We release a BEAT2-derived dataset pairing speech audio and timed transcripts with full-body robot motion, together with training, inference, and evaluation code. The accompanying benchmark covers co-speech characteristics, robot-motion quality, and runtime efficiency.
II Related Work
II-A Humanoid Whole-Body Motion Generation
Recent whole-body tracking systems [11, 12] enable humanoids to execute diverse motion references. Such references can be obtained by retargeting human motion, as in GMR [10] and OmniRetarget [13], or generated from language instructions, as in FRoM-W1 [14] and TextOp [15]. OMG [16] further unifies language, audio, and human-motion conditioning within a shared generative framework.
Within this broader setting, robot co-speech generation focuses on gestures accompanying spoken utterances. Yoon et al. [17] generate transcript-conditioned upper-body gestures and demonstrate execution on NAO. RoboPerform [18] generates humanoid motion from speech audio using a generic text prompt rather than the utterance transcript. RoboGesture [6] combines hierarchical semantic–acoustic conditioning with streaming generation of upper-body and hand gestures. PhysDrift [7] uses separate speech and text encoders for one-step robot-native motion generation. However, these approaches either omit utterance-specific audio–text conditioning, focus on upper-body motion, or do not fully specify the temporal organization of their multimodal features. ECHO-G combines frame-aligned acoustic features with token-level transcript embeddings for full-body robot-space generation, preserving their distinct granularities. We also provide paired audio–text–robot data and code for training, inference, and evaluation.
II-B Holistic Human Co-Speech Motion Generation
BEAT [19] provides multimodal speech–gesture data and introduces CaMN for integrating audio, text, and auxiliary conditions. Building on BEAT2, EMAGE [4] combines adaptive content–rhythm fusion with masked gesture modeling and compositional motion priors for holistic generation. DiffSHEG [20] jointly generates expressions and gestures through diffusion, while GestureLSM [5] combines flow matching, latent shortcut learning, and spatiotemporal modeling of body regions for efficient gesture generation. These methods primarily synthesize human-motion representations.
Evaluation considers distributional fidelity, motion variation, and speech–motion alignment. Yoon et al. [3] introduced Fréchet Gesture Distance (FGD), and EMAGE [4] adopted skeleton-aware features for distributional evaluation. Audio2Gestures [8] examines motion diversity and multimodality, while beat-alignment measures assess temporal correspondence between motion and audio [21, 4]. Our benchmark adapts these evaluation dimensions to robot motion and complements them with measures of robot-motion quality and runtime efficiency.
III Method
\widetilde{b}_{\tau n} derived from motion-frame times u_{\tau} and token-center times m_{n}. Multiplying global weights by the exponentiated log-prior and renormalizing yields the local weights. Matrices are transposed for display; colored and dashed arrows indicate stronger and weaker attention, respectively.As shown in Fig. 2, ECHO-G generates full-body robot-motion references from speech audio and timed transcripts. SGDiT integrates acoustic and linguistic conditions within a rectified-flow model, while a fixed whole-body motion tracker executes the generated joint-position references.
III-A Problem Formulation and Motion Representation
Given speech audio a and its word-timed transcript y, ECHO-G models a conditional distribution over full-body robot-motion sequences:
p_{\theta}(\mathbf{R}\mid a,y),where \theta denotes the generator parameters and \mathbf{R}=[\mathbf{r}_{1},\ldots,\mathbf{r}_{T}]^{\top}\in\mathbb{R}^{T\times D} contains T motion frames. For each frame \tau=1,\ldots,T, we use D=39 features:
\mathbf{r}_{\tau}=\bigl[\boldsymbol{\rho}_{\tau}^{\top},\,\delta\psi_{\tau},\,(\mathbf{v}_{\tau}^{\mathrm{loc}})^{\top},\,\mathbf{q}_{\tau}^{\top}\bigr]^{\top},where \boldsymbol{\rho}_{\tau}\in\mathbb{R}^{6} encodes the base orientation using a 6-D rotation representation, \delta\psi_{\tau}\in\mathbb{R} is the inter-frame yaw increment, \mathbf{v}_{\tau}^{\mathrm{loc}}\in\mathbb{R}^{3} is the base linear velocity in a yaw-aligned local frame, and \mathbf{q}_{\tau}\in\mathbb{R}^{29} contains the robot’s joint angles in a fixed order. Absolute root translation is omitted to make the learning target invariant to global position offsets.
The generator operates on normalized motion features:
\mathbf{x}_{\tau}=(\mathbf{r}_{\tau}-\boldsymbol{\mu})\oslash\boldsymbol{\sigma},where \boldsymbol{\mu},\boldsymbol{\sigma}\in\mathbb{R}^{D} are the feature-wise mean and standard deviation computed from the training split, and \oslash denotes element-wise division. We denote the normalized sequence by \mathbf{X}=[\mathbf{x}_{1},\ldots,\mathbf{x}_{T}]^{\top}. Generated sequences are denormalized before evaluation or execution.
III-B Audio–Text Conditioning
Acoustic features extracted by a frozen speech encoder [22] are linearly interpolated to the T motion frames. The transcript text in y is tokenized into w_{1:N} and encoded by a frozen language model [23]. Separate layer normalization and learned affine projections map the two feature sequences to \mathbf{A}\in\mathbb{R}^{T\times d} and \mathbf{H}\in\mathbb{R}^{N\times d}, respectively, where d is the generator’s hidden dimension and N is the number of transcript tokens. The projected conditions retain their frame-level and token-level organization.
Using tokenizer character offsets, we derive approximate token intervals B=\{(s_{n},e_{n})\}_{n=1}^{N} from the word-level timestamps, where s_{n} and e_{n} are the associated start and end times. Temporal conditioning uses token centers m_{n}=(s_{n}+e_{n})/2 and motion-frame times u_{\tau}=(\tau-1)/f, where f is the motion frame rate. Both times are measured from the utterance onset. The combined condition is c=(\mathbf{A},\mathbf{H},B).
III-C Speech-Grounded Diffusion Transformer
SGDiT maps a noisy normalized motion sequence \mathbf{X}_{t}\in\mathbb{R}^{T\times D} and conditions c to the flow velocity v_{\theta}(\mathbf{X}_{t},t,c). Its transformer blocks combine temporal self-attention, transcript cross-attention, and feed-forward processing. Flow time t\in[0,1] modulates the blocks through adaptive layer normalization [24].
Acoustic conditioning. Each input motion token combines a projected motion frame, its aligned acoustic condition, and a positional embedding:
\mathbf{z}_{\tau}=\mathbf{W}_{r}\mathbf{x}_{t,\tau}+\mathbf{A}_{\tau}+\mathbf{p}_{\tau},where \mathbf{x}_{t,\tau} is frame \tau of \mathbf{X}_{t}, \mathbf{W}_{r}\in\mathbb{R}^{d\times D} is a learned projection, and \mathbf{p}_{\tau}\in\mathbb{R}^{d} is a learned frame-position embedding. Temporal self-attention then exchanges information bidirectionally across the acoustically conditioned motion sequence.
Transcript conditioning. Cross-attention retrieves linguistic context through two paths over the same transcript features. The global path provides content-based access to the full token sequence, while the local path adds a preference for temporally nearby tokens. Updated motion features, after normalization and flow-time modulation, provide queries \mathbf{Q}_{\tau}; the projected transcript features \mathbf{H} provide keys \mathbf{K}_{n} and values \mathbf{V}_{n} shared by both paths. For one attention head, the content scores and global attention are
\displaystyle S_{\tau n} \\ \displaystyle=\gamma\,\hat{\mathbf{Q}}_{\tau}^{\top}\hat{\mathbf{K}}_{n}, \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{g}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}),where hats denote L2-normalized queries and keys, \gamma is a bounded learned logit scale, and masked softmax normalizes over non-padding transcript tokens.
The local path adds a Gaussian temporal prior to the shared content scores, yielding the temporally reweighted attention illustrated in Fig. 3:
\displaystyle b_{\tau n} \\ \displaystyle=-\frac{(u_{\tau}-m_{n}-\delta)^{2}}{2\sigma^{2}}, \\ \displaystyle\widetilde{b}_{\tau n} \\ \displaystyle=\max(b_{\tau n},-\kappa), \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{l}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}+\widetilde{\mathbf{b}}),where \sigma and \delta control the temporal width and offset, and \kappa bounds the log penalty. Both paths use the same token padding mask.
The two distributions are combined using temporal support:
\displaystyle g_{\tau} \\ \displaystyle=\alpha\max_{n\in I}\exp(b_{\tau n}), \\ \displaystyle\Pi_{\tau n} \\ \displaystyle=(1-g_{\tau})\Pi^{\mathrm{g}}_{\tau n}+g_{\tau}\Pi^{\mathrm{l}}_{\tau n},where I contains the non-padding token indices and \alpha is the learned base mixing coefficient. Support uses the unclipped prior, reducing the local contribution when the frame is distant from all offset-adjusted token centers. The mixed weights aggregate the shared values into a text-conditioned update, which is projected and added to the motion features through a gated residual connection. A linear output head produces the final T\times D flow-velocity prediction.
III-D Training Objective
We train SGDiT with rectified flow matching [9, 25]. For a normalized motion–condition pair (\mathbf{X},c), we sample standard Gaussian noise \boldsymbol{\epsilon}\in\mathbb{R}^{T\times D} and a sequence-level flow time t\sim\mathcal{U}(0,1). The interpolated motion and target velocity are
\mathbf{X}_{t}=t\mathbf{X}+(1-t)\boldsymbol{\epsilon},\qquad\mathbf{V}^{\star}=\mathbf{X}-\boldsymbol{\epsilon}.The predicted velocity \hat{\mathbf{V}}=v_{\theta}(\mathbf{X}_{t},t,c) is supervised by matching the target flow and its adjacent-frame differences:
\displaystyle\mathcal{L}_{\mathrm{flow}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\hat{\mathbf{V}},\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{temp}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\Delta_{\tau}\hat{\mathbf{V}},\Delta_{\tau}\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{gen}} \\ \displaystyle=\mathcal{L}_{\mathrm{flow}}+\lambda_{\mathrm{temp}}\mathcal{L}_{\mathrm{temp}}.Here \Delta_{\tau} denotes first differences along the motion-frame axis, and \lambda_{\mathrm{temp}} weights the temporal term. MSE is averaged over feature dimensions and valid frames; the temporal term uses only adjacent pairs of valid frames.
III-E Inference and Tracking
At inference, the output length T is specified by the input clip duration at the motion frame rate. We initialize \mathbf{X}^{(0)}\in\mathbb{R}^{T\times D} with standard Gaussian noise and integrate the learned velocity field from flow time 0 to 1, keeping c fixed. Using K uniform Euler steps, we update
\displaystyle t_{k} \\ \displaystyle=\frac{k}{K}, \\ \displaystyle\mathbf{X}^{(k+1)} \\ \displaystyle=\mathbf{X}^{(k)}+\frac{1}{K}v_{\theta}(\mathbf{X}^{(k)},t_{k},c),for k=0,\ldots,K-1. The final state \hat{\mathbf{X}}=\mathbf{X}^{(K)} is converted to robot-motion references by inverting the normalization in (3). Their joint-angle components are supplied to the fixed SONIC motion tracker [11] as joint-position references and executed alongside speech playback.
IV Experiments and Results
IV-A Experimental Setup
IV-A1 Datasets and Preprocessing
We construct a robot-space co-speech dataset from BEAT2 [4]. The original long-form recordings are segmented into 22,192 utterance-level clips at speech pauses using word-level forced alignment, retaining the corresponding audio, transcript, SMPL-X motion, and word timestamps. Each motion sequence is retargeted to the 29-DoF Unitree G1 using GMR [10]. The retargeted motions are converted to a unified Z-up coordinate system and foot-ground aligned using a clip-wise vertical root offset, followed by recomputation of forward kinematics and motion derivatives. We further canonicalize each sequence by removing its initial global yaw while preserving subsequent root dynamics. Robot motions remain at the native 30 fps throughout processing.
We then filter the resulting speech–robot pairs using both robot-motion quality checks and cross-modal consistency checks. For robot-motion quality, we discard clips that violate criteria on foot contact, self-collision, joint continuity and limits, smoothness, or severe high-frequency motion artifacts. We also retain only pairs whose audio and motion durations differ by at most one motion frame.
We adopt a speaker-held-out split, holding out three English speakers for validation and using the remaining speakers for training. After filtering, the final dataset contains 14,987 training clips and 3,242 validation clips.
IV-A2 Evaluation Metrics
All methods use the same held-out split as the candidate input set, with eligibility determined by their native input and output-length requirements. Prediction–reference comparisons use the common temporal prefix of each eligible pair.
Co-Speech Motion Characteristics. Fréchet Gesture Distance (FGD) [3, 4] measures distributional discrepancy using a shared skeleton-aware G1 joint-motion encoder. Div, MM, and BA are computed from forward-kinematic body positions with fixed base rotation and translation. Diversity (Div) measures the frame-wise L1 deviation from each sequence’s temporal mean pose. For stochastic generators, Multimodality (MM) [8] averages pairwise L1 distances among 20 samples generated for each input condition. Beat Alignment (BA) [21, 4] matches speech onsets to the nearest detected upper-body motion beats. We report absolute gaps to the matched reference statistics for Div and BA.
Robot Motion Quality.
Body jerk [16] is estimated from third-order finite differences of reconstructed world-space body positions, scaled by the cube of the frame rate. We average jerk magnitudes over all valid frame–body pairs and report the absolute gap between generated and reference means (\DeltaJerk) [26]. Foot-ground error measures the vertical distance of the lowest sole-proxy surface from the ground plane. Contact sliding speed measures the maximum horizontal sole-point speed per foot, averaged over detected contact intervals.
Efficiency. End-to-end inference time covers input audio and aligned-transcript reading, online feature encoding, motion generation, and GMR or VAE-based motion mapping when applicable, ending at the robot-motion reference. We report total processing time divided by the total number of output frames. Peak RAM increase is the maximum request-window system memory usage above the corresponding pre-load idle baseline.
IV-A3 Implementation Details
We use frozen wav2vec 2.0 large XLSR and Qwen3.5-4B models to obtain 1024-D acoustic features and 2560-D contextual token embeddings, respectively. SGDiT contains 12 transformer blocks with a hidden dimension of 768, 8 attention heads, a feed-forward dimension of 2048, and learned temporal positional embeddings. The attention parameters \gamma,\sigma,\delta,\alpha are learned separately for each layer and head. Attention parameters are constrained to \gamma\in[1,16], \sigma\in[0.25,2] s, \delta\in[-0.5,0.5] s, and \alpha\in[0,0.5] using sigmoid/tanh mappings, with log-prior clipping threshold \kappa=12.
We train for 63,000 optimizer steps on NVIDIA RTX 4090 hardware using FP32 computation. AdamW uses an initial learning rate of 3\times 10^{-4} with cosine decay and no warm-up, weight decay of 10^{-4}, an effective batch size of 48, and gradient clipping at 1.0. We set \lambda_{\mathrm{temp}}=0.5 and jointly drop the acoustic and text encoder features together with the time-distance inputs with probability 0.1. The exponential moving average (EMA) decay is 0.999.
For physical deployment, motion generation runs on a Jetson AGX Orin, which is also used for the efficiency evaluation of all compared methods. Inference uses EMA weights from the checkpoint with the lowest EMA validation loss. Sampling uses eight Euler steps with classifier-free guidance (CFG) scale 1.0. We generate at most 600 motion frames at 30 fps and limit transcripts to 256 tokenizer tokens. Efficiency is measured with batch size one after warm-up.
IV-A4 Baselines and Ablations
We use the publicly released pretrained checkpoints of EMAGE [4] and GestureLSM [5], and retarget their generated human motions to G1 using GMR [10]. We additionally construct Human-Retargeted using the same audio–text conditioning design, backbone configuration, data split, and optimization settings as Ours, but generate normalized 136-D human motion. After denormalization using human-motion training statistics, a pretrained VAE-based mapping converts the samples to 39-D robot references. This mapping remains frozen during human-motion generator training.
Audio-only and Text-only are trained separately in robot space using only the indicated modality. Both variants share the data split, backbone configuration, and optimization settings with Ours. Text-only retains the supplied clip duration and word timings but does not use acoustic features.
IV-B Quantitative Results
- FGD \downarrow
- \DeltaDiv \downarrow
- MM \uparrow
- \DeltaBA \downarrow
- \DeltaJerk \downarrow
- Foot Err. (m) \downarrow
- C-Slide (m/s) \downarrow
- E2E Time (ms/frame) \downarrow
- Peak RAM \Delta (MB) \downarrow
- No data
- EMAGE+GMR
- 4.976
- 0.749
- 0
- 0.172
- 26.951
- 0.013
- 0.163
- 21.3
- 2661
- GestureLSM+GMR
- 5.008
- 0.561
- 1.016
- 0.158
- 24.444
- 0.010
- 0.169
- 20.5
- 3804
- Human-Retargeted
- 4.725
- 0.408
- 1.498
- 0.161
- 8.001
- 0.003
- 0.050
- 6.56
- 18714
- Ours
- 2.278
- 0.320
- 1.786
- 0.063
- 9.039
- 0.008
- 0.052
- 5.96
- 18542
- Audio-only
- 2.360
- 0.360
- 1.702
- 0.081
- 4.497
- 0.008
- 0.038
- Text-only
- 2.436
- 0.429
- 1.681
- 0.113
- 5.940
- 0.009
- 0.049
- Ours
- 2.278
- 0.320
- 1.786
- 0.063
- 9.039
- 0.008
- 0.052
Table I compares direct robot-space generation with human-motion generation followed by retargeting or learned mapping. ECHO-G achieves the best results on all four co-speech metrics and the lowest processing time per output frame. It also improves all three robot-motion quality metrics over EMAGE+GMR and GestureLSM+GMR. Human-Retargeted yields a smaller Jerk gap, lower foot-ground error, and less contact sliding. These results support direct robot-space generation for co-speech modeling with reduced processing time.
Table II evaluates conditioning modalities within the robot-space generator. Joint audio–text conditioning achieves the lowest FGD, \DeltaDiv, and \DeltaBA and the highest MM, outperforming both unimodal variants on the reported co-speech metrics. Audio-only yields the smallest Jerk gap and contact sliding speed, and matches joint conditioning in foot-ground error at the displayed precision. These results support combining acoustic and linguistic information to improve the evaluated co-speech characteristics.
IV-C Qualitative Results
We visualize motions generated from independently prepared audio–text utterances outside BEAT2. These examples provide qualitative evidence of generalization to speech inputs beyond the source dataset. In Fig. 4, audio-only conditioning produces rhythm-responsive motion, but gestures around semantically salient phrases remain small and less clearly related to the spoken content. Text-only conditioning produces content-related gestures, but their timing is less consistently aligned with the audio. Joint conditioning combines speech-responsive timing with more expansive, content-related gestures and fluid transitions in this example.
IV-D Real-Robot Deployment
We conduct real-robot experiments on a Unitree G1, executing the generated joint-position references through a fixed SONIC motion tracker [11] while playing the corresponding speech audio. Fig. 5 shows a representative trial using an independently prepared utterance outside BEAT2, with the robot accompanying its speech with generated body movements. Videos of additional real-robot trials are provided on the project page.
IV-E User Study
- EMAGE+GMR
- 1.80
- 1.87
- 1.95
- 1.59
- Human-Retargeted
- 2.65
- 2.67
- 2.38
- 2.90
- Unimodal (pooled)
- 3.05
- 3.01
- 3.13
- 3.01
- Ours
- 3.49
- 3.45
- 3.31
- 3.70
We conducted a video-rating study with 45 participants using G1 kinematic renderings in MuJoCo. Each participant completed nine trials, with three randomly selected from each of three criterion-specific pools comprising 33 utterances in total. The criteria were human-likeness, rhythm matching, and motion quality. Each trial presented four videos of the same utterance under matched rendering conditions, with hidden method identities and randomized positions. Participants rated each video on a five-point scale.
Each trial compared ECHO-G, Human-Retargeted, EMAGE+GMR, and an audio-only or text-only variant. Unimodal ratings were pooled over the observed trial allocation. Overall denotes the equally weighted mean of the three criterion scores. As shown in Table III, joint audio–text conditioning receives the highest overall mean rating and the highest mean ratings across all three criteria, providing perceptual support for our method.
V Discussion and Limitations
The experimental results support direct robot-space modeling and joint audio–text conditioning for humanoid co-speech generation. Benchmark comparisons show the benefits of these choices for the reported co-speech characteristics, with the direct generation pipeline also requiring less processing time. Qualitative examples, user ratings, and physical demonstrations provide complementary evidence of expressive gestures, perceived quality, and robot execution. Further improvement is needed to better reconcile expressive behavior with robot-motion consistency.
Several limitations remain. First, the gains in co-speech characteristics do not consistently translate into smaller Jerk gaps or better foot-contact measures, leaving room to improve expressive behavior and robot-motion consistency together. Second, the model learns broad speech–gesture associations, with limited training examples of explicit deictic or instructional gestures. This may constrain instruction-aware gesture generation when an utterance calls for a specific semantic motion. Third, the current system requires complete speech audio and timed transcripts and does not yet support causal streaming generation.
VI Conclusion
We presented ECHO-G for full-body humanoid co-speech generation from speech audio and timed transcripts. SGDiT combines frame-aligned acoustic conditioning with global–local transcript cross-attention to generate robot-motion references. Quantitative evaluation, a video-rating study, and physical demonstrations provide complementary evidence for the framework. The released dataset, benchmark, and code support reproducible research on humanoid co-speech generation. Future work will focus on jointly improving gesture expressiveness and robot-motion consistency, enriching training data for instruction-aware semantic gestures, and extending the framework to causal streaming generation.