ROBOTNESS
进阶arXiv

ECHO-G 让 Unitree G1 边说话边做全身手势,FGD 降至 2.278

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
30 秒速读

这是一篇预印本,提出 ECHO-G 框架,从语音音频和带时间戳的转录直接生成人形机器人全身协同语音动作。在 BEAT2 派生数据集的说话人留出验证集上,ECHO-G 的 FGD 为 2.278,优于 EMAGE+GMR 的 4.976、GestureLSM+GMR 的 5.008 和 Human-Retargeted 的 4.725,并在 Unitree G1 真机上部署。该框架省去人体动作重定向,端到端推理时间降至 5.96 毫秒每帧,并公开数据集与代码。

研究问题

给定语音音频和带时间戳的转录,如何为一个具体的人形机器人直接生成既匹配语音韵律又体现语义内容的全身协同语音动作?

问题

现有方法大多生成人体动作表示,再重定向或映射到机器人,容易引入误差和额外计算;面向机器人的方法常只处理上半身,或未同时充分利用帧级语音特征和词级文本特征,缺少公开的音频-文本-机器人配对数据和全面基准。因此难以兼顾表达性和机器人可执行性。

既有方法

人类协同语音模型如 EMAGE、GestureLSM 先合成人体骨架动作,再通过 GMR 重定向或 VAE 映射到机器人;机器人侧的方法如 RoboGesture 仅做上半身流式生成,PhysDrift 虽然使用语音和文本编码器但未面向全身且不公开对齐数据。这些方法在生成空间或条件粒度上存在不足。

新方法

ECHO-G 采用 Speech-Grounded Diffusion Transformer(SGDiT),用 rectified flow 建模从语音和词级时间戳到全身机器人关节动作的一对多分布。它把帧对齐的 wav2vec 2.0 声学特征直接加到运动 token 上,并用全局-局部交叉注意力检索 Qwen3 词向量,局部注意力加入高斯时间先验;输出 39 维机器人特征,包括 6-D 基座朝向、偏航增量、局部速度、29 个关节角,由固定 SONIC 全身跟踪器执行。训练数据来自 BEAT2,经 GMR 重定向到 Unitree G1 并过滤,最终有 14,987 个训练片段和 3,242 个验证片段。

结果

在 BEAT2 派生数据集的说话人留出验证集上,ECHO-G 的 FGD 为 2.278,ΔDiv 为 0.320,MM 为 1.786,ΔBA 为 0.063;相比之下 EMAGE+GMR 为 4.976、0.749、0、0.172,GestureLSM+GMR 为 5.008、0.561、1.016、0.158,Human-Retargeted 为 4.725、0.408、1.498、0.161。机器人运动质量方面,ECHO-G 的足地误差为 0.008 m,接触滑动速度为 0.052 m/s,优于 EMAGE+GMR 和 GestureLSM+GMR,但 ΔJerk 为 9.039,劣于 Human-Retargeted 的 8.001 和音频单模态的 4.497。端到端推理为 5.96 毫秒每帧,低于 EMAGE+GMR 的 21.3、GestureLSM+GMR 的 20.5 和 Human-Retargeted 的 6.56;峰值内存增量 18,542 MB。消融显示联合条件在 FGD、ΔDiv、ΔBA 和 MM 上优于仅音频或仅文本,但 ΔJerk 和接触滑动速度不如音频单模态。45 人视频评分中,ECHO-G 总体均分 3.49,高于 Human-Retargeted 的 2.65、单模态池化的 3.05 和 EMAGE+GMR 的 1.80。真机部署在 Unitree G1 上完成,但只做了定性演示。

局限

作者承认,协同语音指标的提升并未持续带来更小的 Jerk 差距或更好的足部接触指标;模型对明确指示性或指代性手势的训练样本有限,可能影响指令感知手势生成;当前系统需要完整音频和带时间戳的转录,不支持因果流式生成。此外,实验只覆盖 Unitree G1 和 BEAT2 英文数据,真机部署仅作定性演示,没有报告系统化的真机指标。

产业影响

面向服务、导览、教育或娱乐的人形机器人厂商可能在 1 至 3 年内使用此类技术,让机器人在播报、讲解时自然配合手势。前提是解决流式推理、多语言训练、机器人安全约束以及从离线演示到产品级集成的工程问题。现在可用于离线内容生成和展示。

论文全文

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

本文依据 CC BY 4.0 许可发布,经署名转载,原文见 arXiv:2609.39575(PDF)。

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance–motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio–text–robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Index Terms— Human and Humanoid Motion Analysis and Synthesis, Gesture, Posture and Facial Expressions, Humanoid Robot Systems, co-speech gesture generation, flow matching.

I Introduction

In human communication, gestures complement spoken content, convey emphasis, and organize the temporal structure of speech [1, 2]. Inspired by this coordination, we aim to equip speaking humanoids with body motion that reflects both how an utterance is spoken and what it conveys, while respecting the robot’s embodiment. We study full-body co-speech motion generation from the audio and timed transcript of the robot’s own utterance.

Human co-speech research provides a foundation for learning speech–motion relationships. Representative methods explicitly combine acoustic features with transcript-derived linguistic features to model speech rhythm and content [3, 4, 5]. These methods primarily generate human-motion representations. Extending speech-conditioned generation to humanoids requires robot-specific motion references and an interface for their physical execution. Recent robot-oriented methods address the generation of such references from speech. RoboGesture [6] studies audio-driven streaming generation of upper-body and hand gestures, while PhysDrift [7] explores robot-native generation with speech and text encoders.

These advances suggest a full-body robot co-speech generator should combine densely sampled acoustic cues with token-level linguistic content while preserving their distinct granularities. The one-to-many relationship between utterances and gestures [8] further motivates a generative formulation, and real-robot deployment favors direct prediction in robot space. In addition, existing public releases do not consistently provide paired audio–text–robot training data together with a benchmark spanning co-speech characteristics, robot-motion quality, and runtime efficiency.

Motivated by these considerations, we present ECHO-G, a framework for full-body humanoid co-speech generation from speech audio and timed transcripts (Fig. ). Its Speech-Grounded Diffusion Transformer (SGDiT) adds frame-aligned acoustic features to motion tokens and retrieves token-level linguistic context through global–local cross-attention, preserving the distinct granularities of the two conditions. Trained with rectified flow matching [9], SGDiT models the one-to-many relationship between utterances and full-body robot motion, enabling different motions to be sampled for the same utterance. A fixed pretrained whole-body motion tracker executes the joint-position components of the generated references.

To support training and evaluation in humanoid co-speech generation, we construct a BEAT2-derived robot-space dataset through retargeting and embodiment-specific quality filtering [4, 10]. We publicly release the dataset and code for training, inference, and evaluation, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Pipeline comparisons and modality ablations assess the generation-space and conditioning choices, while video ratings and physical demonstrations provide complementary perceptual and deployment evidence.

Our contributions are threefold:

  • We present ECHO-G, a full-body humanoid co-speech generation framework that jointly uses speech audio and timed transcripts, and demonstrate its deployment on a physical humanoid.
  • We develop SGDiT, a rectified-flow model combining frame-aligned acoustic conditioning with global–local transcript cross-attention for one-to-many full-body robot-motion generation.
  • We release a BEAT2-derived dataset pairing speech audio and timed transcripts with full-body robot motion, together with training, inference, and evaluation code. The accompanying benchmark covers co-speech characteristics, robot-motion quality, and runtime efficiency.

II Related Work

Fig. 2: SGDiT architecture and tracking interface. Frame-aligned acoustic features are combined with noisy motion-frame features to form motion tokens. Contextual transcript embeddings provide a shared key–value memory for global and temporally biased text attention. Their attention distributions are fused within each transformer block before value aggregation and residual injection into the motion stream. Rectified-flow sampling produces robot-motion references, whose joint-position components are passed to a fixed whole-body motion tracker for execution.

II-A Humanoid Whole-Body Motion Generation

Recent whole-body tracking systems [11, 12] enable humanoids to execute diverse motion references. Such references can be obtained by retargeting human motion, as in GMR [10] and OmniRetarget [13], or generated from language instructions, as in FRoM-W1 [14] and TextOp [15]. OMG [16] further unifies language, audio, and human-motion conditioning within a shared generative framework.

Within this broader setting, robot co-speech generation focuses on gestures accompanying spoken utterances. Yoon et al. [17] generate transcript-conditioned upper-body gestures and demonstrate execution on NAO. RoboPerform [18] generates humanoid motion from speech audio using a generic text prompt rather than the utterance transcript. RoboGesture [6] combines hierarchical semantic–acoustic conditioning with streaming generation of upper-body and hand gestures. PhysDrift [7] uses separate speech and text encoders for one-step robot-native motion generation. However, these approaches either omit utterance-specific audio–text conditioning, focus on upper-body motion, or do not fully specify the temporal organization of their multimodal features. ECHO-G combines frame-aligned acoustic features with token-level transcript embeddings for full-body robot-space generation, preserving their distinct granularities. We also provide paired audio–text–robot data and code for training, inference, and evaluation.

II-B Holistic Human Co-Speech Motion Generation

BEAT [19] provides multimodal speech–gesture data and introduces CaMN for integrating audio, text, and auxiliary conditions. Building on BEAT2, EMAGE [4] combines adaptive content–rhythm fusion with masked gesture modeling and compositional motion priors for holistic generation. DiffSHEG [20] jointly generates expressions and gestures through diffusion, while GestureLSM [5] combines flow matching, latent shortcut learning, and spatiotemporal modeling of body regions for efficient gesture generation. These methods primarily synthesize human-motion representations.

Evaluation considers distributional fidelity, motion variation, and speech–motion alignment. Yoon et al. [3] introduced Fréchet Gesture Distance (FGD), and EMAGE [4] adopted skeleton-aware features for distributional evaluation. Audio2Gestures [8] examines motion diversity and multimodality, while beat-alignment measures assess temporal correspondence between motion and audio [21, 4]. Our benchmark adapts these evaluation dimensions to robot motion and complements them with measures of robot-motion quality and runtime efficiency.

III Method

Fig. 3: Temporal conditioning in the local transcript-attention branch. Orange and purple heatmaps show global and local attention weights, respectively. The blue heatmap represents the clipped Gaussian log-prior \widetilde{b}_{\tau n} derived from motion-frame times u_{\tau} and token-center times m_{n}. Multiplying global weights by the exponentiated log-prior and renormalizing yields the local weights. Matrices are transposed for display; colored and dashed arrows indicate stronger and weaker attention, respectively.

As shown in Fig. 2, ECHO-G generates full-body robot-motion references from speech audio and timed transcripts. SGDiT integrates acoustic and linguistic conditions within a rectified-flow model, while a fixed whole-body motion tracker executes the generated joint-position references.

III-A Problem Formulation and Motion Representation

Given speech audio a and its word-timed transcript y, ECHO-G models a conditional distribution over full-body robot-motion sequences:

p_{\theta}(\mathbf{R}\mid a,y),
(1)

where \theta denotes the generator parameters and \mathbf{R}=[\mathbf{r}_{1},\ldots,\mathbf{r}_{T}]^{\top}\in\mathbb{R}^{T\times D} contains T motion frames. For each frame \tau=1,\ldots,T, we use D=39 features:

\mathbf{r}_{\tau}=\bigl[\boldsymbol{\rho}_{\tau}^{\top},\,\delta\psi_{\tau},\,(\mathbf{v}_{\tau}^{\mathrm{loc}})^{\top},\,\mathbf{q}_{\tau}^{\top}\bigr]^{\top},
(2)

where \boldsymbol{\rho}_{\tau}\in\mathbb{R}^{6} encodes the base orientation using a 6-D rotation representation, \delta\psi_{\tau}\in\mathbb{R} is the inter-frame yaw increment, \mathbf{v}_{\tau}^{\mathrm{loc}}\in\mathbb{R}^{3} is the base linear velocity in a yaw-aligned local frame, and \mathbf{q}_{\tau}\in\mathbb{R}^{29} contains the robot’s joint angles in a fixed order. Absolute root translation is omitted to make the learning target invariant to global position offsets.

The generator operates on normalized motion features:

\mathbf{x}_{\tau}=(\mathbf{r}_{\tau}-\boldsymbol{\mu})\oslash\boldsymbol{\sigma},
(3)

where \boldsymbol{\mu},\boldsymbol{\sigma}\in\mathbb{R}^{D} are the feature-wise mean and standard deviation computed from the training split, and \oslash denotes element-wise division. We denote the normalized sequence by \mathbf{X}=[\mathbf{x}_{1},\ldots,\mathbf{x}_{T}]^{\top}. Generated sequences are denormalized before evaluation or execution.

III-B Audio–Text Conditioning

Acoustic features extracted by a frozen speech encoder [22] are linearly interpolated to the T motion frames. The transcript text in y is tokenized into w_{1:N} and encoded by a frozen language model [23]. Separate layer normalization and learned affine projections map the two feature sequences to \mathbf{A}\in\mathbb{R}^{T\times d} and \mathbf{H}\in\mathbb{R}^{N\times d}, respectively, where d is the generator’s hidden dimension and N is the number of transcript tokens. The projected conditions retain their frame-level and token-level organization.

Using tokenizer character offsets, we derive approximate token intervals B=\{(s_{n},e_{n})\}_{n=1}^{N} from the word-level timestamps, where s_{n} and e_{n} are the associated start and end times. Temporal conditioning uses token centers m_{n}=(s_{n}+e_{n})/2 and motion-frame times u_{\tau}=(\tau-1)/f, where f is the motion frame rate. Both times are measured from the utterance onset. The combined condition is c=(\mathbf{A},\mathbf{H},B).

III-C Speech-Grounded Diffusion Transformer

SGDiT maps a noisy normalized motion sequence \mathbf{X}_{t}\in\mathbb{R}^{T\times D} and conditions c to the flow velocity v_{\theta}(\mathbf{X}_{t},t,c). Its transformer blocks combine temporal self-attention, transcript cross-attention, and feed-forward processing. Flow time t\in[0,1] modulates the blocks through adaptive layer normalization [24].

Acoustic conditioning. Each input motion token combines a projected motion frame, its aligned acoustic condition, and a positional embedding:

\mathbf{z}_{\tau}=\mathbf{W}_{r}\mathbf{x}_{t,\tau}+\mathbf{A}_{\tau}+\mathbf{p}_{\tau},
(4)

where \mathbf{x}_{t,\tau} is frame \tau of \mathbf{X}_{t}, \mathbf{W}_{r}\in\mathbb{R}^{d\times D} is a learned projection, and \mathbf{p}_{\tau}\in\mathbb{R}^{d} is a learned frame-position embedding. Temporal self-attention then exchanges information bidirectionally across the acoustically conditioned motion sequence.

Transcript conditioning. Cross-attention retrieves linguistic context through two paths over the same transcript features. The global path provides content-based access to the full token sequence, while the local path adds a preference for temporally nearby tokens. Updated motion features, after normalization and flow-time modulation, provide queries \mathbf{Q}_{\tau}; the projected transcript features \mathbf{H} provide keys \mathbf{K}_{n} and values \mathbf{V}_{n} shared by both paths. For one attention head, the content scores and global attention are

\displaystyle S_{\tau n} \\ \displaystyle=\gamma\,\hat{\mathbf{Q}}_{\tau}^{\top}\hat{\mathbf{K}}_{n}, \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{g}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}),

where hats denote L2-normalized queries and keys, \gamma is a bounded learned logit scale, and masked softmax normalizes over non-padding transcript tokens.

The local path adds a Gaussian temporal prior to the shared content scores, yielding the temporally reweighted attention illustrated in Fig. 3:

\displaystyle b_{\tau n} \\ \displaystyle=-\frac{(u_{\tau}-m_{n}-\delta)^{2}}{2\sigma^{2}}, \\ \displaystyle\widetilde{b}_{\tau n} \\ \displaystyle=\max(b_{\tau n},-\kappa), \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{l}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}+\widetilde{\mathbf{b}}),

where \sigma and \delta control the temporal width and offset, and \kappa bounds the log penalty. Both paths use the same token padding mask.

The two distributions are combined using temporal support:

\displaystyle g_{\tau} \\ \displaystyle=\alpha\max_{n\in I}\exp(b_{\tau n}), \\ \displaystyle\Pi_{\tau n} \\ \displaystyle=(1-g_{\tau})\Pi^{\mathrm{g}}_{\tau n}+g_{\tau}\Pi^{\mathrm{l}}_{\tau n},

where I contains the non-padding token indices and \alpha is the learned base mixing coefficient. Support uses the unclipped prior, reducing the local contribution when the frame is distant from all offset-adjusted token centers. The mixed weights aggregate the shared values into a text-conditioned update, which is projected and added to the motion features through a gated residual connection. A linear output head produces the final T\times D flow-velocity prediction.

III-D Training Objective

We train SGDiT with rectified flow matching [9, 25]. For a normalized motion–condition pair (\mathbf{X},c), we sample standard Gaussian noise \boldsymbol{\epsilon}\in\mathbb{R}^{T\times D} and a sequence-level flow time t\sim\mathcal{U}(0,1). The interpolated motion and target velocity are

\mathbf{X}_{t}=t\mathbf{X}+(1-t)\boldsymbol{\epsilon},\qquad\mathbf{V}^{\star}=\mathbf{X}-\boldsymbol{\epsilon}.
(8)

The predicted velocity \hat{\mathbf{V}}=v_{\theta}(\mathbf{X}_{t},t,c) is supervised by matching the target flow and its adjacent-frame differences:

\displaystyle\mathcal{L}_{\mathrm{flow}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\hat{\mathbf{V}},\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{temp}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\Delta_{\tau}\hat{\mathbf{V}},\Delta_{\tau}\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{gen}} \\ \displaystyle=\mathcal{L}_{\mathrm{flow}}+\lambda_{\mathrm{temp}}\mathcal{L}_{\mathrm{temp}}.

Here \Delta_{\tau} denotes first differences along the motion-frame axis, and \lambda_{\mathrm{temp}} weights the temporal term. MSE is averaged over feature dimensions and valid frames; the temporal term uses only adjacent pairs of valid frames.

Fig. 4: Conditioning comparison on an utterance outside BEAT2. Rows show joint audio–text conditioning (Ours), text-only, and audio-only outputs at matched frame indices. In the selected frames, joint conditioning exhibits broader arm extensions, whereas the unimodal outputs generally keep the hands closer to the torso. Colored arrows and circles highlight selected arm and hand movements.

III-E Inference and Tracking

At inference, the output length T is specified by the input clip duration at the motion frame rate. We initialize \mathbf{X}^{(0)}\in\mathbb{R}^{T\times D} with standard Gaussian noise and integrate the learned velocity field from flow time 0 to 1, keeping c fixed. Using K uniform Euler steps, we update

\displaystyle t_{k} \\ \displaystyle=\frac{k}{K}, \\ \displaystyle\mathbf{X}^{(k+1)} \\ \displaystyle=\mathbf{X}^{(k)}+\frac{1}{K}v_{\theta}(\mathbf{X}^{(k)},t_{k},c),

for k=0,\ldots,K-1. The final state \hat{\mathbf{X}}=\mathbf{X}^{(K)} is converted to robot-motion references by inverting the normalization in (3). Their joint-angle components are supplied to the fixed SONIC motion tracker [11] as joint-position references and executed alongside speech playback.

IV Experiments and Results

IV-A Experimental Setup

IV-A1 Datasets and Preprocessing

We construct a robot-space co-speech dataset from BEAT2 [4]. The original long-form recordings are segmented into 22,192 utterance-level clips at speech pauses using word-level forced alignment, retaining the corresponding audio, transcript, SMPL-X motion, and word timestamps. Each motion sequence is retargeted to the 29-DoF Unitree G1 using GMR [10]. The retargeted motions are converted to a unified Z-up coordinate system and foot-ground aligned using a clip-wise vertical root offset, followed by recomputation of forward kinematics and motion derivatives. We further canonicalize each sequence by removing its initial global yaw while preserving subsequent root dynamics. Robot motions remain at the native 30 fps throughout processing.

We then filter the resulting speech–robot pairs using both robot-motion quality checks and cross-modal consistency checks. For robot-motion quality, we discard clips that violate criteria on foot contact, self-collision, joint continuity and limits, smoothness, or severe high-frequency motion artifacts. We also retain only pairs whose audio and motion durations differ by at most one motion frame.

We adopt a speaker-held-out split, holding out three English speakers for validation and using the remaining speakers for training. After filtering, the final dataset contains 14,987 training clips and 3,242 validation clips.

IV-A2 Evaluation Metrics

All methods use the same held-out split as the candidate input set, with eligibility determined by their native input and output-length requirements. Prediction–reference comparisons use the common temporal prefix of each eligible pair.

Co-Speech Motion Characteristics. Fréchet Gesture Distance (FGD) [3, 4] measures distributional discrepancy using a shared skeleton-aware G1 joint-motion encoder. Div, MM, and BA are computed from forward-kinematic body positions with fixed base rotation and translation. Diversity (Div) measures the frame-wise L1 deviation from each sequence’s temporal mean pose. For stochastic generators, Multimodality (MM) [8] averages pairwise L1 distances among 20 samples generated for each input condition. Beat Alignment (BA) [21, 4] matches speech onsets to the nearest detected upper-body motion beats. We report absolute gaps to the matched reference statistics for Div and BA.

Robot Motion Quality.

Fig. 5: Real-robot execution on an utterance outside BEAT2. ECHO-G’s predicted joint angles are supplied as joint-position references to a fixed SONIC motion tracker while the corresponding speech is played. Frames progress from left to right through the book-recommendation utterance shown below, illustrating changes in arm extension, hand height, and torso posture. Colored arrows highlight selected arm movements, while the orange skeletal overlays outline the body configuration.

Body jerk [16] is estimated from third-order finite differences of reconstructed world-space body positions, scaled by the cube of the frame rate. We average jerk magnitudes over all valid frame–body pairs and report the absolute gap between generated and reference means (\DeltaJerk) [26]. Foot-ground error measures the vertical distance of the lowest sole-proxy surface from the ground plane. Contact sliding speed measures the maximum horizontal sole-point speed per foot, averaged over detected contact intervals.

Efficiency. End-to-end inference time covers input audio and aligned-transcript reading, online feature encoding, motion generation, and GMR or VAE-based motion mapping when applicable, ending at the robot-motion reference. We report total processing time divided by the total number of output frames. Peak RAM increase is the maximum request-window system memory usage above the corresponding pre-load idle baseline.

IV-A3 Implementation Details

We use frozen wav2vec 2.0 large XLSR and Qwen3.5-4B models to obtain 1024-D acoustic features and 2560-D contextual token embeddings, respectively. SGDiT contains 12 transformer blocks with a hidden dimension of 768, 8 attention heads, a feed-forward dimension of 2048, and learned temporal positional embeddings. The attention parameters \gamma,\sigma,\delta,\alpha are learned separately for each layer and head. Attention parameters are constrained to \gamma\in[1,16], \sigma\in[0.25,2] s, \delta\in[-0.5,0.5] s, and \alpha\in[0,0.5] using sigmoid/tanh mappings, with log-prior clipping threshold \kappa=12.

We train for 63,000 optimizer steps on NVIDIA RTX 4090 hardware using FP32 computation. AdamW uses an initial learning rate of 3\times 10^{-4} with cosine decay and no warm-up, weight decay of 10^{-4}, an effective batch size of 48, and gradient clipping at 1.0. We set \lambda_{\mathrm{temp}}=0.5 and jointly drop the acoustic and text encoder features together with the time-distance inputs with probability 0.1. The exponential moving average (EMA) decay is 0.999.

For physical deployment, motion generation runs on a Jetson AGX Orin, which is also used for the efficiency evaluation of all compared methods. Inference uses EMA weights from the checkpoint with the lowest EMA validation loss. Sampling uses eight Euler steps with classifier-free guidance (CFG) scale 1.0. We generate at most 600 motion frames at 30 fps and limit transcripts to 256 tokenizer tokens. Efficiency is measured with batch size one after warm-up.

IV-A4 Baselines and Ablations

We use the publicly released pretrained checkpoints of EMAGE [4] and GestureLSM [5], and retarget their generated human motions to G1 using GMR [10]. We additionally construct Human-Retargeted using the same audio–text conditioning design, backbone configuration, data split, and optimization settings as Ours, but generate normalized 136-D human motion. After denormalization using human-motion training statistics, a pretrained VAE-based mapping converts the samples to 39-D robot references. This mapping remains frozen during human-motion generator training.

Audio-only and Text-only are trained separately in robot space using only the indicated modality. Both variants share the data split, backbone configuration, and optimization settings with Ours. Text-only retains the supplied clip duration and word timings but does not use acoustic features.

IV-B Quantitative Results

TABLE I: Comparison of direct robot-space generation with human-motion generation followed by retargeting or learned mapping on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from matched ground-truth statistics over each method’s evaluated motion range. Best and second-best results are shown in bold and underlined text, respectively.
  • FGD \downarrow
    Co-Speech Motion Characteristics
    \DeltaDiv \downarrow
    MM \uparrow
    \DeltaBA \downarrow
    \DeltaJerk \downarrow
    Robot Motion Quality
    Foot Err. (m) \downarrow
    C-Slide (m/s) \downarrow
    E2E Time (ms/frame) \downarrow
    Efficiency
    Peak RAM \Delta (MB) \downarrow
    无数据
  • EMAGE+GMR
    Co-Speech Motion Characteristics
    4.976
    0.749
    0
    0.172
    Robot Motion Quality
    26.951
    0.013
    0.163
    Efficiency
    21.3
    2661
  • GestureLSM+GMR
    Co-Speech Motion Characteristics
    5.008
    0.561
    1.016
    0.158
    Robot Motion Quality
    24.444
    0.010
    0.169
    Efficiency
    20.5
    3804
  • Human-Retargeted
    Co-Speech Motion Characteristics
    4.725
    0.408
    1.498
    0.161
    Robot Motion Quality
    8.001
    0.003
    0.050
    Efficiency
    6.56
    18714
  • Ours
    Co-Speech Motion Characteristics
    2.278
    0.320
    1.786
    0.063
    Robot Motion Quality
    9.039
    0.008
    0.052
    Efficiency
    5.96
    18542
TABLE II: Ablation of conditioning modalities for direct robot-space generation on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from the shared ground-truth statistics. Best and second-best results are shown in bold and underlined text, respectively. Ties at the displayed precision receive identical highlighting.
  • Audio-only
    FGD \downarrow
    2.360
    \DeltaDiv \downarrow
    0.360
    MM \uparrow
    1.702
    \DeltaBA \downarrow
    0.081
    \DeltaJerk \downarrow
    4.497
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.038
  • Text-only
    FGD \downarrow
    2.436
    \DeltaDiv \downarrow
    0.429
    MM \uparrow
    1.681
    \DeltaBA \downarrow
    0.113
    \DeltaJerk \downarrow
    5.940
    Foot Err. (m) \downarrow
    0.009
    C-Slide (m/s) \downarrow
    0.049
  • Ours
    FGD \downarrow
    2.278
    \DeltaDiv \downarrow
    0.320
    MM \uparrow
    1.786
    \DeltaBA \downarrow
    0.063
    \DeltaJerk \downarrow
    9.039
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.052

Table I compares direct robot-space generation with human-motion generation followed by retargeting or learned mapping. ECHO-G achieves the best results on all four co-speech metrics and the lowest processing time per output frame. It also improves all three robot-motion quality metrics over EMAGE+GMR and GestureLSM+GMR. Human-Retargeted yields a smaller Jerk gap, lower foot-ground error, and less contact sliding. These results support direct robot-space generation for co-speech modeling with reduced processing time.

Table II evaluates conditioning modalities within the robot-space generator. Joint audio–text conditioning achieves the lowest FGD, \DeltaDiv, and \DeltaBA and the highest MM, outperforming both unimodal variants on the reported co-speech metrics. Audio-only yields the smallest Jerk gap and contact sliding speed, and matches joint conditioning in foot-ground error at the displayed precision. These results support combining acoustic and linguistic information to improve the evaluated co-speech characteristics.

IV-C Qualitative Results

We visualize motions generated from independently prepared audio–text utterances outside BEAT2. These examples provide qualitative evidence of generalization to speech inputs beyond the source dataset. In Fig. 4, audio-only conditioning produces rhythm-responsive motion, but gestures around semantically salient phrases remain small and less clearly related to the spoken content. Text-only conditioning produces content-related gestures, but their timing is less consistently aligned with the audio. Joint conditioning combines speech-responsive timing with more expansive, content-related gestures and fluid transitions in this example.

IV-D Real-Robot Deployment

We conduct real-robot experiments on a Unitree G1, executing the generated joint-position references through a fixed SONIC motion tracker [11] while playing the corresponding speech audio. Fig. 5 shows a representative trial using an independently prepared utterance outside BEAT2, with the robot accompanying its speech with generated body movements. Videos of additional real-robot trials are provided on the project page.

IV-E User Study

TABLE III: Mean user-study ratings from 45 participants on a five-point scale. Higher is better. Unimodal scores are pooled as described in the text. Best and second-best means are shown in bold and underlined text, respectively.
  • EMAGE+GMR
    Overall
    1.80
    Human- likeness
    1.87
    Rhythm Matching
    1.95
    Motion Quality
    1.59
  • Human-Retargeted
    Overall
    2.65
    Human- likeness
    2.67
    Rhythm Matching
    2.38
    Motion Quality
    2.90
  • Unimodal (pooled)
    Overall
    3.05
    Human- likeness
    3.01
    Rhythm Matching
    3.13
    Motion Quality
    3.01
  • Ours
    Overall
    3.49
    Human- likeness
    3.45
    Rhythm Matching
    3.31
    Motion Quality
    3.70

We conducted a video-rating study with 45 participants using G1 kinematic renderings in MuJoCo. Each participant completed nine trials, with three randomly selected from each of three criterion-specific pools comprising 33 utterances in total. The criteria were human-likeness, rhythm matching, and motion quality. Each trial presented four videos of the same utterance under matched rendering conditions, with hidden method identities and randomized positions. Participants rated each video on a five-point scale.

Each trial compared ECHO-G, Human-Retargeted, EMAGE+GMR, and an audio-only or text-only variant. Unimodal ratings were pooled over the observed trial allocation. Overall denotes the equally weighted mean of the three criterion scores. As shown in Table III, joint audio–text conditioning receives the highest overall mean rating and the highest mean ratings across all three criteria, providing perceptual support for our method.

V Discussion and Limitations

The experimental results support direct robot-space modeling and joint audio–text conditioning for humanoid co-speech generation. Benchmark comparisons show the benefits of these choices for the reported co-speech characteristics, with the direct generation pipeline also requiring less processing time. Qualitative examples, user ratings, and physical demonstrations provide complementary evidence of expressive gestures, perceived quality, and robot execution. Further improvement is needed to better reconcile expressive behavior with robot-motion consistency.

Several limitations remain. First, the gains in co-speech characteristics do not consistently translate into smaller Jerk gaps or better foot-contact measures, leaving room to improve expressive behavior and robot-motion consistency together. Second, the model learns broad speech–gesture associations, with limited training examples of explicit deictic or instructional gestures. This may constrain instruction-aware gesture generation when an utterance calls for a specific semantic motion. Third, the current system requires complete speech audio and timed transcripts and does not yet support causal streaming generation.

VI Conclusion

We presented ECHO-G for full-body humanoid co-speech generation from speech audio and timed transcripts. SGDiT combines frame-aligned acoustic conditioning with global–local transcript cross-attention to generate robot-motion references. Quantitative evaluation, a video-rating study, and physical demonstrations provide complementary evidence for the framework. The released dataset, benchmark, and code support reproducible research on humanoid co-speech generation. Future work will focus on jointly improving gesture expressiveness and robot-motion consistency, enriching training data for instruction-aware semantic gestures, and extending the framework to causal streaming generation.

References

  1. [1] D. McNeill (1992) Hand and mind: what gestures reveal about thought. University of Chicago Press, Chicago, IL, USA.
  2. [2] A. Kendon (2004) Gesture: visible action as utterance. Cambridge University Press, Cambridge, UK. External Links: Document
  3. [3] Y. Yoon, B. Cha, J. Lee, M. Jang, J. Lee, J. Kim, and G. Lee (2020) Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics 39 (6), pp. 1–16. External Links: Document
  4. [4] H. Liu, Z. Zhu, G. Becherini, Y. Peng, M. Su, Y. Zhou, X. Zhe, N. Iwamoto, B. Zheng, and M. J. Black (2024) EMAGE: towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1144–1154.
  5. [5] P. Liu, L. Song, J. Huang, H. Liu, and C. Xu (2025) GestureLSM: latent shortcut based co-speech gesture generation with spatial-temporal modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10929–10939.
  6. [6] Z. Wang, Z. Ren, P. Shi, Z. Wang, C. Lin, T. Wang, Z. Qi, L. Zhao, H. Wang, and L. Yi (2026) RoboGesture: real-time semantic-aligned co-speech gestures generation for humanoid interaction. In Proceedings of the European Conference on Computer Vision (ECCV),
  7. [7] Z. Liang, X. Xing, M. Yang, W. Zhou, and X. Xu (2026) PhysDrift: bridging the embodiment gap in humanoid co-speech motion generation. arXiv preprint arXiv:2606.19935.
  8. [8] J. Li, D. Kang, W. Pei, X. Zhe, Y. Zhang, Z. He, and L. Bao (2021) Audio2Gestures: generating diverse gestures from speech audio with conditional variational autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11293–11302.
  9. [9] X. Liu, C. Gong, and Q. Liu (2023) Flow straight and fast: learning to generate and transfer data with rectified flow. In International Conference on Learning Representations,
  10. [10] J. P. Araujo, Y. Ze, P. Xu, J. Wu, and C. K. Liu (2025) Retargeting matters: general motion retargeting for humanoid motion tracking. arXiv preprint arXiv:2510.02252.
  11. [11] Z. Luo, Y. Yuan, T. Wang, C. Li, F. Castañeda, S. Chen, Z. Cao, J. Li, D. Minor, Q. Ben, et al. (2026) SONIC: supersizing motion tracking for natural humanoid whole-body control. Science Robotics 11 (117), pp. eaed4592. External Links: Document
  12. [12] M. Chen, K. Wang, B. Zhang, X. Ma, Z. Yang, Y. Ren, Q. Huang, Z. Zhu, Y. Wang, and Z. Su (2026) HoloMotion-1 technical report. arXiv preprint arXiv:2605.15336.
  13. [13] L. Yang, X. Huang, Z. Wu, A. Kanazawa, P. Abbeel, C. Sferrazza, C. K. Liu, R. Duan, and G. Shi (2025) OmniRetarget: interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. arXiv preprint arXiv:2509.26633.
  14. [14] P. Li, Z. Zhuang, Y. Gao, Y. Dong, S. Li, C. Jiang, S. Dou, Z. Xi, E. Zhou, J. Huang, et al. (2026) FRoM-W1: towards general humanoid whole-body control with language instructions. arXiv preprint arXiv:2601.12799.
  15. [15] W. Xie, J. Zheng, J. Han, J. Shi, W. Zhang, C. Bai, and X. Li (2026) TextOp: real-time interactive text-driven humanoid robot motion generation and control. arXiv preprint arXiv:2602.07439.
  16. [16] S. Huang, K. Lee, D. Qiao, G. He, Z. Wang, Y. Li, S. Zhu, and H. Zhao (2026) OMG: omni-modal motion generation for generalist humanoid control. arXiv preprint arXiv:2606.10340.
  17. [17] Y. Yoon, W. Ko, M. Jang, J. Lee, J. Kim, and G. Lee (2019) Robots learn social skills: end-to-end learning of co-speech gesture generation for humanoid robots. In 2019 International Conference on Robotics and Automation (ICRA), pp. 4303–4309. External Links: Document
  18. [18] Z. Li, C. Chi, Y. Wei, B. Zhu, T. Huang, Z. Sun, Y. Peng, P. Wang, Z. Wang, F. Liu, C. Xu, and S. Zhang (2026) Do you have freestyle? expressive humanoid locomotion via audio control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 956–965.
  19. [19] H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng (2022) BEAT: a large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. In European Conference on Computer Vision, pp. 612–630.
  20. [20] J. Chen, Y. Liu, J. Wang, A. Zeng, Y. Li, and Q. Chen (2024) DiffSHEG: a diffusion-based approach for real-time speech-driven holistic 3D expression and gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7352–7361.
  21. [21] R. Li, S. Yang, D. A. Ross, and A. Kanazawa (2021) AI choreographer: music conditioned 3D dance generation with AIST++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13401–13412.
  22. [22] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli (2020) wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, Vol. 33, pp. 12449–12460.
  23. [23] Qwen Team (2026) Qwen3.5: towards native multimodal agents. External Links: Link
  24. [24] W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205.
  25. [25] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) Flow Matching for generative modeling. In International Conference on Learning Representations, External Links: Link
  26. [26] F. Fang, S. Yang, and W. Yang (2026) CoordSpeaker: exploiting gesture captioning for coordinated caption-empowered co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 30761–30771.