ROBOTNESS
中級arXiv

音声と書き起こしを統合した人型ロボット向け発話動作生成「ECHO-G」を発表

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
30秒で読む

ECHO-Gは、音声とタイムコード付き書き起こしの双方を条件として人型ロボットの全身発話動作を直接生成するフレームワークである。BEAT2由来のロボット動作データセットで比較評価し、従来の人間動作生成とリターゲティングの経路に比べFGDやビート同期などのコスピーチ指標と処理時間で優位を示し、Unitree G1での実機動作も確認した。本論文は査読前プレプリントである。

研究課題

発話音声とタイムコード付き書き起こしから、発話内容と韻律の両方に合致し、人型ロボットの身体制約を満たす全身のコスピーチ動作を生成できるか。

問題

人間向けコスピーチ動作生成はSMPL-X等の人間表現を出力するため、ロボットで使うにはリターゲティングや学習済み写像が必要になり、処理時間や動作品質の劣化が生じる。ロボット向けの先行研究は上体や手の動作に限られていたり、発話固有の音声と書き起こしの時間的対応を十分に扱っていなかったりする。さらに音声・テキスト・ロボット動作の対応データと包括的ベンチマークの公開も十分ではない。

従来の手法

EMAGEやGestureLSMは人間の骨格動作を生成し、GMRでUnitree G1へリターゲティングする経路が比較対象となった。ロボット向けではRoboGestureが音声駆動で上体と手の動作を生成し、PhysDriftが音声・テキストエンコーダでロボット動作を生成するが、発話固有の音声とテキストの時間的対応を組み合わせる設計や全身動作への対応が限定的だった。

新しい手法

ECHO-GはSGDiTを核とし、フレーム単位の音響特徴をモーショントークンに付加し、トークン単位の言語特徴にはグローバル経路とガウス時間事前に基づくローカル経路のクロスアテンションを併用する。整流フローマッチングで学習し、同じ発話から複数の全身ロボット動作を直接生成できる。BEAT2をUnitree G1向けに再ターゲティングし品質フィルタをかけた音声・テキスト・ロボット動作データセットと、コスピーチ特性・ロボット動作品質・実行時効率を測るベンチマークを公開した。推論では生成した関節位置参照を固定のSONICモーショントラッカーへ渡して実機実行する。

結果

BEAT2の話者除外データ(学習14,987件、検証3,242件)で、ECHO-GはFGD 2.278、ΔDiv 0.320、MM 1.786、ΔBA 0.063、ΔJerk 9.039、足接地誤差0.008 m、接触滑り速度0.052 m/s、推論時間5.96 ms/フレーム、ピークRAM増加18,542 MBだった。EMAGE+GMRはFGD 4.976、ΔBA 0.172、推論21.3 ms/フレーム、GestureLSM+GMRはFGD 5.008、ΔBA 0.158、推論20.5 ms/フレームで、ECHO-Gが4つのコスピーチ指標すべてと処理時間で最良だった。Human-RetargetedはΔJerk 8.001、足接地誤差0.003 m、接触滑り速度0.050 m/sで、ECHO-Gよりロボット動作品質の一部指標で良好だった。条件除去では、音声のみがFGD 2.360、ΔDiv 0.360、MM 1.702、ΔBA 0.081、テキストのみが2.436、0.429、1.681、0.113だったのに対し、音声+テキストは2.278、0.320、1.786、0.063で、コスピーチ指標で上回った。一方ΔJerkは音声のみが4.497、接触滑り速度は音声のみが0.038 m/sで、音声+テキストの9.039、0.052 m/sより良好だった。45人参加の動画評価では、総合平均がECHO-G 3.49、Human-Retargeted 2.65、EMAGE+GMR 1.80、単一モダリティ統合3.05で、ECHO-Gが最高だった。実機はUnitree G1でSONICトラッカー経由の動作を確認した。

限界

著者らは、コスピーチ特性の改善が必ずしもJerkギャップや足接地・接触滑りの改善に結びつかず、表現力とロボット動作整合性の両立が今後の課題と述べている。また、明示的な指示やダイクティックなジェスチャーの学習例が少なく、特定の意味動作を要求する発話への対応が限られる。現状は完全な音声とタイムコード付き書き起こしを必要とし、因果的なストリーミング生成には未対応である。実機評価は定性的であり、定量的な成功率や試行回数は報告されていない。査読済み会議・論文誌への採録は本論文では報告されていない。

産業への影響

人型ロボットの対話、接客、案内、プレゼンテーションなどの場面で、音声合成と連動した自然な全身ジェスチャー生成に応用できる。論文ではUnitree G1向けデータセットとコードが公開されており、同型ロボットや29自由度構成の小型人型ロボットを扱う企業は現在のプロトタイプ段階でも試用しやすい。ただし異なる自由度や体型のロボットへ展開するには、再ターゲティングとデータセット再構築が必要になる。1〜3年でカスタマーサービスや教育用ロボットへの組み込みが期待されるが、その前にストリーミング生成、明示的な指示的ジェスチャーの学習、ロボット動作品質と表現力の両立が解決される必要がある。

論文全文

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

CC BY 4.0 のもとで公開された論文です。出典を明記して転載しています。原文は arXiv:2609.39575(PDF)。

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance–motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio–text–robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Index Terms— Human and Humanoid Motion Analysis and Synthesis, Gesture, Posture and Facial Expressions, Humanoid Robot Systems, co-speech gesture generation, flow matching.

I Introduction

In human communication, gestures complement spoken content, convey emphasis, and organize the temporal structure of speech [1, 2]. Inspired by this coordination, we aim to equip speaking humanoids with body motion that reflects both how an utterance is spoken and what it conveys, while respecting the robot’s embodiment. We study full-body co-speech motion generation from the audio and timed transcript of the robot’s own utterance.

Human co-speech research provides a foundation for learning speech–motion relationships. Representative methods explicitly combine acoustic features with transcript-derived linguistic features to model speech rhythm and content [3, 4, 5]. These methods primarily generate human-motion representations. Extending speech-conditioned generation to humanoids requires robot-specific motion references and an interface for their physical execution. Recent robot-oriented methods address the generation of such references from speech. RoboGesture [6] studies audio-driven streaming generation of upper-body and hand gestures, while PhysDrift [7] explores robot-native generation with speech and text encoders.

These advances suggest a full-body robot co-speech generator should combine densely sampled acoustic cues with token-level linguistic content while preserving their distinct granularities. The one-to-many relationship between utterances and gestures [8] further motivates a generative formulation, and real-robot deployment favors direct prediction in robot space. In addition, existing public releases do not consistently provide paired audio–text–robot training data together with a benchmark spanning co-speech characteristics, robot-motion quality, and runtime efficiency.

Motivated by these considerations, we present ECHO-G, a framework for full-body humanoid co-speech generation from speech audio and timed transcripts (Fig. ). Its Speech-Grounded Diffusion Transformer (SGDiT) adds frame-aligned acoustic features to motion tokens and retrieves token-level linguistic context through global–local cross-attention, preserving the distinct granularities of the two conditions. Trained with rectified flow matching [9], SGDiT models the one-to-many relationship between utterances and full-body robot motion, enabling different motions to be sampled for the same utterance. A fixed pretrained whole-body motion tracker executes the joint-position components of the generated references.

To support training and evaluation in humanoid co-speech generation, we construct a BEAT2-derived robot-space dataset through retargeting and embodiment-specific quality filtering [4, 10]. We publicly release the dataset and code for training, inference, and evaluation, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Pipeline comparisons and modality ablations assess the generation-space and conditioning choices, while video ratings and physical demonstrations provide complementary perceptual and deployment evidence.

Our contributions are threefold:

  • We present ECHO-G, a full-body humanoid co-speech generation framework that jointly uses speech audio and timed transcripts, and demonstrate its deployment on a physical humanoid.
  • We develop SGDiT, a rectified-flow model combining frame-aligned acoustic conditioning with global–local transcript cross-attention for one-to-many full-body robot-motion generation.
  • We release a BEAT2-derived dataset pairing speech audio and timed transcripts with full-body robot motion, together with training, inference, and evaluation code. The accompanying benchmark covers co-speech characteristics, robot-motion quality, and runtime efficiency.

II Related Work

Fig. 2: SGDiT architecture and tracking interface. Frame-aligned acoustic features are combined with noisy motion-frame features to form motion tokens. Contextual transcript embeddings provide a shared key–value memory for global and temporally biased text attention. Their attention distributions are fused within each transformer block before value aggregation and residual injection into the motion stream. Rectified-flow sampling produces robot-motion references, whose joint-position components are passed to a fixed whole-body motion tracker for execution.

II-A Humanoid Whole-Body Motion Generation

Recent whole-body tracking systems [11, 12] enable humanoids to execute diverse motion references. Such references can be obtained by retargeting human motion, as in GMR [10] and OmniRetarget [13], or generated from language instructions, as in FRoM-W1 [14] and TextOp [15]. OMG [16] further unifies language, audio, and human-motion conditioning within a shared generative framework.

Within this broader setting, robot co-speech generation focuses on gestures accompanying spoken utterances. Yoon et al. [17] generate transcript-conditioned upper-body gestures and demonstrate execution on NAO. RoboPerform [18] generates humanoid motion from speech audio using a generic text prompt rather than the utterance transcript. RoboGesture [6] combines hierarchical semantic–acoustic conditioning with streaming generation of upper-body and hand gestures. PhysDrift [7] uses separate speech and text encoders for one-step robot-native motion generation. However, these approaches either omit utterance-specific audio–text conditioning, focus on upper-body motion, or do not fully specify the temporal organization of their multimodal features. ECHO-G combines frame-aligned acoustic features with token-level transcript embeddings for full-body robot-space generation, preserving their distinct granularities. We also provide paired audio–text–robot data and code for training, inference, and evaluation.

II-B Holistic Human Co-Speech Motion Generation

BEAT [19] provides multimodal speech–gesture data and introduces CaMN for integrating audio, text, and auxiliary conditions. Building on BEAT2, EMAGE [4] combines adaptive content–rhythm fusion with masked gesture modeling and compositional motion priors for holistic generation. DiffSHEG [20] jointly generates expressions and gestures through diffusion, while GestureLSM [5] combines flow matching, latent shortcut learning, and spatiotemporal modeling of body regions for efficient gesture generation. These methods primarily synthesize human-motion representations.

Evaluation considers distributional fidelity, motion variation, and speech–motion alignment. Yoon et al. [3] introduced Fréchet Gesture Distance (FGD), and EMAGE [4] adopted skeleton-aware features for distributional evaluation. Audio2Gestures [8] examines motion diversity and multimodality, while beat-alignment measures assess temporal correspondence between motion and audio [21, 4]. Our benchmark adapts these evaluation dimensions to robot motion and complements them with measures of robot-motion quality and runtime efficiency.

III Method

Fig. 3: Temporal conditioning in the local transcript-attention branch. Orange and purple heatmaps show global and local attention weights, respectively. The blue heatmap represents the clipped Gaussian log-prior \widetilde{b}_{\tau n} derived from motion-frame times u_{\tau} and token-center times m_{n}. Multiplying global weights by the exponentiated log-prior and renormalizing yields the local weights. Matrices are transposed for display; colored and dashed arrows indicate stronger and weaker attention, respectively.

As shown in Fig. 2, ECHO-G generates full-body robot-motion references from speech audio and timed transcripts. SGDiT integrates acoustic and linguistic conditions within a rectified-flow model, while a fixed whole-body motion tracker executes the generated joint-position references.

III-A Problem Formulation and Motion Representation

Given speech audio a and its word-timed transcript y, ECHO-G models a conditional distribution over full-body robot-motion sequences:

p_{\theta}(\mathbf{R}\mid a,y),
(1)

where \theta denotes the generator parameters and \mathbf{R}=[\mathbf{r}_{1},\ldots,\mathbf{r}_{T}]^{\top}\in\mathbb{R}^{T\times D} contains T motion frames. For each frame \tau=1,\ldots,T, we use D=39 features:

\mathbf{r}_{\tau}=\bigl[\boldsymbol{\rho}_{\tau}^{\top},\,\delta\psi_{\tau},\,(\mathbf{v}_{\tau}^{\mathrm{loc}})^{\top},\,\mathbf{q}_{\tau}^{\top}\bigr]^{\top},
(2)

where \boldsymbol{\rho}_{\tau}\in\mathbb{R}^{6} encodes the base orientation using a 6-D rotation representation, \delta\psi_{\tau}\in\mathbb{R} is the inter-frame yaw increment, \mathbf{v}_{\tau}^{\mathrm{loc}}\in\mathbb{R}^{3} is the base linear velocity in a yaw-aligned local frame, and \mathbf{q}_{\tau}\in\mathbb{R}^{29} contains the robot’s joint angles in a fixed order. Absolute root translation is omitted to make the learning target invariant to global position offsets.

The generator operates on normalized motion features:

\mathbf{x}_{\tau}=(\mathbf{r}_{\tau}-\boldsymbol{\mu})\oslash\boldsymbol{\sigma},
(3)

where \boldsymbol{\mu},\boldsymbol{\sigma}\in\mathbb{R}^{D} are the feature-wise mean and standard deviation computed from the training split, and \oslash denotes element-wise division. We denote the normalized sequence by \mathbf{X}=[\mathbf{x}_{1},\ldots,\mathbf{x}_{T}]^{\top}. Generated sequences are denormalized before evaluation or execution.

III-B Audio–Text Conditioning

Acoustic features extracted by a frozen speech encoder [22] are linearly interpolated to the T motion frames. The transcript text in y is tokenized into w_{1:N} and encoded by a frozen language model [23]. Separate layer normalization and learned affine projections map the two feature sequences to \mathbf{A}\in\mathbb{R}^{T\times d} and \mathbf{H}\in\mathbb{R}^{N\times d}, respectively, where d is the generator’s hidden dimension and N is the number of transcript tokens. The projected conditions retain their frame-level and token-level organization.

Using tokenizer character offsets, we derive approximate token intervals B=\{(s_{n},e_{n})\}_{n=1}^{N} from the word-level timestamps, where s_{n} and e_{n} are the associated start and end times. Temporal conditioning uses token centers m_{n}=(s_{n}+e_{n})/2 and motion-frame times u_{\tau}=(\tau-1)/f, where f is the motion frame rate. Both times are measured from the utterance onset. The combined condition is c=(\mathbf{A},\mathbf{H},B).

III-C Speech-Grounded Diffusion Transformer

SGDiT maps a noisy normalized motion sequence \mathbf{X}_{t}\in\mathbb{R}^{T\times D} and conditions c to the flow velocity v_{\theta}(\mathbf{X}_{t},t,c). Its transformer blocks combine temporal self-attention, transcript cross-attention, and feed-forward processing. Flow time t\in[0,1] modulates the blocks through adaptive layer normalization [24].

Acoustic conditioning. Each input motion token combines a projected motion frame, its aligned acoustic condition, and a positional embedding:

\mathbf{z}_{\tau}=\mathbf{W}_{r}\mathbf{x}_{t,\tau}+\mathbf{A}_{\tau}+\mathbf{p}_{\tau},
(4)

where \mathbf{x}_{t,\tau} is frame \tau of \mathbf{X}_{t}, \mathbf{W}_{r}\in\mathbb{R}^{d\times D} is a learned projection, and \mathbf{p}_{\tau}\in\mathbb{R}^{d} is a learned frame-position embedding. Temporal self-attention then exchanges information bidirectionally across the acoustically conditioned motion sequence.

Transcript conditioning. Cross-attention retrieves linguistic context through two paths over the same transcript features. The global path provides content-based access to the full token sequence, while the local path adds a preference for temporally nearby tokens. Updated motion features, after normalization and flow-time modulation, provide queries \mathbf{Q}_{\tau}; the projected transcript features \mathbf{H} provide keys \mathbf{K}_{n} and values \mathbf{V}_{n} shared by both paths. For one attention head, the content scores and global attention are

\displaystyle S_{\tau n} \\ \displaystyle=\gamma\,\hat{\mathbf{Q}}_{\tau}^{\top}\hat{\mathbf{K}}_{n}, \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{g}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}),

where hats denote L2-normalized queries and keys, \gamma is a bounded learned logit scale, and masked softmax normalizes over non-padding transcript tokens.

The local path adds a Gaussian temporal prior to the shared content scores, yielding the temporally reweighted attention illustrated in Fig. 3:

\displaystyle b_{\tau n} \\ \displaystyle=-\frac{(u_{\tau}-m_{n}-\delta)^{2}}{2\sigma^{2}}, \\ \displaystyle\widetilde{b}_{\tau n} \\ \displaystyle=\max(b_{\tau n},-\kappa), \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{l}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}+\widetilde{\mathbf{b}}),

where \sigma and \delta control the temporal width and offset, and \kappa bounds the log penalty. Both paths use the same token padding mask.

The two distributions are combined using temporal support:

\displaystyle g_{\tau} \\ \displaystyle=\alpha\max_{n\in I}\exp(b_{\tau n}), \\ \displaystyle\Pi_{\tau n} \\ \displaystyle=(1-g_{\tau})\Pi^{\mathrm{g}}_{\tau n}+g_{\tau}\Pi^{\mathrm{l}}_{\tau n},

where I contains the non-padding token indices and \alpha is the learned base mixing coefficient. Support uses the unclipped prior, reducing the local contribution when the frame is distant from all offset-adjusted token centers. The mixed weights aggregate the shared values into a text-conditioned update, which is projected and added to the motion features through a gated residual connection. A linear output head produces the final T\times D flow-velocity prediction.

III-D Training Objective

We train SGDiT with rectified flow matching [9, 25]. For a normalized motion–condition pair (\mathbf{X},c), we sample standard Gaussian noise \boldsymbol{\epsilon}\in\mathbb{R}^{T\times D} and a sequence-level flow time t\sim\mathcal{U}(0,1). The interpolated motion and target velocity are

\mathbf{X}_{t}=t\mathbf{X}+(1-t)\boldsymbol{\epsilon},\qquad\mathbf{V}^{\star}=\mathbf{X}-\boldsymbol{\epsilon}.
(8)

The predicted velocity \hat{\mathbf{V}}=v_{\theta}(\mathbf{X}_{t},t,c) is supervised by matching the target flow and its adjacent-frame differences:

\displaystyle\mathcal{L}_{\mathrm{flow}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\hat{\mathbf{V}},\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{temp}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\Delta_{\tau}\hat{\mathbf{V}},\Delta_{\tau}\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{gen}} \\ \displaystyle=\mathcal{L}_{\mathrm{flow}}+\lambda_{\mathrm{temp}}\mathcal{L}_{\mathrm{temp}}.

Here \Delta_{\tau} denotes first differences along the motion-frame axis, and \lambda_{\mathrm{temp}} weights the temporal term. MSE is averaged over feature dimensions and valid frames; the temporal term uses only adjacent pairs of valid frames.

Fig. 4: Conditioning comparison on an utterance outside BEAT2. Rows show joint audio–text conditioning (Ours), text-only, and audio-only outputs at matched frame indices. In the selected frames, joint conditioning exhibits broader arm extensions, whereas the unimodal outputs generally keep the hands closer to the torso. Colored arrows and circles highlight selected arm and hand movements.

III-E Inference and Tracking

At inference, the output length T is specified by the input clip duration at the motion frame rate. We initialize \mathbf{X}^{(0)}\in\mathbb{R}^{T\times D} with standard Gaussian noise and integrate the learned velocity field from flow time 0 to 1, keeping c fixed. Using K uniform Euler steps, we update

\displaystyle t_{k} \\ \displaystyle=\frac{k}{K}, \\ \displaystyle\mathbf{X}^{(k+1)} \\ \displaystyle=\mathbf{X}^{(k)}+\frac{1}{K}v_{\theta}(\mathbf{X}^{(k)},t_{k},c),

for k=0,\ldots,K-1. The final state \hat{\mathbf{X}}=\mathbf{X}^{(K)} is converted to robot-motion references by inverting the normalization in (3). Their joint-angle components are supplied to the fixed SONIC motion tracker [11] as joint-position references and executed alongside speech playback.

IV Experiments and Results

IV-A Experimental Setup

IV-A1 Datasets and Preprocessing

We construct a robot-space co-speech dataset from BEAT2 [4]. The original long-form recordings are segmented into 22,192 utterance-level clips at speech pauses using word-level forced alignment, retaining the corresponding audio, transcript, SMPL-X motion, and word timestamps. Each motion sequence is retargeted to the 29-DoF Unitree G1 using GMR [10]. The retargeted motions are converted to a unified Z-up coordinate system and foot-ground aligned using a clip-wise vertical root offset, followed by recomputation of forward kinematics and motion derivatives. We further canonicalize each sequence by removing its initial global yaw while preserving subsequent root dynamics. Robot motions remain at the native 30 fps throughout processing.

We then filter the resulting speech–robot pairs using both robot-motion quality checks and cross-modal consistency checks. For robot-motion quality, we discard clips that violate criteria on foot contact, self-collision, joint continuity and limits, smoothness, or severe high-frequency motion artifacts. We also retain only pairs whose audio and motion durations differ by at most one motion frame.

We adopt a speaker-held-out split, holding out three English speakers for validation and using the remaining speakers for training. After filtering, the final dataset contains 14,987 training clips and 3,242 validation clips.

IV-A2 Evaluation Metrics

All methods use the same held-out split as the candidate input set, with eligibility determined by their native input and output-length requirements. Prediction–reference comparisons use the common temporal prefix of each eligible pair.

Co-Speech Motion Characteristics. Fréchet Gesture Distance (FGD) [3, 4] measures distributional discrepancy using a shared skeleton-aware G1 joint-motion encoder. Div, MM, and BA are computed from forward-kinematic body positions with fixed base rotation and translation. Diversity (Div) measures the frame-wise L1 deviation from each sequence’s temporal mean pose. For stochastic generators, Multimodality (MM) [8] averages pairwise L1 distances among 20 samples generated for each input condition. Beat Alignment (BA) [21, 4] matches speech onsets to the nearest detected upper-body motion beats. We report absolute gaps to the matched reference statistics for Div and BA.

Robot Motion Quality.

Fig. 5: Real-robot execution on an utterance outside BEAT2. ECHO-G’s predicted joint angles are supplied as joint-position references to a fixed SONIC motion tracker while the corresponding speech is played. Frames progress from left to right through the book-recommendation utterance shown below, illustrating changes in arm extension, hand height, and torso posture. Colored arrows highlight selected arm movements, while the orange skeletal overlays outline the body configuration.

Body jerk [16] is estimated from third-order finite differences of reconstructed world-space body positions, scaled by the cube of the frame rate. We average jerk magnitudes over all valid frame–body pairs and report the absolute gap between generated and reference means (\DeltaJerk) [26]. Foot-ground error measures the vertical distance of the lowest sole-proxy surface from the ground plane. Contact sliding speed measures the maximum horizontal sole-point speed per foot, averaged over detected contact intervals.

Efficiency. End-to-end inference time covers input audio and aligned-transcript reading, online feature encoding, motion generation, and GMR or VAE-based motion mapping when applicable, ending at the robot-motion reference. We report total processing time divided by the total number of output frames. Peak RAM increase is the maximum request-window system memory usage above the corresponding pre-load idle baseline.

IV-A3 Implementation Details

We use frozen wav2vec 2.0 large XLSR and Qwen3.5-4B models to obtain 1024-D acoustic features and 2560-D contextual token embeddings, respectively. SGDiT contains 12 transformer blocks with a hidden dimension of 768, 8 attention heads, a feed-forward dimension of 2048, and learned temporal positional embeddings. The attention parameters \gamma,\sigma,\delta,\alpha are learned separately for each layer and head. Attention parameters are constrained to \gamma\in[1,16], \sigma\in[0.25,2] s, \delta\in[-0.5,0.5] s, and \alpha\in[0,0.5] using sigmoid/tanh mappings, with log-prior clipping threshold \kappa=12.

We train for 63,000 optimizer steps on NVIDIA RTX 4090 hardware using FP32 computation. AdamW uses an initial learning rate of 3\times 10^{-4} with cosine decay and no warm-up, weight decay of 10^{-4}, an effective batch size of 48, and gradient clipping at 1.0. We set \lambda_{\mathrm{temp}}=0.5 and jointly drop the acoustic and text encoder features together with the time-distance inputs with probability 0.1. The exponential moving average (EMA) decay is 0.999.

For physical deployment, motion generation runs on a Jetson AGX Orin, which is also used for the efficiency evaluation of all compared methods. Inference uses EMA weights from the checkpoint with the lowest EMA validation loss. Sampling uses eight Euler steps with classifier-free guidance (CFG) scale 1.0. We generate at most 600 motion frames at 30 fps and limit transcripts to 256 tokenizer tokens. Efficiency is measured with batch size one after warm-up.

IV-A4 Baselines and Ablations

We use the publicly released pretrained checkpoints of EMAGE [4] and GestureLSM [5], and retarget their generated human motions to G1 using GMR [10]. We additionally construct Human-Retargeted using the same audio–text conditioning design, backbone configuration, data split, and optimization settings as Ours, but generate normalized 136-D human motion. After denormalization using human-motion training statistics, a pretrained VAE-based mapping converts the samples to 39-D robot references. This mapping remains frozen during human-motion generator training.

Audio-only and Text-only are trained separately in robot space using only the indicated modality. Both variants share the data split, backbone configuration, and optimization settings with Ours. Text-only retains the supplied clip duration and word timings but does not use acoustic features.

IV-B Quantitative Results

TABLE I: Comparison of direct robot-space generation with human-motion generation followed by retargeting or learned mapping on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from matched ground-truth statistics over each method’s evaluated motion range. Best and second-best results are shown in bold and underlined text, respectively.
  • FGD \downarrow
    Co-Speech Motion Characteristics
    \DeltaDiv \downarrow
    MM \uparrow
    \DeltaBA \downarrow
    \DeltaJerk \downarrow
    Robot Motion Quality
    Foot Err. (m) \downarrow
    C-Slide (m/s) \downarrow
    E2E Time (ms/frame) \downarrow
    Efficiency
    Peak RAM \Delta (MB) \downarrow
    データなし
  • EMAGE+GMR
    Co-Speech Motion Characteristics
    4.976
    0.749
    0
    0.172
    Robot Motion Quality
    26.951
    0.013
    0.163
    Efficiency
    21.3
    2661
  • GestureLSM+GMR
    Co-Speech Motion Characteristics
    5.008
    0.561
    1.016
    0.158
    Robot Motion Quality
    24.444
    0.010
    0.169
    Efficiency
    20.5
    3804
  • Human-Retargeted
    Co-Speech Motion Characteristics
    4.725
    0.408
    1.498
    0.161
    Robot Motion Quality
    8.001
    0.003
    0.050
    Efficiency
    6.56
    18714
  • Ours
    Co-Speech Motion Characteristics
    2.278
    0.320
    1.786
    0.063
    Robot Motion Quality
    9.039
    0.008
    0.052
    Efficiency
    5.96
    18542
TABLE II: Ablation of conditioning modalities for direct robot-space generation on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from the shared ground-truth statistics. Best and second-best results are shown in bold and underlined text, respectively. Ties at the displayed precision receive identical highlighting.
  • Audio-only
    FGD \downarrow
    2.360
    \DeltaDiv \downarrow
    0.360
    MM \uparrow
    1.702
    \DeltaBA \downarrow
    0.081
    \DeltaJerk \downarrow
    4.497
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.038
  • Text-only
    FGD \downarrow
    2.436
    \DeltaDiv \downarrow
    0.429
    MM \uparrow
    1.681
    \DeltaBA \downarrow
    0.113
    \DeltaJerk \downarrow
    5.940
    Foot Err. (m) \downarrow
    0.009
    C-Slide (m/s) \downarrow
    0.049
  • Ours
    FGD \downarrow
    2.278
    \DeltaDiv \downarrow
    0.320
    MM \uparrow
    1.786
    \DeltaBA \downarrow
    0.063
    \DeltaJerk \downarrow
    9.039
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.052

Table I compares direct robot-space generation with human-motion generation followed by retargeting or learned mapping. ECHO-G achieves the best results on all four co-speech metrics and the lowest processing time per output frame. It also improves all three robot-motion quality metrics over EMAGE+GMR and GestureLSM+GMR. Human-Retargeted yields a smaller Jerk gap, lower foot-ground error, and less contact sliding. These results support direct robot-space generation for co-speech modeling with reduced processing time.

Table II evaluates conditioning modalities within the robot-space generator. Joint audio–text conditioning achieves the lowest FGD, \DeltaDiv, and \DeltaBA and the highest MM, outperforming both unimodal variants on the reported co-speech metrics. Audio-only yields the smallest Jerk gap and contact sliding speed, and matches joint conditioning in foot-ground error at the displayed precision. These results support combining acoustic and linguistic information to improve the evaluated co-speech characteristics.

IV-C Qualitative Results

We visualize motions generated from independently prepared audio–text utterances outside BEAT2. These examples provide qualitative evidence of generalization to speech inputs beyond the source dataset. In Fig. 4, audio-only conditioning produces rhythm-responsive motion, but gestures around semantically salient phrases remain small and less clearly related to the spoken content. Text-only conditioning produces content-related gestures, but their timing is less consistently aligned with the audio. Joint conditioning combines speech-responsive timing with more expansive, content-related gestures and fluid transitions in this example.

IV-D Real-Robot Deployment

We conduct real-robot experiments on a Unitree G1, executing the generated joint-position references through a fixed SONIC motion tracker [11] while playing the corresponding speech audio. Fig. 5 shows a representative trial using an independently prepared utterance outside BEAT2, with the robot accompanying its speech with generated body movements. Videos of additional real-robot trials are provided on the project page.

IV-E User Study

TABLE III: Mean user-study ratings from 45 participants on a five-point scale. Higher is better. Unimodal scores are pooled as described in the text. Best and second-best means are shown in bold and underlined text, respectively.
  • EMAGE+GMR
    Overall
    1.80
    Human- likeness
    1.87
    Rhythm Matching
    1.95
    Motion Quality
    1.59
  • Human-Retargeted
    Overall
    2.65
    Human- likeness
    2.67
    Rhythm Matching
    2.38
    Motion Quality
    2.90
  • Unimodal (pooled)
    Overall
    3.05
    Human- likeness
    3.01
    Rhythm Matching
    3.13
    Motion Quality
    3.01
  • Ours
    Overall
    3.49
    Human- likeness
    3.45
    Rhythm Matching
    3.31
    Motion Quality
    3.70

We conducted a video-rating study with 45 participants using G1 kinematic renderings in MuJoCo. Each participant completed nine trials, with three randomly selected from each of three criterion-specific pools comprising 33 utterances in total. The criteria were human-likeness, rhythm matching, and motion quality. Each trial presented four videos of the same utterance under matched rendering conditions, with hidden method identities and randomized positions. Participants rated each video on a five-point scale.

Each trial compared ECHO-G, Human-Retargeted, EMAGE+GMR, and an audio-only or text-only variant. Unimodal ratings were pooled over the observed trial allocation. Overall denotes the equally weighted mean of the three criterion scores. As shown in Table III, joint audio–text conditioning receives the highest overall mean rating and the highest mean ratings across all three criteria, providing perceptual support for our method.

V Discussion and Limitations

The experimental results support direct robot-space modeling and joint audio–text conditioning for humanoid co-speech generation. Benchmark comparisons show the benefits of these choices for the reported co-speech characteristics, with the direct generation pipeline also requiring less processing time. Qualitative examples, user ratings, and physical demonstrations provide complementary evidence of expressive gestures, perceived quality, and robot execution. Further improvement is needed to better reconcile expressive behavior with robot-motion consistency.

Several limitations remain. First, the gains in co-speech characteristics do not consistently translate into smaller Jerk gaps or better foot-contact measures, leaving room to improve expressive behavior and robot-motion consistency together. Second, the model learns broad speech–gesture associations, with limited training examples of explicit deictic or instructional gestures. This may constrain instruction-aware gesture generation when an utterance calls for a specific semantic motion. Third, the current system requires complete speech audio and timed transcripts and does not yet support causal streaming generation.

VI Conclusion

We presented ECHO-G for full-body humanoid co-speech generation from speech audio and timed transcripts. SGDiT combines frame-aligned acoustic conditioning with global–local transcript cross-attention to generate robot-motion references. Quantitative evaluation, a video-rating study, and physical demonstrations provide complementary evidence for the framework. The released dataset, benchmark, and code support reproducible research on humanoid co-speech generation. Future work will focus on jointly improving gesture expressiveness and robot-motion consistency, enriching training data for instruction-aware semantic gestures, and extending the framework to causal streaming generation.

References

  1. [1] D. McNeill (1992) Hand and mind: what gestures reveal about thought. University of Chicago Press, Chicago, IL, USA.
  2. [2] A. Kendon (2004) Gesture: visible action as utterance. Cambridge University Press, Cambridge, UK. External Links: Document
  3. [3] Y. Yoon, B. Cha, J. Lee, M. Jang, J. Lee, J. Kim, and G. Lee (2020) Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics 39 (6), pp. 1–16. External Links: Document
  4. [4] H. Liu, Z. Zhu, G. Becherini, Y. Peng, M. Su, Y. Zhou, X. Zhe, N. Iwamoto, B. Zheng, and M. J. Black (2024) EMAGE: towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1144–1154.
  5. [5] P. Liu, L. Song, J. Huang, H. Liu, and C. Xu (2025) GestureLSM: latent shortcut based co-speech gesture generation with spatial-temporal modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10929–10939.
  6. [6] Z. Wang, Z. Ren, P. Shi, Z. Wang, C. Lin, T. Wang, Z. Qi, L. Zhao, H. Wang, and L. Yi (2026) RoboGesture: real-time semantic-aligned co-speech gestures generation for humanoid interaction. In Proceedings of the European Conference on Computer Vision (ECCV),
  7. [7] Z. Liang, X. Xing, M. Yang, W. Zhou, and X. Xu (2026) PhysDrift: bridging the embodiment gap in humanoid co-speech motion generation. arXiv preprint arXiv:2606.19935.
  8. [8] J. Li, D. Kang, W. Pei, X. Zhe, Y. Zhang, Z. He, and L. Bao (2021) Audio2Gestures: generating diverse gestures from speech audio with conditional variational autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11293–11302.
  9. [9] X. Liu, C. Gong, and Q. Liu (2023) Flow straight and fast: learning to generate and transfer data with rectified flow. In International Conference on Learning Representations,
  10. [10] J. P. Araujo, Y. Ze, P. Xu, J. Wu, and C. K. Liu (2025) Retargeting matters: general motion retargeting for humanoid motion tracking. arXiv preprint arXiv:2510.02252.
  11. [11] Z. Luo, Y. Yuan, T. Wang, C. Li, F. Castañeda, S. Chen, Z. Cao, J. Li, D. Minor, Q. Ben, et al. (2026) SONIC: supersizing motion tracking for natural humanoid whole-body control. Science Robotics 11 (117), pp. eaed4592. External Links: Document
  12. [12] M. Chen, K. Wang, B. Zhang, X. Ma, Z. Yang, Y. Ren, Q. Huang, Z. Zhu, Y. Wang, and Z. Su (2026) HoloMotion-1 technical report. arXiv preprint arXiv:2605.15336.
  13. [13] L. Yang, X. Huang, Z. Wu, A. Kanazawa, P. Abbeel, C. Sferrazza, C. K. Liu, R. Duan, and G. Shi (2025) OmniRetarget: interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. arXiv preprint arXiv:2509.26633.
  14. [14] P. Li, Z. Zhuang, Y. Gao, Y. Dong, S. Li, C. Jiang, S. Dou, Z. Xi, E. Zhou, J. Huang, et al. (2026) FRoM-W1: towards general humanoid whole-body control with language instructions. arXiv preprint arXiv:2601.12799.
  15. [15] W. Xie, J. Zheng, J. Han, J. Shi, W. Zhang, C. Bai, and X. Li (2026) TextOp: real-time interactive text-driven humanoid robot motion generation and control. arXiv preprint arXiv:2602.07439.
  16. [16] S. Huang, K. Lee, D. Qiao, G. He, Z. Wang, Y. Li, S. Zhu, and H. Zhao (2026) OMG: omni-modal motion generation for generalist humanoid control. arXiv preprint arXiv:2606.10340.
  17. [17] Y. Yoon, W. Ko, M. Jang, J. Lee, J. Kim, and G. Lee (2019) Robots learn social skills: end-to-end learning of co-speech gesture generation for humanoid robots. In 2019 International Conference on Robotics and Automation (ICRA), pp. 4303–4309. External Links: Document
  18. [18] Z. Li, C. Chi, Y. Wei, B. Zhu, T. Huang, Z. Sun, Y. Peng, P. Wang, Z. Wang, F. Liu, C. Xu, and S. Zhang (2026) Do you have freestyle? expressive humanoid locomotion via audio control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 956–965.
  19. [19] H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng (2022) BEAT: a large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. In European Conference on Computer Vision, pp. 612–630.
  20. [20] J. Chen, Y. Liu, J. Wang, A. Zeng, Y. Li, and Q. Chen (2024) DiffSHEG: a diffusion-based approach for real-time speech-driven holistic 3D expression and gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7352–7361.
  21. [21] R. Li, S. Yang, D. A. Ross, and A. Kanazawa (2021) AI choreographer: music conditioned 3D dance generation with AIST++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13401–13412.
  22. [22] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli (2020) wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, Vol. 33, pp. 12449–12460.
  23. [23] Qwen Team (2026) Qwen3.5: towards native multimodal agents. External Links: Link
  24. [24] W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205.
  25. [25] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) Flow Matching for generative modeling. In International Conference on Learning Representations, External Links: Link
  26. [26] F. Fang, S. Yang, and W. Yang (2026) CoordSpeaker: exploiting gesture captioning for coordinated caption-empowered co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 30761–30771.