ROBOTNESS
중급arXiv

ECHO-G: 오디오와 시간 대본으로 휴머노이드 전신 코스피치 동작을 생성하는 프레임워크

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
30초 요약

이 논문은 음성 오디오와 시간 정렬된 대본을 함께 조건으로 사용해 휴머노이드 로봇의 전신 코스피치 동작을 생성하는 ECHO-G를 제안한다. BEAT2 기반 로봇 데이터셋에서 기존 인간 동작 생성 후 리타게팅 파이프라인보다 코스피치 지표와 처리 시간이 우수했고, Unitree G1 실물 로봇 배포를 시연했다. 프리프린트다.

연구 질문

음성 오디오와 시간 정렬된 대본에서 휴머노이드 로봇의 전신 코스피치 동작을 로봇 고유 공간에서 직접 생성하고, 서로 다른 시간 해상도의 두 조건을 효과적으로 결합할 수 있는가

문제

기존 인간 코스피치 동작 생성기는 SMPL-X 같은 인간 모션 표현을 만들고 이를 로봇으로 리타게팅하는 방식이라 로봇 실체 제약을 학습하지 못하고 변환 오버헤드가 발생한다. 로봇 대상 연구는 상체나 손에 국한되거나 음성과 언어 조건의 시간적 구조를 제대로 보존하지 못했고, 오디오 텍스트 로봇 쌍 데이터와 평가 벤치마크도 부족했다.

기존 접근

EMAGE와 GestureLSM 등은 인간 모션을 생성한 뒤 GMR로 Unitree G1에 리타게팅했다. RoboGesture는 오디오 기반 스트리밍으로 상체와 손 동작을 만들었고, PhysDrift는 로봇 네이티브 생성을 시도했지만 음성과 텍스트 인코더를 단순 결합했다. 이 방식들은 프레임 단위 음향 신호와 토큰 단위 언어 의미의 해상도 차이를 보존하지 못하거나 전신 동작을 직접 만들지 못했다.

새 접근

ECHO-G는 Speech-Grounded Diffusion Transformer인 SGDiT를 사용한다. wav2vec 2.0 large XLSR로 추출한 1024차원 음향 특징을 모션 프레임에 정렬해 모션 토큰에 더하고, Qwen3.5-4B로 인코딩한 2560차원 대본 토큰 임베딩은 글로벌 경로와 시간 가우시안 사전을 적용한 로컬 경로의 교차 어텐션으로 검색한다. 교차 어텐션 가중치는 프레임과 토큰 중심 시간의 로그 사전을 더해 혼합한다. 정류 흐름 매칭으로 훈련하며 39차원 로봇 모션 특징, 즉 6D 베이스 회전, 요 증분, 로컬 선속도, 29개 관절각을 직접 예측한다. BEAT2에서 22,192개 발화를 분할하고 GMR로 Unitree G1에 리타게팅한 뒤 품질 필터를 거쳐 훈련 14,987개, 검증 3,242개 클립을 공개했다. 추론은 8 오일러 스텝, CFG 스케일 1.0, 최대 600프레임, 256토큰이다. 생성된 관절각은 고정된 SONIC 전신 모션 트래커에 위치 레퍼런스로 전달된다.

결과

BEAT2 화자 분리 평가에서 ECHO-G는 FGD 2.278, ΔDiv 0.320, MM 1.786, ΔBA 0.063을 기록해 EMAGE+GMR(FGD 4.976, ΔDiv 0.749, MM 0, ΔBA 0.172), GestureLSM+GMR(5.008, 0.561, 1.016, 0.158), Human-Retargeted(4.725, 0.408, 1.498, 0.161)보다 코스피치 지표가 좋았다. 처리 시간은 5.96 ms/frame으로 Human-Retargeted 6.56 ms/frame, EMAGE+GMR 21.3 ms/frame, GestureLSM+GMR 20.5 ms/frame보다 낮았다. 로봇 모션 품질에서 ΔJerk는 9.039로 Human-Retargeted 8.001보다 컸고, foot error 0.008 m, contact sliding 0.052 m/s로 Human-Retargeted(0.003 m, 0.050 m/s)보다 나빴지만 EMAGE+GMR과 GestureLSM+GMR보다는 좋았다. modality ablation에서 Ours는 Audio-only와 Text-only보다 FGD, ΔDiv, ΔBA, MM이 우수했고, Audio-only는 ΔJerk 4.497과 contact sliding 0.038 m/s로 더 작았다. 45명 사용자 평가에서 Ours는 overall 3.49, human-likeness 3.45, rhythm matching 3.31, motion quality 3.70으로 Human-Retargeted(2.65, 2.67, 2.38, 2.90)와 pooled unimodal(3.05, 3.01, 3.13, 3.01)보다 높았다. 실물 Unitree G1에서 SONIC 트래커로 생성 관절각을 실행했고, 효율 측정은 Jetson AGX Orin에서 배치 1, warm-up 후 수행했다.

한계

저자들은 코스피치 특성 개선이 로봇 모션 일관성의 Jerk와 foot-contact 지표 개선으로 이어지지 않았다고 밝혔다. 학습 데이터에 명시적 지시적 제스처가 부족해 특정 의미 동작이 필요한 발화에서 한계가 있고, 현재는 전체 오디오와 시간 정렬된 대본이 필요해 인과적 스트리밍 생성을 지원하지 않는다. 영어 BEAT2 화자 분할로만 평가했고, 실물 로봇 평가는 정량적 지표 없이 시연 수준이다.

업계 영향

휴머노이드 서비스 로봇, 안내 로봇, 교육용 로봇 등 음성 응답과 함께 자연스러운 전신 제스처가 필요한 제품에 적용 가능하다. Unitree G1 실증과 Jetson AGX Orin 추론은 하드웨어 통합 가능성을 보여주지만, 실시간 스트리밍 입력과 안전한 전신 동작 검증, 지시적 제스처 데이터 확보가 선행되어야 한다. 상용 적용은 1~3년 뒤가 될 수 있다.

논문 전문

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

CC BY 4.0 라이선스로 공개된 논문입니다. 출처를 밝혀 전재하며, 원문은 arXiv:2609.39575(PDF)에서 볼 수 있습니다.

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance–motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio–text–robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio–text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Index Terms— Human and Humanoid Motion Analysis and Synthesis, Gesture, Posture and Facial Expressions, Humanoid Robot Systems, co-speech gesture generation, flow matching.

I Introduction

In human communication, gestures complement spoken content, convey emphasis, and organize the temporal structure of speech [1, 2]. Inspired by this coordination, we aim to equip speaking humanoids with body motion that reflects both how an utterance is spoken and what it conveys, while respecting the robot’s embodiment. We study full-body co-speech motion generation from the audio and timed transcript of the robot’s own utterance.

Human co-speech research provides a foundation for learning speech–motion relationships. Representative methods explicitly combine acoustic features with transcript-derived linguistic features to model speech rhythm and content [3, 4, 5]. These methods primarily generate human-motion representations. Extending speech-conditioned generation to humanoids requires robot-specific motion references and an interface for their physical execution. Recent robot-oriented methods address the generation of such references from speech. RoboGesture [6] studies audio-driven streaming generation of upper-body and hand gestures, while PhysDrift [7] explores robot-native generation with speech and text encoders.

These advances suggest a full-body robot co-speech generator should combine densely sampled acoustic cues with token-level linguistic content while preserving their distinct granularities. The one-to-many relationship between utterances and gestures [8] further motivates a generative formulation, and real-robot deployment favors direct prediction in robot space. In addition, existing public releases do not consistently provide paired audio–text–robot training data together with a benchmark spanning co-speech characteristics, robot-motion quality, and runtime efficiency.

Motivated by these considerations, we present ECHO-G, a framework for full-body humanoid co-speech generation from speech audio and timed transcripts (Fig. ). Its Speech-Grounded Diffusion Transformer (SGDiT) adds frame-aligned acoustic features to motion tokens and retrieves token-level linguistic context through global–local cross-attention, preserving the distinct granularities of the two conditions. Trained with rectified flow matching [9], SGDiT models the one-to-many relationship between utterances and full-body robot motion, enabling different motions to be sampled for the same utterance. A fixed pretrained whole-body motion tracker executes the joint-position components of the generated references.

To support training and evaluation in humanoid co-speech generation, we construct a BEAT2-derived robot-space dataset through retargeting and embodiment-specific quality filtering [4, 10]. We publicly release the dataset and code for training, inference, and evaluation, together with a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Pipeline comparisons and modality ablations assess the generation-space and conditioning choices, while video ratings and physical demonstrations provide complementary perceptual and deployment evidence.

Our contributions are threefold:

  • We present ECHO-G, a full-body humanoid co-speech generation framework that jointly uses speech audio and timed transcripts, and demonstrate its deployment on a physical humanoid.
  • We develop SGDiT, a rectified-flow model combining frame-aligned acoustic conditioning with global–local transcript cross-attention for one-to-many full-body robot-motion generation.
  • We release a BEAT2-derived dataset pairing speech audio and timed transcripts with full-body robot motion, together with training, inference, and evaluation code. The accompanying benchmark covers co-speech characteristics, robot-motion quality, and runtime efficiency.

II Related Work

Fig. 2: SGDiT architecture and tracking interface. Frame-aligned acoustic features are combined with noisy motion-frame features to form motion tokens. Contextual transcript embeddings provide a shared key–value memory for global and temporally biased text attention. Their attention distributions are fused within each transformer block before value aggregation and residual injection into the motion stream. Rectified-flow sampling produces robot-motion references, whose joint-position components are passed to a fixed whole-body motion tracker for execution.

II-A Humanoid Whole-Body Motion Generation

Recent whole-body tracking systems [11, 12] enable humanoids to execute diverse motion references. Such references can be obtained by retargeting human motion, as in GMR [10] and OmniRetarget [13], or generated from language instructions, as in FRoM-W1 [14] and TextOp [15]. OMG [16] further unifies language, audio, and human-motion conditioning within a shared generative framework.

Within this broader setting, robot co-speech generation focuses on gestures accompanying spoken utterances. Yoon et al. [17] generate transcript-conditioned upper-body gestures and demonstrate execution on NAO. RoboPerform [18] generates humanoid motion from speech audio using a generic text prompt rather than the utterance transcript. RoboGesture [6] combines hierarchical semantic–acoustic conditioning with streaming generation of upper-body and hand gestures. PhysDrift [7] uses separate speech and text encoders for one-step robot-native motion generation. However, these approaches either omit utterance-specific audio–text conditioning, focus on upper-body motion, or do not fully specify the temporal organization of their multimodal features. ECHO-G combines frame-aligned acoustic features with token-level transcript embeddings for full-body robot-space generation, preserving their distinct granularities. We also provide paired audio–text–robot data and code for training, inference, and evaluation.

II-B Holistic Human Co-Speech Motion Generation

BEAT [19] provides multimodal speech–gesture data and introduces CaMN for integrating audio, text, and auxiliary conditions. Building on BEAT2, EMAGE [4] combines adaptive content–rhythm fusion with masked gesture modeling and compositional motion priors for holistic generation. DiffSHEG [20] jointly generates expressions and gestures through diffusion, while GestureLSM [5] combines flow matching, latent shortcut learning, and spatiotemporal modeling of body regions for efficient gesture generation. These methods primarily synthesize human-motion representations.

Evaluation considers distributional fidelity, motion variation, and speech–motion alignment. Yoon et al. [3] introduced Fréchet Gesture Distance (FGD), and EMAGE [4] adopted skeleton-aware features for distributional evaluation. Audio2Gestures [8] examines motion diversity and multimodality, while beat-alignment measures assess temporal correspondence between motion and audio [21, 4]. Our benchmark adapts these evaluation dimensions to robot motion and complements them with measures of robot-motion quality and runtime efficiency.

III Method

Fig. 3: Temporal conditioning in the local transcript-attention branch. Orange and purple heatmaps show global and local attention weights, respectively. The blue heatmap represents the clipped Gaussian log-prior \widetilde{b}_{\tau n} derived from motion-frame times u_{\tau} and token-center times m_{n}. Multiplying global weights by the exponentiated log-prior and renormalizing yields the local weights. Matrices are transposed for display; colored and dashed arrows indicate stronger and weaker attention, respectively.

As shown in Fig. 2, ECHO-G generates full-body robot-motion references from speech audio and timed transcripts. SGDiT integrates acoustic and linguistic conditions within a rectified-flow model, while a fixed whole-body motion tracker executes the generated joint-position references.

III-A Problem Formulation and Motion Representation

Given speech audio a and its word-timed transcript y, ECHO-G models a conditional distribution over full-body robot-motion sequences:

p_{\theta}(\mathbf{R}\mid a,y),
(1)

where \theta denotes the generator parameters and \mathbf{R}=[\mathbf{r}_{1},\ldots,\mathbf{r}_{T}]^{\top}\in\mathbb{R}^{T\times D} contains T motion frames. For each frame \tau=1,\ldots,T, we use D=39 features:

\mathbf{r}_{\tau}=\bigl[\boldsymbol{\rho}_{\tau}^{\top},\,\delta\psi_{\tau},\,(\mathbf{v}_{\tau}^{\mathrm{loc}})^{\top},\,\mathbf{q}_{\tau}^{\top}\bigr]^{\top},
(2)

where \boldsymbol{\rho}_{\tau}\in\mathbb{R}^{6} encodes the base orientation using a 6-D rotation representation, \delta\psi_{\tau}\in\mathbb{R} is the inter-frame yaw increment, \mathbf{v}_{\tau}^{\mathrm{loc}}\in\mathbb{R}^{3} is the base linear velocity in a yaw-aligned local frame, and \mathbf{q}_{\tau}\in\mathbb{R}^{29} contains the robot’s joint angles in a fixed order. Absolute root translation is omitted to make the learning target invariant to global position offsets.

The generator operates on normalized motion features:

\mathbf{x}_{\tau}=(\mathbf{r}_{\tau}-\boldsymbol{\mu})\oslash\boldsymbol{\sigma},
(3)

where \boldsymbol{\mu},\boldsymbol{\sigma}\in\mathbb{R}^{D} are the feature-wise mean and standard deviation computed from the training split, and \oslash denotes element-wise division. We denote the normalized sequence by \mathbf{X}=[\mathbf{x}_{1},\ldots,\mathbf{x}_{T}]^{\top}. Generated sequences are denormalized before evaluation or execution.

III-B Audio–Text Conditioning

Acoustic features extracted by a frozen speech encoder [22] are linearly interpolated to the T motion frames. The transcript text in y is tokenized into w_{1:N} and encoded by a frozen language model [23]. Separate layer normalization and learned affine projections map the two feature sequences to \mathbf{A}\in\mathbb{R}^{T\times d} and \mathbf{H}\in\mathbb{R}^{N\times d}, respectively, where d is the generator’s hidden dimension and N is the number of transcript tokens. The projected conditions retain their frame-level and token-level organization.

Using tokenizer character offsets, we derive approximate token intervals B=\{(s_{n},e_{n})\}_{n=1}^{N} from the word-level timestamps, where s_{n} and e_{n} are the associated start and end times. Temporal conditioning uses token centers m_{n}=(s_{n}+e_{n})/2 and motion-frame times u_{\tau}=(\tau-1)/f, where f is the motion frame rate. Both times are measured from the utterance onset. The combined condition is c=(\mathbf{A},\mathbf{H},B).

III-C Speech-Grounded Diffusion Transformer

SGDiT maps a noisy normalized motion sequence \mathbf{X}_{t}\in\mathbb{R}^{T\times D} and conditions c to the flow velocity v_{\theta}(\mathbf{X}_{t},t,c). Its transformer blocks combine temporal self-attention, transcript cross-attention, and feed-forward processing. Flow time t\in[0,1] modulates the blocks through adaptive layer normalization [24].

Acoustic conditioning. Each input motion token combines a projected motion frame, its aligned acoustic condition, and a positional embedding:

\mathbf{z}_{\tau}=\mathbf{W}_{r}\mathbf{x}_{t,\tau}+\mathbf{A}_{\tau}+\mathbf{p}_{\tau},
(4)

where \mathbf{x}_{t,\tau} is frame \tau of \mathbf{X}_{t}, \mathbf{W}_{r}\in\mathbb{R}^{d\times D} is a learned projection, and \mathbf{p}_{\tau}\in\mathbb{R}^{d} is a learned frame-position embedding. Temporal self-attention then exchanges information bidirectionally across the acoustically conditioned motion sequence.

Transcript conditioning. Cross-attention retrieves linguistic context through two paths over the same transcript features. The global path provides content-based access to the full token sequence, while the local path adds a preference for temporally nearby tokens. Updated motion features, after normalization and flow-time modulation, provide queries \mathbf{Q}_{\tau}; the projected transcript features \mathbf{H} provide keys \mathbf{K}_{n} and values \mathbf{V}_{n} shared by both paths. For one attention head, the content scores and global attention are

\displaystyle S_{\tau n} \\ \displaystyle=\gamma\,\hat{\mathbf{Q}}_{\tau}^{\top}\hat{\mathbf{K}}_{n}, \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{g}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}),

where hats denote L2-normalized queries and keys, \gamma is a bounded learned logit scale, and masked softmax normalizes over non-padding transcript tokens.

The local path adds a Gaussian temporal prior to the shared content scores, yielding the temporally reweighted attention illustrated in Fig. 3:

\displaystyle b_{\tau n} \\ \displaystyle=-\frac{(u_{\tau}-m_{n}-\delta)^{2}}{2\sigma^{2}}, \\ \displaystyle\widetilde{b}_{\tau n} \\ \displaystyle=\max(b_{\tau n},-\kappa), \\ \displaystyle\boldsymbol{\Pi}^{\mathrm{l}} \\ \displaystyle=\operatorname{maskedSoftmax}(\mathbf{S}+\widetilde{\mathbf{b}}),

where \sigma and \delta control the temporal width and offset, and \kappa bounds the log penalty. Both paths use the same token padding mask.

The two distributions are combined using temporal support:

\displaystyle g_{\tau} \\ \displaystyle=\alpha\max_{n\in I}\exp(b_{\tau n}), \\ \displaystyle\Pi_{\tau n} \\ \displaystyle=(1-g_{\tau})\Pi^{\mathrm{g}}_{\tau n}+g_{\tau}\Pi^{\mathrm{l}}_{\tau n},

where I contains the non-padding token indices and \alpha is the learned base mixing coefficient. Support uses the unclipped prior, reducing the local contribution when the frame is distant from all offset-adjusted token centers. The mixed weights aggregate the shared values into a text-conditioned update, which is projected and added to the motion features through a gated residual connection. A linear output head produces the final T\times D flow-velocity prediction.

III-D Training Objective

We train SGDiT with rectified flow matching [9, 25]. For a normalized motion–condition pair (\mathbf{X},c), we sample standard Gaussian noise \boldsymbol{\epsilon}\in\mathbb{R}^{T\times D} and a sequence-level flow time t\sim\mathcal{U}(0,1). The interpolated motion and target velocity are

\mathbf{X}_{t}=t\mathbf{X}+(1-t)\boldsymbol{\epsilon},\qquad\mathbf{V}^{\star}=\mathbf{X}-\boldsymbol{\epsilon}.
(8)

The predicted velocity \hat{\mathbf{V}}=v_{\theta}(\mathbf{X}_{t},t,c) is supervised by matching the target flow and its adjacent-frame differences:

\displaystyle\mathcal{L}_{\mathrm{flow}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\hat{\mathbf{V}},\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{temp}} \\ \displaystyle=\mathbb{E}\bigl[\operatorname{MSE}(\Delta_{\tau}\hat{\mathbf{V}},\Delta_{\tau}\mathbf{V}^{\star})\bigr], \\ \displaystyle\mathcal{L}_{\mathrm{gen}} \\ \displaystyle=\mathcal{L}_{\mathrm{flow}}+\lambda_{\mathrm{temp}}\mathcal{L}_{\mathrm{temp}}.

Here \Delta_{\tau} denotes first differences along the motion-frame axis, and \lambda_{\mathrm{temp}} weights the temporal term. MSE is averaged over feature dimensions and valid frames; the temporal term uses only adjacent pairs of valid frames.

Fig. 4: Conditioning comparison on an utterance outside BEAT2. Rows show joint audio–text conditioning (Ours), text-only, and audio-only outputs at matched frame indices. In the selected frames, joint conditioning exhibits broader arm extensions, whereas the unimodal outputs generally keep the hands closer to the torso. Colored arrows and circles highlight selected arm and hand movements.

III-E Inference and Tracking

At inference, the output length T is specified by the input clip duration at the motion frame rate. We initialize \mathbf{X}^{(0)}\in\mathbb{R}^{T\times D} with standard Gaussian noise and integrate the learned velocity field from flow time 0 to 1, keeping c fixed. Using K uniform Euler steps, we update

\displaystyle t_{k} \\ \displaystyle=\frac{k}{K}, \\ \displaystyle\mathbf{X}^{(k+1)} \\ \displaystyle=\mathbf{X}^{(k)}+\frac{1}{K}v_{\theta}(\mathbf{X}^{(k)},t_{k},c),

for k=0,\ldots,K-1. The final state \hat{\mathbf{X}}=\mathbf{X}^{(K)} is converted to robot-motion references by inverting the normalization in (3). Their joint-angle components are supplied to the fixed SONIC motion tracker [11] as joint-position references and executed alongside speech playback.

IV Experiments and Results

IV-A Experimental Setup

IV-A1 Datasets and Preprocessing

We construct a robot-space co-speech dataset from BEAT2 [4]. The original long-form recordings are segmented into 22,192 utterance-level clips at speech pauses using word-level forced alignment, retaining the corresponding audio, transcript, SMPL-X motion, and word timestamps. Each motion sequence is retargeted to the 29-DoF Unitree G1 using GMR [10]. The retargeted motions are converted to a unified Z-up coordinate system and foot-ground aligned using a clip-wise vertical root offset, followed by recomputation of forward kinematics and motion derivatives. We further canonicalize each sequence by removing its initial global yaw while preserving subsequent root dynamics. Robot motions remain at the native 30 fps throughout processing.

We then filter the resulting speech–robot pairs using both robot-motion quality checks and cross-modal consistency checks. For robot-motion quality, we discard clips that violate criteria on foot contact, self-collision, joint continuity and limits, smoothness, or severe high-frequency motion artifacts. We also retain only pairs whose audio and motion durations differ by at most one motion frame.

We adopt a speaker-held-out split, holding out three English speakers for validation and using the remaining speakers for training. After filtering, the final dataset contains 14,987 training clips and 3,242 validation clips.

IV-A2 Evaluation Metrics

All methods use the same held-out split as the candidate input set, with eligibility determined by their native input and output-length requirements. Prediction–reference comparisons use the common temporal prefix of each eligible pair.

Co-Speech Motion Characteristics. Fréchet Gesture Distance (FGD) [3, 4] measures distributional discrepancy using a shared skeleton-aware G1 joint-motion encoder. Div, MM, and BA are computed from forward-kinematic body positions with fixed base rotation and translation. Diversity (Div) measures the frame-wise L1 deviation from each sequence’s temporal mean pose. For stochastic generators, Multimodality (MM) [8] averages pairwise L1 distances among 20 samples generated for each input condition. Beat Alignment (BA) [21, 4] matches speech onsets to the nearest detected upper-body motion beats. We report absolute gaps to the matched reference statistics for Div and BA.

Robot Motion Quality.

Fig. 5: Real-robot execution on an utterance outside BEAT2. ECHO-G’s predicted joint angles are supplied as joint-position references to a fixed SONIC motion tracker while the corresponding speech is played. Frames progress from left to right through the book-recommendation utterance shown below, illustrating changes in arm extension, hand height, and torso posture. Colored arrows highlight selected arm movements, while the orange skeletal overlays outline the body configuration.

Body jerk [16] is estimated from third-order finite differences of reconstructed world-space body positions, scaled by the cube of the frame rate. We average jerk magnitudes over all valid frame–body pairs and report the absolute gap between generated and reference means (\DeltaJerk) [26]. Foot-ground error measures the vertical distance of the lowest sole-proxy surface from the ground plane. Contact sliding speed measures the maximum horizontal sole-point speed per foot, averaged over detected contact intervals.

Efficiency. End-to-end inference time covers input audio and aligned-transcript reading, online feature encoding, motion generation, and GMR or VAE-based motion mapping when applicable, ending at the robot-motion reference. We report total processing time divided by the total number of output frames. Peak RAM increase is the maximum request-window system memory usage above the corresponding pre-load idle baseline.

IV-A3 Implementation Details

We use frozen wav2vec 2.0 large XLSR and Qwen3.5-4B models to obtain 1024-D acoustic features and 2560-D contextual token embeddings, respectively. SGDiT contains 12 transformer blocks with a hidden dimension of 768, 8 attention heads, a feed-forward dimension of 2048, and learned temporal positional embeddings. The attention parameters \gamma,\sigma,\delta,\alpha are learned separately for each layer and head. Attention parameters are constrained to \gamma\in[1,16], \sigma\in[0.25,2] s, \delta\in[-0.5,0.5] s, and \alpha\in[0,0.5] using sigmoid/tanh mappings, with log-prior clipping threshold \kappa=12.

We train for 63,000 optimizer steps on NVIDIA RTX 4090 hardware using FP32 computation. AdamW uses an initial learning rate of 3\times 10^{-4} with cosine decay and no warm-up, weight decay of 10^{-4}, an effective batch size of 48, and gradient clipping at 1.0. We set \lambda_{\mathrm{temp}}=0.5 and jointly drop the acoustic and text encoder features together with the time-distance inputs with probability 0.1. The exponential moving average (EMA) decay is 0.999.

For physical deployment, motion generation runs on a Jetson AGX Orin, which is also used for the efficiency evaluation of all compared methods. Inference uses EMA weights from the checkpoint with the lowest EMA validation loss. Sampling uses eight Euler steps with classifier-free guidance (CFG) scale 1.0. We generate at most 600 motion frames at 30 fps and limit transcripts to 256 tokenizer tokens. Efficiency is measured with batch size one after warm-up.

IV-A4 Baselines and Ablations

We use the publicly released pretrained checkpoints of EMAGE [4] and GestureLSM [5], and retarget their generated human motions to G1 using GMR [10]. We additionally construct Human-Retargeted using the same audio–text conditioning design, backbone configuration, data split, and optimization settings as Ours, but generate normalized 136-D human motion. After denormalization using human-motion training statistics, a pretrained VAE-based mapping converts the samples to 39-D robot references. This mapping remains frozen during human-motion generator training.

Audio-only and Text-only are trained separately in robot space using only the indicated modality. Both variants share the data split, backbone configuration, and optimization settings with Ours. Text-only retains the supplied clip duration and word timings but does not use acoustic features.

IV-B Quantitative Results

TABLE I: Comparison of direct robot-space generation with human-motion generation followed by retargeting or learned mapping on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from matched ground-truth statistics over each method’s evaluated motion range. Best and second-best results are shown in bold and underlined text, respectively.
  • FGD \downarrow
    Co-Speech Motion Characteristics
    \DeltaDiv \downarrow
    MM \uparrow
    \DeltaBA \downarrow
    \DeltaJerk \downarrow
    Robot Motion Quality
    Foot Err. (m) \downarrow
    C-Slide (m/s) \downarrow
    E2E Time (ms/frame) \downarrow
    Efficiency
    Peak RAM \Delta (MB) \downarrow
    자료 없음
  • EMAGE+GMR
    Co-Speech Motion Characteristics
    4.976
    0.749
    0
    0.172
    Robot Motion Quality
    26.951
    0.013
    0.163
    Efficiency
    21.3
    2661
  • GestureLSM+GMR
    Co-Speech Motion Characteristics
    5.008
    0.561
    1.016
    0.158
    Robot Motion Quality
    24.444
    0.010
    0.169
    Efficiency
    20.5
    3804
  • Human-Retargeted
    Co-Speech Motion Characteristics
    4.725
    0.408
    1.498
    0.161
    Robot Motion Quality
    8.001
    0.003
    0.050
    Efficiency
    6.56
    18714
  • Ours
    Co-Speech Motion Characteristics
    2.278
    0.320
    1.786
    0.063
    Robot Motion Quality
    9.039
    0.008
    0.052
    Efficiency
    5.96
    18542
TABLE II: Ablation of conditioning modalities for direct robot-space generation on the BEAT2 speaker-held-out evaluation data. \DeltaDiv, \DeltaBA, and \DeltaJerk denote absolute deviations from the shared ground-truth statistics. Best and second-best results are shown in bold and underlined text, respectively. Ties at the displayed precision receive identical highlighting.
  • Audio-only
    FGD \downarrow
    2.360
    \DeltaDiv \downarrow
    0.360
    MM \uparrow
    1.702
    \DeltaBA \downarrow
    0.081
    \DeltaJerk \downarrow
    4.497
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.038
  • Text-only
    FGD \downarrow
    2.436
    \DeltaDiv \downarrow
    0.429
    MM \uparrow
    1.681
    \DeltaBA \downarrow
    0.113
    \DeltaJerk \downarrow
    5.940
    Foot Err. (m) \downarrow
    0.009
    C-Slide (m/s) \downarrow
    0.049
  • Ours
    FGD \downarrow
    2.278
    \DeltaDiv \downarrow
    0.320
    MM \uparrow
    1.786
    \DeltaBA \downarrow
    0.063
    \DeltaJerk \downarrow
    9.039
    Foot Err. (m) \downarrow
    0.008
    C-Slide (m/s) \downarrow
    0.052

Table I compares direct robot-space generation with human-motion generation followed by retargeting or learned mapping. ECHO-G achieves the best results on all four co-speech metrics and the lowest processing time per output frame. It also improves all three robot-motion quality metrics over EMAGE+GMR and GestureLSM+GMR. Human-Retargeted yields a smaller Jerk gap, lower foot-ground error, and less contact sliding. These results support direct robot-space generation for co-speech modeling with reduced processing time.

Table II evaluates conditioning modalities within the robot-space generator. Joint audio–text conditioning achieves the lowest FGD, \DeltaDiv, and \DeltaBA and the highest MM, outperforming both unimodal variants on the reported co-speech metrics. Audio-only yields the smallest Jerk gap and contact sliding speed, and matches joint conditioning in foot-ground error at the displayed precision. These results support combining acoustic and linguistic information to improve the evaluated co-speech characteristics.

IV-C Qualitative Results

We visualize motions generated from independently prepared audio–text utterances outside BEAT2. These examples provide qualitative evidence of generalization to speech inputs beyond the source dataset. In Fig. 4, audio-only conditioning produces rhythm-responsive motion, but gestures around semantically salient phrases remain small and less clearly related to the spoken content. Text-only conditioning produces content-related gestures, but their timing is less consistently aligned with the audio. Joint conditioning combines speech-responsive timing with more expansive, content-related gestures and fluid transitions in this example.

IV-D Real-Robot Deployment

We conduct real-robot experiments on a Unitree G1, executing the generated joint-position references through a fixed SONIC motion tracker [11] while playing the corresponding speech audio. Fig. 5 shows a representative trial using an independently prepared utterance outside BEAT2, with the robot accompanying its speech with generated body movements. Videos of additional real-robot trials are provided on the project page.

IV-E User Study

TABLE III: Mean user-study ratings from 45 participants on a five-point scale. Higher is better. Unimodal scores are pooled as described in the text. Best and second-best means are shown in bold and underlined text, respectively.
  • EMAGE+GMR
    Overall
    1.80
    Human- likeness
    1.87
    Rhythm Matching
    1.95
    Motion Quality
    1.59
  • Human-Retargeted
    Overall
    2.65
    Human- likeness
    2.67
    Rhythm Matching
    2.38
    Motion Quality
    2.90
  • Unimodal (pooled)
    Overall
    3.05
    Human- likeness
    3.01
    Rhythm Matching
    3.13
    Motion Quality
    3.01
  • Ours
    Overall
    3.49
    Human- likeness
    3.45
    Rhythm Matching
    3.31
    Motion Quality
    3.70

We conducted a video-rating study with 45 participants using G1 kinematic renderings in MuJoCo. Each participant completed nine trials, with three randomly selected from each of three criterion-specific pools comprising 33 utterances in total. The criteria were human-likeness, rhythm matching, and motion quality. Each trial presented four videos of the same utterance under matched rendering conditions, with hidden method identities and randomized positions. Participants rated each video on a five-point scale.

Each trial compared ECHO-G, Human-Retargeted, EMAGE+GMR, and an audio-only or text-only variant. Unimodal ratings were pooled over the observed trial allocation. Overall denotes the equally weighted mean of the three criterion scores. As shown in Table III, joint audio–text conditioning receives the highest overall mean rating and the highest mean ratings across all three criteria, providing perceptual support for our method.

V Discussion and Limitations

The experimental results support direct robot-space modeling and joint audio–text conditioning for humanoid co-speech generation. Benchmark comparisons show the benefits of these choices for the reported co-speech characteristics, with the direct generation pipeline also requiring less processing time. Qualitative examples, user ratings, and physical demonstrations provide complementary evidence of expressive gestures, perceived quality, and robot execution. Further improvement is needed to better reconcile expressive behavior with robot-motion consistency.

Several limitations remain. First, the gains in co-speech characteristics do not consistently translate into smaller Jerk gaps or better foot-contact measures, leaving room to improve expressive behavior and robot-motion consistency together. Second, the model learns broad speech–gesture associations, with limited training examples of explicit deictic or instructional gestures. This may constrain instruction-aware gesture generation when an utterance calls for a specific semantic motion. Third, the current system requires complete speech audio and timed transcripts and does not yet support causal streaming generation.

VI Conclusion

We presented ECHO-G for full-body humanoid co-speech generation from speech audio and timed transcripts. SGDiT combines frame-aligned acoustic conditioning with global–local transcript cross-attention to generate robot-motion references. Quantitative evaluation, a video-rating study, and physical demonstrations provide complementary evidence for the framework. The released dataset, benchmark, and code support reproducible research on humanoid co-speech generation. Future work will focus on jointly improving gesture expressiveness and robot-motion consistency, enriching training data for instruction-aware semantic gestures, and extending the framework to causal streaming generation.

References

  1. [1] D. McNeill (1992) Hand and mind: what gestures reveal about thought. University of Chicago Press, Chicago, IL, USA.
  2. [2] A. Kendon (2004) Gesture: visible action as utterance. Cambridge University Press, Cambridge, UK. External Links: Document
  3. [3] Y. Yoon, B. Cha, J. Lee, M. Jang, J. Lee, J. Kim, and G. Lee (2020) Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics 39 (6), pp. 1–16. External Links: Document
  4. [4] H. Liu, Z. Zhu, G. Becherini, Y. Peng, M. Su, Y. Zhou, X. Zhe, N. Iwamoto, B. Zheng, and M. J. Black (2024) EMAGE: towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1144–1154.
  5. [5] P. Liu, L. Song, J. Huang, H. Liu, and C. Xu (2025) GestureLSM: latent shortcut based co-speech gesture generation with spatial-temporal modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10929–10939.
  6. [6] Z. Wang, Z. Ren, P. Shi, Z. Wang, C. Lin, T. Wang, Z. Qi, L. Zhao, H. Wang, and L. Yi (2026) RoboGesture: real-time semantic-aligned co-speech gestures generation for humanoid interaction. In Proceedings of the European Conference on Computer Vision (ECCV),
  7. [7] Z. Liang, X. Xing, M. Yang, W. Zhou, and X. Xu (2026) PhysDrift: bridging the embodiment gap in humanoid co-speech motion generation. arXiv preprint arXiv:2606.19935.
  8. [8] J. Li, D. Kang, W. Pei, X. Zhe, Y. Zhang, Z. He, and L. Bao (2021) Audio2Gestures: generating diverse gestures from speech audio with conditional variational autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11293–11302.
  9. [9] X. Liu, C. Gong, and Q. Liu (2023) Flow straight and fast: learning to generate and transfer data with rectified flow. In International Conference on Learning Representations,
  10. [10] J. P. Araujo, Y. Ze, P. Xu, J. Wu, and C. K. Liu (2025) Retargeting matters: general motion retargeting for humanoid motion tracking. arXiv preprint arXiv:2510.02252.
  11. [11] Z. Luo, Y. Yuan, T. Wang, C. Li, F. Castañeda, S. Chen, Z. Cao, J. Li, D. Minor, Q. Ben, et al. (2026) SONIC: supersizing motion tracking for natural humanoid whole-body control. Science Robotics 11 (117), pp. eaed4592. External Links: Document
  12. [12] M. Chen, K. Wang, B. Zhang, X. Ma, Z. Yang, Y. Ren, Q. Huang, Z. Zhu, Y. Wang, and Z. Su (2026) HoloMotion-1 technical report. arXiv preprint arXiv:2605.15336.
  13. [13] L. Yang, X. Huang, Z. Wu, A. Kanazawa, P. Abbeel, C. Sferrazza, C. K. Liu, R. Duan, and G. Shi (2025) OmniRetarget: interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. arXiv preprint arXiv:2509.26633.
  14. [14] P. Li, Z. Zhuang, Y. Gao, Y. Dong, S. Li, C. Jiang, S. Dou, Z. Xi, E. Zhou, J. Huang, et al. (2026) FRoM-W1: towards general humanoid whole-body control with language instructions. arXiv preprint arXiv:2601.12799.
  15. [15] W. Xie, J. Zheng, J. Han, J. Shi, W. Zhang, C. Bai, and X. Li (2026) TextOp: real-time interactive text-driven humanoid robot motion generation and control. arXiv preprint arXiv:2602.07439.
  16. [16] S. Huang, K. Lee, D. Qiao, G. He, Z. Wang, Y. Li, S. Zhu, and H. Zhao (2026) OMG: omni-modal motion generation for generalist humanoid control. arXiv preprint arXiv:2606.10340.
  17. [17] Y. Yoon, W. Ko, M. Jang, J. Lee, J. Kim, and G. Lee (2019) Robots learn social skills: end-to-end learning of co-speech gesture generation for humanoid robots. In 2019 International Conference on Robotics and Automation (ICRA), pp. 4303–4309. External Links: Document
  18. [18] Z. Li, C. Chi, Y. Wei, B. Zhu, T. Huang, Z. Sun, Y. Peng, P. Wang, Z. Wang, F. Liu, C. Xu, and S. Zhang (2026) Do you have freestyle? expressive humanoid locomotion via audio control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 956–965.
  19. [19] H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng (2022) BEAT: a large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. In European Conference on Computer Vision, pp. 612–630.
  20. [20] J. Chen, Y. Liu, J. Wang, A. Zeng, Y. Li, and Q. Chen (2024) DiffSHEG: a diffusion-based approach for real-time speech-driven holistic 3D expression and gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7352–7361.
  21. [21] R. Li, S. Yang, D. A. Ross, and A. Kanazawa (2021) AI choreographer: music conditioned 3D dance generation with AIST++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13401–13412.
  22. [22] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli (2020) wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, Vol. 33, pp. 12449–12460.
  23. [23] Qwen Team (2026) Qwen3.5: towards native multimodal agents. External Links: Link
  24. [24] W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205.
  25. [25] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) Flow Matching for generative modeling. In International Conference on Learning Representations, External Links: Link
  26. [26] F. Fang, S. Yang, and W. Yang (2026) CoordSpeaker: exploiting gesture captioning for coordinated caption-empowered co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 30761–30771.