A Dual-System Vision-Language-Action Architecture for Whole-Upper-Body Humanoid Control
日本語版は未提供のため、英語原文で表示しています。
30秒で読む
A slow VLM planner (7–9 Hz) and a fast visuomotor policy (200 Hz) control a full humanoid upper body from language.
研究課題
How to get language understanding and fast reactive control in one deployable stack?
問題
Large VLMs are too slow for 200 Hz torque control.
従来の手法
Scripted skills or single-rate policies.
新しい手法
Two systems communicating via a latent vector; trained on ~500 h of teleop.
結果
Multi-robot collaboration and picking of thousands of novel household items.
限界
No public benchmark; single-embodiment; teleop data cost.
産業への影響
The architecture pattern is now widely copied; watch actuator bandwidth as the next bottleneck.