A Dual-System Vision-Language-Action Architecture for Whole-Upper-Body Humanoid Control
아직 한국어판이 없어 영어 원문으로 표시.
30초 요약
A slow VLM planner (7–9 Hz) and a fast visuomotor policy (200 Hz) control a full humanoid upper body from language.
연구 질문
How to get language understanding and fast reactive control in one deployable stack?
문제
Large VLMs are too slow for 200 Hz torque control.
기존 접근
Scripted skills or single-rate policies.
새 접근
Two systems communicating via a latent vector; trained on ~500 h of teleop.
결과
Multi-robot collaboration and picking of thousands of novel household items.
한계
No public benchmark; single-embodiment; teleop data cost.
업계 영향
The architecture pattern is now widely copied; watch actuator bandwidth as the next bottleneck.