A Dual-System Vision-Language-Action Architecture for Whole-Upper-Body Humanoid Control
In 30 seconds
A slow VLM planner (7–9 Hz) and a fast visuomotor policy (200 Hz) control a full humanoid upper body from language.
Research question
How to get language understanding and fast reactive control in one deployable stack?
Problem
Large VLMs are too slow for 200 Hz torque control.
Previous approach
Scripted skills or single-rate policies.
New approach
Two systems communicating via a latent vector; trained on ~500 h of teleop.
Results
Multi-robot collaboration and picking of thousands of novel household items.
Limitations
No public benchmark; single-embodiment; teleop data cost.
Industry impact
The architecture pattern is now widely copied; watch actuator bandwidth as the next bottleneck.