MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
日本語版は未提供のため、英語原文で表示しています。
要旨より
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.
要旨より。当社による要約は作成中です。