MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
From the abstract
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.
From the abstract. Our summary is in progress.