MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
Noch nicht auf Deutsch verfügbar: Anzeige im englischen Original.
Aus dem Abstract
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.
Aus dem Abstract. Unsere Zusammenfassung folgt.