MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
暂无中文版,显示英文原文。
摘自论文摘要
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.
摘自论文摘要,摘要生成中。