MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
아직 한국어판이 없어 영어 원문으로 표시.
초록 발췌
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.
초록에서 가져왔습니다. 요약을 준비 중입니다.