World-action modelExpert
Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang
This preprint introduces Magic-W0, a robot manipulation foundation model that couples continuous action generation with explicit latent predictions of current 3D geometry, action-induced 3D motion, and future semantics. On RoboDojo-Sim it posts the highest average score among world-action models at 27.10, and on five real-robot tasks it reaches 94.6% average success after fine-tuning, beating π0.5 at 91.8%. The result matters because it shifts policies from blind imitation toward predicting how candidate actions will change the physical world, with gains in generalization and long-horizon manipulation.