Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence
This preprint introduces Magic-W0, a robot manipulation foundation model that couples continuous action generation with explicit latent predictions of current 3D geometry, action-induced 3D motion, and future semantics. On RoboDojo-Sim it posts the highest average score among world-action models at 27.10, and on five real-robot tasks it reaches 94.6% average success after fine-tuning, beating π0.5 at 91.8%. The result matters because it shifts policies from blind imitation toward predicting how candidate actions will change the physical world, with gains in generalization and long-horizon manipulation.