PhysWAM: Physically Consistent World Action Model for Autonomous Driving
From the abstract
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer.
From the abstract. Our summary is in progress.