ROBOTNESS
ExpertarXiv

DiffWAM: A Fast and Efficient Navigation World Action Model

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou
In 30 seconds

DiffWAM is a preprint that introduces a UAV navigation world-action model. It recovers camera trajectories directly from the predictive features of a frozen video model, avoiding future-video decoding and multi-frame geometric reconstruction at deployment. The paper reports a trajectory RMSE of 0.3492 m and adds FastDreamer, an execution system that overlaps computation with flight for continuous UAV control.

Research question

Can the motion implicit in future visual prediction be recovered directly from the predictive representations of a frozen video model, rather than by generating future videos and reconstructing geometry?

Problem

Using pretrained video foundation models for embodied UAV navigation usually requires costly future-video synthesis and geometric reconstruction, which is too expensive and slow for real-time continuous flight.

Previous approach

Earlier methods converted video-model priors into UAV motion by generating future video rollouts and performing geometric reconstruction, then inferring motion from those outputs. Their shortcoming is the high computational cost and latency of decoding future frames and reconstructing multi-frame geometry during deployment.

New approach

DiffWAM is a geometry-conditioned navigation world-action model that maps multi-level predictive features from a frozen video model directly to continuous camera trajectories. A Grid-Motion module preserves spatial-temporal motion associations, and a Latent2Pose module grounds those features with first-frame geometry to produce metrically meaningful 3D motion. During deployment it does not decode future video or reconstruct multi-frame geometry; those are used only for offline supervision. FastDreamer overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution.

Results

The abstract reports a trajectory RMSE of 0.3492 m. It also begins to state an endpoint success metric but the abstract is truncated, so that value is not fully reported. No baselines, trial counts, or simulation/real-world split are given in the abstract.

Limitations

The abstract does not list explicit limitations. Evident limitations include a single reported trajectory error, an incomplete endpoint success metric, no baseline comparison, no stated number of trials or hardware platform, and no code or compute details. As a preprint, the results have not yet been peer-reviewed.

Industry impact

UAV autonomy providers, drone platform makers, and firms building real-time visual navigation stacks could potentially use this to reduce compute and latency when applying video foundation models to flight control. Near-term (1–3 years) research prototypes are plausible if the method is validated beyond the abstract; production use in 3+ years would require peer-reviewed real-world flight results, open code/weights, integration with flight controllers, and safety certification.

The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.