Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
From the abstract
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.
From the abstract. Our summary is in progress.