Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
暂无中文版,显示英文原文。
摘自论文摘要
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.
摘自论文摘要,摘要生成中。