Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
From the abstract
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach.
From the abstract. Our summary is in progress.