In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks
From the abstract
We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow.
From the abstract. Our summary is in progress.