ROBOTNESS
IntermediatearXiv

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
In 30 seconds

This preprint introduces RoboCoach, a world-model-guided coaching loop that decides which subtask demonstrations to collect and which reusable skill expert adapters to update. On real robots, 150 added subtask demonstrations lift complete-task success from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, and coached skills transfer to four unseen compositions where a baseline scores 0%. The work matters because it shows world models can direct scarce real-world supervision toward the specific reusable skills that fail, rather than requiring expensive end-to-end data.

Research question

Can an action-conditioned world model guide which subtask demonstrations to acquire and which reusable skill expert adapters to update in order to improve long-horizon robot manipulation with limited additional supervision?

Problem

Long-horizon manipulation reuses skills across many task compositions, but end-to-end demonstrations are expensive and local execution errors propagate across later subtasks. Existing active and corrective imitation learning can select additional supervision, but they do not connect imagined failures to a specific reusable skill-level component that should receive the update.

Previous approach

Prior work combined generalist vision-language-action policies, active/corrective imitation learning, action-conditioned world models for policy evaluation, and modular LoRA adapters. The shortcoming is that these pieces are usually applied separately: world models simulate outcomes but do not directly produce a concrete skill-level request for real demonstrations, and active learning often updates a single shared policy rather than the specific reusable skill that failed.

New approach

RoboCoach introduces a Route–Imagine–Diagnose–Improve loop. A shared action-conditioned world model called CoachWorld rolls out the active skill expert in closed loop; a progress judge detects the first subtask that times out; aggregated imagined timeouts are ranked by a task-balanced scorecard; and the top pairs receive targeted human demonstrations that update only the corresponding expert LoRA adapters while the shared VLA backbone stays frozen.

Results

CoachWorld achieves LPIPS 0.0996 and FVD 51.71 on DROID-180, and LPIPS 0.1712 on Lab Franka-180, outperforming Ctrl-World, Cosmos 3, and OSCAR-2B on most metrics. Across 22 task–policy pairs, imagined and deployed success have Pearson r=0.820 and Spearman ρ=0.840. The RoboMeter progress judge achieves 0.414 s switch MAE, 82.73% [email protected], and 88.21% outcome F1 on reference video; under CoachWorld-generated rollouts these are 0.815 s, 62.73%, and 83.84%. In simulation after three coaching rounds, RoboCoach reaches 71.2% on LIBERO and 68.0% on RoboTwin. Applying the same targeted demonstrations to selected experts instead of a shared global adapter adds +3.4 points on LIBERO and +13.2 points on RoboTwin; world-model targeting beats random expert selection by +2.8 and +12.8 points. On real robots, 150 added subtask demonstrations per platform raise Franka success from 13.3% to 75.0% and AgileX from 40.0% to 83.8%, versus 30.0% and 47.5% for uniform acquisition with a shared adapter. On four held-out compositions, RoboCoach averages 35.0% across 20 trials per route (65%, 15%, 25%, 35% per route) while the shared-policy baseline scores 0%.

Limitations

Authors note that occlusion and out-of-view contact/collision degrade CoachWorld prediction, seed agreement cannot remove systematic world-model bias, and the expert library is predefined by skill semantics. The evaluation uses only two real-robot platforms, a small set of held-out compositions (four routes, 20 trials each), and fixed demonstration budgets. The paper does not report a code release.

Industry impact

Robot OEMs, warehouse and logistics automation vendors, and generalist VLA providers could use this to prioritize data collection for long-horizon manipulation and update reusable skill libraries with far fewer human demonstrations. In 1–3 years, offline world-model diagnosis could route human data collection in labs or factories. In 3+ years, it could support fleet self-improvement if world-model and progress-judge reliability are proven across varied environments, occlusion, and safety-critical tasks.

The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.