π0.5: a Vision-Language-Action Model with Open-World Generalization
Co-training a VLA on heterogeneous data (multi-robot, web, high-level subtask labels) lets a mobile manipulator clean unseen homes end-to-end.
Can a single policy generalise to entirely new homes without per-site data?
Prior VLAs overfit to the scenes and embodiments in their training set.
Fine-tune per task or per site; limited transfer across embodiments.
Hierarchical inference (semantic subtask prediction → low-level action) with a mixture of data sources.
Multi-minute tasks in unseen homes; large gains over ablations removing web or cross-embodiment data.
Still needs teleoperated data; failure recovery is limited; evaluation homes chosen by the authors.
Reinforces the case that data mixture, not just scale, drives generalisation — relevant to any humanoid company's data strategy.