ROBOTNESS
ExpertarXiv120 citationsSample brief

π0.5: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence teamPhysical Intelligence
In 30 seconds

Co-training a VLA on heterogeneous data (multi-robot, web, high-level subtask labels) lets a mobile manipulator clean unseen homes end-to-end.

Research question

Can a single policy generalise to entirely new homes without per-site data?

Problem

Prior VLAs overfit to the scenes and embodiments in their training set.

Previous approach

Fine-tune per task or per site; limited transfer across embodiments.

New approach

Hierarchical inference (semantic subtask prediction → low-level action) with a mixture of data sources.

Results

Multi-minute tasks in unseen homes; large gains over ablations removing web or cross-embodiment data.

Limitations

Still needs teleoperated data; failure recovery is limited; evaluation homes chosen by the authors.

Industry impact

Reinforces the case that data mixture, not just scale, drives generalisation — relevant to any humanoid company's data strategy.