ROBOTNESS
专业arXiv被引 120 次示例解读

π0.5: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence teamPhysical Intelligence
暂无中文版,显示英文原文。
30 秒速读

Co-training a VLA on heterogeneous data (multi-robot, web, high-level subtask labels) lets a mobile manipulator clean unseen homes end-to-end.

研究问题

Can a single policy generalise to entirely new homes without per-site data?

问题

Prior VLAs overfit to the scenes and embodiments in their training set.

既有方法

Fine-tune per task or per site; limited transfer across embodiments.

新方法

Hierarchical inference (semantic subtask prediction → low-level action) with a mixture of data sources.

结果

Multi-minute tasks in unseen homes; large gains over ablations removing web or cross-embodiment data.

局限

Still needs teleoperated data; failure recovery is limited; evaluation homes chosen by the authors.

产业影响

Reinforces the case that data mixture, not just scale, drives generalisation — relevant to any humanoid company's data strategy.