π0.5: a Vision-Language-Action Model with Open-World Generalization
Co-training a VLA on heterogeneous data (multi-robot, web, high-level subtask labels) lets a mobile manipulator clean unseen homes end-to-end.
New robotics and physical-AI papers, with what each one means for the industry.
Co-training a VLA on heterogeneous data (multi-robot, web, high-level subtask labels) lets a mobile manipulator clean unseen homes end-to-end.
A slow VLM planner (7–9 Hz) and a fast visuomotor policy (200 Hz) control a full humanoid upper body from language.
RL policies trained in GPU simulation transfer zero-shot to a commodity humanoid on rough terrain.
Survey of sensor modalities, coverage and learning methods for in-hand manipulation.
A video world model trained on robot egocentric footage predicts outcomes of actions for evaluation and planning.
Cobot adoption in Korean SMEs is limited by integration cost more than by unit price.
Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.