MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
日本語版は未提供のため、英語原文で表示しています。
要旨より
Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.
要旨より。当社による要約は作成中です。