When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models
日本語版は未提供のため、英語原文で表示しています。
要旨より
Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.
要旨より。当社による要約は作成中です。