When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models
暂无中文版,显示英文原文。
摘自论文摘要
Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.
摘自论文摘要,摘要生成中。