When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models
From the abstract
Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.
From the abstract. Our summary is in progress.