Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
From the abstract
Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress.
From the abstract. Our summary is in progress.