Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
日本語版は未提供のため、英語原文で表示しています。
要旨より
Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress.
要旨より。当社による要約は作成中です。