Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model
From the abstract
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning.
From the abstract. Our summary is in progress.
The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.