Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model
This preprint introduces Optimus-R, a memory-centric VLA framework that turns robotic manipulation adaptation into explicit query-skill memory tuning instead of repeated parameter updates. With only 30% of training data it raises LIBERO average success to 88.6% (vs 72.1% for π0.5) and reaches 33.3% real-world success from 20 demos per task; in lifelong learning it reduces forgetting to 10.0% vs 15.0%. The approach matters for low-data robot adaptation and retaining prior skills.