ROBOTNESS
研究

论文

机器人与具身智能领域的最新论文,以及每篇对产业的意义。

20 篇论文
暂无中文版,显示英文原文。
按技术筛选:VLA 视觉语言动作模型,清除
arXiv
VLA专业

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA专业

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Safe VLA专业

FailBank:将运行时反馈变为 VLA 策略的持久自我进化

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

FailBank 是一个四阶段自进化框架,将运行时 CBF 屏蔽的修正建议转化为学习记录,用于持续更新视觉-语言-动作模型,而无需在部署时使用屏蔽。在 VLA-Arena 静态障碍套件上,FailBank 在两个骨干模型上分别将任务成功率提高 8.5 和 6.9 个百分点,同时将策略诱导累积成本降低 35.6% 和 23.8%。与运行时屏蔽 AEGIS 相比,成功率分别提高 25.4 和 9.5 个百分点。本文为预印本,尚未经过同行评审。

arXiv
Memory-centric VLA进阶

Optimus-R:记忆中心VLA框架实现低数据适应与终身学习

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

本文提出预印本方法Optimus-R,一个记忆中心的视觉语言动作模型框架,将机器人适应任务转化为显式的查询技能记忆调优,而非反复更新模型参数。在LIBERO、CALVIN、RoboTwin 2.0仿真基准及真实机器人上,Optimus-R在低数据适应、跨域迁移和终身学习中均优于基线π0.5,显著缓解灾难性遗忘。该框架通过内联记忆接口、查询技能记忆库和轻量桥接适配机制,为机器人持续学习提供了高效路径。

arXiv
VLA专业

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Grounding专业

GroundingPI:以视觉基元构建4B接地基础模型,34项基准平均73.68%超越GPT-6 Astra

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

这是一篇预印本论文。作者提出 GroundingPI,一个基于点和边界框等视觉基元的 4B 参数接地基础模型,在 34 个接地基准上达到 73.68% 平均得分,超过更大的 GPT-6 Astra(71.54%)。该模型作为视觉骨干用于机器人操作和自动驾驶时,在 RoboTwin 2.0 全部四个分布外设置中优于所有主流骨干,相对最强基线最高提升 24.8%,并在 RoboCasa-GR1 上以 50% 演示数据超过其他基线用 75% 数据的表现。这项工作表明专门的接地预训练可成为物理智能的感知基础,尤其适合系统1的快速执行。

arXiv
Aerial Manipulation专业

合成监督与场景图组合实现跨站点空中操作,MPAR对齐执行进度

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

该预印本提出统一框架,将局部全身VLA行为组合为场景级空中操作。在仿真中完成21/50个跨站点任务(42.0%),物理平台验证了2/5次取水浇花任务。关键贡献包括无需真实平台演示的合成训练、MPAR对齐异步动作块,以及场景图引导的可行交接与拓扑转移。

arXiv
VLA专业

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
VLA进阶

ChunkTrust:利用动作专家证据自适应调整机器人策略的执行时域

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

机器人AI模型每次预测一段未来动作,执行多少步之后再重新观察环境,过去大多由工程师手动设为固定值。清华大学、北京智源人工智能研究院(BAAI)等机构发布的这篇预印本提出一个即插即用模块,在运行中自动决定这一步数,无需重新训练,就提升了Physical Intelligence和英伟达现有模型在仿真和真实双臂机器人上的成功率。

arXiv
Real-time VLA专业

两阶段非均匀Flow去噪把VLA推理从61.6毫秒压到22.0毫秒,真机叠衣验证实时执行方法

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

这是一篇arXiv预印本,尚未经过同行评审。作者以π0.5为基线,在双臂真机上测量了从相机、本体感知到命令响应的端到端延迟,并提出两阶段非均匀Flow Matching去噪,把推理步数从10步减到2步,模型推理时间从61.557毫秒降到21.956毫秒。在30次叠衣真机试验中,训练型方法Legato和免训练方法Temporal Smoothing表现最好;把两阶段去噪与这两种方法结合,推理成本大幅下降,但任务成功率有小幅回落。

arXiv
VLA专业

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
VLA专业

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
VLA专业

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

arXiv
VLA专业

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai, Fangming Liu

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies.

arXiv
VLA专业

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself.

arXiv
VLA专业

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.

arXiv
VLA专业

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo.

论文来自 arXiv 机器人领域。尚未生成摘要的论文,显示摘要开头部分。