ROBOTNESS
研究

论文

机器人与具身智能领域的最新论文,以及每篇对产业的意义。

106 篇论文
暂无中文版,显示英文原文。
arXiv
Perception专业

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RL专业

PL-MPC 在模型预测控制中同时改进 critic 监督、终端估计与策略蒸馏,真机 unseen 尺寸成功率超 TD-M(PC)2

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

PL-MPC 是 arXiv 预印本,未注明经同行评审。该方法在 TD-M(PC)2 上同步修改 critic 监督、MPC 终端值估计和 planner 到策略的转移,不改世界模型与 MPPI。在 HumanoidBench 的 balance-hard 上 TAR 从 98±18 升至 387±255,hurdle 从 199±13 升至 466±200;在 KUKA IIWA14 真机 wrench–nut 对齐中,训练尺寸和两个未见尺寸的观测成功率均高于 TD-M(PC)2。

arXiv
Trajectory Correction进阶

MaDE:冻结后处理算子将神经轨迹预测强制到学习动力学流形,显著降低残差但位移误差上升

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

论文提出Markovian Dynamics Enforcer(MaDE),一个无需真值控制、时间不变的后处理算子,可为任意神经轨迹预测器附加动力学一致性与不等式约束修正。在inD真实车辆轨迹上,MaDE将一步运动学自行车模型残差从原始预测器的0.1703至0.1714降至0.0071至0.0072,但平均位移误差增大1.57至1.83倍。该工作为arXiv预印本,未经同行评审。

arXiv
Terrain traversability专业

经验驱动的持续学习预测四足机器人地形可通过性,验证门控降低遗忘23.1%

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

这篇预印本提出一个持续学习管线,利用冻结的DINOv3视觉特征和证据回归器从接触前图像预测五种足端交互指标,并通过有界重放和验证门控持续更新。在Unitree Go2真机顺序经过三种新地形时,带历史验证门控的方法比仅重放方法将最终锚点负对数似然退化降低23.1%,同时新地形适应相近。该预印本尚未经同行评审。

arXiv
Real-time VLA专业

两阶段非均匀Flow去噪把VLA推理从61.6毫秒压到22.0毫秒,真机叠衣验证实时执行方法

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

这是一篇arXiv预印本,尚未经过同行评审。作者以π0.5为基线,在双臂真机上测量了从相机、本体感知到命令响应的端到端延迟,并提出两阶段非均匀Flow Matching去噪,把推理步数从10步减到2步,模型推理时间从61.557毫秒降到21.956毫秒。在30次叠衣真机试验中,训练型方法Legato和免训练方法Temporal Smoothing表现最好;把两阶段去噪与这两种方法结合,推理成本大幅下降,但任务成功率有小幅回落。

arXiv
Text-to-3D Policy进阶

T3DP:细粒度语言-行为对齐提升文本到3D策略的未见过规格泛化

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

本文提出T3DP框架,通过在指令token和局部行为表示之间建立双向细粒度对齐,让3D扩散策略能执行训练中未出现的精细行为规格。在Meta-World、ManiSkill和RoboTwin上,T3DP较全局语言-行为对齐平均成功率提升11.0至14.2个百分点;在真实机器人上平均成功率从47.5%升至65.0%。该工作为预印本,尚未经过同行评审。

arXiv
Safe RL-MPC专业

RL引导PAC-NMPC:固定翼无人机视觉导航实现概率安全与远见

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

本文提出将强化学习训练的actor-critic和传感器预测模型嵌入PAC-NMPC框架,为未知环境中基于感知的导航提供有限时间概率安全保证和长期性能。在固定翼无人机上,该方法在仿真中取得90%成功率,在硬件试验中取得80%成功率,均优于纯RL和传统规划基线。本文为预印本,未经同行评审。

arXiv
VLA专业

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
Navigation专业

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Zhihao Zheng, Mooi Choo Chuah

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences.

arXiv
Manipulation进阶

RoboCoach 让世界模型主动当教练:想象失败后定向补数据,真机成功率升至 75.0% 和 83.8%

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach 是一个预印本框架,用世界模型中的想象执行来诊断长程机器人任务中最先失败的子任务,并据此请求对应技能专家的演示数据。在 Franka 和 AgileX 真机上,仅追加 150 条子任务演示,成功率分别从 13.3% 升至 75.0%、从 40.0% 升至 83.8%,显著超过统一采集加共享适配器的基线。该工作表明世界模型可以成为主动教练,把稀缺的真实数据导向可复用的技能模块。

arXiv
Navigation专业

Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

Max Burns, Maisha Khanum, Monroe Kennedy, Steven H. Collins

Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model.

arXiv
Perception专业

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul, Muhayy Ud Din, Irfan Hussain

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly.

arXiv
Manipulation专业

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.

arXiv
VLA专业

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
Manipulation专业

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions.

arXiv
Humanoid专业

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations.

arXiv
Learning专业

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning.

arXiv
Navigation专业

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

Xiaoyu Yang, Sen Han, Da Li, Nan Wu

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer.

arXiv
Robotics专业

Cascaded consensus splitting for multi-branch contingency games

Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire, Laurent Pautet, David Filliat, David Fridovich-Keil, Georgios Bakirtzis

Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While contingency games capture intent uncertainty, existing formulations rely on a single branching time, oversimplifying interactions in which different agents' intentions are revealed at different times.

arXiv
VLA专业

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

论文来自 arXiv 机器人领域。尚未生成摘要的论文,显示摘要开头部分。