ROBOTNESS
研究

论文

机器人与具身智能领域的最新论文,以及每篇对产业的意义。

106 篇论文
暂无中文版,显示英文原文。
arXiv
Perception专业

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation专业

DiffWAM:直接从冻结视频模型预测表征解码无人机连续轨迹的世界动作模型

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAM 是 arXiv 预印本提出的导航世界动作模型,它从冻结视频模型的中间预测表征中直接生成连续相机轨迹,部署时无需完成未来视频生成和几何重建。在自建 DiffWAM-1000 基准上,DiffWAM 轨迹 RMSE 为 0.3492 米,端点成功率为 74.40%;机载 DiffWAM-Flash 在 NVIDIA Jetson AGX Thor 上模型管道延迟为 1.08 秒。该工作表明预测视频表征可被高效地落地为连续三维运动,为生成再重建的导航管线提供了直接替代方案。

arXiv
VLA专业

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Robotics专业

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
Manipulation专业

Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments. However, their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems are unavailable or obstructed.

arXiv
Learning专业

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic Actuators专业

拓扑优化多向软气动执行器 TO-mdiSPAs:仿真实现三臂夹持器多方向弯曲

Swagatam Islam Sarkar, Prabhat Kumar

该预印本提出一种基于拓扑优化的多向软气动执行器(mdiSPA)系统设计方法,采用三场鲁棒公式和达西定律建模压力载荷,并用MMA求解。优化得到的非传统气腔在Abaqus仿真中表现出多方向弯曲,由三臂构成的夹持器可沿不同方向自适应抓取。研究基于数值模拟,尚未进行实物制造或实验验证。

arXiv
Navigation专业

Active Mapping of Underwater Litter Using Camera-Sonar Fusion

David Rete, Patrick Boros, Lucian Busoniu

Marine litter is a growing threat to the underwater ecosystem, driving demand for autonomous survey methods that can locate debris efficiently over large areas. Existing survey methods typically follow predefined paths or operate with a single sensing modality, typically a camera (with image quality suffering in poor-visibility conditions) or sonar (usually noisy and low-resolution).

arXiv
Manipulation专业

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions.

arXiv
Grounding专业

GroundingPI:以视觉基元构建4B接地基础模型,34项基准平均73.68%超越GPT-6 Astra

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

这是一篇预印本论文。作者提出 GroundingPI,一个基于点和边界框等视觉基元的 4B 参数接地基础模型,在 34 个接地基准上达到 73.68% 平均得分,超过更大的 GPT-6 Astra(71.54%)。该模型作为视觉骨干用于机器人操作和自动驾驶时,在 RoboTwin 2.0 全部四个分布外设置中优于所有主流骨干,相对最强基线最高提升 24.8%,并在 RoboCasa-GR1 上以 50% 演示数据超过其他基线用 75% 数据的表现。这项工作表明专门的接地预训练可成为物理智能的感知基础,尤其适合系统1的快速执行。

arXiv
Aerial Manipulation专业

合成监督与场景图组合实现跨站点空中操作,MPAR对齐执行进度

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

该预印本提出统一框架,将局部全身VLA行为组合为场景级空中操作。在仿真中完成21/50个跨站点任务(42.0%),物理平台验证了2/5次取水浇花任务。关键贡献包括无需真实平台演示的合成训练、MPAR对齐异步动作块,以及场景图引导的可行交接与拓扑转移。

arXiv
VLA专业

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
Navigation专业

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.

arXiv
Floorplan SLAM专业

MVP-SLAM:多相机视觉惯性SLAM用楼层平面图在线纠正施工场景漂移

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

MVP-SLAM是一篇预印本,提出一种多相机视觉惯性SLAM系统,利用建筑楼层平面图在线校正漂移。在Hilti–Trimble SLAM Challenge 2026中,它在Localization任务以0.29米平均RMSE排名第2,SLAM任务以0.24米排名第5,是所有在线运行、整合平面图并定位到楼层平面图中的方法中成绩最好的。这项工作表明,仅靠相机和IMU也能在施工场景利用平面图先验在线持续抑制累积误差。

arXiv
Distributed Manipulation专业

膜耦合Delta阵列突破执行器间距限制:8×8机器人阵列操控15至90毫米物体

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

这篇预印本提出一种由8×8个3自由度Delta机器人和弹性膜组成的分布式操控系统,通过将64个离散接触点变成连续表面,消除了执行器间距对物体尺寸的下限。在硬件上,该系统可操控15至90毫米的物体,跨度是执行器间距43.3毫米的六倍范围;强化学习策略在仿真中命令19个Delta邻域达到95.1%成功率,并在未调整的情况下以76%成功率迁移到真机。

arXiv
Manipulation专业

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State Estimation专业

融合倾斜测量的线性推力约束卡尔曼滤波提升多无人机敏捷飞行状态估计

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

这是一篇预印本。作者针对视觉多无人机飞行中仅用位置测量导致高阶状态估计延迟的问题,提出融合视觉检测器提供的倾斜信息,并设计了一种线性推力约束卡尔曼滤波器。在三个数据集上,姿态感知估计将速度和加速度误差平均降低40%和57%,闭环仿真中支持超过2g的横向机动跟踪。

arXiv
VLA进阶

ChunkTrust:利用动作专家证据自适应调整机器人策略的执行时域

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

机器人AI模型每次预测一段未来动作,执行多少步之后再重新观察环境,过去大多由工程师手动设为固定值。清华大学、北京智源人工智能研究院(BAAI)等机构发布的这篇预印本提出一个即插即用模块,在运行中自动决定这一步数,无需重新训练,就提升了Physical Intelligence和英伟达现有模型在仿真和真实双臂机器人上的成功率。

论文来自 arXiv 机器人领域。尚未生成摘要的论文,显示摘要开头部分。