ROBOTNESS
研究

论文

机器人与具身智能领域的最新论文,以及每篇对产业的意义。

96 篇论文
暂无中文版,显示英文原文。
arXiv
Navigation专业

Centralized Multi-UAV Exploration and 3D Reconstruction Using Single-UAV Planners

João Félix Mendes, Meysam Basiri, Rodrigo Ventura

Extending single Unmanned Aerial Vehicles (UAVs) exploration methods to multi-UAV teams can improve coverage speed and robustness, but introduces challenges such as consistent mapping, safe navigation, and deployment strategy. In this work, we present a centralized multi-UAV exploration framework that enables the use of existing single-UAV sampling-based planners in a multi-UAV setting.

arXiv
Humanoid专业

StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry

Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model.

arXiv
Visual Grounding专业

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou

Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.

arXiv
Tool Co-Design专业

联合优化工具与策略,机器人粉末称量真机误差降低45%

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

这篇arXiv预印本提出一种工具形态与控制策略联合设计框架,外层用BOHB搜索勺形工具参数,内层为每个形态训练SAC策略,并用几何相似度热启动加速搜索。真机实验中,优化后的工具在7种粉末上的平均称量误差为2.23±3.00毫克,相比标准工具4.09±6.82毫克降低45%。该结果来自未经同行评审的预印本。

arXiv
VLA专业

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA专业

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Navigation专业

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.

arXiv
Safe VLA专业

FailBank:将运行时反馈变为 VLA 策略的持久自我进化

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

FailBank 是一个四阶段自进化框架,将运行时 CBF 屏蔽的修正建议转化为学习记录,用于持续更新视觉-语言-动作模型,而无需在部署时使用屏蔽。在 VLA-Arena 静态障碍套件上,FailBank 在两个骨干模型上分别将任务成功率提高 8.5 和 6.9 个百分点,同时将策略诱导累积成本降低 35.6% 和 23.8%。与运行时屏蔽 AEGIS 相比,成功率分别提高 25.4 和 9.5 个百分点。本文为预印本,尚未经过同行评审。

arXiv
Manipulation专业

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning?

arXiv
Safety专业

NEUPRO用神经符号规则让机器人安全推理可解释

Zihan Ye, Jiayi Liu, Puze Liu, Jiayun Li, Georgia Chalvatzaki, Jan Peters, Kristian Kersting

预印本论文提出NEUPRO,一个神经符号框架,将安全规范表示为一阶逻辑规则,并通过可微分推理器从图像中学习安全谓词。在真实机器人数据集REASON上,NEUPRO在七个任务的安全分类准确率达到0.92±0.02,显著优于VLM基线,并能解释安全违规原因。该工作为机器人安全提供了可解释且可迁移的表示,但尚未经同行评审。

arXiv
Manipulation专业

Tactile Curiosity Drives Robot Interaction

Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge.

arXiv
Manipulation专业

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai, Weinan Zhang, Xinghang Li, Huaping Liu

World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud.

arXiv
Safety专业

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances.

arXiv
Perception专业

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation专业

DiffWAM:直接从冻结视频模型预测表征解码无人机连续轨迹的世界动作模型

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAM 是 arXiv 预印本提出的导航世界动作模型,它从冻结视频模型的中间预测表征中直接生成连续相机轨迹,部署时无需完成未来视频生成和几何重建。在自建 DiffWAM-1000 基准上,DiffWAM 轨迹 RMSE 为 0.3492 米,端点成功率为 74.40%;机载 DiffWAM-Flash 在 NVIDIA Jetson AGX Thor 上模型管道延迟为 1.08 秒。该工作表明预测视频表征可被高效地落地为连续三维运动,为生成再重建的导航管线提供了直接替代方案。

arXiv
VLA专业

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

论文来自 arXiv 机器人领域。尚未生成摘要的论文,显示摘要开头部分。