ROBOTNESS
Research

Papers

New robotics and physical-AI papers, with what each one means for the industry.

106 papers
arXiv
PerceptionExpert

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV NavigationExpert

DiffWAM: A Fast and Efficient Navigation World Action Model

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAM is a preprint that introduces a UAV navigation world-action model. It recovers camera trajectories directly from the predictive features of a frozen video model, avoiding future-video decoding and multi-frame geometric reconstruction at deployment. The paper reports a trajectory RMSE of 0.3492 m and adds FastDreamer, an execution system that overlaps computation with flight for continuous UAV control.

arXiv
VLAExpert

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
RoboticsExpert

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
ManipulationExpert

Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments. However, their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems are unavailable or obstructed.

arXiv
LearningExpert

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic ActuatorsExpert

TO-mdiSPAs: Topology Optimization of multi-directional Soft Pneumatic Actuators

Swagatam Islam Sarkar, Prabhat Kumar

This preprint introduces a topology optimization method for designing multi-directional soft pneumatic actuators using Darcy-law pressure modeling and a robust three-field formulation. Optimized chambers are assembled into a simulated three-arm gripper; Abaqus tests at 0.15 MPa show single-arm tip displacements up to about 15.8 mm, with bending directions tracking the vector sum of activated chamber pressures. If experimentally validated, it could replace heuristic soft actuator design with systematic multi-axis bending for soft grippers.

arXiv
NavigationExpert

Active Mapping of Underwater Litter Using Camera-Sonar Fusion

David Rete, Patrick Boros, Lucian Busoniu

Marine litter is a growing threat to the underwater ecosystem, driving demand for autonomous survey methods that can locate debris efficiently over large areas. Existing survey methods typically follow predefined paths or operate with a single sensing modality, typically a camera (with image quality suffering in poor-visibility conditions) or sonar (usually noisy and low-resolution).

arXiv
ManipulationExpert

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions.

arXiv
GroundingExpert

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

This preprint introduces GroundingPI, a 4B grounding foundation model that predicts object points and boxes as quantized token coordinates rather than using a general-purpose vision-language backbone. Across 34 grounding benchmarks it averages 73.68%, ahead of a larger GPT-6 Astra baseline at 71.54%, and as a visual backbone it improves downstream manipulation (RoboTwin 2.0, RoboCasa-GR1) and autonomous driving (nuScenes L2 0.296 m). The result matters because it argues grounding is a distinct perceptual layer that can make embodied foundation models more precise.

arXiv
Aerial ManipulationExpert

From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

This preprint introduces a framework for scene-scale aerial manipulation on an articulated uncrewed aerial manipulator, combining synthetic whole-body VLA training, measured-progress-aligned trajectory realization, and scene-graph-guided mission composition. In simulation, local skills achieve 39/60 successes under oracle handoff, and full multi-site missions achieve 21/50 (42.0%), while a monolithic whole-task VLA baseline achieves 0/50. Physical trials validate representative tasks without platform demonstrations, reducing risk and data cost for aerial manipulation.

arXiv
VLAExpert

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
NavigationExpert

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.

arXiv
Floorplan SLAMExpert

MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

This preprint introduces MVP-SLAM, an online visual-inertial SLAM system for construction sites that uses an as-planned floor plan to correct drift using only two back-to-back fisheye cameras and an IMU. It ranked 2nd of 22 teams in the Localization task of the Hilti–Trimble SLAM Challenge 2026 with 0.29 m mean RMSE, and 5th of 62 teams in SLAM with 0.24 m, the best among systems that integrate floor plans online. The work matters because it enables drift-bounded, plan-aligned localization during active site traversal without depth sensors.

arXiv
Distributed ManipulationExpert

Making Waves: A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

This preprint presents an 8×8 array of 64 three-degree-of-freedom delta robots coupled by a stretchable membrane, creating a continuous surface that can manipulate objects both larger and smaller than the 43.3 mm actuator spacing. On real hardware, a learned policy transferred from simulation and succeeded on 76% of 50 trials across objects from 30 to 90 mm, while local hand-designed primitives routed 15 mm cubes at 7.6 mm/s. The work matters because it removes the centre-to-centre spacing lower bound on object size for distributed manipulation.

arXiv
ManipulationExpert

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State EstimationExpert

Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

This preprint introduces vision-based pose-aware state estimators for neighboring multirotor UAVs, integrating visual tilt measurements to infer thrust direction, including a new linear thrust-constrained Kalman filter. Across two real-world and one simulated dataset, pose-aware methods cut mean velocity and acceleration errors by 40% and 57% versus position-only filters, remove a ~300 ms acceleration delay, and enable simulated follower tracking of >2 g lateral maneuvers where position-only estimation fails. This matters for collision avoidance and coordination in agile multi-UAV operations.

arXiv
VLAIntermediate

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

Robot AI models plan a short burst of movements at a time, and how many of those moves the robot carries out before it looks again is usually fixed by hand. This preprint from Tsinghua University, BAAI and partners adds a plug-in that picks that number while the robot works, raising success rates of existing Physical Intelligence and NVIDIA models in simulation and on a real two-arm robot without retraining them.

Topic pages

Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.