ROBOTNESS
Research

Papers

New robotics and physical-AI papers, with what each one means for the industry.

96 papers
arXiv
RoboticsExpert

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
ManipulationExpert

Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments. However, their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems are unavailable or obstructed.

arXiv
LearningExpert

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic ActuatorsExpert

TO-mdiSPAs: Topology Optimization of multi-directional Soft Pneumatic Actuators

Swagatam Islam Sarkar, Prabhat Kumar

This preprint introduces a topology optimization method for designing multi-directional soft pneumatic actuators using Darcy-law pressure modeling and a robust three-field formulation. Optimized chambers are assembled into a simulated three-arm gripper; Abaqus tests at 0.15 MPa show single-arm tip displacements up to about 15.8 mm, with bending directions tracking the vector sum of activated chamber pressures. If experimentally validated, it could replace heuristic soft actuator design with systematic multi-axis bending for soft grippers.

arXiv
NavigationExpert

Active Mapping of Underwater Litter Using Camera-Sonar Fusion

David Rete, Patrick Boros, Lucian Busoniu

Marine litter is a growing threat to the underwater ecosystem, driving demand for autonomous survey methods that can locate debris efficiently over large areas. Existing survey methods typically follow predefined paths or operate with a single sensing modality, typically a camera (with image quality suffering in poor-visibility conditions) or sonar (usually noisy and low-resolution).

arXiv
World-action modelExpert

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

This preprint introduces Magic-W0, a robot manipulation foundation model that couples continuous action generation with explicit latent predictions of current 3D geometry, action-induced 3D motion, and future semantics. On RoboDojo-Sim it posts the highest average score among world-action models at 27.10, and on five real-robot tasks it reaches 94.6% average success after fine-tuning, beating π0.5 at 91.8%. The result matters because it shifts policies from blind imitation toward predicting how candidate actions will change the physical world, with gains in generalization and long-horizon manipulation.

arXiv
GroundingExpert

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

This preprint introduces GroundingPI, a 4B grounding foundation model that predicts object points and boxes as quantized token coordinates rather than using a general-purpose vision-language backbone. Across 34 grounding benchmarks it averages 73.68%, ahead of a larger GPT-6 Astra baseline at 71.54%, and as a visual backbone it improves downstream manipulation (RoboTwin 2.0, RoboCasa-GR1) and autonomous driving (nuScenes L2 0.296 m). The result matters because it argues grounding is a distinct perceptual layer that can make embodied foundation models more precise.

arXiv
Aerial ManipulationExpert

From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

This preprint introduces a framework for scene-scale aerial manipulation on an articulated uncrewed aerial manipulator, combining synthetic whole-body VLA training, measured-progress-aligned trajectory realization, and scene-graph-guided mission composition. In simulation, local skills achieve 39/60 successes under oracle handoff, and full multi-site missions achieve 21/50 (42.0%), while a monolithic whole-task VLA baseline achieves 0/50. Physical trials validate representative tasks without platform demonstrations, reducing risk and data cost for aerial manipulation.

arXiv
VLAExpert

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
NavigationExpert

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.

arXiv
Floorplan SLAMExpert

MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

This preprint introduces MVP-SLAM, an online visual-inertial SLAM system for construction sites that uses an as-planned floor plan to correct drift using only two back-to-back fisheye cameras and an IMU. It ranked 2nd of 22 teams in the Localization task of the Hilti–Trimble SLAM Challenge 2026 with 0.29 m mean RMSE, and 5th of 62 teams in SLAM with 0.24 m, the best among systems that integrate floor plans online. The work matters because it enables drift-bounded, plan-aligned localization during active site traversal without depth sensors.

arXiv
Distributed ManipulationExpert

Making Waves: A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

This preprint presents an 8×8 array of 64 three-degree-of-freedom delta robots coupled by a stretchable membrane, creating a continuous surface that can manipulate objects both larger and smaller than the 43.3 mm actuator spacing. On real hardware, a learned policy transferred from simulation and succeeded on 76% of 50 trials across objects from 30 to 90 mm, while local hand-designed primitives routed 15 mm cubes at 7.6 mm/s. The work matters because it removes the centre-to-centre spacing lower bound on object size for distributed manipulation.

arXiv
ManipulationExpert

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State EstimationExpert

Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

This preprint introduces vision-based pose-aware state estimators for neighboring multirotor UAVs, integrating visual tilt measurements to infer thrust direction, including a new linear thrust-constrained Kalman filter. Across two real-world and one simulated dataset, pose-aware methods cut mean velocity and acceleration errors by 40% and 57% versus position-only filters, remove a ~300 ms acceleration delay, and enable simulated follower tracking of >2 g lateral maneuvers where position-only estimation fails. This matters for collision avoidance and coordination in agile multi-UAV operations.

arXiv
PerceptionExpert

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RLExpert

Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

This preprint extends policy-constrained TD-MPC into PL-MPC by changing critic training targets, MPPI terminal values, and actor distillation without altering the world model or planner. On HumanoidBench, it lifts total average return from 98±18 to 387±255 on balance-hard and from 199±13 to 466±200 on hurdle; in zero-shot sim-to-real wrench–nut alignment on a KUKA IIWA14 it achieves 74.2% vs 61.3% success on the training object size. The result matters because it shows targeted interactions in the planning–learning loop can improve hard robot control tasks beyond the usual planner–policy alignment.

arXiv
NavigationExpert

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls.

arXiv
Terrain traversabilityExpert

Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

This preprint presents a quadruped traversability system that predicts five foot-contact outcomes from pre-contact camera images using a frozen DINOv3 backbone and an evidential regressor, then continually updates the model through replay with a historical validation gate. On a sequential real-robot stream across three unseen surface types, the gate cut anchor negative log-likelihood degradation by 23.1% versus replay without the gate while keeping new-terrain adaptation nearly equal. It matters for legged robots that must operate safely on unfamiliar ground without forgetting earlier experience.

Topic pages

Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.