ROBOTNESS
Research

Papers

New robotics and physical-AI papers, with what each one means for the industry.

106 papers
arXiv
PerceptionExpert

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RLExpert

Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

This preprint extends policy-constrained TD-MPC into PL-MPC by changing critic training targets, MPPI terminal values, and actor distillation without altering the world model or planner. On HumanoidBench, it lifts total average return from 98±18 to 387±255 on balance-hard and from 199±13 to 466±200 on hurdle; in zero-shot sim-to-real wrench–nut alignment on a KUKA IIWA14 it achieves 74.2% vs 61.3% success on the training object size. The result matters because it shows targeted interactions in the planning–learning loop can improve hard robot control tasks beyond the usual planner–policy alignment.

arXiv
NavigationExpert

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls.

arXiv
Terrain traversabilityExpert

Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

This preprint presents a quadruped traversability system that predicts five foot-contact outcomes from pre-contact camera images using a frozen DINOv3 backbone and an evidential regressor, then continually updates the model through replay with a historical validation gate. On a sequential real-robot stream across three unseen surface types, the gate cut anchor negative log-likelihood degradation by 23.1% versus replay without the gate while keeping new-terrain adaptation nearly equal. It matters for legged robots that must operate safely on unfamiliar ground without forgetting earlier experience.

arXiv
Real-time VLAExpert

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

This preprint reports a two-stage non-uniform Flow Matching denoiser that cuts the π0.5 VLA's action-generation steps from 10 to 2 and model-inference time from 61.6 ms to 22.0 ms, a 2.8× speedup. The authors also built a distributed real-time VLA execution framework and benchmarked six action-scheduling methods on a 180-trial bimanual physical garment-folding task; Legato achieved 96.7% success as the best training-based method, while Temporal Smoothing led training-free methods at 76.7%. Combining the fast sampler with those methods gave 1.68–2.80× inference speedups at a 6.7–10.0 percentage-point success loss, highlighting a practical speed–quality tradeoff for real-time robot policies.

arXiv
Text-to-3D PolicyIntermediate

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

This preprint introduces T3DP, a text-to-3D policy framework that aligns instruction tokens with local behavior segments from demonstrations to improve generalization to unseen fine-grained behavior specifications. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0 to +14.2 points; on real-robot tasks it raises average success from 47.5% to 65.0% (+17.5 points). The result matters for robot systems that must follow precise language such as target positions, displacements, or articulated states without exhaustive demonstration coverage.

arXiv
Safe RL-MPCExpert

RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

This preprint introduces AC-PAC-NMPC, a hybrid controller that places RL-trained actor, critic, and sensor-prediction models inside a sampling-based stochastic NMPC framework to provide finite-time probabilistic safety bounds and long-horizon vision-based navigation. On a fixed-wing UAV with an onboard depth camera, it raised real-robot success from 40% (both RL actor and best PAC-NMPC baseline) to 80% with the lowest cost; in simulation it reached 90% success versus 85% for the strongest baseline. This matters for agile robots that need learned long-range behavior without sacrificing formal collision-probability guarantees.

arXiv
VLAExpert

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
ManipulationIntermediate

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

This preprint introduces RoboCoach, a world-model-guided coaching loop that decides which subtask demonstrations to collect and which reusable skill expert adapters to update. On real robots, 150 added subtask demonstrations lift complete-task success from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, and coached skills transfer to four unseen compositions where a baseline scores 0%. The work matters because it shows world models can direct scarce real-world supervision toward the specific reusable skills that fail, rather than requiring expensive end-to-end data.

arXiv
NavigationExpert

Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

Max Burns, Maisha Khanum, Monroe Kennedy, Steven H. Collins

Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model.

arXiv
ManipulationExpert

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.

arXiv
VLAExpert

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
ManipulationExpert

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions.

arXiv
HumanoidExpert

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations.

arXiv
LearningExpert

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning.

arXiv
NavigationExpert

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

Xiaoyu Yang, Sen Han, Da Li, Nan Wu

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer.

arXiv
RoboticsExpert

Cascaded consensus splitting for multi-branch contingency games

Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire, Laurent Pautet, David Filliat, David Fridovich-Keil, Georgios Bakirtzis

Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While contingency games capture intent uncertainty, existing formulations rely on a single branching time, oversimplifying interactions in which different agents' intentions are revealed at different times.

arXiv
VLAExpert

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

Topic pages

Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.