ROBOTNESS
Forschung

Studien

Neue Studien zu Robotik und Physical AI, jeweils mit ihrer Bedeutung für die Branche.

106 Studien
Noch nicht auf Deutsch verfügbar: Anzeige im englischen Original.
arXiv
PerceptionExperten

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RLExperten

PL-MPC erweitert TD-M(PC)2 um Multi-Step-TD, adaptive Terminalwerte und returngewichtete Aktordestillation

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

Das Preprint stellt PL-MPC vor, eine Erweiterung des policy-constrained TD-M(PC)2-Rückgrats, die Critic-Supervision, MPPI-Terminalwertbewertung und Planer-zu-Policy-Transfer verändert. Auf HumanoidBench steigt der Total Average Return bei balance-hard von 98±18 auf 387±255 und bei hurdle von 199±13 auf 466±200; auf vier DMControl-Aufgaben bleibt die Methode konkurrenzfähig. In einem Sim-to-Real-Transfer mit einem KUKA IIWA14 erreicht PL-MPC höhere Erfolgsraten als TD-M(PC)2 auf der Trainingsgröße und zwei ungesehenen Objektgrößen.

arXiv
NavigationExperten

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls.

arXiv
Terrain traversabilityExperten

Erfahrungsgetriebenes Continual Learning für Terrain-Traversabilität bei Quadruped-Robotern mit Validierungs-Gate

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

Das Preprint stellt eine Pipeline vor, die für einen Unitree Go2 aus Bildern vor dem Kontakt mit einem eingefrorenen DINOv3-Backbone fünf propriozeptive Interaktionsindikatoren vorhersagt und deren Unsicherheit über eine evidentiale Regression schätzt. In einem sequenziellen Hardware-Versuch über drei unbekannte Untergründe senkt das Validierungs-Gate die finale Anchor-NLL-Verschlechterung um 23,1 Prozent gegenüber Replay ohne Gate bei ähnlicher Anpassung an neue Terrains. Damit wird ein Weg gezeigt, kontinuierliches Lernen aus Robotererfahrung mit kontrollierter Beibehaltung früherer Terrain-Modelle zu verbinden.

arXiv
VLAExperten

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain.

arXiv
Text-to-3D PolicyFortgeschritten

T3DP: Feingranulare Sprach-Verhaltens-Ausrichtung für Text-to-3D-Policies

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

Der Preprint stellt T3DP vor, ein Framework, das feingranulare Sprach-Verhaltens-Ausrichtung nutzt, um textgesteuerte 3D-Diffusionsrichtlinien für ungesehene Verhaltensspezifikationen zu verbessern. In Simulationen (Meta-World, ManiSkill, RoboTwin) erreicht T3DP durchschnittlich 11,0 bis 14,2 Prozentpunkte mehr Erfolg als globale Ausrichtung; auf einem realen PiPer-X-Roboter steigt der Durchschnitt von 47,5 auf 65,0 Prozent. Die lokale Zuordnung von Instruktionstoken zu Verhaltenssegmenten erhält spezifikationsrelevante Unterschiede, die globale Verfahren oft verwischen.

arXiv
NavigationExperten

RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models.

arXiv
VLAExperten

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
NavigationExperten

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Zhihao Zheng, Mooi Choo Chuah

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences.

arXiv
ManipulationFortgeschritten

RoboCoach: Weltmodelle als aktive Coaches für kompositionale Roboterfähigkeiten

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach ist ein Preprint und beschreibt ein Framework, das mit dem aktionskonditionierten Weltmodell CoachWorld imaginierte Fehlschläge in langfristigen Manipulationsaufgaben lokalisiert und daraus gezielte Demonstrationsanfragen sowie LoRA-Updates für die zuständigen Skill-Experten ableitet. In Simulation und auf zwei echten Roboterplattformen steigt die Erfolgsrate mit 150 zusätzlichen Teilaufgaben-Demonstrationen von 13,3 % auf 75,0 % beim Franka und von 40,0 % auf 83,8 % beim AgileX, während eine gleichmäßige Datenerfassung mit gemeinsamem Adapter nur 30,0 % bzw. 47,5 % erreicht. Die Arbeit zeigt, dass Weltmodelle nicht nur simulieren, sondern als aktive Coaches entscheiden können, welche wiederverwendbare Fähigkeit als Nächstes gelernt werden soll.

arXiv
NavigationExperten

Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

Max Burns, Maisha Khanum, Monroe Kennedy, Steven H. Collins

Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model.

arXiv
PerceptionExperten

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul, Muhayy Ud Din, Irfan Hussain

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly.

arXiv
ManipulationExperten

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.

arXiv
VLAExperten

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
ManipulationExperten

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions.

arXiv
HumanoidExperten

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations.

arXiv
LearningExperten

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning.

arXiv
NavigationExperten

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

Xiaoyu Yang, Sen Han, Da Li, Nan Wu

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer.

arXiv
RoboticsExperten

Cascaded consensus splitting for multi-branch contingency games

Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire, Laurent Pautet, David Filliat, David Fridovich-Keil, Georgios Bakirtzis

Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While contingency games capture intent uncertainty, existing formulations rely on a single branching time, oversimplifying interactions in which different agents' intentions are revealed at different times.

arXiv
VLAExperten

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

Die Studien stammen aus den Robotik-Feeds von arXiv. Solange unsere Zusammenfassung fehlt, erscheinen die ersten Zeilen des Abstracts.