ROBOTNESS
Forschung

Studien

Neue Studien zu Robotik und Physical AI, jeweils mit ihrer Bedeutung für die Branche.

106 Studien
Noch nicht auf Deutsch verfügbar: Anzeige im englischen Original.
arXiv
NavigationExperten

Centralized Multi-UAV Exploration and 3D Reconstruction Using Single-UAV Planners

João Félix Mendes, Meysam Basiri, Rodrigo Ventura

Extending single Unmanned Aerial Vehicles (UAVs) exploration methods to multi-UAV teams can improve coverage speed and robustness, but introduces challenges such as consistent mapping, safe navigation, and deployment strategy. In this work, we present a centralized multi-UAV exploration framework that enables the use of existing single-UAV sampling-based planners in a multi-UAV setting.

arXiv
HumanoidExperten

StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry

Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model.

arXiv
Visual GroundingExperten

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou

Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.

arXiv
Co-speech HumanoidFortgeschritten

ECHO-G: Vollkörper-Co-Speech-Bewegungsgenerierung für humanoide Roboter aus Audiospur und zeitlich zugeordnetem Transkript

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

ECHO-G ist ein als Preprint veröffentlichtes Framework zur Vollkörper-Co-Speech-Bewegungsgenerierung für humanoide Roboter. Es kombiniert frame-genaue akustische Merkmale mit token-basiertem Transkriptkontext in einem Diffusion-Transformer und erzeugt direkt Bewegungsreferenzen im Roboterkoordinatenraum. Auf einem aus BEAT2 abgeleiteten G1-Datensatz erreicht der Ansatz die besten Co-Speech-Kennzahlen und die niedrigste Inferenzzeit pro Frame; in einem Nutzerstudium mit 45 Personen wird die gemeinsame Audio-Text-Konditionierung bevorzugt.

arXiv
ManipulationExperten

Tool-Policy Co-Design for Powder Weighing in Laboratory Automation

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

Autonomous powder weighing is one of many bottlenecks in laboratory automation due to the complex, non-linear dynamics of heterogeneous materials. Robot chemists performing this task utilise standard tools shaped for the dexterity of human hands, whose fixed geometry sets the dynamics that the control policy needs to regulate.

arXiv
VLAExperten

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLAExperten

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
NavigationExperten

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.

arXiv
VLAExperten

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress.

arXiv
ManipulationExperten

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning?

arXiv
UAV traffic monitoringFortgeschritten

Vorhersage statt Erkennung: Drohnengestützte Stauprognose verbessert adaptive Ampelsteuerung

Samira Hayat, Christian Raffelsberger

Das Paper ist ein Preprint, der eine Multiagenten-Simulation vorstellt, in der Drohnen Verkehrsstaus an Knotenpunkten erkennen und vorhersagen, um adaptive Ampelschaltungen auszulösen. Zentrale Ergebnisse sind, dass die Leistung ab einer Drohnenzahl etwa gleich der Knotenanzahl abflacht und dass eine vorausschauende Signaladaption die Staudauer ungefähr doppelt so stark reduziert wie eine reaktive Erkennung. Die Vorhersagegenauigkeit bleibt jedoch bei maximal 55 Prozent begrenzt, was auf Verbesserungsbedarf bei den Vorläufersignalen hinweist.

arXiv
VLAExperten

Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning.

arXiv
SafetyExperten

NEUPRO lernt interpretierbare Sicherheitsregeln aus Bildern für Robotersteuerung

Zihan Ye, Jiayi Liu, Puze Liu, Jiayun Li, Georgia Chalvatzaki, Jan Peters, Kristian Kersting

Der vorliegende Preprint stellt NEUPRO vor, ein neuro-symbolisches Framework, das Sicherheitsanforderungen als symbolische Regeln formuliert und mit differenzierbarem Reasoning aus Bildern grounded. Auf dem neu veröffentlichten realen Robotik-Datensatz REASON erreicht NEUPRO über alle Aufgaben eine Sicherheitsklassifikationsgenauigkeit von 0,92 ± 0,02 und liegt damit deutlich über VLM-Baselines. Der Ansatz liefert explizite Erklärungen für Sicherheitsverletzungen und ist durch graphbasiertes Reasoning 9,1-mal schneller im Training als eine tensorbasierte Variante.

arXiv
ManipulationExperten

Tactile Curiosity Drives Robot Interaction

Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge.

arXiv
ManipulationExperten

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai, Weinan Zhang, Xinghang Li, Huaping Liu

World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud.

arXiv
SafetyExperten

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances.

Die Studien stammen aus den Robotik-Feeds von arXiv. Solange unsere Zusammenfassung fehlt, erscheinen die ersten Zeilen des Abstracts.