ROBOTNESS
研究

論文

ロボティクスとフィジカルAIの最新論文と、それぞれが産業にもたらす意味。

論文106本
日本語版は未提供のため、英語原文で表示しています。
arXiv
Perception上級

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RL上級

PL-MPCがTD-M(PC)2にMTDとATEとRADを追加、HumanoidBenchの高難度タスクで平均報酬を大幅改善

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

プレプリントとして公開された本研究では、TD-M(PC)2を基盤に、批評家の学習目標を多段階化するMTD、計画時の終端価値に不確実性ペナルティを加えるATE、実現報酬で重み付けした模倣を行うRADを組み込んだPL-MPCを提案した。HumanoidBenchのbalance-hardでTotal Average Returnが98±18から387±255へ、hurdleで199±13から466±200へ向上し、実機のKUKA IIWA14によるレンチとナットの位置合わせでもTD-M(PC)2を上回る成功率を報告した。

arXiv
Navigation上級

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls.

arXiv
Terrain traversability上級

検証ゲート付き継続学習で四脚ロボットの地形走破性を予測、履歴忘却を23.1%低減

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

本論文は、四脚ロボットが未踏地形で得た接触指標を、DINOv3の視覚特徴から不確実性付きで予測する継続学習パイプラインを提案したプレプリントである。逐次的な実機データストリームで、履歴検証ゲート付きリプレイはゲートなしリプレイに比べ最終アンカーNLL劣化を23.1%相対低減し、新地形への適応はほぼ同等だった。この手法は、地形の見た目ではなくロボット自身の接触応答に基づく走破性評価を可能にする点で重要である。

arXiv
VLA上級

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain.

arXiv
Text-to-3D Policy中級

T3DPが未見仕様の汎化を改善、言語と3D行動の局所アライメントで成功率向上

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

本プレプリントは、3Dビジオモータポリシーが学習時に見ていない細粒度な行動仕様を言語から実行する問題に取り組み、T3DPと呼ぶ枠組みを提案した。T3DPは指示のトークンと行動セグメントの局所的な対応関係を双方向に学習し、類似仕様を混同しない言語表現を獲得する。シミュレーション3ベンチマークで平均成功率をグローバルアライメント比11.0〜14.2ポイント、実機で47.5%から65.0%へ17.5ポイント改善した。

arXiv
Navigation上級

RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models.

arXiv
VLA上級

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
Navigation上級

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Zhihao Zheng, Mooi Choo Chuah

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences.

arXiv
Manipulation中級

RoboCoach:世界モデルを能動的コーチに用い、再利用可能な技能専門家を選択的に改善

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoachは、世界モデル内で長いタスクを模擬実行し、最初に失敗する部分作業と対応技能専門家を特定して追加実演を要求する枠組みである。このプレプリントでは、2つのシミュレーション環境と2台の実機で評価し、実機では150件の部分作業実演を追加するだけでFrankaの成功率を13.3%から75.0%に、AgileXを40.0%から83.8%に改善したと報告している。さらに、更新した専門家は未学習の4つの構成タスクで平均35.0%の成功率を示し、均一取得の共有方策ベースラインの0%を上回った。

arXiv
Navigation上級

Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

Max Burns, Maisha Khanum, Monroe Kennedy, Steven H. Collins

Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model.

arXiv
Perception上級

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul, Muhayy Ud Din, Irfan Hussain

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly.

arXiv
Manipulation上級

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.

arXiv
VLA上級

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
Manipulation上級

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions.

arXiv
Humanoid上級

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations.

arXiv
Learning上級

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning.

arXiv
Navigation上級

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

Xiaoyu Yang, Sen Han, Da Li, Nan Wu

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer.

arXiv
Robotics上級

Cascaded consensus splitting for multi-branch contingency games

Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire, Laurent Pautet, David Filliat, David Fridovich-Keil, Georgios Bakirtzis

Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While contingency games capture intent uncertainty, existing formulations rely on a single branching time, oversimplifying interactions in which different agents' intentions are revealed at different times.

arXiv
VLA上級

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

論文は arXiv のロボティクス分野から取得しています。当社の要約が未作成の場合は、要旨の冒頭を表示します。