ROBOTNESS
연구

논문

새로 나온 로봇, 피지컬 AI 논문과, 각 논문이 업계에 갖는 의미.

논문 106편
아직 한국어판이 없어 영어 원문으로 표시.
arXiv
Perception전문가

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RL전문가

PL-MPC, 계획-학습 폐루프를 함께 개선해 HumanoidBench와 실물 렌치-너트 정렬 성능 향상

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

PL-MPC는 정책 제약형 TD-M(PC)2 백본에서 비평가 지도, MPPI 종단 가치 평가, 플래너-정책 전이를 동시에 바꾼 방법이다. HumanoidBench에서는 balance-hard의 TAR(Total Average Return)가 98±18에서 387±255로, hurdle은 199±13에서 466±200으로 올랐고, KUKA IIWA14 실물 실험에서는 학습 물체 크기와 미학습 크기 모두에서 TD-M(PC)2보다 높은 성공률을 보였다. 아직 동료 심사를 거치지 않은 프리프린트다.

arXiv
Navigation전문가

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls.

arXiv
Terrain traversability전문가

경험 기반 지속 학습 CL-P로 쿼드러패드 로봇의 지형 통과성을 예측한다

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

프리프린트로 공개된 이 논문은 쿼드러패드 로봇이 접촉 전 이미지로부터 평면 발 미끄러짐, 평균 수직 지면 반력, 견인 지수, 이동 비용, 접지 하중률 등 5가지 상호작용 지표를 예측하고, 검증 게이트를 적용한 지속 학습으로 새로운 지형에 적응하면서 과거 성능 저하를 줄이는 파이프라인을 제안했다. 하드웨어 순차 실험에서 과거 앵커 NLL 성능 저하를 리플레이 단독 대비 23.1% 줄였으며, 시뮬레이션 내비게이션에서도 안정적으로 동작했다.

arXiv
VLA전문가

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain.

arXiv
Text-to-3D Policy중급

Text-to-3D Policy(T3DP), 토큰-행동 정렬로 보지 못한 세부 지시를 수행

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

T3DP는 언어 명령의 토큰과 행동 세그먼트를 지역적으로 정렬해 훈련에서 보지 못한 세부 사양을 수행하는 텍스트-3D 정책이다. Meta-World, ManiSkill, RoboTwin에서 전역 정렬 대비 평균 성공률을 11.0~14.2%포인트 높였고, 실물 PiperX 로봇에서는 평균 성공률을 47.5%에서 65.0%로 17.5%포인트 끌어올렸다. 이 논문은 아직 동료 심사를 거치지 않은 프리프린트이며, 같은 장면에서 목표 위치나 이동량이 다른 작업을 언어로 세밀하게 지시하는 데 중요하다.

arXiv
Navigation전문가

RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models.

arXiv
VLA전문가

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
Navigation전문가

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Zhihao Zheng, Mooi Choo Chuah

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences.

arXiv
Manipulation중급

RoboCoach, 세계 모델로 상상한 실패를 다음 로봇 기술 학습 신호로 바꾼다

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

프리프린트로 공개된 RoboCoach는 행동 조건 세계 모델 CoachWorld에서 로봇 실행을 굴려 실패한 하위 작업을 찾고, 해당 기술 전문가에만 추가 데모를 요청해 LoRA 어댑터를 갱신하는 프레임워크다. 22개 작업-정책 쌍에서 상상 실행과 실제 실행 성공률 간 스피어만 상관이 0.840이었고, 실물 로봇에서 플랫폼당 150개 하위 작업 데모만 추가해 Franka는 13.3%에서 75.0%로, AgileX는 40.0%에서 83.8%로 성공률이 올랐다. 이는 세계 모델이 데이터를 더 만드는 것을 넘어 어떤 기술을 가르칠지 결정하는 능동 코치가 될 수 있음을 보여준다.

arXiv
Navigation전문가

Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

Max Burns, Maisha Khanum, Monroe Kennedy, Steven H. Collins

Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model.

arXiv
Perception전문가

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul, Muhayy Ud Din, Irfan Hussain

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly.

arXiv
Manipulation전문가

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit.

arXiv
VLA전문가

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
Manipulation전문가

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions.

arXiv
Humanoid전문가

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations.

arXiv
Learning전문가

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning.

arXiv
Navigation전문가

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

Xiaoyu Yang, Sen Han, Da Li, Nan Wu

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer.

arXiv
Robotics전문가

Cascaded consensus splitting for multi-branch contingency games

Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire, Laurent Pautet, David Filliat, David Fridovich-Keil, Georgios Bakirtzis

Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While contingency games capture intent uncertainty, existing formulations rely on a single branching time, oversimplifying interactions in which different agents' intentions are revealed at different times.

arXiv
VLA전문가

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

논문은 arXiv 로봇 피드에서 가져옵니다. 요약이 아직 작성되지 않은 논문은 초록의 첫 부분을 보여줍니다.