ROBOTNESS
연구

논문

새로 나온 로봇, 피지컬 AI 논문과, 각 논문이 업계에 갖는 의미.

논문 106편
아직 한국어판이 없어 영어 원문으로 표시.
arXiv
Perception전문가

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation전문가

DiffWAM, 미래 영상 생성 없이 비디오 예측 표현을 UAV 궤적으로 직접 전환

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAM은 동결된 비디오 기초 모델의 중간 예측 표현을 첫 프레임 기하 정보와 결합해 연속 카메라 궤적으로 직접 변환하는 내비게이션 세계-행동 모델이다. DiffWAM-1000 1,000개 평가에서 궤적 RMSE 0.3492m, 엔드포인트 성공률 74.40%를 기록했고, DiffWAM-Flash는 Jetson AGX Thor에서 모델 파이프라인 지연 1.080초를 보였다. 이 연구는 arXiv 프리프린트이며, 미래 비디오 생성과 다중 프레임 복원 없이 예측 표현만으로 구조적 UAV 비행을 수행할 수 있음을 보였다.

arXiv
VLA전문가

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Robotics전문가

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
Manipulation전문가

Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments. However, their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems are unavailable or obstructed.

arXiv
Learning전문가

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic Actuators전문가

토폴로지 최적화로 다방향 굽힘 구현한 소프트 공압 액추에이터 TO-mdiSPAs

Swagatam Islam Sarkar, Prabhat Kumar

이 논문은 Darcy 법칙과 3필드 강건 토폴로지 최적화를 이용해 여러 방향으로 굽힐 수 있는 소프트 공압 액추에이터 mdiSPA의 챔버 구조를 설계했다. 최적화된 챔버는 기존 사각형 챔버와 다른 형상을 보였고, Abaqus 시뮬레이션에서 세 개의 암을 조합한 그리퍼가 입력 압력 조합에 따라 다양한 방향으로 변형했다. 다만 아직 동료 심사를 거치지 않은 프리프린트이며, 실물 제작과 실험 검증은 남아 있다.

arXiv
Navigation전문가

Active Mapping of Underwater Litter Using Camera-Sonar Fusion

David Rete, Patrick Boros, Lucian Busoniu

Marine litter is a growing threat to the underwater ecosystem, driving demand for autonomous survey methods that can locate debris efficiently over large areas. Existing survey methods typically follow predefined paths or operate with a single sensing modality, typically a camera (with image quality suffering in poor-visibility conditions) or sonar (usually noisy and low-resolution).

arXiv
Manipulation전문가

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions.

arXiv
Grounding전문가

GroundingPI, 점과 박스 기반 4B 그라운딩 모델로 물리 지능 지각 강화

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

XPeng 등 공동 연구진이 점과 박스를 공용 어휘의 양자화 좌표로 생성하는 4B 파라미터 그라운딩 기반 모델 GroundingPI를 공개했다. 이 프리프린트는 34개 벤치마크 평균 73.68%로 GPT-6 Astra(71.54%)를 넘었고, RoboTwin 2.0의 네 가지 OOD 설정 모두에서 비교 백본 중 1위를 기록했으며, RoboCasa-GR1에서는 50% 데모만으로 다른 모델의 75% 성능을 웃돌았다. 범용 VLM의 정밀 지각 한계를 줄여 실행 계층을 위한 지각 기반 모델의 가능성을 보였다.

arXiv
Aerial Manipulation전문가

Scene Graph와 MPAR로 로컬 전신 VLA를 장면 규모 공중 조작으로 확장

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

이 프리프린트는 관절형 무인 항공 조작기(UAM)에서 로컬 전신 비전-언어-행동(VLA) 정책을 장면 규모 임무로 확장하는 통합 프레임워크를 제안했다. 합성 데모로 VLA를 훈련하고 측정 진행 정렬(MPAR)로 비동기 행동 청크를 연속 궤적으로 실현하며, Scene Graph로 언어 목표를 실행 가능한 핸드오프 상태로 연결한다. 시뮬레이션 다중 사이트 임무 50회 중 21회(42.0%)를 완료했고 물리 UAM에서도 검증해 합성 데이터 기반 공중 조작 가능성을 보였다.

arXiv
VLA전문가

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
Navigation전문가

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.

arXiv
Floorplan SLAM전문가

MVP-SLAM, 평면도 사전정보로 건설 현장 드리프트를 온라인 보정하는 다중 카메라 비주얼-관성 SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

룩셈부르크대 연구팀이 arXiv 프리프린트로 공개한 MVP-SLAM은 두 대의 반대 방향 어안 카메라와 IMU만으로 실내 건설 현장에서 이동 궤적을 추정하면서, 바닥 평면도의 벽과 매칭해 누적 드리프트를 온라인으로 보정하는 비주얼-관성 SLAM 시스템이다. Hilti–Trimble SLAM Challenge 2026에서 Localization 과제 평균 RMSE 0.29 m로 22개 팀 중 2위, SLAM 과제 0.24 m로 62개 팀 중 5위를 기록했으며, 온라인으로 평면도를 통합하고 평면도 좌표계에서 위치를 추정하는 시스템 가운데 두 과제 모두 1위였다. 평면도를 쓰지 않는 자체 비교군 대비 Localization 오차를 4.9배 줄였다.

arXiv
Distributed Manipulation전문가

Membrane-Coupled Delta Array가 43.3mm 액추에이터 간격보다 작은 15mm 물체까지 조작

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

연구팀은 3자유도 델타 로봇 64대(8×8)의 끝단을 신축성 원단으로 연결해 192자유도 연속 표면을 만들었다. 이 표면은 중심 간격 43.3mm보다 작은 15mm 물체부터 90mm 물체까지 하나의 플랫폼에서 조작했고, 19개 델타 주변에 저차 DCT 계수로 명령하는 정책이 전체 배열 명령보다 배치 오차를 절반으로 줄였다. 실물 전이 실험에서는 적응 없이 76% 성공률을 보였다. 이 논문은 arXiv 프리프린트다.

arXiv
Manipulation전문가

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State Estimation전문가

비전 기반 다중 UAV 민첩 비행을 위한 상태 추정 재검토

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

체코 프라하 공과대학 연구진이 비전 기반 다중 UAV 비행에서 이웃 기체의 속도와 가속도를 정확히 추정하기 위해 시각 검출기가 제공하는 기울기 정보를 상태 추정에 통합했다. arXiv 프리프린트로 공개된 이 연구는 위치만 쓰는 4종 추정기와 자세 인식 5종 추정기를 실제 데이터 2종과 고충실도 시뮬레이션 데이터 1종에서 비교했으며, 자세 인식 방식이 속도 오차 평균 40%, 가속도 오차 평균 57%를 줄였다. 제안한 선형 추력 구속 칼만 필터가 가장 우수했고, 위치만 쓰는 필터는 약 300ms의 고정 지연을 보인 반면 자세 인식 추정기는 카메라 프레임률 수준의 물리적 응답 한계 근처에서 동작했다.

arXiv
VLA중급

ChunkTrust: 행동 전문가 신호로 로봇 정책의 실행 구간을 조절하는 방법

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

로봇 AI 모델은 앞으로의 동작을 한 묶음씩 예측하는데, 그중 몇 개를 실행한 뒤 다시 주변을 살필지는 대개 사람이 고정값으로 정해 왔다. 칭화대와 베이징즈위안인공지능연구원(BAAI) 등이 낸 이번 프리프린트는 이 값을 작업 중에 스스로 고르는 부가 모듈을 제안했고, 피지컬인텔리전스와 엔비디아의 기존 모델을 다시 학습시키지 않고도 시뮬레이션과 실제 양팔 로봇에서 성공률을 끌어올렸다.

논문은 arXiv 로봇 피드에서 가져옵니다. 요약이 아직 작성되지 않은 논문은 초록의 첫 부분을 보여줍니다.