ROBOTNESS
연구

논문

새로 나온 로봇, 피지컬 AI 논문과, 각 논문이 업계에 갖는 의미.

논문 96편
아직 한국어판이 없어 영어 원문으로 표시.
arXiv
Navigation전문가

Centralized Multi-UAV Exploration and 3D Reconstruction Using Single-UAV Planners

João Félix Mendes, Meysam Basiri, Rodrigo Ventura

Extending single Unmanned Aerial Vehicles (UAVs) exploration methods to multi-UAV teams can improve coverage speed and robustness, but introduces challenges such as consistent mapping, safe navigation, and deployment strategy. In this work, we present a centralized multi-UAV exploration framework that enables the use of existing single-UAV sampling-based planners in a multi-UAV setting.

arXiv
Humanoid전문가

StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry

Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model.

arXiv
Visual Grounding전문가

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou

Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.

arXiv
Tool Co-Design전문가

실험실 자동화 분말 계량을 위한 Tool-Policy Co-Design 프레임워크

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

연구진은 로봇 팔이 쓰는 분말 분배 도구의 형태와 강화학습 제어 정책을 함께 최적화하는 Tool-Policy Co-Design 프레임워크를 제안했다. BOHB 기반 외부 루프와 SAC 기반 내부 루프로 도구 깊이, 폭, 림 스파이크 구조를 탐색하고, 유사한 기하 구조의 정책을 재사용해 탐색 속도를 높였다. 실물 로봇 실험에서 공동 설계한 도구는 표준 도구 대비 실제 계량 오차를 45% 줄였고, 이 논문은 동료 심사를 거치지 않은 프리프린트다.

arXiv
VLA전문가

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA전문가

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Navigation전문가

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.

arXiv
Safe VLA전문가

FailBank, 런타임 CBF 피드백을 LoRA 정책 개선으로 바꾸는 4단계 자기진화 프레임워크

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

FailBank는 런타임 CBF 교정 신호를 관찰 전용으로 수집해 VLA 정책을 LoRA로 반복 개선하는 프레임워크다. VLA-Arena 정적 장애물 평가에서 π0.5 백본 기준 평균 성공률을 66.0%에서 74.5%로 8.5%p 높이고 정책 유발 누적 비용을 35.6% 줄였다. 런타임 차폐 AEGIS와 비교하면 성공률을 25.4%p 높였으며, 이 연구는 프리프린트이고 실물 로봇 실험은 보고되지 않았다.

arXiv
Robotics전문가

Game-Guided Skill Discovery through Self-Play for Playable Agent Control

Seungeun Rho, Jeonghwan Kim, Xue Bin Peng, Sehoon Ha

We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions.

arXiv
Manipulation전문가

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning?

arXiv
Safety전문가

NEUPRO, 해석 가능한 로봇 안전 규칙을 시각 입력에서 학습하는 신경-기호 프레임워크

Zihan Ye, Jiayi Liu, Puze Liu, Jiayun Li, Georgia Chalvatzaki, Jan Peters, Kristian Kersting

이 논문은 로봇 안전 요구사항을 해석 가능한 1차 논리 규칙으로 표현하고, 시각 입력에서 안전 관련 술어를 학습하는 NEUPRO를 제안한다. REASON이라는 첫 실물 로봇 해석 가능 안전 벤치마크에서 NEUPRO는 과제 평균 0.92의 안전 분류 정확도를 보였고, 텐서 기반 추론기 대비 학습 9.1배, 추론 1.6배 빠르며 메모리 사용이 14.3배 낮았다. 이 연구는 아직 동료 심사를 거치지 않은 프리프린트다.

arXiv
Manipulation전문가

Tactile Curiosity Drives Robot Interaction

Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge.

arXiv
Manipulation전문가

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai, Weinan Zhang, Xinghang Li, Huaping Liu

World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud.

arXiv
Safety전문가

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances.

arXiv
Perception전문가

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation전문가

DiffWAM, 미래 영상 생성 없이 비디오 예측 표현을 UAV 궤적으로 직접 전환

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAM은 동결된 비디오 기초 모델의 중간 예측 표현을 첫 프레임 기하 정보와 결합해 연속 카메라 궤적으로 직접 변환하는 내비게이션 세계-행동 모델이다. DiffWAM-1000 1,000개 평가에서 궤적 RMSE 0.3492m, 엔드포인트 성공률 74.40%를 기록했고, DiffWAM-Flash는 Jetson AGX Thor에서 모델 파이프라인 지연 1.080초를 보였다. 이 연구는 arXiv 프리프린트이며, 미래 비디오 생성과 다중 프레임 복원 없이 예측 표현만으로 구조적 UAV 비행을 수행할 수 있음을 보였다.

arXiv
VLA전문가

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

논문은 arXiv 로봇 피드에서 가져옵니다. 요약이 아직 작성되지 않은 논문은 초록의 첫 부분을 보여줍니다.