ROBOTNESS
연구

논문

새로 나온 로봇, 피지컬 AI 논문과, 각 논문이 업계에 갖는 의미.

논문 20편
아직 한국어판이 없어 영어 원문으로 표시.
기술 필터 적용 중: VLA(비전 언어 행동) 모델, 해제
arXiv
VLA전문가

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA전문가

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Safe VLA전문가

FailBank, 런타임 CBF 피드백을 LoRA 정책 개선으로 바꾸는 4단계 자기진화 프레임워크

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

FailBank는 런타임 CBF 교정 신호를 관찰 전용으로 수집해 VLA 정책을 LoRA로 반복 개선하는 프레임워크다. VLA-Arena 정적 장애물 평가에서 π0.5 백본 기준 평균 성공률을 66.0%에서 74.5%로 8.5%p 높이고 정책 유발 누적 비용을 35.6% 줄였다. 런타임 차폐 AEGIS와 비교하면 성공률을 25.4%p 높였으며, 이 연구는 프리프린트이고 실물 로봇 실험은 보고되지 않았다.

arXiv
Memory-centric VLA중급

Optimus-R, 쿼리-스킬 메모리 뱅크로 VLA 적응 비용과 망각을 낮추다

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

하얼빈공업대학(선전)과 Pengcheng Laboratory 연구팀이 공개한 프리프린트에서 VLA 모델의 새 작업 적응을 메모리 중심으로 바꾼 Optimus-R을 제안했다. Optimus-R은 학습 가능한 메모리 토큰을 VLA prefix 스트림에 삽입해 제어 관련 쿼리와 스킬 표현을 추출하고, 이를 쿼리-스킬 메모리 뱅크에 저장해 재사용한다. 실험에서 30% 학습 데이터만으로 LIBERO 평균 성공률을 16.5%p, CALVIN 평균 완료 길이를 0.42, RoboTwin 2.0 Hard 성공률을 5.0%p 높였고, 실물 로봇 20개 데모 크로스도메인 적응에서 33.3% 성공률로 π0.5 대비 13.9%p 우위를 보였다.

arXiv
VLA전문가

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Grounding전문가

GroundingPI, 점과 박스 기반 4B 그라운딩 모델로 물리 지능 지각 강화

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

XPeng 등 공동 연구진이 점과 박스를 공용 어휘의 양자화 좌표로 생성하는 4B 파라미터 그라운딩 기반 모델 GroundingPI를 공개했다. 이 프리프린트는 34개 벤치마크 평균 73.68%로 GPT-6 Astra(71.54%)를 넘었고, RoboTwin 2.0의 네 가지 OOD 설정 모두에서 비교 백본 중 1위를 기록했으며, RoboCasa-GR1에서는 50% 데모만으로 다른 모델의 75% 성능을 웃돌았다. 범용 VLM의 정밀 지각 한계를 줄여 실행 계층을 위한 지각 기반 모델의 가능성을 보였다.

arXiv
Aerial Manipulation전문가

Scene Graph와 MPAR로 로컬 전신 VLA를 장면 규모 공중 조작으로 확장

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

이 프리프린트는 관절형 무인 항공 조작기(UAM)에서 로컬 전신 비전-언어-행동(VLA) 정책을 장면 규모 임무로 확장하는 통합 프레임워크를 제안했다. 합성 데모로 VLA를 훈련하고 측정 진행 정렬(MPAR)로 비동기 행동 청크를 연속 궤적으로 실현하며, Scene Graph로 언어 목표를 실행 가능한 핸드오프 상태로 연결한다. 시뮬레이션 다중 사이트 임무 50회 중 21회(42.0%)를 완료했고 물리 UAM에서도 검증해 합성 데이터 기반 공중 조작 가능성을 보였다.

arXiv
VLA전문가

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
VLA중급

ChunkTrust: 행동 전문가 신호로 로봇 정책의 실행 구간을 조절하는 방법

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

로봇 AI 모델은 앞으로의 동작을 한 묶음씩 예측하는데, 그중 몇 개를 실행한 뒤 다시 주변을 살필지는 대개 사람이 고정값으로 정해 왔다. 칭화대와 베이징즈위안인공지능연구원(BAAI) 등이 낸 이번 프리프린트는 이 값을 작업 중에 스스로 고르는 부가 모듈을 제안했고, 피지컬인텔리전스와 엔비디아의 기존 모델을 다시 학습시키지 않고도 시뮬레이션과 실제 양팔 로봇에서 성공률을 끌어올렸다.

arXiv
Real-time VLA전문가

실시간 VLA를 위한 단계 인지 2단계 Flow 디노이징과 시스템 수준 평가

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

이 논문은 프리프린트로, VLA 모델 추론과 로봇 실행 사이 시간 불일치를 줄이기 위해 2단계 비균일 Flow Matching 디노이징과 분산 실시간 실행 프레임워크를 제안했다. π0.5 기준 모델에서 NFE를 10에서 2로 줄여 모델 추론 시간을 61.557 ms에서 21.956 ms로 2.804배 단축했고, 실물 양팔 옷 접기 과제에서 여섯 가지 실시간 실행 방법을 비교해 Legato가 학습 기반, Temporal Smoothing이 비학습 기반 중 가장 우수했다. 이 결과는 모델 추론 효율과 로봇 시스템 타이밍을 함께 최적화해야 함을 보여준다.

arXiv
VLA전문가

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
VLA전문가

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
VLA전문가

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

arXiv
VLA전문가

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai, Fangming Liu

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies.

arXiv
VLA전문가

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself.

arXiv
VLA전문가

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.

arXiv
VLA전문가

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo.

논문은 arXiv 로봇 피드에서 가져옵니다. 요약이 아직 작성되지 않은 논문은 초록의 첫 부분을 보여줍니다.