ROBOTNESS
연구

논문

새로 나온 로봇, 피지컬 AI 논문과, 각 논문이 업계에 갖는 의미.

논문 15편
아직 한국어판이 없어 영어 원문으로 표시.
arXiv
VLA전문가

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA전문가

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
VLA전문가

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
VLA전문가

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
VLA중급

ChunkTrust: 행동 전문가 신호로 로봇 정책의 실행 구간을 조절하는 방법

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

로봇 AI 모델은 앞으로의 동작을 한 묶음씩 예측하는데, 그중 몇 개를 실행한 뒤 다시 주변을 살필지는 대개 사람이 고정값으로 정해 왔다. 칭화대와 베이징즈위안인공지능연구원(BAAI) 등이 낸 이번 프리프린트는 이 값을 작업 중에 스스로 고르는 부가 모듈을 제안했고, 피지컬인텔리전스와 엔비디아의 기존 모델을 다시 학습시키지 않고도 시뮬레이션과 실제 양팔 로봇에서 성공률을 끌어올렸다.

arXiv
VLA전문가

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
VLA전문가

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
VLA전문가

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

arXiv
VLA전문가

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai, Fangming Liu

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies.

arXiv
VLA전문가

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself.

arXiv
VLA전문가

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.

arXiv
VLA전문가

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo.

논문은 arXiv 로봇 피드에서 가져옵니다. 요약이 아직 작성되지 않은 논문은 초록의 첫 부분을 보여줍니다.