ROBOTNESS
Research

Papers

New robotics and physical-AI papers, with what each one means for the industry.

20 papers
Filtered by technology vla · clear
arXiv
VLAExpert

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLAExpert

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Safe VLAExpert

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

This preprint presents FailBank, a framework that turns an observe-only safety teacher's runtime corrections into training records for vision-language-action policies. On VLA-Arena static-obstacle tasks, it raises task success by 8.5 and 6.9 percentage points over base policies for two VLA backbones while cutting policy-induced cumulative cost by 35.6% and 23.8%. The method matters because it lets robot policies learn a persistent balance between completing a task and avoiding unintended contact, instead of relying on temporary runtime shields.

arXiv
Memory-centric VLAIntermediate

Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

This preprint introduces Optimus-R, a memory-centric VLA framework that turns robotic manipulation adaptation into explicit query-skill memory tuning instead of repeated parameter updates. With only 30% of training data it raises LIBERO average success to 88.6% (vs 72.1% for π0.5) and reaches 33.3% real-world success from 20 demos per task; in lifelong learning it reduces forgetting to 10.0% vs 15.0%. The approach matters for low-data robot adaptation and retaining prior skills.

arXiv
VLAExpert

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
GroundingExpert

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

This preprint introduces GroundingPI, a 4B grounding foundation model that predicts object points and boxes as quantized token coordinates rather than using a general-purpose vision-language backbone. Across 34 grounding benchmarks it averages 73.68%, ahead of a larger GPT-6 Astra baseline at 71.54%, and as a visual backbone it improves downstream manipulation (RoboTwin 2.0, RoboCasa-GR1) and autonomous driving (nuScenes L2 0.296 m). The result matters because it argues grounding is a distinct perceptual layer that can make embodied foundation models more precise.

arXiv
Aerial ManipulationExpert

From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

This preprint introduces a framework for scene-scale aerial manipulation on an articulated uncrewed aerial manipulator, combining synthetic whole-body VLA training, measured-progress-aligned trajectory realization, and scene-graph-guided mission composition. In simulation, local skills achieve 39/60 successes under oracle handoff, and full multi-site missions achieve 21/50 (42.0%), while a monolithic whole-task VLA baseline achieves 0/50. Physical trials validate representative tasks without platform demonstrations, reducing risk and data cost for aerial manipulation.

arXiv
VLAExpert

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
VLAIntermediate

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

Robot AI models plan a short burst of movements at a time, and how many of those moves the robot carries out before it looks again is usually fixed by hand. This preprint from Tsinghua University, BAAI and partners adds a plug-in that picks that number while the robot works, raising success rates of existing Physical Intelligence and NVIDIA models in simulation and on a real two-arm robot without retraining them.

arXiv
Real-time VLAExpert

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

This preprint reports a two-stage non-uniform Flow Matching denoiser that cuts the π0.5 VLA's action-generation steps from 10 to 2 and model-inference time from 61.6 ms to 22.0 ms, a 2.8× speedup. The authors also built a distributed real-time VLA execution framework and benchmarked six action-scheduling methods on a 180-trial bimanual physical garment-folding task; Legato achieved 96.7% success as the best training-based method, while Temporal Smoothing led training-free methods at 76.7%. Combining the fast sampler with those methods gave 1.68–2.80× inference speedups at a 6.7–10.0 percentage-point success loss, highlighting a practical speed–quality tradeoff for real-time robot policies.

arXiv
VLAExpert

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
VLAExpert

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
VLAExpert

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

arXiv
VLAExpert

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai, Fangming Liu

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies.

arXiv
VLAExpert

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself.

arXiv
VLAExpert

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.

arXiv
VLAExpert

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo.

Topic pages

Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.