ROBOTNESS
研究

論文

ロボティクスとフィジカルAIの最新論文と、それぞれが産業にもたらす意味。

論文20本
日本語版は未提供のため、英語原文で表示しています。
技術「VLA(視覚言語行動)モデル」で絞り込み中、解除
arXiv
VLA上級

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA上級

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Safe VLA上級

FailBank、ランタイム安全フィードバックを自己進化に変換しVLAの成功率と安全性を同時改善

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

米ノートルダム大学の研究チームは、視覚言語行動モデル(VLA)の実行時安全フィードバックを永続的なポリシー改善に変換するFailBankを開発した。固定CBF教師による観測のみの収集、成果認識付き選別、蓄積型失敗バンク、ガード付きLoRA更新を組み合わせ、VLA-Arena静的障害物スイートで成功率向上とポリシー起因コスト低減を同時に達成した。本論文はプレプリントである。

arXiv
Memory-centric VLA中級

Optimus-R、VLA適応をメモリ中心に再設計し少データ性能を大幅改善

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

Optimus-Rは、視覚言語行動モデル(VLA)のタスク適応をパラメータ更新ではなく明示的なメモリ操作として定式化するフレームワークである。インラインメモリインタフェースで制御に関連するクエリとスキル表現を抽出し、クエリスキルメモリバンクに外部化することで、限られたデモでの効率的なスキル学習と破滅的忘却の緩和を実現した。本論文はプレプリントであり、LIBEROの30%データで平均成功率を16.5ポイント改善したほか、実機20デモで33.3%の成功率を報告している。

arXiv
VLA上級

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Grounding上級

GroundingPI、点とボックスで物理的知能向けの知覚基盤を確立

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

GroundingPIは、点とボックスを量子化座標として出力する4Bパラメータの視覚グラウンディング基盤モデルである。34件のグラウンディングベンチマークで平均73.68%を達成し、より大規模なGPT-6 Astraの71.54%を上回った。ロボット操作と自動運転の下流タスクで既存バックボーンを上回る転移性能を示し、物理的知能のための知覚基盤として有望である。本論文はarXivのみで公開されたプレプリントである。

arXiv
Aerial Manipulation上級

局所全身VLAをシーン規模空中マニピュレーションへ拡張、MPARで実測進捗に整合

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

本研究は、多関節空中マニピュレータ向けに、合成データで学習した局所全身VLA行動をシーン規模ミッションに合成するフレームワークを提案した。主要な結果として、シミュレーションでのマルチサイト任務成功率42.0%、500ms遅延下でのMPARによるテイクオーバー位相誤差の中央値0.212秒削減、実機での多関節UAM検証を報告した。本論文は査読済み採録が明記されていないプレプリントである。

arXiv
VLA上級

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
VLA中級

ChunkTrust:アクションエキスパートの手がかりでロボット方策の実行ホライズンを適応的に決める

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

ロボットAIモデルは先の動作をひとまとまりで予測するが、そのうち何ステップを実行してから周囲を見直すかは、これまで技術者が固定値で決めることが多かった。清華大学や北京智源人工知能研究院(BAAI)などによる今回のプレプリントは、この値を動作中に自動で選ぶ追加モジュールを示し、Physical IntelligenceとNVIDIAの既存モデルを再学習せずに、シミュレーションと実機の双腕ロボットで成功率を高めた。

arXiv
Real-time VLA上級

VLA推論を10ステップから2ステップへ、段階依存の非一様Flow Denoisingで61.6msを22.0msに短縮

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

このプレプリントは、VLAのモデル推論とロボット実行系の遅延を実測し、Flow Matchingの速度場が終端付近で方向補正に集中することを示した。その上で標準の10NFEから2NFEへ削減する2段階非一様デノイジングを提案し、推論時間を61.557msから21.956msに短縮した。双腕Tシャツ折りタスクで6種類のリアルタイム実行手法を比較し、学習ベースではLegato、学習不要ではTemporal Smoothingが最高で、2段階推論を組み合わせると推論コストは大幅に減るが成功率は小幅に低下した。

arXiv
VLA上級

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

arXiv
VLA上級

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost.

arXiv
VLA上級

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Timothy K Johnsen, Marco Levorato

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control.

arXiv
VLA上級

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai, Fangming Liu

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies.

arXiv
VLA上級

Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li, Yuxuan Feng, Jinyu Gu

Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself.

arXiv
VLA上級

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.

arXiv
VLA上級

Rho: A Foundation for Efficiently Adaptable VLA Models

Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo.

論文は arXiv のロボティクス分野から取得しています。当社の要約が未作成の場合は、要旨の冒頭を表示します。