ROBOTNESS
研究

論文

ロボティクスとフィジカルAIの最新論文と、それぞれが産業にもたらす意味。

論文96本
日本語版は未提供のため、英語原文で表示しています。
arXiv
Navigation上級

Centralized Multi-UAV Exploration and 3D Reconstruction Using Single-UAV Planners

João Félix Mendes, Meysam Basiri, Rodrigo Ventura

Extending single Unmanned Aerial Vehicles (UAVs) exploration methods to multi-UAV teams can improve coverage speed and robustness, but introduces challenges such as consistent mapping, safe navigation, and deployment strategy. In this work, we present a centralized multi-UAV exploration framework that enables the use of existing single-UAV sampling-based planners in a multi-UAV setting.

arXiv
Humanoid上級

StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry

Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model.

arXiv
Visual Grounding上級

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou

Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.

arXiv
Tool Co-Design上級

粉末計量ロボットのツール形態と制御方策を共設計、実機誤差を標準ツール比45%削減

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

本論文は、実験室自動化における粉末計量のために、吐出ツールの形状と強化学習制御方策を同時に最適化するバイレベル共設計フレームワークを提案した。BOHBによる形態探索と幾何学的類似度に基づく方策のウォームスタートを組み合わせ、7種類の粉体でシミュレーションと実機評価を行った。共設計ツールは標準ツールに比べて実世界の計量誤差を45%削減し、未学習の材料でも優位を示した。本論文は査読済み会議録の記載がないプレプリントである。

arXiv
VLA上級

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

arXiv
VLA上級

Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal, Vijay Raghunathan

A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy.

arXiv
Navigation上級

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.

arXiv
Safe VLA上級

FailBank、ランタイム安全フィードバックを自己進化に変換しVLAの成功率と安全性を同時改善

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

米ノートルダム大学の研究チームは、視覚言語行動モデル(VLA)の実行時安全フィードバックを永続的なポリシー改善に変換するFailBankを開発した。固定CBF教師による観測のみの収集、成果認識付き選別、蓄積型失敗バンク、ガード付きLoRA更新を組み合わせ、VLA-Arena静的障害物スイートで成功率向上とポリシー起因コスト低減を同時に達成した。本論文はプレプリントである。

arXiv
Manipulation上級

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning?

arXiv
Safety上級

NEUPROが画像から安全述語を学習、解釈可能なロボット安全制御を実現

Zihan Ye, Jiayi Liu, Puze Liu, Jiayun Li, Georgia Chalvatzaki, Jan Peters, Kristian Kersting

本論文はプレプリントである。研究者らは、人間が与える記号的ルールと微分可能推論を組み合わせたNEUPROを提案し、生画像から安全に関わる述語を学習してロボットの安全違反を解釈可能に判定した。実ロボット画像を含むREASONデータセットで7タスク平均精度0.92±0.02を達成し、従来のブラックボックスコスト定式化に代わる再利用可能な安全推論の可能性を示した。

arXiv
Manipulation上級

Tactile Curiosity Drives Robot Interaction

Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge.

arXiv
Manipulation上級

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai, Weinan Zhang, Xinghang Li, Huaping Liu

World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud.

arXiv
Safety上級

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances.

arXiv
Perception上級

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation上級

DiffWAM、動画基盤モデルの予測表現から直接UAV軌道を生成

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAMは、凍結した動画基盤モデルの中間予測表現と初期フレームの幾何情報から、連続するカメラ軌道を直接復元するナビゲーション世界行動モデルである。本論文はプレプリントであり、DiffWAM-1000で軌道RMSE 0.3492 m、終点成功率74.40%を報告した。将来ビデオの生成と再構築をデプロイ時に省く経路として、UAVナビゲーションの低遅延化と実機展開可能性を示した。

arXiv
VLA上級

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

論文は arXiv のロボティクス分野から取得しています。当社の要約が未作成の場合は、要旨の冒頭を表示します。