ROBOTNESS
研究

論文

ロボティクスとフィジカルAIの最新論文と、それぞれが産業にもたらす意味。

論文106本
日本語版は未提供のため、英語原文で表示しています。
arXiv
Perception上級

Non-Invasive Inspection of Water Canals Using Dronar

Michael Zielinski, Zhizhan Wang, Benjamin Dymond, Reza Razavian, Zhongwang Dou

Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona, metro area. Over time, the concrete canals can experience a range of issues, including canal lining deformation, cracked concrete, and sediment buildup on the canal floor.

arXiv
UAV Navigation上級

DiffWAM、動画基盤モデルの予測表現から直接UAV軌道を生成

Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

DiffWAMは、凍結した動画基盤モデルの中間予測表現と初期フレームの幾何情報から、連続するカメラ軌道を直接復元するナビゲーション世界行動モデルである。本論文はプレプリントであり、DiffWAM-1000で軌道RMSE 0.3492 m、終点成功率74.40%を報告した。将来ビデオの生成と再構築をデプロイ時に省く経路として、UAVナビゲーションの低遅延化と実機展開可能性を示した。

arXiv
VLA上級

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires.

arXiv
Robotics上級

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
Manipulation上級

Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments. However, their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems are unavailable or obstructed.

arXiv
Learning上級

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic Actuators上級

TO-mdiSPAs:トポロジー最適化による多方向ソフト空気圧アクチュエータの設計

Swagatam Islam Sarkar, Prabhat Kumar

本研究は、任意の3次元方向へ曲がるソフト空気圧アクチュエータ(SPA)をトポロジー最適化で自動設計する手法TO-mdiSPAsを提案した。Darcy則と排水項で設計依存の圧力荷重をモデル化し、3フィールドのロバスト最適化を適用した。数値シミュレーションでは、最適化した3本アーム構成が複数方向への屈曲と把持動作を示した。査読前のプレプリントである。

arXiv
Navigation上級

Active Mapping of Underwater Litter Using Camera-Sonar Fusion

David Rete, Patrick Boros, Lucian Busoniu

Marine litter is a growing threat to the underwater ecosystem, driving demand for autonomous survey methods that can locate debris efficiently over large areas. Existing survey methods typically follow predefined paths or operate with a single sensing modality, typically a camera (with image quality suffering in poor-visibility conditions) or sonar (usually noisy and low-resolution).

arXiv
Manipulation上級

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions.

arXiv
Grounding上級

GroundingPI、点とボックスで物理的知能向けの知覚基盤を確立

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

GroundingPIは、点とボックスを量子化座標として出力する4Bパラメータの視覚グラウンディング基盤モデルである。34件のグラウンディングベンチマークで平均73.68%を達成し、より大規模なGPT-6 Astraの71.54%を上回った。ロボット操作と自動運転の下流タスクで既存バックボーンを上回る転移性能を示し、物理的知能のための知覚基盤として有望である。本論文はarXivのみで公開されたプレプリントである。

arXiv
Aerial Manipulation上級

局所全身VLAをシーン規模空中マニピュレーションへ拡張、MPARで実測進捗に整合

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

本研究は、多関節空中マニピュレータ向けに、合成データで学習した局所全身VLA行動をシーン規模ミッションに合成するフレームワークを提案した。主要な結果として、シミュレーションでのマルチサイト任務成功率42.0%、500ms遅延下でのMPARによるテイクオーバー位相誤差の中央値0.212秒削減、実機での多関節UAM検証を報告した。本論文は査読済み採録が明記されていないプレプリントである。

arXiv
VLA上級

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
Navigation上級

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.

arXiv
Floorplan SLAM上級

MVP-SLAM、床プラン事前情報でドリフト補正する複数カメラ視覚慣性SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

ルクセンブルク大学の研究チームが、建設現場向けに2台の魚眼カメラとIMUのみで動作し、設計図の床プランを逐次統合してドリフトを補正するオンラインSLAM「MVP-SLAM」を発表した(査読前プレプリント)。Hilti–Trimble SLAM Challenge 2026のLocalizationタスクで22チーム中2位(平均RMSE 0.29m)、SLAMタスクで62チーム中5位(0.24m)となり、オンラインで床プランを統合し位置特定する手法としては両タスクで首位だった。床プランをカメラのみの視覚慣性SLAMにオンラインで組み込み、軌跡構築中に補正できる点が重要である。

arXiv
Distributed Manipulation上級

膜結合8x8デルタアレイがアクチュエータ間隔以下の物体を操作、DCT方策で実機成功率76%

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

8x8の3自由度デルタロボット64台の先端を伸縮性膜で結合し、離散接触を連続面に変換した。15mmから90mmまで、アクチュエータ中心間距離43.3mmをまたいで単一面上で操作でき、低次離散コサイン変換係数に作用する方策は実機に適応なしで76%の成功率を示した。本論文は査読前のプレプリントである。

arXiv
Manipulation上級

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State Estimation上級

機敏な視覚ベース複数UAV飛行に向けた状態推定の再考

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

本論文はプレプリントである。視覚検出器が与える機体傾斜情報を推力方向の観測として統合し、隣接UAVの速度と加速度を低遅延で推定する手法を提案した。2件の実データセットと1件のフォトリアリスティックシミュレーションデータセットで評価し、位置のみの推定に比べ速度誤差平均40%、加速度誤差平均57%を削減した。閉ループのリーダー・フォロワーシミュレーションでは、提案手法が2gを超える横方向機動の追従を可能にした。

arXiv
VLA中級

ChunkTrust:アクションエキスパートの手がかりでロボット方策の実行ホライズンを適応的に決める

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

ロボットAIモデルは先の動作をひとまとまりで予測するが、そのうち何ステップを実行してから周囲を見直すかは、これまで技術者が固定値で決めることが多かった。清華大学や北京智源人工知能研究院(BAAI)などによる今回のプレプリントは、この値を動作中に自動で選ぶ追加モジュールを示し、Physical IntelligenceとNVIDIAの既存モデルを再学習せずに、シミュレーションと実機の双腕ロボットで成功率を高めた。

論文は arXiv のロボティクス分野から取得しています。当社の要約が未作成の場合は、要旨の冒頭を表示します。