ROBOTNESS
研究

論文

ロボティクスとフィジカルAIの最新論文と、それぞれが産業にもたらす意味。

論文93本
日本語版は未提供のため、英語原文で表示しています。
arXiv
Robotics上級

Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?

Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu

Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline.

arXiv
Continuum sensing上級

磁気とIMUを融合したモジュール式連続体ロボットの自己姿勢推定、外部カメラ不要で16.7Hz更新

Zheng Cao, Guo Ning Sue, Xiangyun Bu, David Quinn, Junzhe Hu, Carmel Majidi

本論文は、分散配置したIMUと基板コイルによる磁気計測を組み合わせ、外部カメラに依存せず連続体ロボットの三次元形状を推定するモジュール式センシング方式を開発した。4シナリオ20試行でバックボーンRMSE 10.0±3.2 mm、先端RMSE 14.9±6.5 mm、更新レート16.7 Hzを達成し、先端高さの閉ループ制御も実証した。これはプレプリントであり、査読済みとは報告されていない。

arXiv
Learning上級

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs.

arXiv
Soft Pneumatic Actuators上級

TO-mdiSPAs:トポロジー最適化による多方向ソフト空気圧アクチュエータの設計

Swagatam Islam Sarkar, Prabhat Kumar

本研究は、任意の3次元方向へ曲がるソフト空気圧アクチュエータ(SPA)をトポロジー最適化で自動設計する手法TO-mdiSPAsを提案した。Darcy則と排水項で設計依存の圧力荷重をモデル化し、3フィールドのロバスト最適化を適用した。数値シミュレーションでは、最適化した3本アーム構成が複数方向への屈曲と把持動作を示した。査読前のプレプリントである。

arXiv
World-action model上級

Magic-W0、構造化世界・行動基盤モデルがRoboDojoでWAM最高スコア、実機成功率94.6%

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

Magic-W0は、ロボット操作のための構造化世界・行動基盤モデルである。現在状態、行動による遷移、未来状態を明示的にモデル化し行動生成と双方向に結合し、RoboDojo-SimでWAM中最高の平均Score 27.10、実機5タスクで平均成功率94.6%を達成した。本論文は査読前のプレプリントである。

arXiv
Grounding上級

GroundingPI、点とボックスで物理的知能向けの知覚基盤を確立

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang

GroundingPIは、点とボックスを量子化座標として出力する4Bパラメータの視覚グラウンディング基盤モデルである。34件のグラウンディングベンチマークで平均73.68%を達成し、より大規模なGPT-6 Astraの71.54%を上回った。ロボット操作と自動運転の下流タスクで既存バックボーンを上回る転移性能を示し、物理的知能のための知覚基盤として有望である。本論文はarXivのみで公開されたプレプリントである。

arXiv
Aerial Manipulation上級

局所全身VLAをシーン規模空中マニピュレーションへ拡張、MPARで実測進捗に整合

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie

本研究は、多関節空中マニピュレータ向けに、合成データで学習した局所全身VLA行動をシーン規模ミッションに合成するフレームワークを提案した。主要な結果として、シミュレーションでのマルチサイト任務成功率42.0%、500ms遅延下でのMPARによるテイクオーバー位相誤差の中央値0.212秒削減、実機での多関節UAM検証を報告した。本論文は査読済み採録が明記されていないプレプリントである。

arXiv
VLA上級

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy.

arXiv
Floorplan SLAM上級

MVP-SLAM、床プラン事前情報でドリフト補正する複数カメラ視覚慣性SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

ルクセンブルク大学の研究チームが、建設現場向けに2台の魚眼カメラとIMUのみで動作し、設計図の床プランを逐次統合してドリフトを補正するオンラインSLAM「MVP-SLAM」を発表した(査読前プレプリント)。Hilti–Trimble SLAM Challenge 2026のLocalizationタスクで22チーム中2位(平均RMSE 0.29m)、SLAMタスクで62チーム中5位(0.24m)となり、オンラインで床プランを統合し位置特定する手法としては両タスクで首位だった。床プランをカメラのみの視覚慣性SLAMにオンラインで組み込み、軌跡構築中に補正できる点が重要である。

arXiv
Distributed Manipulation上級

膜結合8x8デルタアレイがアクチュエータ間隔以下の物体を操作、DCT方策で実機成功率76%

Bailey Dacre, Andrés Faíña, Oliver Kroemer, Zeynep Temel

8x8の3自由度デルタロボット64台の先端を伸縮性膜で結合し、離散接触を連続面に変換した。15mmから90mmまで、アクチュエータ中心間距離43.3mmをまたいで単一面上で操作でき、低次離散コサイン変換係数に作用する方策は実機に適応なしで76%の成功率を示した。本論文は査読前のプレプリントである。

arXiv
Manipulation上級

Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors

Antonio Franchi, Amr Afifi

Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load.

arXiv
Multi-UAV State Estimation上級

機敏な視覚ベース複数UAV飛行に向けた状態推定の再考

Michal Pliska, Matouš Vrba, Ondřej Víta, Martin Jiroušek, Viktor Walter, Martin Saska

本論文はプレプリントである。視覚検出器が与える機体傾斜情報を推力方向の観測として統合し、隣接UAVの速度と加速度を低遅延で推定する手法を提案した。2件の実データセットと1件のフォトリアリスティックシミュレーションデータセットで評価し、位置のみの推定に比べ速度誤差平均40%、加速度誤差平均57%を削減した。閉ループのリーダー・フォロワーシミュレーションでは、提案手法が2gを超える横方向機動の追従を可能にした。

arXiv
Perception上級

GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration

João Félix Mendes, Rodrigo Ventura, Meysam Basiri

Autonomous exploration demands that robots continuously evaluate candidate viewpoints based on their expected information gain and execution cost. Sampling-based planners estimate this gain by volumetric raycasting and, due to its computational cost, evaluate candidates under an assumption of mutual independence, ignoring the overlap between viewpoints along the same path.

arXiv
Model-based RL上級

PL-MPCがTD-M(PC)2にMTDとATEとRADを追加、HumanoidBenchの高難度タスクで平均報酬を大幅改善

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

プレプリントとして公開された本研究では、TD-M(PC)2を基盤に、批評家の学習目標を多段階化するMTD、計画時の終端価値に不確実性ペナルティを加えるATE、実現報酬で重み付けした模倣を行うRADを組み込んだPL-MPCを提案した。HumanoidBenchのbalance-hardでTotal Average Returnが98±18から387±255へ、hurdleで199±13から466±200へ向上し、実機のKUKA IIWA14によるレンチとナットの位置合わせでもTD-M(PC)2を上回る成功率を報告した。

arXiv
Terrain traversability上級

検証ゲート付き継続学習で四脚ロボットの地形走破性を予測、履歴忘却を23.1%低減

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

本論文は、四脚ロボットが未踏地形で得た接触指標を、DINOv3の視覚特徴から不確実性付きで予測する継続学習パイプラインを提案したプレプリントである。逐次的な実機データストリームで、履歴検証ゲート付きリプレイはゲートなしリプレイに比べ最終アンカーNLL劣化を23.1%相対低減し、新地形への適応はほぼ同等だった。この手法は、地形の見た目ではなくロボット自身の接触応答に基づく走破性評価を可能にする点で重要である。

arXiv
Real-time VLA上級

VLA推論を10ステップから2ステップへ、段階依存の非一様Flow Denoisingで61.6msを22.0msに短縮

Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang

このプレプリントは、VLAのモデル推論とロボット実行系の遅延を実測し、Flow Matchingの速度場が終端付近で方向補正に集中することを示した。その上で標準の10NFEから2NFEへ削減する2段階非一様デノイジングを提案し、推論時間を61.557msから21.956msに短縮した。双腕Tシャツ折りタスクで6種類のリアルタイム実行手法を比較し、学習ベースではLegato、学習不要ではTemporal Smoothingが最高で、2段階推論を組み合わせると推論コストは大幅に減るが成功率は小幅に低下した。

arXiv
Safe RL-MPC上級

RL誘導型PAC-NMPCが未知環境の視覚ベース航法で確率的安全を実現

Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore

強化学習で訓練したアクタークリティックとセンサ予測モデルをPAC-NMPCに統合し、未知環境での知覚ベース航法に有限時間の確率的衝突回避保証を与える手法を提案した。固定翼機の実機実験では成功率80パーセントを達成し、RL単独方策や地図とA*を用いるPAC-NMPCの40パーセントを上回った。本論文はarXivプレプリントであり、査読済みではない。

arXiv
VLA上級

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, Yukang Zhuo, Zheng Chong, Yunshuang Nie, Jian Zhang, Weijia Liufu, Qingman Wu, Heming Xu, Bingchang Song, Dantong Wu, Zhiyuan Wang, Hang Xu, Jianhua Han

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert.

論文は arXiv のロボティクス分野から取得しています。当社の要約が未作成の場合は、要旨の冒頭を表示します。