ROBOTNESS
Research

Papers

New robotics and physical-AI papers, with what each one means for the industry.

8 papers
arXiv
Co-speech HumanoidIntermediate

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

ECHO-G is a preprint framework that generates full-body speech-synchronized motion for humanoid robots from audio and timed transcripts. It models motion directly in robot space with a Speech-Grounded Diffusion Transformer and reports a Fréchet Gesture Distance of 2.278 versus 4.976 for EMAGE+GMR, with 5.96 ms/frame inference time. This matters for making expressive, deployable co-speech gestures on humanoids without human-motion retargeting.

arXiv
UAV traffic monitoringIntermediate

Prediction is Better than Detection: Traffic Congestion Control using Drones

Samira Hayat, Christian Raffelsberger

This preprint introduces a Mesa-based multi-agent simulation where UAVs patrol road junctions to detect and predict traffic jams and trigger adaptive signal control. Across sweeps of fleet size, traffic level, and network size, detection and prediction rates plateau once drone count roughly matches junction count; prediction-triggered signal changes cut jam duration by up to 8 ticks versus up to 5 ticks for detection-triggered changes, roughly doubling benefit. It suggests onboard prediction rather than larger fleets is the higher-value path for aerial congestion management.

arXiv
Memory-centric VLAIntermediate

Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

This preprint introduces Optimus-R, a memory-centric VLA framework that turns robotic manipulation adaptation into explicit query-skill memory tuning instead of repeated parameter updates. With only 30% of training data it raises LIBERO average success to 88.6% (vs 72.1% for π0.5) and reaches 33.3% real-world success from 20 demos per task; in lifelong learning it reduces forgetting to 10.0% vs 15.0%. The approach matters for low-data robot adaptation and retaining prior skills.

arXiv
VLAIntermediate

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong

Robot AI models plan a short burst of movements at a time, and how many of those moves the robot carries out before it looks again is usually fixed by hand. This preprint from Tsinghua University, BAAI and partners adds a plug-in that picks that number while the robot works, raising success rates of existing Physical Intelligence and NVIDIA models in simulation and on a real two-arm robot without retraining them.

arXiv
Text-to-3D PolicyIntermediate

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

This preprint introduces T3DP, a text-to-3D policy framework that aligns instruction tokens with local behavior segments from demonstrations to improve generalization to unseen fine-grained behavior specifications. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0 to +14.2 points; on real-robot tasks it raises average success from 47.5% to 65.0% (+17.5 points). The result matters for robot systems that must follow precise language such as target positions, displacements, or articulated states without exhaustive demonstration coverage.

arXiv
ManipulationIntermediate

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

This preprint introduces RoboCoach, a world-model-guided coaching loop that decides which subtask demonstrations to collect and which reusable skill expert adapters to update. On real robots, 150 added subtask demonstrations lift complete-task success from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, and coached skills transfer to four unseen compositions where a baseline scores 0%. The work matters because it shows world models can direct scarce real-world supervision toward the specific reusable skills that fail, rather than requiring expensive end-to-end data.

Topic pages

Papers come from arXiv robotics feeds. Where our summary is not written yet, you see the opening lines of the abstract.