Yuan Gao 0024

dblp:76/2452-24 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0003-3324-4418ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Configuration-Adaptive Visual Relative Localization for Spherical Modular Self-Reconfigurable Robots
abstract
Spherical Modular Self-reconfigurable Robots (SMSRs) have been popular in recent years. Their Self-reconfigurable nature allows them to adapt to different en-vironments and tasks, and achieve what a single module could not achieve. To collaborate with each other, relative localization between each module and assembly is crucial. Existing relative localization methods either have low accuracy, which is unsuit-able for short-distance collaborations, or are designed for fixed-shape robots, whose visual features remain static over time. This paper proposes the first visual relative localization method for SMSRs. We first detect and identify individual modules of SMSRs, and adopt visual tracking to improve the detection and identification robustness. Using an optimization-based method, tracking result is then fused with odometry to estimate the relative pose between assemblies. To deal with the non-convexity of the optimization problem, we adopt semi-definite relaxation to transform it into a convex form. The proposed method is validated and analysed in real-world experiments. The overall localization performance and the performance under time-varying configuration are evaluated. The result shows that the relative position estimation accuracy reaches 2%, and the orientation estimation accuracy reaches 6.64°, and that our method surpasses the state-of-the-art methods.
Qiu Zheng, Yuxiao Tu, Yuan Gao 0024, Guanqi Liang, Tin Lun Lam
ICRA4
2025 Inertial Parameters Identification for Floating-Base Multibody Systems Using Spinning Trajectories
abstract
Inertial parameter identification is crucial for accurate robot control, but existing methods for fixed-base manipulators are insufficient for floating-base systems. To address this, we propose the Decomposed Inertia Identification (DII) framework, which utilizes inertia transfer theory and the Recursive Parameter Null Space Algorithm (RPNA) to decompose base parameters into fixed-base and residual subsets. This approach reduces optimization complexity and enables symbolic parameter identification. Inspired by animal spinning behaviors, we use spinning trajectories to excite leg dynamics, overcoming high-DoFs challenges. The Covariance Matrix Adaptation Evolution Strategy (CMA-ES) optimizes parameters under physical consistency constraints. The method was validated on a Unitree Go1 quadruped robot, achieving 98.1% parameter convergence within 200 iterations during spinning locomotion (0.5–2 rad/s yaw velocity). Updating inertial parameters reduced tracking errors by 63% in body posture control and improved straight-line locomotion accuracy by 89% under payload variations on leg. The DII framework bridges fixed- and floating-base systems, enabling the application of mature fixed-base methodologies to floating-base robots and advancing self-model identification for real-world applications.
Hongwu Zhu, Yongyuan Xu, Ziyi Zhou 0004, Yuan Gao 0024, Ning Ding 0003
INDIN4
2025 Multimodal Deformation Estimation of Soft Pneumatic Gripper During Operation
abstract
Soft pneumatic robots are gaining significant attention due to their compliance and adaptability in unstructured environments. While emerging dual-chamber soft pneumatic robots can achieve complex 3D deformations beyond conventional single-axis bending, real-time proprioception remains challenging due to the high degrees of freedom and the complex interaction between chambers. To address this issue, we propose a multimodal learning-based sensing method that combines camera and inertial measurement unit (IMU) and then extracts full-body shape information using deep learning algorithms. Our method enhances proprioception by effectively processing high-dimensional sensor data, providing real-time feedback on the gripper shape. The average error of key points was found to be 3.67mm (Var 8.39) for our method, while the error was 4.36mm (Var 10.47) when a camera was used alone, or 9.32mm (Var 21.29) when an IMU was used alone. Our multimodal learning-based shape estimation and reconstruction empower soft pneumatic grippers to be seamlessly integrated into the embodied AI framework, significantly improving their reliability and thus paving the way for applications in service robotics, ehabilitation robotics, and human-robot collaborations.
Changheng Cai, Fei Xiao 0014, Marcellus Vanza, Taoyang Wang, Fangbing Zhou, Xuanyang Xu, Jian Zhu 0005, Yuan Gao 0024
IROS8
2025 Understanding Users' Perceptions and Expectations toward a Social Balloon Robot via an Exploratory Study
Tianyi Xia, Manqiu Liao, Yuan Gao 0024, Chun Yu, Yuntao Wang 0001, Yuanchun Shi
UIST9
2025 OC-HMAS: Dynamic Self-Organization and Self-Correction in Heterogeneous Multiagent Systems Using Multimodal Large Models
abstract
Heterogeneous multiagent systems (HMASs) leverage diverse agent capabilities to address complex tasks in dynamic environments, yet traditional approaches face limitations in autonomy and generalization when adapting to evolving scenarios. To overcome these challenges, we propose OC-HMAS, an IoT-integrated framework that synergizes self-organization and self-correction through multimodal perception. The system processes RGB images, LiDAR point clouds, and instance segmentation maps for real-time environmental awareness, while vision-language models and large language models (LLMs) jointly enable context-aware task decomposition, role allocation, and adaptive planning. Integrated path optimization and obstacle avoidance mechanisms further ensure operational safety and scalability across logistics, inspection, and search-and-rescue operations. Experimental validation demonstrates the framework’s superiority over SMRC-LLM, with 5.15% higher success rates and 14.2% faster task completion in logistics, alongside 4.69% accuracy gains and 12.9% time reduction in inspection scenarios. These results validate its enhanced adaptability in IoT-augmented environments, establishing a new benchmark for autonomous HMAS deployment.
Ping Feng, Tingting Yang 0001, Mingyang Liang, Lin Wang 0015, Yuan Gao 0024
IEEE Internet Things J.5
2025 Unlocking Drone Perception in Low AGL Heights: Progressive Semi-Supervised Learning for Ground-to-Aerial Perception Knowledge Transfer
abstract
We explore the novel challenge of drone perception across varying low AGL (above ground level) heights, a task essential for dynamic tasks, unlike the fixed ground viewpoint in autonomous driving. Supervised learning for this incurs high annotation costs, and current semi-supervised methods struggle with viewpoint differences. In this paper, we introduce ground-to-aerial perception knowledge transfer and propose a progressive semi-supervised learning framework for drone perception using only labeled data from the ground viewpoint and unlabeled data from flying viewpoints. The framework hinges on four key components: 1) a dense viewpoint sampling strategy, segmenting the vertical flight height range into evenly distributed intervals; 2) nearest neighbor pseudo-labeling, inferring labels of the nearest neighbor viewpoint using a model learned on the preceding viewpoint; 3) MixView, generating augmented images among different viewpoints to mitigate viewpoint differences; and 4) a progressive distillation strategy, gradually learning until reaching the maximum flying height. To validate our approach, we create both synthesized and real-world datasets. Extensive experimental analyses reveal a remarkable relative accuracy improvement of 25.7% and 16.9% for the synthesized dataset and the real world, respectively. Code and datasets are available on https://github.com/FreeformRobotics/Progressive-Self-Distillation-for-Ground-to-Aerial-Perception-Knowledge-Transfer.
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Yuan Gao 0024, Tin Lun Lam
IEEE Trans. Intell. Transp. Syst.5
2024 PepperPose: Full-Body Pose Estimation with a Companion Robot
abstract
Accurate full-body pose estimation across diverse actions in a user-friendly and location-agnostic manner paves the way for interactive applications in realms like sports, fitness, and healthcare. This task becomes challenging in real-world scenarios due to factors like the user’s dynamic positioning, the diversity of actions, and the varying acceptability of the pose-capturing system. In this context, we present PepperPose, a novel companion robot system tailored for optimized pose estimation. Unlike traditional methods, PepperPose actively tracks the user and refines its viewpoint, facilitating enhanced pose accuracy across different locations and actions. This allows users to enjoy a seamless action-sensing experience. Our evaluation, involving 30 participants undertaking daily functioning and exercise actions in a home-like space, underscores the robot’s promising capabilities. Moreover, we demonstrate the opportunities that PepperPose presents for human-robot interaction, its current limitations, and future developments.
Lingxiao Zhong, Chun Yu, Yuntao Wang 0001, Yuan Gao 0024, Tin Lun Lam, Yuanchun Shi
CHI7
2024 Meta-Reinforcement Learning Based Cooperative Surface Inspection of 3D Uncertain Structures using Multi-robot Systems
abstract
This paper presents a decentralized cooperative motion planning approach for surface inspection of 3D structures which includes uncertainties like size, number, shape, position, using multi-robot systems (MRS). Given that most of existing works mainly focus on surface inspection of single and fully known 3D structures, our motivation is two-fold: first, 3D structures separately distributed in 3D environments are complex, therefore the use of MRS intuitively can facilitate an inspection by fully taking advantage of sensors with different capabilities. Second, performing the aforementioned tasks when considering uncertainties is a complicated and time-consuming process because we need to explore, figure out the size and shape of 3D structures and then plan surface-inspection path. To overcome these challenges, we present a meta-learning approach that provides a decentralized planner for each robot to improve the exploration and surface inspection capabilities. The experimental results demonstrate our method can outperform other methods by approximately 10.5%-27% on success rate and 70%-75% on inspection speed.
Yuan Gao 0024, Junjie Hu 0003, Fuqin Deng, Tin Lun Lam
ICRA2
2024 Vision-Language Model-based Physical Reasoning for Robot Liquid Perception
abstract
There is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an open challenge as they can only process textual input. Recent advancements in large vision-language models (LVLMs) have enabled a more comprehensive understanding of the physical world by incorporating visual input, which provides richer contextual information than language alone. In this work, we proposed a novel paradigm that leveraged GPT-4V(ision), the state-of-the-art LVLM by OpenAI, to enable embodied agents to perceive liquid objects via image-based environmental feedback. Specifically, we exploited the physical understanding of GPT-4V to interpret the visual representation (e.g., time-series plot) of non-visual feedback (e.g., F/T sensor data), indirectly enabling multimodal perception beyond vision and language using images as proxies. We evaluated our method using 10 common household liquids with containers of various geometry and material. Without any training or fine-tuning, we demonstrated that our method can enable the robot to indirectly perceive the physical response of liquids and estimate their viscosity. We also showed that by jointly reasoning over the visual and physical attributes learned through interactions, our method could recognize liquid objects in the absence of strong visual cues (e.g., container labels with legible text or symbols), increasing the accuracy from 69.0%—achieved by the best-performing vision-only variant—to 86.0%.
Wenqiang Lai, Tianwei Zhang 0002, Tin Lun Lam, Yuan Gao 0024
IROS4
2023 Boosting LightWeight Depth Estimation via Knowledge Distillation
Junjie Hu 0003, Chenyou Fan, Hualie Jiang, Xiyue Guo, Yuan Gao 0024, Xiangyong Lu, Tin Lun Lam
KSEM (1)5
2023 Asymmetric Self-Play-Enabled Intelligent Heterogeneous Multirobot Catching System Using Deep Multiagent Reinforcement Learning
abstract
Aiming to develop a more robust and intelligent heterogeneous system for adversarial catching in security and rescue tasks, in this article, we discuss the specialities of applying asymmetric self-play and curriculum learning techniques to deal with the increasing heterogeneity and number of different robots in modern heterogeneous multirobot systems (HMRS). Our method, based on actor-critic multiagent reinforcement learning, provides a framework that can enable cooperative behaviors among heterogeneous multirobot teams. This leads to the development of an HMRS for complex catching scenarios that involve several robot teams and real-world constraints. We conduct simulated experiments to evaluate different mechanisms' influence on our method's performance, and real-world experiments to assess our system's performance in complex real-world catching problems. In addition, a bridging study is conducted to compare our method with a state-of-the-art method called S2M2 in heterogeneous catching problems, and our method performs better in adversarial settings. As a result, we show that the proposed framework, through fusing asymmetric self-play and curriculum learning during training, is able to successfully complete the HMRS catching task under realistic constraints in both simulation and the real world, thus providing a direction for future large-scale intelligent security & rescue HMRS.
Yuan Gao 0024, Xi Chen 0051, Junjie Hu 0003, Fuqin Deng, Tin Lun Lam
IEEE Trans. Robotics1
2022 Abnormal Occupancy Grid Map Recognition using Attention Network
abstract
The occupancy grid map is a critical component of autonomous positioning and navigation in the mobile robotic system, as many other systems' performance depends heavily on it. To guarantee the quality of the occupancy grid maps, researchers previously had to perform tedious manual recognition for a long time. This work focuses on automatic abnormal occupancy grid map recognition using the residual neural network with novel attention mechanism modules. We propose an effective channel and spatial Residual Squeeze-and-Excitation (csRSE) attention module, which contains a residual block for producing hierarchical features, followed by both channel SE (cSE) block and spatial SE (sSE) block for the sufficient information extraction along the channel and spatial pathways. To further summarize the occupancy grid map characteristics and experiments with our csRSE attention modules, we constructed a dataset called occupancy grid map dataset (OGMD) for our experiments. On this OGMD test dataset, we tested a few variants of our proposed structure and compared them with other attention mechanisms. Our experimental results show that the proposed attention network can infer the abnormal map with state-of-the-art (SOTA) accuracy of 96.23% for abnormal occupancy grid map recognition.
Fuqin Deng, Mingjian Liang, Ningbo Yi, Yuan Gao 0024, Tin Lun Lam
ICRA7
2022 AB-Mapper: Attention and BicNet based Multi-agent Path Planning for Dynamic Environment
abstract
Multi-agent path finding in dynamic environments is of great academic and practical value for multi-robot systems in the real world. To improve the effectiveness and efficiency of the learning process during path planning in dynamic environments, we introduce an algorithm called Attention and BicNet based Multi-agent path planning with effective reinforcement (AB-Mapper) under the actor-critic reinforcement learning framework. In this framework, on one hand, we design an actor-network that can utilize the BicNet with communication function to achieve the intra-team coordination. On the other hand, we propose a critic network that can selectively allocate attention weights to surrounding agents. This attention mechanism allows an individual agent to automatically learn a better evaluation of actions by considering the behaviours of its surrounding agents. Compared with the SOTA method Mapper in crowded environments with dynamic obstacles, our AB-Mapper is more effective (90.27±0.06% vs. 61.65±13.90% in terms of mean success rate) in solving the general multi-agent path finding problem.
Huifeng Guan, Yuan Gao 0024, Fuqin Deng, Tin Lun Lam
IROS2
2021 FEANet: Feature-Enhanced Attention Network for RGB-Thermal Real-time Semantic Segmentation
abstract
The RGB-Thermal (RGB-T) information for semantic segmentation has been extensively explored in recent years. However, most existing RGB-T semantic segmentation usually compromises spatial resolution to achieve real-time inference speed, which leads to poor performance. To better extract detail spatial information, we propose a two-stage Feature-Enhanced Attention Network (FEANet) for the RGB-T semantic segmentation task. Specifically, we introduce a Feature-Enhanced Attention Module (FEAM) to excavate and enhance multi-level features from both the channel and spatial views. Benefited from the proposed FEAM module, our FEANet can preserve the spatial information and shift more attention to high-resolution features from the fused RGB-T images. Extensive experiments on the urban scene dataset demonstrate that our FEANet outperforms other state-of-the-art (SOTA) RGB-T methods in terms of objective metrics and subjective visual comparison (+2.6% in global mAcc and +0.8% in global mIoU). For the 480 × 640 RGB-T test images, our FEANet can run with a real-time speed on an NVIDIA GeForce RTX 2080 Ti card.
Fuqin Deng, Mingjian Liang, Hongmin Wang, Yuan Gao 0024, Junjie Hu 0003, Xiyue Guo, Tin Lun Lam
IROS6
2012 What Does Touch Tell Us about Emotions in Touchscreen-Based Gameplay?
abstract
The increasing number of people playing games on touch-screen mobile phones raises the question of whether touch behaviors reflect players’ emotional states. This prospect would not only be a valuable evaluation indicator for game designers, but also for real-time personalization of the game experience. Psychology studies on acted touch behavior show the existence of discriminative affective profiles. In this article, finger-stroke features during gameplay on an iPod were extracted and their discriminative power analyzed. Machine learning algorithms were used to build systems for automatically discriminating between four emotional states (Excited, Relaxed, Frustrated, Bored), two levels of arousal and two levels of valence. Accuracy reached between 69% and 77% for the four emotional states, and higher results (~89%) were obtained for discriminating between two levels of arousal and two levels of valence. We conclude by discussing the factors relevant to the generalization of the results to applications other than games.
Yuan Gao 0024, Nadia Bianchi-Berthouze, Hongying Meng
ACM Trans. Comput. Hum. Interact.1