VLDB 2026 Research / reviewers in the wild / expert
Qing Gao 0002
dblp:16/4671-2
· DBLP profile ↗
23ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-5395-6175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Triple Spectral Fusion for Sensor-Based Human Activity RecognitionabstractThe field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identify daily activities. Despite the advancements in learning-based methods, it is challenging to perform information fusion from the temporal perspective due to the complexities in fusing heterogeneous sensor data and establishing long-term context correlations. This paper proposes a novel triple spectral fusion framework tailored for HAR. First, we develop an adaptive complementary filtering technique for noise suppression and organize each IMU's sensors into posture and motion modality nodes. Given that IMU nodes form a dynamic heterogeneous graph, we then apply adaptive filtering within the graph Fourier domain to merge both homogeneous and heterogeneous node information. Furthermore, an adaptive wavelet frequency selection approach is implemented to suppress context redundancy and shorten the length of features. This approach enhances both timestamp-based graph aggregation and the correlation of long-term contexts. Our framework uses adaptive filtering in the Fourier, graph Fourier, and wavelet domains, enabling effective multi-sensor fusion and context correlation. Extensive experiments on ten benchmark datasets demonstrate the superior performance of our framework. Ye Zhang 0037, Longguang Wang, Qing Gao 0002, Chaocan Xiang, Mohammed Bennamoun, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | M3-HOI: Multi-modal mining network for video-based human-object interaction recognition
Bohong Wu, Qing Gao 0002 |
Pattern Recognit. | 2 |
| 2025 | 3D Whole-Body Pose Estimation Using Graph High-Resolution Network for Humanoid Robot TeleoperationabstractIn the realm of robotics, teleoperation plays a pivotal role in performing high-risk or intricate tasks, and obtaining precise 3D whole-body pose is crucial for this purpose. Traditional two-stage methods have limitations in estimating different body parts, leading to complex systems and higher estimation errors. In order to address these issues,the paper introduces a novel framework called Graph High-Resolution Network (GraphHRNet) for accurate 3D whole-body pose estimation, which is essential for the teleoperation of humanoid robots. GraphHRNet effectively captures global structural information and local details by integrating a High-Resolution Module and a Multi-branch Regression Module. The High-Resolution Module utilizes an enhanced graph convolution kernel to fuse multi-scale features, capturing global information, while the Multi-branch Regression Module focuses on refining and predicting accurate 3D coordinates for intricate body parts such as hands and face. Experimental results on the H3WB dataset demonstrate that GraphHRNet surpasses state-of-the-art (SOTA) methods in 3D whole-body pose estimation, significantly improving performance. Furthermore, the paper explores the potential application of this approach in a tele-operation system for humanoid robots, providing an intuitive and high-fidelity solution for remotely executing complex tasks. The code have been publicly available at https://github.com/Z-mingyu/GraphHRNet.git Qing Gao 0002, Yuanchuan Lai, Ye Zhang 0037, Tao Chang, Yulan Guo |
ICRA | 2 |
| 2025 | Fine-Grained Action Recognition Using Cross-Modal Attention Network for Human-Robot Sign Language InteractionabstractWith the growing demand for barrier-free communication among the deaf and mute, human-robot sign language interaction has gradually gained attention as an auxiliary tool. Action recognition serves as a crucial information source for robots to understand human behavior, enabling robots to recognize signs and achieve natural interaction with deaf people through it. However, existing Human-Robot Interaction (HRI) technologies based on action recognition mainly focus on coarse-grained human movements, failing to capture and respond to nuanced actions in real-world scenarios. Additionally, multi-modal action recognition often employs early or late fusion methods to integrate various modalities, lacking the exploration of relationships between modalities, resulting in the loss of some correlated information. To enable robots to adapt to diverse scenarios for a more nuanced understanding of human behaviors, we propose a fine-grained action recognition framework using Cross-Modal Attention network (CMA) based on RGB and skeleton. Firstly, holistic features including face, hand, and body are extracted by a pose estimator, effectively representing intricate human actions. Subsequently, to fully leverage the extracted fine-grained features, skeleton is represented as heatmap volumes. Finally, a Cross-Attention Interaction (CAI) module is designed to explore the intrinsic connections between RGB and skeleton, facilitating mutual learning of their respective advantageous features in the deep layers of feature extraction, thereby achieving information interaction. Simultaneously, HRI experiments are conducted on the large-scale fine-grained action dataset, WLASL2000. In this HRI system, the robotic arm responds by performing sign language aligned with the human actions identified by CMA, showcasing the practicality and effectiveness of our proposed model in real-world scenarios. Qing Gao 0002, Xianfeng Cheng, Xuerui Li, Zhaojie Ju |
SMC | 2 |
| 2025 | SignRobot: Sign Language Recognition for Robot Interaction Based on Dual-Stream Multi-Fusion with Frame Enhancement NetworkabstractDeaf and hard-of-hearing individuals often rely on signs for communication, but limited translation resources restrict their daily needs. Robots equipped with sign language recognition and interaction capabilities can assist in bridging this gap. To enhance sign language translation resources, we design a robotic system called SignRobot, which accurately recognizes and responds to signs. To improve recognition performance, we develop a sign language recognition network based on dual-stream multi-fusion and frame enhancement (MFE-Net), using RGB and heatmaps as inputs. Specifically, an frame enhancement module with parallel spatial and motion guidance is introduced to emphasize key spatial regions and movement changes. By employing a multi-fusion strategy, we achieve feature-level interaction and adaptive late fusion between modalities, improving accuracy and robustness. Experimental results show that MFE-Net surpasses state-of-the-art methods on the PHOENIX14 and PHOENIX14-T datasets. Additionally, our SignRobot successfully demonstrates sign recognition and robotic responses in sign language, representing a promising advancement in robot-assisted communication for deaf people. Qing Gao 0002, Yuanchuan Lai, Yang Zhang 0028, Zhaojie Ju |
SMC | 2 |
| 2025 | 3D skeleton aware driver behavior recognition framework for autonomous driving system
Rongtian Huo, Junkang Chen, Ye Zhang 0037, Qing Gao 0002 |
Neurocomputing | 4 |
| 2025 | Robust Depth Estimation Under Sensor Degradations: A Multi-Sensor Fusion PerspectiveabstractThe significance of depth estimation has spurred recent endeavors to enhance it through Multi-Sensor Fusion (MSF). However, prevailing MSF methods exhibit limitations concerning accuracy and resilience when confronted with sensor degradations. While certain forms of degradation, such as suboptimal lighting and adverse weather conditions, can be mitigated by collecting pertinent data in data-driven learning, this approach proves ineffective for Out-of-Distribution (OOD) sensor degradations. In this paper, we propose a novel approach termed Combinable and Separable Multi-Sensor Fusion (CSMSF) designed to bolster depth estimation robustness against multiple sensor degradations. CSMSF hinges on four core principles: i) improved performance is achieved with an increased number of valid sensors, ii) a single valid sensor can independently enable its own depth estimation, iii) maintaining a judicious equilibrium between accuracy and model complexity, and iv) autonomous diagnosis of sensor observation failure. Leveraging these advantages, CSMSF identifies and rejects degraded sensors, allowing autonomous selection of valid sensors for scene depth estimation. The experimental results demonstrate the superior robustness of the proposed CSMSF, underscoring its efficacy in addressing challenges associated with sensor degradations across diverse environmental conditions. Junjie Hu 0003, Chenyou Fan, Mete Ozay, Qing Gao 0002, Yulan Guo, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Holistic-Based Cross-Attention Modal Fusion Network for Video Sign Language RecognitionabstractAs a bridge between the deaf people and the outside, sign language primarily involves hand movements, complemented by intricate facial and body expressions. To enhance the engagement of sign language users in today's prevalent online social activities, it is necessary to incorporate video sign language recognition (SLR) technology into social multimedia. However, current multimodal video SLR methods predominantly rely on limited data, leading to poor robustness and susceptibility to overfitting. Additionally, most works employed simple concatenation, failing to explore effective interaction among different modalities. To address these issues, we propose a cross-attention modal fusion network (CAMFuse) based on red-green-blue (RGB) and skeleton to achieve more robust video SLR. First, CAMFuse introduces a comprehensive bimodal framework that considers coarse-grained features from body and fine-grained features from hands and face. Second, departing from previous skeleton-based video SLR methods represented through graphs, CAMFuse adopts heatmap volumes to reduce storage and promote subsequent interaction. Last, a space-time cross-attention fusion (ST-CAF) module is applied to the deeper feature extraction stages of RGB and skeleton, aiming to mine complementary relationships between them and learn the excellent information from each other, reducing the independence among modalities. Experimental results demonstrate the effectiveness of our proposed CAMFuse, outperforming the state-of-the-art methods on the popular isolated video sign language datasets WLASL-2000 (53.12%) and AUTSL (96.18%) dataset. Qing Gao 0002, Haixing Mai, Zhaojie Ju |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Differentiable Prior-Driven Data Augmentation for Sensor-Based Human Activity RecognitionabstractSensor-based human activity recognition (HAR) usually suffers from the problem of insufficient annotated data, due to the difficulty in labeling the intuitive signals of wearable sensors. To this end, recent advances have adopted handcrafted operations or generative models for data augmentation. The handcrafted operations are driven by some physical priors of human activities, e.g., action distortion and strength fluctuations. However, these approaches may face challenges in maintaining semantic data properties. Although the generative models have better data adaptability, it is difficult for them to incorporate important action priors into data generation. This article proposes a differentiable prior-driven data augmentation framework for HAR. First, we embed the handcrafted augmentation operations into a differentiable module, which adaptively selects and optimizes the operations to be combined together. Then, we construct a generative module to add controllable perturbations to the data derived by the handcrafted operations and further improve the diversity of data augmentation. By integrating the handcrafted operation module and the generative module into one learnable framework, the generalization performance of the recognition models is enhanced effectively. Extensive experimental results with three different classifiers on five public datasets demonstrate the effectiveness of the proposed framework. Project page:https://github.com/crocodilegogogo/DriveData-Under-Review. Ye Zhang 0037, Qing Gao 0002, Qingtang Ding, Boyang Li 0007, Yulan Guo |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Lifelong-MonoDepth: Lifelong Learning for Multidomain Monocular Metric Depth EstimationabstractWith the rapid advancements in autonomous driving and robot navigation, there is a growing demand for lifelong learning (LL) models capable of estimating metric (absolute) depth. LL approaches potentially offer significant cost savings in terms of model training, data storage, and collection. However, the quality of RGB images and depth maps is sensor-dependent, and depth maps in the real world exhibit domain-specific characteristics, leading to variations in depth ranges. These challenges limit existing methods to LL scenarios with small domain gaps and relative depth map estimation. To facilitate lifelong metric depth learning, we identify three crucial technical challenges that require attention: 1) developing a model capable of addressing the depth scale variation through scale-aware depth learning; 2) devising an effective learning strategy to handle significant domain gaps; and 3) creating an automated solution for domain-aware depth inference in practical applications. Based on the aforementioned considerations, in this article, we present 1) a lightweight multihead framework that effectively tackles the depth scale imbalance; 2) an uncertainty-aware LL solution that adeptly handles significant domain gaps; and 3) an online domain-specific predictor selection method for real-time inference. Through extensive numerical studies, we show that the proposed method can achieve good efficiency, stability, and plasticity, leading the benchmarks by 8%-15%. The code is available at https://github.com/FreeformRobotics/Lifelong-MonoDepth. Junjie Hu 0003, Chenyou Fan, Liguang Zhou, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | TMS-Net: A multi-feature multi-stream multi-level information sharing network for skeleton-based sign language recognition
Yuquan Leng, Junkang Chen, Yang Zhang 0028, Qing Gao 0002 |
Neurocomputing | 6 |
| 2024 | AIP-Net: An anchor-free instance-level human part detection network
Ye Zhang 0037, Yuquan Leng, Qing Gao 0002 |
Neurocomputing | 4 |
| 2024 | SML: A Skeleton-based multi-feature learning method for sign language recognition
Yuquan Leng, Zengrong Lin, Xuerui Li, Qing Gao 0002 |
Knowl. Based Syst. | 6 |
| 2024 | Sharing-Net: Lightweight feedforward network for skeleton-based action recognition based on information sharing mechanism
Qing Gao 0002, Zhaojie Ju, Yulan Guo |
Pattern Recognit. | 2 |
| 2023 | Deep Depth Completion From Extremely Sparse Data: A SurveyabstractDepth completion aims at predicting dense pixel-wise depth from an extremely sparse map captured from a depth sensor, e.g., LiDARs. It plays an essential role in various applications such as autonomous driving, 3D reconstruction, augmented reality, and robot navigation. Recent successes on the task have been demonstrated and dominated by deep learning based solutions. In this article, for the first time, we provide a comprehensive literature review that helps readers better grasp the research trends and clearly understand the current advances. We investigate the related studies from the design aspects of network architectures, loss functions, benchmark datasets, and learning strategies with a proposal of a novel taxonomy that categorizes existing methods. Besides, we present a quantitative comparison of model performance on three widely used benchmarks, including indoor and outdoor datasets. Finally, we discuss the challenges of prior works and provide readers with some insights for future research directions. Junjie Hu 0003, Chenyu Bao, Mete Ozay, Chenyou Fan, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Oropharynx Visual Detection by Using a Multi-Attention Single-Shot Multibox Detector for Human-Robot Collaborative Oropharynx SamplingabstractThe pandemic of COVID-19 has increased the demand for the oropharynx sampling robots. For an automatic oropharynx sampling, detection and localization of the oropharynx objects are essential. First, in response to the small-object and real-time needs of visual oropharynx detection, a lightweight multi-attention single-shot multibox detector (MASSD) method is designed. This method can effectively improve the detection accuracy of oropharynx sampling regions, especially small regions, while ensuring sufficient speed by introducing spatial attention, channel attention, and feature fusion mechanisms into the single-shot multibox detector. Second, the proposed MASSD is applied to an oropharyngeal swab (OP-swab) robot system to detect oropharynx sampling regions and conduct autonomous sampling. In the experiment, training and validation based on a custom oropharynx dataset verify the effectiveness and efficiency of the proposed MASSD. The detection accuracy can reach 81.3% of mean average [email protected]:0.95 at 104 frames per second and the application experiment on the OP-swab robot system performs oropharynx sampling with 100% success accuracy in human–robot collaboration strategy. Qing Gao 0002, Yongquan Chen, Zhaojie Ju |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | An Efficient RGB-D Hand Gesture Detection Framework for Dexterous Robot Hand-Arm Teleoperation SystemabstractAiming at the problems of accurate and fast hand gesture detection and teleoperation mapping in the hand-based visual teleoperation of dexterous robots, an efficient hand gesture detection framework based on deep learning is proposed in this article. It can achieve an accurate and fast hand gesture detection and teleoperation of dexterous robots based on an anchor-free network architecture by using an RGB-D camera. First, an RGB-D early-fusion method based on the HSV space is proposed, effectively reducing background interference and enhancing hand information. Second, a hand gesture classification network (HandClasNet) is proposed to realize hand detection and localization by detecting the center and corner points of hands, and a HandClasNet is proposed to realize gesture recognition by using a parallel EfficientNet structure. Then, a dexterous robot hand-arm teleoperation system based on the hand gesture detection framework is designed to realize the hand-based teleoperation of a dexterous robot. Our method achieves high accuracy with fast speed on public and custom hand datasets and outperforms some state-of-the-art methods. In addition, the application of the proposed method in the hand-based teleoperation system can control the grasping of various objects by a dexterous hand-arm system in real time and accurately, which verifies the efficiency of our method. Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Chuliang Chi |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | Parallel Dual-Hand Detection by Using Hand and Body Features for Robot TeleoperationabstractVisual hand-based robot teleoperation provides a powerful guarantee for robots to complete complex tasks. However, detection and distinction of dual hands on images are difficult because of the small differences between left and right hands. To solve this problem, a parallel dual-hand detection and distinction method that combines the features of hands with the relationship features between the dual hands and body pose is proposed to achieve robust and accurate dual-hand detection. This parallel dual-hand detection method includes a hand detection module, a body pose estimation module, and a fusion module. In the hand detection module, a hand detector that realizes fast and accurate hand detection by detecting the center and corner points of hands is designed. In the body pose estimation module, a body pose estimator with dual-hand positions is proposed. The fusion module is designed to fuse hand detection and dual-hand estimation results to achieve distinction between left and right hands. Finally, the parallel dual-hand detection method is applied to a bimanual robot teleoperation system by using a designed dual-hand teleoperation framework. The proposed parallel dual-hand detection method can achieve 98.54% mAP of hand detection with 18 frames per second on a custom dual-hand detection dataset, and the bimanual robot teleoperation method can achieve 95.4% average accuracy for teleoperation tasks. Experimental results show the high accuracy and speed of our proposed parallel dual-hand detection method and its practicability in bimanual robot teleoperation. Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Shiwu Lai |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | Mouth Cavity Visual Analysis Based on Deep Learning for Oropharyngeal Swab Robot SamplingabstractThe visual analysis of the mouth cavity plays a significant role in the pathogen specimen sampling and disease diagnosis of the mouth cavity. Aiming at performance defects of general detectors based on deep learning in detecting mouth cavity components, this article proposes a mouth cavity analysis network (MCNet), which is an instance segmentation method with spatial features, and a mouth cavity dataset (MCData), which is the first available dataset for mouth cavity detecting and segmentation. First, given the lack of a mouth cavity image dataset, the MCData for detecting and segmenting key parts in the mouth cavity was developed for model training and testing. Second, the MCNet was designed based on the mask region-based convolutional neural network. To improve the performance of feature extraction, a parallel multiattention module was designed. Besides, to solve low detection accuracy of small-sized objects, a multiscale region proposal network structure was designed. Then, the mouth cavity spatial structure features were introduced, and the detection confidence could be refined to increase the detection accuracy. The MCNet achieved 81.5% detection accuracy and 78.1% segmentation accuracy (intersection over union = 0.50:0.95) on the MCData. Comparative experiments with the MCData showed that the proposed MCNet outperformed state-of-the-art approaches with the task of mouth cavity instance segmentation. In addition, the MCNet has been used in an oropharyngeal swab robot for COVID-19 oropharyngeal sampling. Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Tianwei Zhang 0002, Yuquan Leng |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2022 | Improvement of Unconstrained Appearance-Based Gaze Tracking with LSTMabstractGaze tracking is not only an important research direction in computer vision but also an important non-verbal clue in human life. What is important is that the direction of gaze can be used as a reference for judging a person’s intentions. In order to improve the accuracy of predicting gaze direction, a model of 3D gaze tracking based on bidirectional Long Short-Term Memory (LSTM) is proposed in this paper. The backbone network of the model is ResNet and its variants. The output of the model is the angular error of gaze direction. To improve the accuracy of the model prediction, the attention mechanism is adopted in this work. The ablation experiments are conducted on the selected Gaze360, which is a dataset with sufficiently large and diverse data. The angular error of the proposed model decreases from 13.5° to 12.6°. Guoxu Li, Lihong Dai, Qing Gao 0002, Hongwei Gao 0002, Zhaojie Ju |
SMC | 3 |
| 2022 | Deep Temporal Model-Based Identity-Aware Hand Detection for Space Human-Robot InteractionabstractHand detection is a crucial technology for space human-robot interaction (SHRI), and the awareness of hand identities is particularly critical. However, most advanced works have three limitations: 1) the low detection accuracy of small-size objects; 2) insufficient temporal feature modeling between frames in videos; and 3) the inability of real-time detection. In the article, a temporal detector (called TA-RSSD) is proposed based on the SSD and spatiotemporal long short-term memory (ST-LSTM) for real-time detection in SHRI applications. Next, based on the online tubelet analysis, a real-time identity-awareness module is designed for multiple hand object identification. Several notable properties are described as follows: 1) the hybrid structure of the Resnet-101 and the SSD improves the detection accuracy of small objects; 2) three-level feature pyramidal structure retains rich semantic information without losing detailed information; 3) a group of the redesigned temporal attentional LSTM (TA-LSTM) is utilized for three-level feature map modeling, which effectively achieves background suppression and scale suppression; 4) low-level attention maps are used to eliminate in-class similarity between hand objects, which improves the accuracy of identity awareness; and 5) a novel association training scheme enhances the temporal coherence between frames. The proposed model is evaluated on the SHRI-VID dataset (collected according to the task requirements), the AU-AIR dataset, and the ImageNet-VID benchmark. Extensive ablation studies and comparisons on detection and identity-awareness capacities show the superiority of the proposed model. Finally, a set of actual testing is conducted on a space robot, and the results show that the proposed model achieves a real-time speed and high accuracy. Hongwei Gao 0002, Dalin Zhou, Jinguo Liu, Qing Gao 0002, Zhaojie Ju |
IEEE Trans. Cybern. | 5 |
| 2021 | Hand gesture recognition using multimodal data fusion and multiscale parallel convolutional neural network for human-robot interactionabstractAbstract Hand gesture recognition plays an important role in human–robot interaction. The accuracy and reliability of hand gesture recognition are the keys to gesture‐based human–robot interaction tasks. To solve this problem, a method based on multimodal data fusion and multiscale parallel convolutional neural network (CNN) is proposed in this paper to improve the accuracy and reliability of hand gesture recognition. First of all, data fusion is conducted on the sEMG signal, the RGB image, and the depth image of hand gestures. Then, the fused images are generated to two different scale images by downsampling, which are respectively input into two subnetworks of the parallel CNN to obtain two hand gesture recognition results. After that, hand gesture recognition results of the parallel CNN are combined to obtain the final hand gesture recognition result. Finally, experiments are carried out on a self‐made database containing 10 common hand gestures, which verify the effectiveness and superiority of the proposed method for hand gesture recognition. In addition, the proposed method is applied to a seven‐degree‐of‐freedom bionic manipulator to achieve robotic manipulation with hand gestures. Qing Gao 0002, Jinguo Liu, Zhaojie Ju |
Expert Syst. J. Knowl. Eng. | 1 |
| 2020 | Robust real-time hand detection and localization for space human-robot interaction based on deep learning
Qing Gao 0002, Jinguo Liu, Zhaojie Ju |
Neurocomputing | 1 |