Hongwei Gao 0002

dblp:12/2302-2 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-7666-2970ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Spatiotemporal View-Reset Deep Learning With Attentional GRUs for Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted significant attention. However, the skeleton spatial invariance and temporal context modeling are recent challenges for most existing methods. This work proposes a view-reset network, which integrates two important branches: spatial view-reset module (SVRM) and temporal attention module (TAM). In the SVRM, the skeleton model from different perspectives is reset in a unified coordinate system, eliminating the influence of viewpoint changes. In the TAM, the perception of temporal features is jointly enhanced by weighting each frame’s importance in the context. Furthermore, the pretrained residual network (ResNet) is used for prediction. The sample size is increased through data augmentation to improve the robustness of the model. The SVRM, TAM, and ResNet form an end-to-end learning network. The ablation study proved that the model could record the key skeletons and frames in the sequence and then reset the human body to a new position, making it easy for learning. The proposed model is evaluated on four challenging benchmarks based on the performance of the cross-view evaluation metrics. Experiments prove that the proposed model has superior performance and surpasses many state-of-the-art algorithms, with an increase of 1.91% over the top ten on the NTU RGB+D 60 dataset.
Xuna Wang, Hongwei Gao 0002, Zide Liu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.3
2025 GenBEV: Generative Model With Semantic Compensation for Bird's Eye View Segmentation
abstract
Bird’s-Eye View (BEV) semantic segmentation is a key technology for constructing high-precision maps in low-cost visual navigation systems. The main challenge lies in effectively transforming image features into BEV features while preserving rich BEV visual information. Recent works have shown that generative models hold great promise in advancing BEV segmentation. However, these methods primarily focus on producing BEV features using prior knowledge, often overlooking key challenges such as feature shift, confusion, and forgetting during the BEV feature generation process. In this paper, we propose GenBEV, a generative model with semantic compensation that formally addresses inaccuracies and confusion in BEV feature generation. GenBEV leverages the synergistic benefits of data fusion consistency and noise-reduction training to enhance the diversity and reliability of the generated information. This improvement boosts the robustness and generalization of BEV segmentation across diverse scenarios, including those involving complex objects and low-quality images. Specifically, we design an adaptive cross-feature encoder to reduce diffusion variability. During decoding, we integrate the context of BEV features with noisy features to construct semantic embeddings. We show the effectiveness of GenBEV on the nuScenes, KITTI Raw, and KITTI 3D Object datasets. GenBEV achieves segmentation scores of 29.5%, 68.8%, and 39.7%, respectively, surpassing current methods by up to 3.6%, 2.4%, and 2.7%. To the best of our knowledge, GenBEV is the first to address the problem of BEV feature falsification in generative architectures.
Weiming Fan, Yuping Guo, Hong Lyu, Hongwei Gao 0002, Changting Lin, Xu Cheng 0003
IEEE Trans. Intell. Transp. Syst.5
2022 Improvement of Unconstrained Appearance-Based Gaze Tracking with LSTM
abstract
Gaze tracking is not only an important research direction in computer vision but also an important non-verbal clue in human life. What is important is that the direction of gaze can be used as a reference for judging a person’s intentions. In order to improve the accuracy of predicting gaze direction, a model of 3D gaze tracking based on bidirectional Long Short-Term Memory (LSTM) is proposed in this paper. The backbone network of the model is ResNet and its variants. The output of the model is the angular error of gaze direction. To improve the accuracy of the model prediction, the attention mechanism is adopted in this work. The ablation experiments are conducted on the selected Gaze360, which is a dataset with sufficiently large and diverse data. The angular error of the proposed model decreases from 13.5° to 12.6°.
Guoxu Li, Lihong Dai, Qing Gao 0002, Hongwei Gao 0002, Zhaojie Ju
SMC4
2022 Deep Temporal Model-Based Identity-Aware Hand Detection for Space Human-Robot Interaction
abstract
Hand detection is a crucial technology for space human-robot interaction (SHRI), and the awareness of hand identities is particularly critical. However, most advanced works have three limitations: 1) the low detection accuracy of small-size objects; 2) insufficient temporal feature modeling between frames in videos; and 3) the inability of real-time detection. In the article, a temporal detector (called TA-RSSD) is proposed based on the SSD and spatiotemporal long short-term memory (ST-LSTM) for real-time detection in SHRI applications. Next, based on the online tubelet analysis, a real-time identity-awareness module is designed for multiple hand object identification. Several notable properties are described as follows: 1) the hybrid structure of the Resnet-101 and the SSD improves the detection accuracy of small objects; 2) three-level feature pyramidal structure retains rich semantic information without losing detailed information; 3) a group of the redesigned temporal attentional LSTM (TA-LSTM) is utilized for three-level feature map modeling, which effectively achieves background suppression and scale suppression; 4) low-level attention maps are used to eliminate in-class similarity between hand objects, which improves the accuracy of identity awareness; and 5) a novel association training scheme enhances the temporal coherence between frames. The proposed model is evaluated on the SHRI-VID dataset (collected according to the task requirements), the AU-AIR dataset, and the ImageNet-VID benchmark. Extensive ablation studies and comparisons on detection and identity-awareness capacities show the superiority of the proposed model. Finally, a set of actual testing is conducted on a space robot, and the results show that the proposed model achieves a real-time speed and high accuracy.
Hongwei Gao 0002, Dalin Zhou, Jinguo Liu, Qing Gao 0002, Zhaojie Ju
IEEE Trans. Cybern.2
2022 Deep Object Detector With Attentional Spatiotemporal LSTM for Space Human-Robot Interaction
abstract
Global temporal information and local semantic information are essential cues for high-performance online object detection in videos. However, despite their promising detection accuracy in most cases, most state-of-the-art approaches have following two limitations: invalid background/scale suppression and inadequate temporal information mining between frames. Many jobs currently focus on temporal information learning based on a single frame. In this article, we propose an attentional global–local information learning network; this is one of the first attempts to fully use both types of information between frames. Attention maps are creatively utilized to transfer temporal contexts between frames. This also effectively alleviates the adverse effects of scale changes. Furthermore, empowered by a detailed framework, a proposed detector effectively uses multilevel feature extraction. Given these contributions, the proposed detector achieves state-of-the-art performance on challenging benchmarks. Finally, practical experiments are conducted on a space human–robot interaction platform.
Hongwei Gao 0002, Yongquan Chen, Dalin Zhou, Jinguo Liu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.2
2013 PSO-Based SIFT False Matches Elimination for Zooming Image
Hongwei Gao 0002, Dai Peng, Ben Niu 0002, Bin Li 0001
ICIC (2)1
2011 Restoration of Epipolar Line Based on Multi-population Cooperative Particle Swarm Optimization
Hongwei Gao 0002, Jinguo Liu, Fuguo Chen, Ben Niu 0002
ICIC (3)1
2010 An Improved Image Rectification Algorithm Based on Particle Swarm Optimization
Hongwei Gao 0002, Ben Niu 0002, Bin Li 0001, Yang Yu 0002
ICIC (1)1
2009 An Improved Two-Stage Camera Calibration Method Based on Particle Swarm Optimization
Hongwei Gao 0002, Ben Niu 0002, Yang Yu 0002
ICIC (2)1