Guangshuai Gao

dblp:195/5764 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
14since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2027 DAFR-Net: Lightweight deformable adaptive feature reconstruction network for tiny object detection in remote sensing images
Guangshuai Gao, Yunqi Shang, Jiangtao Xi
Expert Syst. Appl.1
2026 DRLA-Net: Divergence-Aware and Ranking Cost-Based Label Assignment for Arbitrary-Oriented Object Detection in Remote Sensing Images
abstract
Remote sensing object detection faces several critical challenges, including 1) arbitrary orientations, 2) objects with high aspect ratios, and 3) a large number of tiny objects. To address these issues, we propose a novel arbitrary-oriented object detection framework based on a Divergence-aware and Ranking cost-based Label Assignment strategy, namely DRLA-Net. Specifically, we design a coarse-to-fine label assignment strategy (DRLA) that models rotated bounding boxes as 2D Gaussian distributions to preserve spatial and directional consistency, while incorporating scale-aware analysis to improve robustness against extreme aspect ratios. To further enhance distributional alignment, we introduce a divergence-aware similarity loss by integrating the Kullback–Leibler Divergence (KLD) with the Generalized Jensen–Shannon Divergence (GJSD), achieving more stable and symmetric optimization. Furthermore, we formulate a unified ranking cost that jointly considers semantic confidence, spatial alignment, and Gaussian similarity, thereby enabling dynamic selection of high-quality positive samples and alleviating the conflict between classification confidence and localization quality. The proposed strategy is applied across the feature pyramid network, which significantly improves the localization of tiny objects with arbitrary shapes and orientations. Extensive experiments on two large-scale rotated object detection datasets, DIOR-R and DOTA-v2.0, demonstrate the effectiveness and superiority of the proposed method.
Guangshuai Gao, Yunqi Shang, Minghong Wei, Zhenduo Guo
IEEE Geosci. Remote. Sens. Lett.1
2026 Supervised Anomaly Detection for Generalized Deepfake Detection
abstract
Generalization to unseen forgeries remains a significant challenge in deepfake detection. Existing methods often focus on detecting specific forgery patterns, such as blending boundaries, identity leakage, etc. However, their overfitting to known fake samples during training restricts their generalization to unseen manipulations. To address this limitation, we propose Supervised Anomaly Detection for Generalized Deepfake Detection (SAD-GDD), a novel framework that treats real data as normal and fake data as anomalies, leveraging anomaly detection to improve generalization. At its core, we propose a Supervised Anomaly Detection Strategy (SADS), which constructs a refined Multivariate Gaussian Distribution (MGD) for real data and applies supervised guidance for fake data to strengthen the discrimination between real and fake. To further enhance performance, we design a two-stream network comprising a High-Frequency Texture Branch (HFT-B) and an RGB Branch (RGB-B), both optimized with SADS. Additionally, we propose a Multi-head Cross-Attention Enhancement Block (MCAEB) to enhance the capability of RGB-B in capturing high-frequency features. Extensive experiments show that SAD-GDD achieves superior generalization compared to state-of-the-art methods across multiple widely used datasets. The code will be made publicly available upon publication.
Guangshuai Gao
IEEE Signal Process. Lett.2
2025 PACTFormer: Peak-Aware Cross-Temporal Transformer for Temporal Action Detection
Zhewen Zhou, Chunlei Li 0002, Guangshuai Gao
PRCV (7)5
2025 Enhanced Foreground-Background Discrimination for Weakly Supervised Semantic Segmentation
abstract
ABSTRACT Weakly supervised semantic segmentation (WSSS) methods are extensively studied due to the availability of image‐level annotations. Relying on class activation maps (CAMs) derived from original classification networks often suffers from issues such as inaccurate object localization, incomplete object regions, and the inclusion of confusing background pixels. To address these issues, we propose a two‐stage method that enhances the foreground–background discriminative ability in a global context (FB‐DGC). Specifically, a cross‐domain feature calibration module (CFCM) is first proposed to calibrate foreground and background salient features using global spatial location information, thereby expanding foreground features while mitigating the impact of inaccurate localization in class activation regions. A class‐specific distance module (CSDM) is further adopted to facilitate the separation of foreground–background features, thereby enhancing the activation of target regions, which alleviates the over‐smoothing of features produced by the network and mitigates issues associated with confused features. In addition, an adaptive edge feature extraction (AEFE) strategy is proposed to identify target features in candidate boundary regions and capture missed features, compensating for drawbacks in recognising the co‐occurrence of multiple targets. The proposed method is extensively evaluated on the challenging PASCAL VOC 2012 and MS COCO 2014 datasets, demonstrating its feasibility and superiority.
Zhoufeng Liu, Bingrui Li, Miao Yu 0005, Guangshuai Gao, Chunlei Li 0002
IET Comput. Vis.4
2024 Deepfake Detection Via Separable Self-Consistency Learning
abstract
Deepfake detection technologies have been developed rapidly in recent years, due to the potential severe security threats induced by the realistic deep facial forgeries. Among the existing deepfake detection methods, self-supervised methods have drawn significant attentions from researchers, because of their better generalization ability against the deep forgeries produced via unseen deepfake techniques. Unfortunately, existing state-of-the-art self-supervised approaches have not properly considered that different pairs of patches from different regions actually give different contributions. Thus, their learned representations are coarse and the generalization performances are less decent. In this paper, we propose a new self-supervised deepfake detection method, named deepfake detection via separable self-consistency learning (SSCLDFD), to improve the generalization ability of deepfake detection. Specifically, to effectively extract detection features, we construct a multi-scale Texture Enhanced Feature Extraction Network (TEFEN), by forming a Central-Difference based Convolution Module (CDCM) to enhance the texture information, which contain rich forgery cues. Since different pairs of patches from different regions (i.e. background and facial regions) tend to give various consistencies, we propose a separable self-consistency loss to explicitly constrain the representation learning. Extensive experiments demonstrate that our SSCL-DFD can give superior generalization performances compared to the state-of-the-art methods.
Yunhong Wang 0001, Wenqi Zhuo, Guangshuai Gao, Yuanfang Guo
ICIP5
2024 DSLA: A Distance-Sensitive Label Assignment Strategy for Oriented Object Detection in Remote Sensing Images
Minghong Wei, Haobin Xiang, Guangshuai Gao, Chunlei Li 0002
ICPR (17)4
2024 Global information aware network with global interaction graph attention for infrared small target detection
abstract
Abstract Detecting small targets in infrared images is crucial for ground surveillance and air traffic control. However, distinguishing small infrared targets from similar backgrounds is challenging due to their lack of structural and textural characteristics. To address these challenges, this study proposes a novel global information‐aware network with global interaction graph attention (GIGA) for infrared small target detection. The GIGA consists of a global interaction layer (GILayer), graph attention weights (GAW), and a global relational learning (GRL) module. Specifically, the GILayer dynamically learns global inter‐pixel relationships of small target images by enhancing the dependencies between feature dimensions. The GAW component calculates pixel‐by‐pixel similarity across the entire feature map using graph attention mechanisms, while the GRL module retains critical similarity features in the feature extraction network, thereby facilitating small target detection. Additionally, the multi‐scale context fusion module utilises self‐attention and dilation convolution to complement richer feature details at different scales. Experimental results on both natural and synthetic datasets demonstrate the proposed method's superiority over other state‐of‐the‐art conventional and deep learning approaches in infrared small target detection.
Ruimin Yang, Guangshuai Gao, Chunlei Li 0002
IET Image Process.3
2024 YOLC: You Only Look Clusters for Tiny Object Detection in Aerial Images
abstract
Detecting objects from aerial images poses significant challenges due to the following factors: 1) Aerial images typically have very large sizes, generally with millions or even hundreds of millions of pixels, while computational resources are limited. 2) Small object size leads to insufficient information for effective detection. 3) Non-uniform object distribution leads to computational resource wastage. To address these issues, we propose YOLC (You Only Look Clusters), an efficient and effective framework that builds on an anchor-free object detector, CenterNet. To overcome the challenges posed by large-scale images and non-uniform object distribution, we introduce a Local Scale Module (LSM) that adaptively searches cluster regions for zooming in for accurate detection. Additionally, we modify the regression loss using Gaussian Wasserstein distance (GWD) to obtain high-quality bounding boxes. Deformable convolution and refinement methods are employed in the detection head to enhance the detection of small objects. We perform extensive experiments on two aerial image datasets, including Visdrone2019 and UAVDT, to demonstrate the effectiveness and superiority of our proposed approach.
Guangshuai Gao, Ziyue Huang 0001, Qingjie Liu 0001, Yunhong Wang 0001
IEEE Trans. Intell. Transp. Syst.2
2023 A Key Feature-Enhanced Network for Remote Sensing Object Detection
abstract
Currently, Remote sensing object detection methods based on deep learning have been widely used in military investigation, ocean monitoring, urban planning and post-disaster relief. However, many difficulties, such as complex backgrounds, indistinguishable object appearances, large-scale variations, and non-uniform distributions, limit the accuracy of detectors. To deal with these problems, a key feature-enhanced network (KFENet) is proposed to improve the detection accuracy. Specifically, a multi-scale feature fusion (MSFF) module is given to obtain richer feature expression, which can explore inherent structural information of images at different scales. Next, an effective activation detection head (EAD-Head) is designed to adaptively capture key features required for classification and regression tasks, respectively. Finally, we implement our strategy within the framework of YOLOX. Experiments on DIOR and RSOD datasets show that the proposed algorithm outperforms state of the art methods.
Yundong Liu, Haonan Kang, Guangshuai Gao, Chunlei Li 0002, Zhoufeng Liu
ICIP4
2022 PSGCNet: A Pyramidal Scale and Global Context Guided Network for Dense Object Counting in Remote-Sensing Images
abstract
Object counting, which aims to count the accurate number of object instances in images, has been attracting more and more attention. However, challenges such as large-scale variation, complex background interference, and nonuniform density distribution greatly limit the counting accuracy, particularly striking in remote-sensing imagery. To mitigate the above issues, this article proposes a novel framework for dense object counting in remote-sensing images, which incorporates a pyramidal scale module (PSM) and a global context module (GCM), dubbed PSGCNet, where PSM is used to adaptively capture multi-scale information and GCM is to guide the model to select suitable scales generated from PSM. Moreover, a reliable supervision manner improved from Bayesian and counting loss (BCL) is utilized to learn the density probability and then compute the count expectation at each annotation. It can relieve nonuniform density distribution to a certain extent. Extensive experiments on four remote-sensing counting datasets demonstrate the effectiveness of the proposed method and its superiority compared with state of the arts. Additionally, experiments extended on four commonly used crowd counting datasets further validate the generalization ability of the model. Code is available athttps://github.com/gaoguangshuai/psgcnet.
Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 MRDet: A Multihead Network for Accurate Rotated Object Detection in Aerial Images
abstract
Objects in aerial images usually have arbitrary orientations and are densely located over the ground, making them extremely challenge to be detected. Many of the recent developed methods attempt to solve these issues by estimating an extra orientation parameter and placing dense anchors, which will result in high model complexity and computational costs. In this article, we propose an arbitrary-oriented region proposal network (AO-RPN) to generate oriented proposals transformed from horizontal anchors. The AO-RPN is very efficient with only a few amounts of parameters increase than the original RPN. Furthermore, to obtain accurate bounding boxes, we decouple the detection task into multiple subtasks and propose a multihead network to accomplish them. Each head is specially designed to learn the features optimal for the corresponding task, which allows our network to detect objects accurately. We name it multihead rotated object detector (MRDet). We evaluate the performance of the proposed MRDet on two challenging benchmarks, i.e., DOTA and HRSC2016, and compare it with several state-of-the-art methods. Our method achieves very promising results, which clearly demonstrates its effectiveness. Code has been available athttps://github.com/qinr/MRDet.
Ran Qin, Qingjie Liu 0001, Guangshuai Gao, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Co-Saliency Detection With Co-Attention Fully Convolutional Network
abstract
Co-saliency detection aims to detect common salient objects from a group of relevant images. Some attempts have been made with the Fully Convolutional Network (FCN) framework and achieve satisfactory detection results. However, due to stacking convolution layers and pooling operation, the boundary details tend to be lost. In addition, existing models often utilize the extracted features without discrimination, leading to redundancy in representation since actually not all features are helpful to the final prediction and some even bring distraction. In this paper, we propose a co-attention module embedded FCN framework, called as Co-Attention FCN (CA-FCN). Specifically, the co-attention module is plugged into the high-level convolution layers of FCN, which can assign larger attention weights on the common salient objects and smaller ones on the background and uncommon distractors to boost final detection performance. Extensive experiments on three popular co-saliency benchmark datasets demonstrate the superiority of the proposed CA-FCN, which outperforms state-of-the-arts in most cases. Besides, the effectiveness of our new co-attention module is also validated with ablation studies.
Guangshuai Gao, Wenting Zhao 0007, Qingjie Liu 0001, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Counting From Sky: A Large-Scale Data Set for Remote Sensing Object Counting and a Benchmark Method
abstract
Object counting, whose aim is to estimate the number of objects from a given image, is an important and challenging computation task. Significant efforts have been devoted to addressing this problem and achieved great progress, yet counting the number of ground objects from remote sensing images is barely studied. In this article, we are interested in counting dense objects from remote sensing images. Compared with object counting in a natural scene, this task is challenging in the following factors: large-scale variation, complex cluttered background, and orientation arbitrariness. More importantly, the scarcity of data severely limits the development of research in this field. To address these issues, we first construct a large-scale object counting data set with remote sensing images, which contains four important geographic objects: buildings, crowded ships in harbors, and large vehicles and small vehicles in parking lots. We then benchmark the data set by designing a novel neural network that can generate a density map of an input image. The proposed network consists of three parts, namely attention module, scale pyramid module, and deformable convolution module (DCM) to attack the aforementioned challenging factors. Extensive experiments are performed on the proposed data set and one crowd counting data set, which demonstrates the challenges of the proposed data set and the superiority and effectiveness of our method compared with state-of-the-art methods.
Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Counting Dense Objects in Remote Sensing Images
abstract
Estimating accurate number of interested objects from a given image is a challenging yet important task. Significant efforts have been made to address this problem and achieve great progress, yet counting number of ground objects from remote sensing images is barely studied. In this paper, we are interested in counting dense objects from remote sensing images. Compared with object counting in natural scene, this task is challenging in following factors: large scale variation, complex cluttered background and orientation arbitrariness. More importantly, the scarcity of data severely limits the development of research in this field. To address these issues, we first construct a large-scale object counting dataset based on remote sensing images, which contains four kinds of objects: buildings, crowded ships in harbor, large-vehicles and small-vehicles in parking lot. We then benchmark the dataset by designing a novel neural network which can generate density map of an input image. The proposed network consists of three parts namely convolution block attention module (CBAM), scale pyramid module (SPM) and deformable convolution module (DCM). Experiments on the proposed dataset and comparisons with state of the art methods demonstrate the challenging of the proposed dataset, and superiority and effectiveness of our method.
Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001
ICASSP1
2020 Unsupervised Conditional Disentangle Network For Image Dehazing
abstract
Image dehazing aims to restore the blurry image information caused by the ambiguities of unknown scene radiance and transmission. Instead of using paired images or depth information, we propose an Unsupervised Conditional Disentangle Network (UCDN) using unpaired dataset. Our approach enforces the constraint by introducing physical-based disentanglement. Unlike other unsupervised dehazing models, our approach adapts the multi-concentration of fog and outperforms on the dataset with different concentrations. Extensive experiments on synthesized dataset demonstrate that our approach can surpass state-of-the-arts. Meanwhile, through benchmarking on our collected natural hazy dataset, our approach can generate more perceptually appealing dehazing results.
Yizhou Jin, Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001
ICIP2
2019 Robust low-rank decomposition of multi-channel feature matrices for fabric defect detection
Chunlei Li 0002, Chaodie Liu, Guangshuai Gao, Zhoufeng Liu, Yu-Ping Wang 0002
Multim. Tools Appl.3
2017 Fabric Defect Detection Algorithm Based on Multi-channel Feature Extraction and Joint Low-Rank Decomposition
Chaodie Liu, Guangshuai Gao, Zhoufeng Liu, Chunlei Li 0002
ICIG (1)2