Ruigang Fu

dblp:174/9736 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-1279-0615ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2025 Learning a Perspective-Invariant Descriptor for Remote Sensing Image Matching
abstract
Image descriptors are crucial in remote sensing image matching tasks. However, the presence of nonlinear transformation and dimensional collapse inherent in the perspective imaging process often poses challenges to achieving accurate matches. Existing descriptors lack a theoretical analysis of the perspective distortion process and fail to mine the patterns hidden in the perspective imaging process, consequently limiting their efficacy in remote sensing image matching. To uncover the underlying patterns in the image and devise a perspective-invariant descriptor, this paper proposes a perspective-invariant descriptor network (PIDNet). In our approach, we first analyze the remote sensing imaging process and demonstrate that it can be described in a new, conceptually simple linear space named the perspective distortion space. Second, we extract the bases from this space via the intersection-over-union (IoU) metric. As a result, each element in the space can be linearly expressed by the bases. Finally, we utilize these bases to design and learn a perspective-invariant descriptor. The core idea of our descriptor is based on the fact that each base corresponds to a unique imaging viewpoint. Therefore, any imaging viewpoint can be linearly represented as a combination of the bases. To implement our PIDNet, we propose a perspective sampling network module (PSNM) based on the spatial transform networks (STN) since no modules are available for our image sampling process. Furthermore, we introduce a perspective convolutional layer (PCLayer) to extract intermediate covariant features. Then, we concatenate the covariant features to learn a perspective-invariant descriptor. Experimental results on three datasets, including single-modal and multi-modal images, demonstrate the superior performance of PIDNet compared to state-of-the-art methods. Our source code will be publicly available at PIDNet.
Jia Wang 0054, Zhiguo Qu, Lingshuang Kong, Encai Liu, Ruigang Fu
IEEE Trans. Circuits Syst. Video Technol.7
2024 RPSC: Robust Pseudo-Labeling for Semantic Clustering
abstract
Clustering methods achieve performance improvement by jointly learning representation and cluster assignment. However, they do not consider the confidence of pseudo-labels which are not optimal as supervised information, resulting into error accumulation. To address this issue, we propose a Robust Pseudo-labeling for Semantic Clustering (RPSC) approach, which includes two stages. In the first stage (RPSC-Self), we design a semantic pseudo-labeling scheme by using the consistency of samples, i.e., samples with same semantics should be close to each other in the embedding space. To exploit robust semantic pseudo-labels for self-supervised learning, we propose a soft contrastive loss (SCL) which encourage the model to believe high-confidence sematic pseudo-labels and be less driven by low-confidence pseudo-labels. In the second stage (RPSC-Semi), we first determine the semantic pseudo-label of a sample based on the distance between itself and cluster centers, followed by screening out reliable semantic pseudo-label by exploiting the consistency. These reliable pseudo-labels are used as supervised information in the pseudo-semi-supervised learning algorithm to further improve the performance. Experimental results show that RPSC outperforms 18 competitive clustering algorithms significantly on six challenging image benchmarks. In particular, RPSC achieves an accuracy of 0.688 on ImageNet-Dogs, which is an up to 24% improvement, compared with the second-best method. Meanwhile, we conduct ablation studies to investigate effects of different augmented strategies on RPSC as well as contributions of terms in SCL to clustering performance. Besides, experimental results indicate that SCL can be easily integrated into existing clustering methods and bring performance improvement.
Sihang Liu 0007, Wenming Cao 0002, Ruigang Fu, Kaixiang Yang 0001, Zhiwen Yu 0002
AAAI3
2024 Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection
abstract
Visible-infrared (RGB-IR) image fusion has shown great potentials in object detection based on unmanned aerial ve-hicles (UAVs). However, the weakly misalignment problem between multimodal image pairs limits its performance in object detection. Most existing methods often ignore the modality gap and emphasize a strict alignment, resulting in an upper bound of alignment quality and an increase of implementation costs. To address these challenges, we propose a novel method named Offset-guided Adaptive Feature Alignment (OAFA), which could adaptively adjust the relative positions between multimodal features. Considering the impact of modality gap on the cross-modality spa-tial matching, a Cross-modality Spatial Offset Modeling (CSOM) module is designed to establish a common sub-space to estimate the precise feature-level offsets. Then, an Offset-guided Deformable Alignment and Fusion (ODAF) module is utilized to implicitly capture optimal fusion po-sitions for detection task rather than conducting a strict alignment. Comprehensive experiments demonstrate that our method not only achieves state-of-the-art performance in the UAVs-based object detection task but also shows strong robustness to the weakly misalignment problem.
Chen Chen 0152, Jiahao Qi, Kangcheng Bin, Ruigang Fu, Xikun Hu, Ping Zhong 0001
CVPR5
2024 Lighten CARAFE: Dynamic Lightweight Upsampling with Guided Reassemble Kernels
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yinghui Gao, Ping Zhong 0001
ICPR (4)1
2024 Content-Aware Feature Upsampling for Voxel-Based 3D Semantic Segmentation
Ruigang Fu, Qingyong Hu, Ping Zhong 0001
ICPR (30)2
2023 Enhanced Cross-Domain Dim and Small Infrared Target Detection via Content-Decoupled Feature Alignment
abstract
The detection and recognition of dim and small infrared (IR) targets across domains pose two formidable challenges: distributional discrepancies in samples and scarcity or absence of annotated instances in the target domain. While current unsupervised domain adaptive object detection methods can somewhat alleviate the performance degradation caused by these issues, they fail to address the differences in semantic content between the background environments in different application scenarios. This results in a semantic gap that impedes the algorithm’s adaptability. We propose a content-decoupled unsupervised domain adaptive method to mitigate the adverse impact of the semantic gap on domain adaptation. Specifically, we introduce an unsupervised task-guided branch in the one-stage detector that executes style transfer in real-time during training, directing the feature representation via shared parameters. A novel content-decoupled feature alignment module aligns semantically unrelated features in the source and target domains while preserving the intrinsic semantics of dim and small IR targets. The experiments on a cross-band IR dataset reveal that the proposed method enhances the mean average precision by 18.3% compared to the baseline detector without increasing computational burden.
Yu Zhang 0060, Yan Zhang 0147, Zhiguang Shi, Ruigang Fu, Di Liu 0018, Yi Zhang 0128
IEEE Trans. Geosci. Remote. Sens.4
2022 Multi-aperture optical imaging systems and their mathematical light field acquisition models
abstract
Inspired by the compound eyes of insects, many multi-aperture optical imaging systems have been proposed to improve the imaging quality, e.g., to yield a high-resolution image or an image with a large field-of-view. Previous research has reviewed existing multi-aperture optical imaging systems, but few papers emphasize the light field acquisition model which is essential to bridge the gap between configuration design and application. In this paper, we review typical multi-aperture optical imaging systems (i.e., artificial compound eye, light field camera, and camera array), and then summarize general mathematical light field acquisition models for different configurations. These mathematical models provide methods for calculating the key indexes of a specific multi-aperture optical imaging system, such as the field-of-view and sub-image overlap ratio. The mathematical tools simplify the quantitative design and evaluation of imaging systems for researchers.
Qiming Qi, Ruigang Fu, Zhengzheng Shao, Hongqi Fan
Frontiers Inf. Technol. Electron. Eng.2
2022 Remote Sensing Object Detection Based on Receptive Field Expansion Block
abstract
Due to the rapid development of deep learning techniques and the collection of large-scale remote sensing datasets, convolutional neural networks (CNNs) have made significant progress in remote sensing object detection. However, due to the diversity of objects in remote sensing images, multiscale object detection is still a challenging task. In this letter, a novel object detection framework based on feature pyramid network (FPN) is proposed to improve the detection performance of multiscale objects. First, a receptive field expansion block (RFEB) is designed and added on the top of the backbone to expand the receptive field of FPN adaptively. In this way, the context information around each object is well captured. Then, the features obtained via RFEB are delivered to feature maps at all pyramid levels, remedying the drawback of FPN that semantic information captured by deep layers is gradually diluted when transmitted to lower layers. Third, since the classic backbone of FPN, which produces large receptive fields based on large downsampling factors, may limit the effectiveness of RFEB, the backbone of the original FPN is modified using dilated convolution to ease the resolution drop of feature maps while maintaining a large receptive field. As a feature extractor, the proposed framework can be easily deployed in other FPN-based methods. The experiments on the benchmark for object DetectIon in Optical Remote sensing images (DIOR) dataset demonstrate the proposed method’s superiority over considered state-of-the-art baseline methods in terms of detection accuracy.
Xiaohu Dong, Ruigang Fu, Yinghui Gao, Yao Qin 0002, Yuanxin Ye
IEEE Geosci. Remote. Sens. Lett.2
2022 Remote Sensing Object Detection Based on Gated Context-Aware Module
abstract
Recently, deep learning algorithms, especially feature pyramid network (FPN), have achieved significant progress in object detection of natural scene images. However, due to the complex scenes of remote sensing images and the diversity of remote sensing objects, FPN still faces the following drawback when applied to remote sensing object detection. Specifically, in the original FPN, the features of each proposal are extracted by RoIAlign. However, these features have limited effective receptive fields, making FPN lack of crucial contextual information to accurately classify and locate objects, as well as filter some background noises that possess similar appearance with objects. To alleviate the above problem, in this letter, we propose a gated context aware module (G-CAM), and replace the original RoIAlign in FPN with the proposed G-CAM to adaptively incorporate the useful local context surrounding each proposal and the global context of the whole image into FPN, enabling FPN to effectively detect objects in remote sensing images Extensive experiments have been conducted on the DIOR and RSOD datasets, which validates that the proposed method achieves superior performance to the considered state-of-the-art methods in terms of detection accuracy.
Xiaohu Dong, Yao Qin 0002, Ruigang Fu, Yinghui Gao, Yuanxin Ye
IEEE Geosci. Remote. Sens. Lett.3
2022 Multiscale Deformable Attention and Multilevel Features Aggregation for Remote Sensing Object Detection
abstract
In this letter, a novel object detection method based on feature pyramid network (FPN) is proposed to improve the detection performance of remote sensing objects. First, since the information in the background regions may interfere with object detection, a novel multi-scale deformable attention module (MSDAM) is designed and added on the top of the backbone of FPN to make the network suppress the background features while highlight the target features. The proposed MSDAM generates attention maps from feature maps with multi-scale deformable receptive fields, thus can fit remote sensing objects of various shapes and sizes better and predict more precise attention maps for remote sensing images. Second, in the original FPN, each proposal is predicted based on feature grids pooled from only one feature level. This process is suboptimal as the information discarded in other feature levels and the global contextual information are also meaningful to object detection. Thus, a multi-level features aggregation module (MLFAM) is proposed to aggregate the multi-level outputs of FPN and the global context of the whole image, generating more powerful pyramidal representations for the subsequent object detection. The experiments conducted on the DIOR and RSOD datasets demonstrate the superiority of the proposed method over the considered state-of-the-art baseline methods in terms of detection accuracy.
Xiaohu Dong, Yao Qin 0002, Ruigang Fu, Yinghui Gao, Yuanxin Ye
IEEE Geosci. Remote. Sens. Lett.3
2022 Learning Nonlocal Quadrature Contrast for Detection and Recognition of Infrared Rotary-Wing UAV Targets in Complex Background
abstract
Traditional data-driven algorithms suffer from data reliance, hyperparameter sensitivity, and faint characteristics in infrared (IR) “low, slow, and small” unmanned aerial vehicle target detection and recognition, resulting in performance degradation in complex backgrounds. Inspired by model-driven methods, this article proposes a learnable feature modulation module that uses prior knowledge to enhance feature representation. Specifically, this method converts the local contrast measure into a nonlocal quadrature difference measure in deep feature space, considering feature points that break semantic continuity as the potential target locations through a self-attentive approach. On this basis, considering the scale changes of aircraft targets during radial approach to IR detectors, a multiscale single-stage detector is designed by effective receptive field calculation. In this network structure, a bidirectional serial feature modulation method is used to fully retain the multiscale features of the target and ensure adaptability to point, spot, and area targets while satisfying real-time requirements. The ablation studies verify the effectiveness of each component and help determine the optimal parameter configuration. Finally, comparison experiments with state-of-the-art methods are conducted on a 10k scale IR dataset. The experimental results show that the detection accuracy of this method is better than that of other baseline methods while ensuring real-time performance, especially in highly complex and low-contrast scenes, achieving superior higher target detection accuracy.
Yu Zhang 0060, Yan Zhang 0147, Ruigang Fu, Zhiguang Shi, Di Liu 0018
IEEE Trans. Geosci. Remote. Sens.3
2020 Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao
BMVC1
2018 CNN with coarse-to-fine layer for hierarchical classification
abstract
Most of the traditional convolution neural network (CNN)‐based classification models are flat classifiers, which have an underlying assumption that all classes are equally difficult to distinguish. However, visual separability between different object categories is highly uneven in the real world. Recently, hierarchical classification has been proven effective for CNNs, more and more attempts have been made to exploit category hierarchies in CNN models. In this study, the authors propose a novel hierarchical CNN architecture, called coarse‐to‐fine CNN. It is simple, with a proposed coarse‐to‐fine layer on the top of a generic CNN. The coarse‐to‐fine layer is inspired by the Bayesian equation, where the coarse prediction can affect the fine prediction directly. Arbitrary CNNs can perform the hierarchical classification by adding the proposed layer. The training of a coarse‐to‐fine CNN is end‐to‐end, it can be optimised by typical stochastic gradient descent. In the test phase, it outputs multiple hierarchical predictions simultaneously. Experimental results on the benchmark datasets MNIST, CIFAR‐10, and CIFAR‐100 show clear advantages over the compared baselines.
Ruigang Fu, Yinghui Gao
IET Comput. Vis.1
2018 Visualizing and analyzing convolution neural networks with gradient information
Ruigang Fu, Yinghui Gao
Neurocomputing1
2016 Fully automatic figure-ground segmentation algorithm based on deep convolutional neural network and GrabCut
abstract
Figure‐ground segmentation is used to extract the foreground from the background, where the foreground is usually defined as the region containing the most meaningful object of the image. In fact, the algorithms that take advantage of human–computer interaction often attain better performance and they are based on the ‘one‐to‐one’ model. In this study, the authors present a novel algorithm for figure‐ground segmentation based on the GrabCut algorithm, which is a common segmentation algorithm that is user interactive. However, instead of a real user, they attempt to use a pre‐trained deep convolutional neural network to interact with GrabCut for completing its job successfully. Weizmann's segmentation evaluation database is used as the test dataset and the results show that their algorithm works well for figure‐ground segmentation. While the previous automatic segmentation algorithms are required to rank their segments empirically in order to find the position of the foreground after the segmentation, their algorithm is fully automatic.
Ruigang Fu, Yinghui Gao
IET Image Process.1