Fenlong Jiang

dblp:246/6895 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-3714-0600ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficiency-Aware Federated Learning for Image Classification via Model Adaptation
Yiheng Lu, Ziyu Guan, Maoguo Gong, Wei Zhao 0019, Zhuping Hu, Fenlong Jiang
IEEE Trans. Circuits Syst. Video Technol.6
2026 Uncertainty-Aware Local Bayesian Framework for Hyperspectral Image Classification With Noisy Labels
abstract
Various deep learning-based methods have greatly improved hyperspectral image (HSI) classification performance, but these models are sensitive to noisy training labels. Human annotation on remote sensing images inevitably introduced label noise, which degrades the model prediction confidence. Understanding the spatial characteristics and distribution of such annotation errors is crucial for both diagnosing dataset annotation failures and guiding effective robust learning strategies. Current noisy label learning methods pay limited attention to visualizing noise label distributions, and these approaches often exhibit poor compatibility with noise-free models. Leveraging the relationship between the prediction uncertainty and label noise, we propose a Local Bayesian Framework (LBF) for noisy HSI classification and noise labels awareness. LBF adapts standard CNN, GCN, or Transformer backbones via local Bayesian adaptation (LBA) to evaluate prediction uncertainty and employs an uncertainty-monitoring optimization strategy (U-MOS) for training. Without major architectural changes, LBF delivers accurate uncertainty maps that highlight noisy regions, suppresses overfitting to corrupted labels, and consistently improves classification robustness across four benchmark HSI datasets.
Mingyang Zhang 0002, Ziqi Di, Hao Liu 0123, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Circuits Syst. Video Technol.4
2025 A nonlocal superpatch-based reweighted low-rank representation method for hyperspectral unmixing
Maoguo Gong, Xiangming Jiang, Tao Zhan 0005, Fenlong Jiang
Knowl. Based Syst.5
2025 Multi-scale hierarchical feature fusion network for change detection
Hanhong Zheng, Mingyang Zhang 0002, Maoguo Gong, A. K. Qin 0001, Tongfei Liu, Fenlong Jiang
Pattern Recognit.6
2025 Spatial-Spectral Aggregation Transformer With Diffusion Prior for Hyperspectral Image Super-Resolution
abstract
Constrained by imaging systems, hyperspectral images (HSIs) always have a low spatial resolution. Deep learning-based HSI super-resolution methods have achieved impressive results through learning the nonlinear mapping between low-resolution (LR) and high-resolution (HR) images. However, most of them take the LR image or its upsampled version through bicubic interpolation as input, leading to low-quality features and limited details captured by the network. As a powerful generative model, diffusion model has the ability to learn both contextual semantics and textual details from distinct timesteps, enabling the effective exploration of spatial-spectral distributions in high-dimensional data. In this paper, we propose a novel method that extracts high-quality prior information from original images to assist in super-resolution through pretraining a diffusion model. Specifically, we first train a diffusion model using original HSI patches in a self-supervised manner and then obtain prior features from the pretrained denoising U-Net decoder. To efficiently incorporate the prior features into the super-resolution model, we propose an adaptive fusion module based on spatial and spectral attention mechanisms, which enhances features in both dimensions while preserving the original characteristics. Additionally, to leverage the complementarity of spatial and spectral information, we design a spatial-spectral aggregation Transformer module that incorporates an adaptive interaction module to facilitate information exchange across different dimensions, thereby enhancing the representation capability. Extensive experiments on three public hyperspectral datasets demonstrate that the proposed method achieves excellent super-resolution performance and outperforms the state-of-the-art methods in terms of quantitative quality and visual results.
Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yu Zhou 0051, Fenlong Jiang, Yue Wu 0004
IEEE Trans. Circuits Syst. Video Technol.7
2025 A General Uncertainty-Guided Bayesian Adaptation Framework for Building Change Detection
abstract
Existing deep learning-based building change detection (BCD) methods are often hindered by sample imbalance and imagery noise, which leads to inaccurate predictions, particularly for building edges and minority changed class regions. To overcome these limitations, we propose a novel Uncertainty-guided Bayesian Adaptation (UBA) framework, designed as a plug-and-play module to enhance existing BCD methods. The UBA framework consists of two core components. First, a Local Bayesian Adaptation strategy (LBs) pragmatically adapts the output layer of any BCD network, enabling efficient prediction uncertainty estimation and decomposition. We demonstrate that the decomposed aleatoric and epistemic uncertainty terms semantically highlight building edges and minority changed class regions, respectively. Based on this insight, we propose an Uncertainty-Weighted Optimization Mechanism (U-Wom) that leverages these uncertainty maps to dynamically re-weight the loss function. This mechanism guides the model to focus its learning on these challenging, fine-grained regions. Extensive experiments on several widely-used BCD datasets show that the UBA framework consistently and significantly improves the performance of various state-of-the-art methods.
Ziqi Di, Mingyang Zhang 0002, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.3
2025 Change Masked Modality Alignment Network for Multimodal Change Detection
abstract
Using multimodal remote sensing images for change detection (CD) can significantly improve the feasibility and reliability in challenging environments. However, the differences in imaging mechanisms make multimodal images highly heterogeneous. A key challenge for multimodal CD (MCD) is that the heterogeneity of the modalities and changes in ground objects are intertwined during processing. To address this issue, this article proposes a change masked modality alignment network (CMMAN), which uses a multitask framework consisting of one CD branch and two image modal transformation (IMT) branches. Specifically, to ensure a unified feature space, bi-temporal multimodal images are first input into the same Swin-Transformer-based encoder. The extracted features are then fed simultaneously into the CD branch and separately into the two IMT branches. In the CD branch, the decoder is also designed based on the Swin-Transformer, and a weakly modality-correlated feature enhancement (WMCFE) module is introduced to mitigate the interference of modality heterogeneity on CD. For the two IMT branches, both employ a generative adversarial network (GAN) to transform between modalities, and the distributions of features from different modalities are aligned through simultaneous optimization. Uniquely, the change probability map predicted by the CD branch is utilized to mask the change regions in IMT, further decoupling ground object changes and modal heterogeneity. Experimental results on multiple public datasets demonstrate that the proposed CMMAN significantly improves MCD performance and shows good compatibility and portability with various common backbone networks.
Fenlong Jiang, Husheng Wu, Dan Feng 0002, Yu Zhou 0051, Mingyang Zhang 0002, Maoguo Gong, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.1
2025 D3PM: Dual-Stream Denoising Diffusion Probabilistic Model for Change Detection in Multimodal Remote Sensing Images
abstract
Detecting land cover changes from multi-temporal and multi-modal remote sensing images acquired by different sensors at the same location is a complex yet highly valuable task. Recently, diffusion models, exemplified by the Denoising Diffusion Probabilistic Model (DDPM), have garnered significant attention for their remarkable performance and straightforward architecture. These models excel in image generation, distribution modeling, and feature extraction, making them highly promising for advancing Multimodal Change Detection (MCD). In this paper, we propose a Dual-stream Denoising Diffusion Probabilistic Model (D3PM) to address the challenges of MCD. Specifically, D3PM leverages DDPM to design two distinct processing streams, one for each image modality. The first stream employs an unconditional DDPM, whose denoising encoder-decoder network can achieve robust feature extraction. The second stream employs a conditional DDPM to facilitate modal translation, enabling the extracted features to align with the characteristics of the other modality, thereby improving cross-modality comparability. To further enhance performance, we constructed a CD task branch based on the decoder features of the two DDPMs across multiple denoising time steps. Additionally, we designed a collaborative learning optimization strategy with asynchronous time steps, fostering cross-task knowledge sharing and mutual enhancement while preserving the integrity of individual task learning. Experimental results on multiple public datasets demonstrate the effectiveness and superiority of the proposed D3PM, which achieves efficient modal transformation and alignment, mitigates modal heterogeneity interference, and significantly improves detection performance.
Fenlong Jiang, Xinlong Huo, Mingyang Zhang 0002, Maoguo Gong, Yan Pu, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.1
2025 Adaptive Center-Focused Hybrid Attention Network for Change Detection in Hyperspectral Images
abstract
Hyperspectral images (HSIs) capture extensive spatial and spectral information, facilitating detailed change detection (CD) of complex land covers. However, the high correlation among spectral data can lead to information redundancy, increasing processing dimensions and introducing irrelevant or detrimental data to CD. To address these challenges, we propose an adaptive center-focused hybrid attention network (ACFHAN) for CD in HSIs. This network adaptively emphasizes the spatial regions and spectral channels most pertinent to CD while suppressing irrelevant information. The architecture establishes an end-to-end mapping from the two HSIs to the change results, featuring multiple center-focused hybrid attention blocks (CFHABs). Each CFHAB integrates two different attention modules, including an adaptive spatial–spectral hybrid self-attention (S2HSA) module that dynamically adjusts spatial–spectral feature weights and a center-focused attention (CFA) module that enhances the area most relevant to the center pixel to be classified. Additionally, to tackle the challenges of expensive labeling, we further designed a multiscale superpixel-based data augmentation method which combines traditional unsupervised and supervised methods to provide sufficient low-cost but high-confidence labeled data for CD. Experimental results across various HSI CD datasets validate the effectiveness of our proposed method.
Fenlong Jiang, Shining Zhang, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.1
2025 BARNet: Boundary-Aware Refinement Network for Weakly Supervised Change Detection
abstract
Remote sensing change detection plays a critical role in urban land management and disaster assessment. However, most existing methods rely on expensive and time-consuming pixel-level labels, limiting their practical applicability. Weakly supervised change detection methods, such as only using image-level labels, offer the potential to reduce annotation costs while maintaining robust detection performance. However, this coarse supervisory information often makes it difficult to accurately capture fine-grained details, resulting in poor pixel-level detection accuracy. To overcome these challenges, we propose a novel Boundary-Aware Refinement Network (BARNet) for weakly supervised change detection, which utilizes a two-stage framework that first generates pixel-level pseudo labels via image-level CD activation maps, then subsequently trains a pixel-level CD network using these generated pseudo labels. Specifically, the first stage adopted a teacher-student distillation image-level CD network, which integrated a multi-scale boundary feature attention module, along with activation ambiguity loss and contrastive learning loss as feature separation constraints, to generate high-quality pseudo labels. In the second stage, these pseudo labels are used to provide deep supervision a pixel-level CD network, where the boundary-aware decoupling module further refines boundary information, leading to more precise segmentation of change areas. Extensive experiments on three public datasets demonstrate that BARNet not only achieves state-of-the-art performance in the weakly supervised change detection domain but also shows competitive performance with existing fully supervised methods, significantly reducing annotation costs while maintaining detection accuracy. With its strong performance, BARNet demonstrates great potential for practical applications in scenarios with limited supervision.
Fenlong Jiang, Zikang Zhong, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.1
2025 Meta-Collaborative Learning for Arbitrarily Scaled Hyperspectral Image Super-Resolution
abstract
Deep learning-based methods for hyperspectral image super-resolution (SR) have achieved significant success in recent years. These methods typically consist of feature extraction module (FEM) and upsampling module. However, due to structural limitations of the upsampling module, most current methods focus on training separate models for different scale factors, which ignores the exploration of potential feature interdependence among different scale factors. In response to these challenges, we introduce a novel framework, called “meta-collaborative learning for arbitrarily scaled hyperspectral image super-resolution” (MCArb). Specifically, MCArb integrates a collaborative learning framework with a meta-learning-based 3-D upsampling module (3DMetaUM) and a scale-aware feature adaptation module (SAFAM). It enables training multiple SR tasks at different scale factors within a single network at the same time. This strategy is able not only to process arbitrary-scale-factor SR for hyperspectral images but also to harness the latent feature interdependence among different scales. In this study, we applied the MCArb framework to transform three deep learning-based hyperspectral image SR networks to MCArb methods, resulting in significant performance enhancements across five hyperspectral datasets. These improvements showcase the proposed MCArb framework’s ability to enhance feature extraction efficiency and capitalize on latent interscale correlations. This code is available athttps://github.com/ShuangWu-XDU/MCArb_HSI_SR.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Yu Zhou 0051, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.5
2025 Commonality Feature Representation Learning for Unsupervised Multimodal Change Detection
abstract
The main challenge of multimodal change detection (MCD) is that multimodal bitemporal images (MBIs) cannot be compared directly to identify changes. To overcome this problem, this paper proposes a novel commonality feature representation learning (CFRL) and constructs a CFRL-based unsupervised MCD framework. The CFRL is composed of a Siamese-based encoder and two decoders. First, the Siamese-based encoder can map original MBIs in the same feature space for extracting the representative features of each modality. Then, the two decoders are used to reconstruct the original MBIs by regressing themselves, respectively. Meanwhile, we swap the decoders to reconstruct the pseudo-MBIs to conduct modality alignment. Subsequently, all reconstructed images are input to the Siamese-based encoder again to map them in a same feature space, by which representative features are obtained. On this basis, latent commonality features between MBIs can be extracted by minimizing the distance between these representative features. These latent commonality features are comparable and can be used to identify changes. Notably, the proposed CFRL can be performed simultaneously in two modalities corresponding to MBIs. Therefore, two change magnitude images (CMIs) can be generated simultaneously by measuring the difference between the commonality features of MBIs. Finally, a simple threshold algorithm or a clustering algorithm can be employed to divide CMIs into binary change maps. Extensive experiments on six publicly available MCD datasets show that the proposed CFRL-based framework can achieve superior performance compared with other state-of-the-art approaches.
Tongfei Liu, Mingyang Zhang 0002, Maoguo Gong, Qingfu Zhang 0001, Fenlong Jiang, Hanhong Zheng, Di Lu 0004
IEEE Trans. Image Process.5
2025 SPCNet: Deep Self-Paced Curriculum Network Incorporated With Inductive Bias
abstract
The vulnerability to poor local optimum and the memorization of noise data limit the generalizability and reliability of massively parameterized convolutional neural networks (CNNs) on complex real-world data. Self-paced curriculum learning (SPCL), which models the easy-to-hard learning progression from human beings, is considered as a potential savior. In spite of the fact that numerous SPCL solutions have been explored, it still confronts two main challenges exactly in solving deep networks. By virtue of various designed regularizers, existing weighting schemes independent of the learning objective heavily rely on the prior knowledge. In addition, alternative optimization strategy (AOS) enables the tedious iterative training procedure, thus there is still not an efficient framework that integrates the SPCL paradigm well with networks. This article delivers a novel insight that attention mechanism allows for adaptive enhancement in the contribution of diverse instance information to the gradient propagation. Accordingly, we propose a general-purpose deep SPCL paradigm that incorporates the preferences of implicit regularizer for different samples into the network structure with inductive bias, which in turn is formalized as the self-paced curriculum network (SPCNet). Our proposal allows simultaneous online difficulty estimation, adaptive sample selection, and model updating in an end-to-end manner, which significantly facilitates the collaboration of SPCL to deep networks. Experiments on image classification and scene classification tasks demonstrate that our approach surpasses the state-of-the-art schemes and obtains superior performance.
Yue Zhao 0024, Maoguo Gong, Mingyang Zhang 0002, A. K. Qin 0001, Fenlong Jiang, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.5
2024 Unsupervised Domain Adaptation for Cross-Scene Hyperspectral Image Classification Based on Decoupled Contrastive Learning
abstract
Recent studies have highlighted the effectiveness of deep domain adaptation (DA) techniques in addressing cross-scene hyperspectral image (HSI) classification challenges. However, most of the existing DA methods often prioritize aligning data distributions while overlooking the intrinsic separability between source and target domain data. In this paper, we propose a decoupled contrastive learning based unsupervised domain adaptation (DCLUDA) method for HSI classification. Unlike conventional adversarial DA methods, our method introduces a unique DA loss specifically designed to minimize class confusion in the target domain. This not only simplifies model training but also enhances class discriminability. Moreover, we employ a decoupled contrastive learning strategy on both domains to enhance data separability within each domain. Finally, we propose a sample selection strategy based on confident learning to select high-confidence samples from the target domain for fine-tuning the DA model. Experiments on two cross-scene HSI classification tasks shown that our proposed DCLUDA outperforms several existing DA methods.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Xiangming Jiang, Yu Zhou 0051, Dan Feng 0002
IJCNN4
2024 Adversarial Feature Equilibrium Network for Multimodal Change Detection in Heterogeneous Remote Sensing Images
abstract
Change detection (CD) methods have been crucial in exploring geo-environmental science. With the advancement of remote sensing (RS) technology, multimodal images acquired from different platforms and sensors are widely used for CD tasks. As an emerging task, multimodal CD (MCD) aims to achieve more comprehensive and precise detection of land cover changes through complementary information in multimodal images. However, there are significant differences between modalities, particularly in heterogeneous images. How to deal with modal differences while effectively integrating change information remains a challenge in MCD. In this article, we propose a novel adversarial feature equilibrium network (AFENet), which establishes an additional adversarial optimization to solve the equilibrium problem between modal differences and land cover changes. Our AFENet aligns the features and reduces the modal gap through a multiscale adversarial domain adaptation (MADA) approach. Meanwhile, a divergence-aware contrastive module (DCM) is designed as a regularization term for adversarial optimization. DCM affects the sensitivity of feature extractors by constraining the mutual information between changed and unchanged pixels. In this case, AFENet can maintain the consistency of feature representation while maximizing the discriminability of change targets. The features extracted from AFENet will then be integrated by our multistream feature fusion (MFF) module and utilized to generate change maps. The effectiveness of our approach is demonstrated on two scene-level multimodal RS datasets. Compared with existing methods, our AFENet achieves state-of-the-art (SOTA) performance on both datasets and outperforms the second-best$F1$score by 4.64% and 1.1%, respectively.
Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Tianqi Gao, Fenlong Jiang
IEEE Trans. Geosci. Remote. Sens.6
2024 Spectral Knowledge Transfer for Remote Sensing Change Detection
abstract
Change detection (CD) in multispectral remote sensing (RS) imagery suffers from low spectral resolution which can lead to degraded recognition of change information from land cover objects. Considering that natural hyperspectral imagery (HSI) is much higher in spectral resolution and more accessible, using it to enhance the spectral information of RS multispectral imagery for CD can improve performance. To achieve this, we propose a spectral knowledge transfer (SKT) framework to allow the creation of pseudo-hyperspectral RS images from the available RS multispectral ones without the need for the real pairs of RS multispectral and hyperspectral images, typically required by existing RS spectral enhancement methods. Specifically, an autoencoder is first trained based on the available pairs of natural HSI and its multispectral counterparts and then calibrated via the available RS multispectral images. The finally obtained decoder module is used to generate the pseudo-hyperspectral image from an input RS multispectral image. We further propose a multispectrum collaborative CD (MCCD) framework that leverages both the real multispectral images and the pseudo hyperspectral images generated from them in a collaborative way to achieve performance improvement. Extensive experiments on two large-scale RS CD datasets and eight existing deep learning-based CD methods demonstrate the stronger efficacy of the proposed method.
Hanhong Zheng, Mingyang Zhang 0002, Maoguo Gong, A. K. Qin 0001, Tongfei Liu, Fenlong Jiang
IEEE Trans. Geosci. Remote. Sens.7
2023 Breaking Hardware Boundaries of IoT Devices via Inverse Feature Completion
abstract
Privacy-preserving collaborative learning enables resource-constrained edge devices (e.g., Internet of Things (IoT) devices and smartphones) to build a knowledge-shared model while keeping individual data locally, achieving privacy preservation by designing an effective communication protocol. However, the learning paradigm raises high requirements for aligned input features of models, which is hard to realize in complicated IoT scenarios with various monitoring indicators. In this article, we propose a novel collaborative learning framework that is tolerant of IoT devices with unaligned feature spaces. Local bilevel optimizations for both model parameters and input features are performed iteratively in the training phase, in which the internal correlations of local sensor data provide additional guidance for the feature inference and completion. The scheme breaks hardware boundaries among various IoT devices in collaboration with the assistance of model inversion inference, which gains a new perspective on the utilization of model confidentiality and requires minimal modifications to the existing collaborative learning process. The framework achieves significant improvement compared with state-of-the-art methods, as we demonstrate through extensive simulations on real-world data sets.
Yuan Gao 0019, Yew-Soon Ong, Maoguo Gong, Fenlong Jiang, Yuanqiao Zhang, Shanfeng Wang
IEEE Internet Things J.4
2023 Context-content collaborative network for building extraction from high-resolution imagery
Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Qingfu Zhang 0001, Di Lu 0004, Hanhong Zheng, Fenlong Jiang
Knowl. Based Syst.7
2023 Self-Supervised Global-Local Contrastive Learning for Fine-Grained Change Detection in VHR Images
abstract
Self-supervised contrastive learning (CL) can learn high-quality feature representations that are beneficial to downstream tasks without labeled data. However, most CL methods are for image-level tasks. For the fine-grained change detection (FCD) tasks, such as change or change trend detection of some specific ground objects, it is usually necessary to perform pixel-level discriminative analysis. Therefore, feature representations learned by image-level CL may have limited effects on FCD. To address this problem, we propose a self-supervised global–local contrastive learning (GLCL) framework, which extends the instance discrimination task to the pixel level. GLCL follows the current mainstream CL paradigm and consists of four parts, including data augmentation to generate different views of the input, an encoder network for feature extraction, a global CL head, and a local CL head to perform image-level and pixel-level instance discrimination tasks, respectively. Through GLCL, features belonging to different perspectives of the same instance will be pulled closer, while features of different instances will be alienated, which can enhance the discriminativeness of feature representations from both global and local perspectives, thereby facilitating downstream FCD tasks. In addition, GLCL makes a targeted structural adaptation to FCD, i.e., the encoder network is undertaken by the common backbone networks of FCD, which can accelerate the deployment on downstream FCD tasks. Experimental results on several real datasets show that compared with other parameter initialization methods, the FCD models pretrained by GLCL can obtain better detection performance.
Fenlong Jiang, Maoguo Gong, Hanhong Zheng, Tongfei Liu, Mingyang Zhang 0002, Jia Liu 0020
IEEE Trans. Geosci. Remote. Sens.1
2022 Landslide Inventory Mapping Method Based on Adaptive Histogram-Mean Distance With Bitemporal VHR Aerial Images
abstract
Landslide inventory mapping (LIM) on the basis of change detection techniques has potential significance for landslide disaster analysis. In this letter, a novel LIM approach based on the adaptive histogram-mean distance (AHMD) is proposed, which adaptively considers spatial contextual information of different landslide regions to improve the detection performance. First, to adapt the shape, size, and distribution of various landslides, an adaptive region around a pixel is extracted by a novel adaptive region extension algorithm without parameter setting. Second, the pixels within the adaptive region are taken to construct the spectral frequency histograms, and then, the adaptive histogram mean (AHM) is developed as the feature of a histogram. Third, the AHMD is defined based on the bin-to-bin (B2B) distance to measure change magnitude between the pairwise AHMs. Finally, LIM can be obtained by a supervised threshold method called double-window flexible pace search (DFPS). Experimental results tested on two real datasets with a very high spatial resolution (VHR) demonstrate the outperformance of the proposed AHMD approach with seven comparative methods.
Tongfei Liu, Maoguo Gong, Fenlong Jiang, Yuanqiao Zhang, Hao Li 0009
IEEE Geosci. Remote. Sens. Lett.3
2022 HFA-Net: High frequency attention siamese network for building change detection in VHR remote sensing images
Hanhong Zheng, Maoguo Gong, Tongfei Liu, Fenlong Jiang, Tao Zhan 0005, Di Lu 0004, Mingyang Zhang 0002
Pattern Recognit.4
2022 Deep Image Inpainting With Enhanced Normalization and Contextual Attention
abstract
Deep learning-based image inpainting has been widely studied, leading to great success. However, many methods adopt convolution and normalization operations, which will bring up some issues to affect the performance. The vanilla normalization cannot distinguish the pixels in corrupted regions from the other valid pixels, resulting in the mean and variance shifts. In addition, the limited receptive field of convolution makes it unable to capture long-range valid information directly. In order to tackle these challenges, we propose a novel deep generative model for image inpainting with two key modules, namely, the channel and spatially adaptive batch normalization (CSA-BN) module, and the selective latent-space-mapping-based contextual attention (SLSM-CA) layer. We replace the vanilla normalization with the CSA-BN module. By channel and spatially adaptive denormalization, the CSA-BN module can mitigate the spatial mean and variance shifts in each channel in a targeted way. In addition, we also integrate the SLSM-CA layer into our model to capture the long-range correlations explicitly. By introducing dual-branch attention and a feature selection module, the SLSM-CA layer can selectively utilize the multi-scale background information to improve prediction quality. What’s more, it introduces the latent spaces to achieve the low-rank approximations of attention matrices and to reduce computational costs. Extensive quantitative and qualitative evaluations demonstrate the superiority of the proposed method compared with state-of-the-art methods.
Jia Liu 0020, Maoguo Gong, Zedong Tang, A. K. Qin 0001, Hao Li 0009, Fenlong Jiang
IEEE Trans. Circuits Syst. Video Technol.6
2022 A Spectral and Spatial Attention Network for Change Detection in Hyperspectral Images
abstract
Hyperspectral images (HSIs) contain rich spectral signatures that reveal more image details and, thus, enable the detection of less noticeable changes on the ground. However, HSI-based change detection (CD) is susceptible to a large amount of irrelevant or noisy spectral and spatial information due to massive spectral bands. To address these issues, we propose a novel spectral and spatial attention network (S2AN) for HSI-based CD, which is capable to suppress CD-irrelevant spectral and spatial information via adaptive spectral and spatial attention mechanisms. S2AN takes as input the image patch from the difference map between two HSIs and outputs the status of change for the patch. Specifically, S2AN is composed of several repeated attention blocks, each of which contains the spectral attention (SpeA) module for directly calculating the attention score for each input channel, the Gaussian spatial attention (GSpaA) module that first constructs an adaptive Gaussian distribution and then samples it to derive the attention scores for each spatial position, and the convolutional feature extraction (CFE) module for extracting features from the attention-weighted input. It is worth mentioning that, in addition to the advantage of the attention, GSpaA also reduces the sensitivity of patch size for patch-based methods. To effectively train S2AN when facing insufficient labeled data, a semisupervised strategy that combines supervised and unsupervised methods to augment labeled training data is proposed. Experiments on several HSI datasets in comparison to existing methods show the superiority of S2AN.
Maoguo Gong, Fenlong Jiang, A. K. Qin 0001, Tongfei Liu, Tao Zhan 0005, Di Lu 0004, Hanhong Zheng, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Building Change Detection for VHR Remote Sensing Images via Local-Global Pyramid Network and Cross-Task Transfer Learning Strategy
abstract
Building change detection (BCD) for very-high-spatial-resolution (VHR) remote sensing images is very important and challenging in the field of remote sensing, as the building is one of the most significant and valuable man-made ground targets. This article proposes a local–global pyramid network (LGPNet) that combines a local feature pyramid module (LFPM) and a global spatial pyramid module (GSPM) for various building feature extraction. The LFPM is constructed using the convolutional kernel with three different pyramid scales, and then, the local pyramid features are obtained by adding features of each scale. In the GSPM, the global spatial pyramid features are extracted by adaptive average pooling to acquire global contextual information from different fields of view on deep features. The LFPM and the GSPM work in a parallel and complementary manner to capture discriminative features of various buildings. In addition to the LFPM and the GSPM, the proposed LGPNet also employs two general attention mechanisms, i.e., the position attention module and the channel attention module, which can select and emphasize adaptively some building features with high semantic responses. Besides, in order to mitigate the influence of other ground targets to a certain extent, a cross-task transfer learning strategy is introduced to make the LGPNet focus on the building, which significantly improves the performance of our method. Extensive experiments on two public available BCD datasets show that the proposed LGPNet can achieve significant improvement compared with eight other state-of-the-art methods. The source code and the pretrained model will be released athttps://github.com/TongfeiLiu/LGPNet.
Tongfei Liu, Maoguo Gong, Di Lu 0004, Qingfu Zhang 0001, Hanhong Zheng, Fenlong Jiang, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.6
2022 Self-Supervised Monocular Depth Estimation With Multiscale Perception
abstract
Extracting 3D information from a single optical image is very attractive. Recently emerging self-supervised methods can learn depth representations without using ground truth depth maps as training data by transforming the depth prediction task into an image synthesis task. However, existing methods rely on a differentiable bilinear sampler for image synthesis, which results in each pixel in a synthetic image being derived from only four pixels in the source image and causes each pixel in the depth map to perceive only a few pixels in the source image. In addition, when calculating the photometric error between a synthetic image and its corresponding target image, existing methods only consider the photometric error within a small neighborhood of each single pixel and therefore ignore correlations between larger areas, which causes the model to tend to fall into the local optima for small patches. In order to extend the perceptual area of the depth map over the source image, we propose a novel multi-scale method that downsamples the predicted depth map and performs image synthesis at different resolutions, which enables each pixel in the depth map to perceive more pixels in the source image and improves the performance of the model. As for the locality of photometric error, we propose a structural similarity (SSIM) pyramid loss to allow the model to sense the difference between images in multiple areas of different sizes. Experimental results show that our method achieves superior performance on both outdoor and indoor benchmarks.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Mingyang Zhang 0002, Fenlong Jiang, Hongyu Zhao 0007
IEEE Trans. Image Process.5
2021 SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate Curvature
abstract
The bottleneck of computation burden limits the widespread use of the 2nd order optimization algorithms for training deep neural networks. In this paper, we present a computationally efficient approximation for natural gradient descent, named Swift Kronecker-Factored Approximate Curvature (SKFAC), which combines Kronecker factorization and a fast low-rank matrix inversion technique. Our research aims at both fully connected and convolutional layers. For the fully connected layers, by utilizing the low-rank property of Kronecker factors of Fisher information matrix, our method only requires inverting a small matrix to approximate the curvature with desirable accuracy. For convolutional layers, we propose a way with two strategies to save computational efforts without affecting the empirical performance by reducing across the spatial dimension or receptive fields of feature maps. Specifically, we propose two effective dimension reduction methods for this purpose: Spatial Subsampling and Reduce Sum. Experimental results of training several deep neural networks on Cifar-10 and ImageNet-1k datasets demonstrate that SKFAC can capture the main curvature and yield comparative performance to K-FAC. The proposed method bridges the wall-clock time gap between the 1st and 2nd order algorithms.
Zedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 0009, Yue Wu 0004, Fan Yu 0004, Zidong Wang 0010, Min Wang 0037
CVPR2
2020 A Semisupervised GAN-Based Multiple Change Detection Framework in Multi-Spectral Images
abstract
Effectively highlighting multiple changes in the earth surface from multi-temporal remote sensing images is a meaningful but challenging task. In order to reduce costs and ensure the performance, it is advisable to employ a semisupervised strategy to achieve this goal. As a discriminative joint classification task, semisupervised change detection aims to extract useful and discriminative features from a large amount of unlabeled data in addition to limited labeled samples. The discriminator of a well-trained generative adversarial network (GAN) is just right for this. Therefore, in this letter, we proposed a semisupervised GAN-based multiple change detection framework for multi-spectral images. First, the GAN is trained by all data without any prior information. Then, we combine two identical trained discriminators to construct a dual-pipeline joint classifier. Finally, the classifier is fine-tuned by a very small amount of labeled data to detect multiple changes. The superior performance of the proposed model over both real multi-spectral data sets demonstrates its robustness and effectiveness.
Fenlong Jiang, Maoguo Gong, Tao Zhan 0005, Xiaolong Fan
IEEE Geosci. Remote. Sens. Lett.1
2020 Structured self-attention architecture for graph-level representation learning
Xiaolong Fan, Maoguo Gong, Yu Xie 0009, Fenlong Jiang, Hao Li 0009
Pattern Recognit.4
2019 Multipopulation Optimization for Multitask Optimization
abstract
Currently, the most of multitask evolutionary algorithms views multiple tasks as factors influencing the evolution of individuals. However, this consideration causes difficulty to assign fitness to individuals, because an individual which performs well on one task can have a bad performance on another task. To avoid this difficulty, this paper proposes a novel multipopulation technique for multitask optimization (MPMTO). The novelty of MPMTO is that it can solve the multiple tasks via a simple and straightforward method by corresponding each population to a task. By this way, the fitness assignment issue can be addressed by just assigning the objective value of the corresponding task to individuals. MPMTO is a general technique so that existing population-based optimization algorithms can be used in each population. This paper uses differential evolutionary algorithm in each population and develops a multipopulation multitask differential evolutionary optimization (mMTDE) based on the proposed multipopulation technique. mMTDE features that each population can use the other populations as the additional knowledge source to create an overlapping population, allowing the populations share information. By this way, the population can improve the efficacy and accuracy of solving multiple tasks. Moreover, the successful inter-task offspring can immigrate back to the corresponding population to fully utilize the inter-task knowledge. We have compared the proposed method with other state-of-the-art methods on benchmark multitask problems. The experimental results show the superiority of the proposed method which could utilizes efficiently the searching knowledge of multiple tasks.
Zedong Tang, Maoguo Gong, Fenlong Jiang, Hao Li 0009, Yue Wu 0004
CEC3