Mingyang Zhang 0002

dblp:76/4874-2 · DBLP profile ↗
← Back
71ranked-venue papers
11as first author
57since 2021 · last 2026
0000-0002-9768-516XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 4 first-author · 25 since 2021Artificial intelligence and machine learning · 31 · 5 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View Consistency
abstract
Human Novel View Synthesis (HNVS) aims to synthesize photorealistic human images from novel viewpoints given observations from known views. Despite significant advances achieved by existing methods such as NeRF, diffusion models, and 3DGS, they still face substantial challenges in achieving stable modeling from a single image. In this paper, we introduce Dual-Constraint Human Gaussian Splatting (DcSplat), a novel, simple, and efficient 3D Gaussian-based framework for single-view 3D human reconstruction. To address occlusion-induced texture missing and depth ambiguities, we introduce two key components: a Latent Multi-View Consistency Constraint Mechanism and a Geometric Constraint Module. The former employs a Latent-space Appearance Transformer (LatentFormer) to learn semantically coherent, view-consistent appearance priors via SMPL-guided pseudo-view fusion. The latter refines noisy SMPL-based depth through a U-Net-like structure conditioned on latent appearance features. These two modules are jointly optimized to generate high-quality Gaussian parameters in a unified latent space. Extensive experiments demonstrate that DcSplat outperforms existing SOTA methods in both geometry and texture quality, while achieving fast inference and lower computational cost.
Tengfei Xiao, Yue Wu 0004, Yongzhe Yuan, Can Qin, Hao Li 0009, Mingyang Zhang 0002
AAAI7
2026 Federated Cross-Device Heterogeneous Few-Shot Adaptation for Edge IoT Systems
abstract
The deployment of federated learning in real-world IoT ecosystems presents intrinsic challenges stemming from hardware asymmetry and sample scarcity, the existing related approaches generally homogenize model architectures and assume abundant labeled data, resulting in an inability to achieve fast generalization on devices with varying computational capabilities and dynamic task conditions. To address the aforementioned challenges, we propose a novel federated cross-device heterogeneous few-shot adaptation (Fed-CHFSA) method for IoT systems. In Fed-CHFSA, collaborating with other devices, each edge device obtains a personalized model that can not only adapt well to the category distribution of respective local data but also recognize unseen categories without data leakage. Specifically, we designed a fine-grained personalized aggregation (FPA) module and an information entropy-driven adaptive feature constraint (EAFC) module for the devices possessing a small amount of labeled data in the model aggregation and training phases of Fed-CHFSA, respectively. In each round of global communication, the edge device performs a certain epoch of personalized training locally under the normalization of EAFC in the feature space. Subsequently, the central server follows the FPA to finely aggregate the received model updates parameter-wise, and redistribute the updated global model to participating devices. After multiple rounds of global communication, every edge device acquires an optimal model more adaptable to local data and more generalized to unseen categories. Compared with existing FL and PFL algorithms on three benchmark few-shot learning (FSL) datasets, the proposed Fed-CHFSA framework achieves the best performance. The effectiveness of FPA and EAFC is also demonstrated by extensive ablation experiments.
Jianzhao Li, Yiting Liu 0004, Boya Deng, Maoguo Gong, Zedong Tang, Mingyang Zhang 0002, Yourun Zhang, Zhuping Hu
IEEE Internet Things J.6
2026 A spatial-spectral-frequency interactive network for multimodal remote sensing classification
abstract
Deep learning-based methods have achieved significant success in remote sensing Earth observation data analysis. Numerous feature fusion techniques address multimodal remote sensing image classification by integrating global and local features. However, these techniques often struggle to extract structural and detail features from heterogeneous and redundant multimodal images, particularly in label-scarce scenarios. With the goal of introducing frequency domain learning to model key and sparse detail features, this paper introduces the spatial–spectral-frequency interaction network (S 2 Fin), which integrates pairwise fusion modules across the spatial, spectral, and frequency domains. Specifically, we propose a high-frequency sparse enhancement transformer to refine spectral signatures by adaptively enhancing discriminative high-frequency components. For spatial-frequency interaction, we present a depth-wise strategy: the adaptive frequency channel module fuses low-frequency structural information with enhanced details in shallow layers, while the high-frequency resonance mask amplifies modality-consistent regions in deep layers using phase similarity. In addition, a spatial–spectral attention fusion module bridges the gap between spectral and spatial branches at intermediate depths. Extensive experiments on four benchmark datasets demonstrate that S 2 Fin exhibits good robustness and generalization, and its performance significantly outperforms state-of-the-art methods in few-sample settings. The code is available at https://github.com/HaoLiu-XDU/SSFin .
Hao Liu 0123, Yunhao Gao, Wei Li 0032, Mingyang Zhang 0002, Maoguo Gong, Lorenzo Bruzzone
Pattern Recognit.4
2026 Uncertainty-Aware Local Bayesian Framework for Hyperspectral Image Classification With Noisy Labels
abstract
Various deep learning-based methods have greatly improved hyperspectral image (HSI) classification performance, but these models are sensitive to noisy training labels. Human annotation on remote sensing images inevitably introduced label noise, which degrades the model prediction confidence. Understanding the spatial characteristics and distribution of such annotation errors is crucial for both diagnosing dataset annotation failures and guiding effective robust learning strategies. Current noisy label learning methods pay limited attention to visualizing noise label distributions, and these approaches often exhibit poor compatibility with noise-free models. Leveraging the relationship between the prediction uncertainty and label noise, we propose a Local Bayesian Framework (LBF) for noisy HSI classification and noise labels awareness. LBF adapts standard CNN, GCN, or Transformer backbones via local Bayesian adaptation (LBA) to evaluate prediction uncertainty and employs an uncertainty-monitoring optimization strategy (U-MOS) for training. Without major architectural changes, LBF delivers accurate uncertainty maps that highlight noisy regions, suppresses overfitting to corrupted labels, and consistently improves classification robustness across four benchmark HSI datasets.
Mingyang Zhang 0002, Ziqi Di, Hao Liu 0123, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Circuits Syst. Video Technol.1
2025 MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency
abstract
3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clouds is very expensive and time-consuming. In addition, these works lack effective self-supervised signals, and existing self-supervised signals often fail to capture sufficiently rich change information. To solve this problem, we assume that the powerful representation of 3D objects should model the consistency information of unchanged regions and distinguish different objects. Based on this assumption, we propose a new unsupervised framework called MUCD to learn change information of multi-temporal point clouds through bidirectional optimization of change segmentor and feature extractor. The training of network is divided into two stages. We first design a foreknowledge point contrastive loss based on the characteristics of the 3DCD task to initialize the feature extractor, and then propose a masked consistency loss to further learn the shared geometric information of unchanged regions in the multi-temporal point clouds, utilizing it as a free and powerful supervised signal to train a change segmentor. In the inference stage, only the segmentor is used to take multi-temporal point clouds as input and produce change segmentation result. Extensive experiments are conducted on SLPCCD and Urb3DCD, two real-world datasets of streets and urban buildings, to verify that our proposed unsupervised method is highly competitive and even outperforms supervised methods in scenes where semantic information changes occur, exhibiting better performance in generalization ability and robustness.
Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001, Qiguang Miao
AAAI6
2025 AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors
abstract
When the current physical adversarial patches cannot deceive thermal infrared detectors, the existing techniques implement adversarial attacks from scratch, such as digital patch generation, material production, and physical deployment. Besides, it is difficult to finely regulate infrared radiation. To address these issues, this paper designs an adversarial thermal display (AdvDisplay ) by assembling thermoelectric coolers (TECs) as an array. Specifically, to reduce the gap between patches in the physical and digital worlds and decrease the power of AdvDisplay device, heat transfer loss and electric power loss are designed to guide the patch optimization. In addition, a precise temperature control scheme for AdvDisplay is proposed based on proportional-integral-derivative (PID) control. Due to the accurate temperature regulation and the reusability of AdvDisplay , our method is able to improve the attack success rate and the efficiency of physical deployments. Extensive experimental results indicate that the proposed method possesses superior adversarial effectiveness compared to other methods and demonstrates strong robustness in physical attacks.
Hao Li 0009, Fanggao Wan, Yue Wu 0004, Mingyang Zhang 0002, Maoguo Gong
AAAI5
2025 PointTruss: K-Truss for Point Cloud Registration
abstract
Point cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both robustness and efficiency. While the k-core reduces computational complexity, which only considers node degree and ignores higher-order topological structures such as triangles, limiting its effectiveness in complex scenarios. To overcome these limitations, we introduce the $k$-truss from graph theory into point cloud registration, leveraging triangle support as a constraint for inlier selection. We further propose a consensus voting-based low-scale sampling strategy to efficiently extract the structural skeleton of the point cloud prior to $k$-truss decomposition. Additionally, we design a spatial distribution score that balances coverage and uniformity of inliers, preventing selections that concentrate on sparse local clusters. Extensive experiments on KITTI, 3DMatch, and 3DLoMatch demonstrate that our method consistently outperforms both traditional and learning-based approaches in various indoor and outdoor scenarios, achieving state-of-the-art results.
Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Qiguang Miao, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001
NeurIPS7
2025 Multi-scale hierarchical feature fusion network for change detection
Hanhong Zheng, Mingyang Zhang 0002, Maoguo Gong, A. K. Qin 0001, Tongfei Liu, Fenlong Jiang
Pattern Recognit.2
2025 Equivariance-Based Markov Decision Process for Unsupervised Point Cloud Registration
abstract
Unsupervised point cloud registration is crucial in 3D computer vision. However, most unsupervised methods struggle to construct effective optimization objectives and reliable unsupervised signals to enhance the performance of the model. To address these issues, with the observation of the significant alignment between the registration process and the Markov Decision Process (MDP), we model point cloud registration as MDP, which can provide more reliable unsupervised signals through the reward. We propose a colored noise based cross-entropy method, which introduces colored noise into sampling process, regulating the power spectral density of the action sequence and expanding the search space, improving the registration effect. Particularly, to strengthen constraints on MDP and training in the transformation space, we utilize equivariance theory to construct transformation equivariant constraint as a new optimization objective and derive equivariant constraint solutions for optimization, providing more reliable unsupervised signals. Extensive experiments demonstrate the superior performance of our method on benchmark datasets.
Yue Wu 0004, Jiayi Lei, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Mingyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.8
2025 Spatial-Spectral Aggregation Transformer With Diffusion Prior for Hyperspectral Image Super-Resolution
abstract
Constrained by imaging systems, hyperspectral images (HSIs) always have a low spatial resolution. Deep learning-based HSI super-resolution methods have achieved impressive results through learning the nonlinear mapping between low-resolution (LR) and high-resolution (HR) images. However, most of them take the LR image or its upsampled version through bicubic interpolation as input, leading to low-quality features and limited details captured by the network. As a powerful generative model, diffusion model has the ability to learn both contextual semantics and textual details from distinct timesteps, enabling the effective exploration of spatial-spectral distributions in high-dimensional data. In this paper, we propose a novel method that extracts high-quality prior information from original images to assist in super-resolution through pretraining a diffusion model. Specifically, we first train a diffusion model using original HSI patches in a self-supervised manner and then obtain prior features from the pretrained denoising U-Net decoder. To efficiently incorporate the prior features into the super-resolution model, we propose an adaptive fusion module based on spatial and spectral attention mechanisms, which enhances features in both dimensions while preserving the original characteristics. Additionally, to leverage the complementarity of spatial and spectral information, we design a spatial-spectral aggregation Transformer module that incorporates an adaptive interaction module to facilitate information exchange across different dimensions, thereby enhancing the representation capability. Extensive experiments on three public hyperspectral datasets demonstrate that the proposed method achieves excellent super-resolution performance and outperforms the state-of-the-art methods in terms of quantitative quality and visual results.
Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yu Zhou 0051, Fenlong Jiang, Yue Wu 0004
IEEE Trans. Circuits Syst. Video Technol.1
2025 A General Uncertainty-Guided Bayesian Adaptation Framework for Building Change Detection
abstract
Existing deep learning-based building change detection (BCD) methods are often hindered by sample imbalance and imagery noise, which leads to inaccurate predictions, particularly for building edges and minority changed class regions. To overcome these limitations, we propose a novel Uncertainty-guided Bayesian Adaptation (UBA) framework, designed as a plug-and-play module to enhance existing BCD methods. The UBA framework consists of two core components. First, a Local Bayesian Adaptation strategy (LBs) pragmatically adapts the output layer of any BCD network, enabling efficient prediction uncertainty estimation and decomposition. We demonstrate that the decomposed aleatoric and epistemic uncertainty terms semantically highlight building edges and minority changed class regions, respectively. Based on this insight, we propose an Uncertainty-Weighted Optimization Mechanism (U-Wom) that leverages these uncertainty maps to dynamically re-weight the loss function. This mechanism guides the model to focus its learning on these challenging, fine-grained regions. Extensive experiments on several widely-used BCD datasets show that the UBA framework consistently and significantly improves the performance of various state-of-the-art methods.
Ziqi Di, Mingyang Zhang 0002, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.2
2025 Scale-Aware Pruning Framework for Remote Sensing Object Detection via Multifeature Representation
abstract
With the rapid advancements in computer vision, high-resolution remote sensing imagery has become a crucial data source for object detection. Nevertheless, effectively utilizing limited computational resources and reducing the burden on satellite edge devices remains a significant challenge. To effectively reduce model complexity while maintaining its representational capacity, this article proposes a scale-aware pruning framework (SAPF) to enhance remote sensing object detection ability. First, this article classifies the convolutional layers in object detection models into two categories: layers with a single-scale feature representation and layers with a multiscale feature representation. For convolutional layers with single-scale features, we utilize singular value decomposition (SVD) to quantify feature importance and assess filter redundancy to enhance model efficiency. By removing less critical filters, this pruning criteria aims to reduce the model size and computational load without compromising performance. However, convolutional layers with multiscale features are crucial for optimizing feature extraction and balancing information capture across various scales. To address this, this article evaluates the similarity between convolutional layers with different scales to determine the contribution of various scale features in multiscale fusion. Surprisingly, the SAPF can reduce the FLOPs and parameters, as well as ensure the representational ability obviously when the YOLO v5s and Faster-RCNN are adopted to classify the NWPU VHR-10, RSOD, and SIMD datasets. This means we can save the training computation resources for the model. Additionally, SAPF can significantly improve the efficiency of the model in object detection to ensure its real-time performance.
Zhuping Hu, Maoguo Gong, Yue Zhao 0024, Mingyang Zhang 0002, Yiheng Lu, Jianzhao Li, Yan Pu, Zhao Wang 0011
IEEE Trans. Geosci. Remote. Sens.4
2025 Change Masked Modality Alignment Network for Multimodal Change Detection
abstract
Using multimodal remote sensing images for change detection (CD) can significantly improve the feasibility and reliability in challenging environments. However, the differences in imaging mechanisms make multimodal images highly heterogeneous. A key challenge for multimodal CD (MCD) is that the heterogeneity of the modalities and changes in ground objects are intertwined during processing. To address this issue, this article proposes a change masked modality alignment network (CMMAN), which uses a multitask framework consisting of one CD branch and two image modal transformation (IMT) branches. Specifically, to ensure a unified feature space, bi-temporal multimodal images are first input into the same Swin-Transformer-based encoder. The extracted features are then fed simultaneously into the CD branch and separately into the two IMT branches. In the CD branch, the decoder is also designed based on the Swin-Transformer, and a weakly modality-correlated feature enhancement (WMCFE) module is introduced to mitigate the interference of modality heterogeneity on CD. For the two IMT branches, both employ a generative adversarial network (GAN) to transform between modalities, and the distributions of features from different modalities are aligned through simultaneous optimization. Uniquely, the change probability map predicted by the CD branch is utilized to mask the change regions in IMT, further decoupling ground object changes and modal heterogeneity. Experimental results on multiple public datasets demonstrate that the proposed CMMAN significantly improves MCD performance and shows good compatibility and portability with various common backbone networks.
Fenlong Jiang, Husheng Wu, Dan Feng 0002, Yu Zhou 0051, Mingyang Zhang 0002, Maoguo Gong, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.6
2025 D3PM: Dual-Stream Denoising Diffusion Probabilistic Model for Change Detection in Multimodal Remote Sensing Images
abstract
Detecting land cover changes from multi-temporal and multi-modal remote sensing images acquired by different sensors at the same location is a complex yet highly valuable task. Recently, diffusion models, exemplified by the Denoising Diffusion Probabilistic Model (DDPM), have garnered significant attention for their remarkable performance and straightforward architecture. These models excel in image generation, distribution modeling, and feature extraction, making them highly promising for advancing Multimodal Change Detection (MCD). In this paper, we propose a Dual-stream Denoising Diffusion Probabilistic Model (D3PM) to address the challenges of MCD. Specifically, D3PM leverages DDPM to design two distinct processing streams, one for each image modality. The first stream employs an unconditional DDPM, whose denoising encoder-decoder network can achieve robust feature extraction. The second stream employs a conditional DDPM to facilitate modal translation, enabling the extracted features to align with the characteristics of the other modality, thereby improving cross-modality comparability. To further enhance performance, we constructed a CD task branch based on the decoder features of the two DDPMs across multiple denoising time steps. Additionally, we designed a collaborative learning optimization strategy with asynchronous time steps, fostering cross-task knowledge sharing and mutual enhancement while preserving the integrity of individual task learning. Experimental results on multiple public datasets demonstrate the effectiveness and superiority of the proposed D3PM, which achieves efficient modal transformation and alignment, mitigates modal heterogeneity interference, and significantly improves detection performance.
Fenlong Jiang, Xinlong Huo, Mingyang Zhang 0002, Maoguo Gong, Yan Pu, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.3
2025 Adaptive Center-Focused Hybrid Attention Network for Change Detection in Hyperspectral Images
abstract
Hyperspectral images (HSIs) capture extensive spatial and spectral information, facilitating detailed change detection (CD) of complex land covers. However, the high correlation among spectral data can lead to information redundancy, increasing processing dimensions and introducing irrelevant or detrimental data to CD. To address these challenges, we propose an adaptive center-focused hybrid attention network (ACFHAN) for CD in HSIs. This network adaptively emphasizes the spatial regions and spectral channels most pertinent to CD while suppressing irrelevant information. The architecture establishes an end-to-end mapping from the two HSIs to the change results, featuring multiple center-focused hybrid attention blocks (CFHABs). Each CFHAB integrates two different attention modules, including an adaptive spatial–spectral hybrid self-attention (S2HSA) module that dynamically adjusts spatial–spectral feature weights and a center-focused attention (CFA) module that enhances the area most relevant to the center pixel to be classified. Additionally, to tackle the challenges of expensive labeling, we further designed a multiscale superpixel-based data augmentation method which combines traditional unsupervised and supervised methods to provide sufficient low-cost but high-confidence labeled data for CD. Experimental results across various HSI CD datasets validate the effectiveness of our proposed method.
Fenlong Jiang, Shining Zhang, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.3
2025 BARNet: Boundary-Aware Refinement Network for Weakly Supervised Change Detection
abstract
Remote sensing change detection plays a critical role in urban land management and disaster assessment. However, most existing methods rely on expensive and time-consuming pixel-level labels, limiting their practical applicability. Weakly supervised change detection methods, such as only using image-level labels, offer the potential to reduce annotation costs while maintaining robust detection performance. However, this coarse supervisory information often makes it difficult to accurately capture fine-grained details, resulting in poor pixel-level detection accuracy. To overcome these challenges, we propose a novel Boundary-Aware Refinement Network (BARNet) for weakly supervised change detection, which utilizes a two-stage framework that first generates pixel-level pseudo labels via image-level CD activation maps, then subsequently trains a pixel-level CD network using these generated pseudo labels. Specifically, the first stage adopted a teacher-student distillation image-level CD network, which integrated a multi-scale boundary feature attention module, along with activation ambiguity loss and contrastive learning loss as feature separation constraints, to generate high-quality pseudo labels. In the second stage, these pseudo labels are used to provide deep supervision a pixel-level CD network, where the boundary-aware decoupling module further refines boundary information, leading to more precise segmentation of change areas. Extensive experiments on three public datasets demonstrate that BARNet not only achieves state-of-the-art performance in the weakly supervised change detection domain but also shows competitive performance with existing fully supervised methods, significantly reducing annotation costs while maintaining detection accuracy. With its strong performance, BARNet demonstrates great potential for practical applications in scenarios with limited supervision.
Fenlong Jiang, Zikang Zhong, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.3
2025 Layer-Interaction Adaptive Pruning for Remote Sensing Scene Classification
abstract
The advancement of Convolutional Neural Networks (CNNs) has enhanced remote sensing scene classification on satellites. However, the increased computational complexity of CNNs will impose a substantial burden on satellite hardware, thereby hindering the practical application. Although various pruning techniques have been developed to reduce the scale of CNNs by assessing the importance of model parameters, these weight-based methods often lead to a degradation in model performance when the original model fails to provide meaningful parameters at under-trained conditions. In this paper, we introduce a novel Layer-interaction Adaptive Pruning (LiAP) method designed to streamline under-trained models. Unlike conventional approaches, LiAP evaluates the importance of neurons based on the distribution of eigenvalues rather than individual parameters. Specifically, this is achieved by projecting the weight matrix of each convolutional layer into the eigenspace, where the eigenvalues are utilized to assess the redundancy of filters. Then each space will be assigned a fined score based on the eigenvalue distribution to provide a robust measure of filter importance. Compared to weight-based pruning methods, LiAP maintains consistent evaluation accuracy between well-trained and under-trained models, as the eigenspace is inherently robust to variations in the weight parameter space. We conducted extensive experiments on VGG-16 and ResNet-50 architectures using datasets such as AID, NWPU-RESISC45, PatternNet, and WHU-RS19 over various data partitions. Notably, our method achieved State-of-the-Art (SOTA) reductions in both FLOPs and parameters for VGG-16 across the aforementioned datasets. These results underscore the efficacy of LiAP in enhancing the efficiency and applicability of CNNs in satellite-based remote sensing tasks.
Yiheng Lu, Zhuping Hu, Wei Zhao 0019, Ziyu Guan, Maoguo Gong, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.6
2025 Collaborative Frequency-Aware Transformer for Unsupervised Multimodal Change Detection in Heterogeneous Remote Sensing Images
abstract
Multimodal change detection (MCD), as an emerging task, aims at recognizing change regions from bi-temporal remote sensing images (RSI) of different modalities. Inspired by the success of the self-attention mechanism in transformer, attempts have been made to solve MCD through the transformer variants. However, transformer-based network optimization requires high-quality training samples. In addition, due to the significant differences in the data distribution, semantic information, and feature representation of multimodal data, transformer-based methods have obvious deficiencies in local feature representation and spatial consistency, especially when dealing with heterogeneous images. To address the above challenges, we propose a collaborative frequency-aware transformer for MCD (CFAT-MCD). As an unsupervised framework, CFAT-MCD is capable of learning more fine-grained patterns of land cover change through a few pseudo-labels. The CFAT is designed to enhance spatial consistency and align the features on a multi-scale basis, which can effectively mitigate the effects of modal differences. In addition, we propose a window-based spatial-frequency collaborative representation (SFCR) module to introduce frequency information into the spatial domain and improve the discriminability of spatial features. Extensive experiments on public datasets and quantitative analyses have validated the superior detection performance of our approach and the effectiveness of each module.
Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Jianzhao Li, Hanhong Zheng, Yue Zhao 0024
IEEE Trans. Geosci. Remote. Sens.4
2025 Bidirectional Stacking Ensemble Curriculum Learning for Hyperspectral Image Imbalanced Classification With Noisy Labels
abstract
Hyperspectral imaging has demonstrated substantial advantages in enhancing classification performance in remote sensing applications due to its abundant spectral information. To address the challenges of label noise and class imbalance in hyperspectral image (HSI) classification, we propose an end-to-end Feature-Guided Network (FGN) for HSI. Instead of merely combining spatial and channel attention, FGN leverages feature-level attention interactions to enhance contextual understanding, leading to better feature extraction, especially for underrepresented classes. Furthermore, a bidirectional loss for curriculum learning (CL) is proposed to rank the HSI training data in a descending or ascending order. The top and bottom loss regularizers are designed to make the proposed model suitable for noisy and imbalanced HSI data distributions. In the phase of selecting pace parameter, a stacking ensemble curriculum learning (SECL) model is established to avoid that the outliers and noisy HSI data are involved into the CL training process. A novel instruction matrix based on sample weights is designed for base classifiers. The outputs of the base models, combined with the expected labels, form the input-output pairs for training the second-level classifier. Experiments conducted on multiple hyperspectral imbalanced datasets with noisy labels demonstrate the superior performance of our method.
Yixin Wang 0009, Hao Li 0009, Maoguo Gong, Yue Wu 0004, Peiran Gong, A. K. Qin 0001, Lining Xing 0001, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.8
2025 Meta-Collaborative Learning for Arbitrarily Scaled Hyperspectral Image Super-Resolution
abstract
Deep learning-based methods for hyperspectral image super-resolution (SR) have achieved significant success in recent years. These methods typically consist of feature extraction module (FEM) and upsampling module. However, due to structural limitations of the upsampling module, most current methods focus on training separate models for different scale factors, which ignores the exploration of potential feature interdependence among different scale factors. In response to these challenges, we introduce a novel framework, called “meta-collaborative learning for arbitrarily scaled hyperspectral image super-resolution” (MCArb). Specifically, MCArb integrates a collaborative learning framework with a meta-learning-based 3-D upsampling module (3DMetaUM) and a scale-aware feature adaptation module (SAFAM). It enables training multiple SR tasks at different scale factors within a single network at the same time. This strategy is able not only to process arbitrary-scale-factor SR for hyperspectral images but also to harness the latent feature interdependence among different scales. In this study, we applied the MCArb framework to transform three deep learning-based hyperspectral image SR networks to MCArb methods, resulting in significant performance enhancements across five hyperspectral datasets. These improvements showcase the proposed MCArb framework’s ability to enhance feature extraction efficiency and capitalize on latent interscale correlations. This code is available athttps://github.com/ShuangWu-XDU/MCArb_HSI_SR.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Yu Zhou 0051, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.1
2025 Commonality Feature Representation Learning for Unsupervised Multimodal Change Detection
abstract
The main challenge of multimodal change detection (MCD) is that multimodal bitemporal images (MBIs) cannot be compared directly to identify changes. To overcome this problem, this paper proposes a novel commonality feature representation learning (CFRL) and constructs a CFRL-based unsupervised MCD framework. The CFRL is composed of a Siamese-based encoder and two decoders. First, the Siamese-based encoder can map original MBIs in the same feature space for extracting the representative features of each modality. Then, the two decoders are used to reconstruct the original MBIs by regressing themselves, respectively. Meanwhile, we swap the decoders to reconstruct the pseudo-MBIs to conduct modality alignment. Subsequently, all reconstructed images are input to the Siamese-based encoder again to map them in a same feature space, by which representative features are obtained. On this basis, latent commonality features between MBIs can be extracted by minimizing the distance between these representative features. These latent commonality features are comparable and can be used to identify changes. Notably, the proposed CFRL can be performed simultaneously in two modalities corresponding to MBIs. Therefore, two change magnitude images (CMIs) can be generated simultaneously by measuring the difference between the commonality features of MBIs. Finally, a simple threshold algorithm or a clustering algorithm can be employed to divide CMIs into binary change maps. Extensive experiments on six publicly available MCD datasets show that the proposed CFRL-based framework can achieve superior performance compared with other state-of-the-art approaches.
Tongfei Liu, Mingyang Zhang 0002, Maoguo Gong, Qingfu Zhang 0001, Fenlong Jiang, Hanhong Zheng, Di Lu 0004
IEEE Trans. Image Process.2
2025 Diffusion Model-Based Visual Compensation Guidance and Visual Difference Analysis for No-Reference Image Quality Assessment
abstract
Existing free-energy guided No-Reference Image Quality Assessment (NR-IQA) methods continue to face challenges in effectively restoring complexly distorted images. The features guiding the main network for quality assessment lack interpretability, and efficiently leveraging high-level feature information remains a significant challenge. As a novel class of state-of-the-art (SOTA) generative model, the diffusion model exhibits the capability to model intricate relationships, enhancing image restoration effectiveness. Moreover, the intermediate variables in the denoising iteration process exhibit clearer and more interpretable meanings for high-level visual information guidance. In view of these, we pioneer the exploration of the diffusion model into the domain of NR-IQA. We design a novel diffusion model for enhancing images with various types of distortions, resulting in higher quality and more interpretable high-level visual information. Our experiments demonstrate that the diffusion model establishes a clear mapping relationship between image reconstruction and image quality scores, which the network learns to guide quality assessment. Finally, to fully leverage high-level visual information, we design two complementary visual branches to collaboratively perform quality evaluation. Extensive experiments are conducted on seven public NR-IQA datasets, and the results demonstrate that the proposed model outperforms SOTA methods for NR-IQA. The codes will be available at https://github.com/handsomewzy/DiffV2IQA.
Zhaoyang Wang 0003, Bo Hu 0008, Mingyang Zhang 0002, Jie Li 0001, Leida Li, Maoguo Gong, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 CCGIB: A Cross-Channel Graph Information Bottleneck Principle
abstract
The empirical studies of most existing graph neural networks (GNNs) broadly take the original node feature and adjacency relationship as single-channel input, ignoring the rich information of multiple graph channels. To circumvent this issue, the multichannel graph analysis framework has been developed to fuse graph information across channels. How to model and integrate shared (i.e., consistency) and channel-specific (i.e., complementarity) information is a key issue in multichannel graph analysis. In this article, we propose a cross-channel graph information bottleneck (CCGIB) principle to maximize the agreement for common representations and the disagreement for channel-specific representations. Under this principle, we formulate the consistency and complementarity information bottleneck (IB) objectives. To enable optimization, a viable approach involves deriving variational lower bound and variational upper bound (VarUB) of mutual information terms, subsequently focusing on optimizing these variational bounds to find the approximate solutions. However, obtaining the lower bounds of cross-channel mutual information objectives proves challenging through direct utilization of variational approximation, primarily due to the independence of the distributions. To address this challenge, we leverage the inherent property of joint distributions and subsequently derive variational bounds to effectively optimize these information objectives. Extensive experiments on graph benchmark datasets demonstrate the superior effectiveness of the proposed method.
Xiaolong Fan, Maoguo Gong, Yue Wu 0004, Mingyang Zhang 0002, Hao Li 0009, Xiangming Jiang
IEEE Trans. Neural Networks Learn. Syst.4
2025 Few-Shot Learning With Enhancements to Data Augmentation and Feature Extraction
abstract
The few-shot image classification task is to enable a model to identify novel classes by using only a few labeled samples as references. In general, the more knowledge a model has, the more robust it is when facing novel situations. Although directly introducing large amounts of new training data to acquire more knowledge is an attractive solution, it violates the purpose of few-shot learning with respect to reducing dependence on big data. Another viable option is to enable the model to accumulate knowledge more effectively from existing data, i.e., improve the utilization of existing data. In this article, we propose a new data augmentation method called self-mixup (SM) to assemble different augmented instances of the same image, which facilitates the model to more effectively accumulate knowledge from limited training data. In addition to the utilization of data, few-shot learning faces another challenge related to feature extraction. Specifically, existing metric-based few-shot classification methods rely on comparing the extracted features of the novel classes, but the widely adopted downsampling structures in various networks can lead to feature degradation due to the violation of the sampling theorem, and the degraded features are not conducive to robust classification. To alleviate this problem, we propose a calibration-adaptive downsampling (CADS) that calibrates and utilizes the characteristics of different features, which can facilitate robust feature extraction and benefit classification. By improving data utilization and feature extraction, our method shows superior performance on four widely adopted few-shot classification datasets.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.5
2025 SPCNet: Deep Self-Paced Curriculum Network Incorporated With Inductive Bias
abstract
The vulnerability to poor local optimum and the memorization of noise data limit the generalizability and reliability of massively parameterized convolutional neural networks (CNNs) on complex real-world data. Self-paced curriculum learning (SPCL), which models the easy-to-hard learning progression from human beings, is considered as a potential savior. In spite of the fact that numerous SPCL solutions have been explored, it still confronts two main challenges exactly in solving deep networks. By virtue of various designed regularizers, existing weighting schemes independent of the learning objective heavily rely on the prior knowledge. In addition, alternative optimization strategy (AOS) enables the tedious iterative training procedure, thus there is still not an efficient framework that integrates the SPCL paradigm well with networks. This article delivers a novel insight that attention mechanism allows for adaptive enhancement in the contribution of diverse instance information to the gradient propagation. Accordingly, we propose a general-purpose deep SPCL paradigm that incorporates the preferences of implicit regularizer for different samples into the network structure with inductive bias, which in turn is formalized as the self-paced curriculum network (SPCNet). Our proposal allows simultaneous online difficulty estimation, adaptive sample selection, and model updating in an end-to-end manner, which significantly facilitates the collaboration of SPCL to deep networks. Experiments on image classification and scene classification tasks demonstrate that our approach surpasses the state-of-the-art schemes and obtains superior performance.
Yue Zhao 0024, Maoguo Gong, Mingyang Zhang 0002, A. K. Qin 0001, Fenlong Jiang, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.3
2024 Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution Network
abstract
Existing hyperspectral image (HSI) super-resolution (SR) methods struggle to effectively capture the complex spectral-spatial relationships and low-level details, while diffusion models represent a promising generative model known for their exceptional performance in modeling complex relations and learning high and low-level visual features. The direct application of diffusion models to HSI SR is hampered by challenges such as difficulties in model convergence and protracted inference time. In this work, we introduce a novel Group-Autoencoder (GAE) framework that synergistically combines with the diffusion model to construct a highly effective HSI SR model (DMGASR). Our proposed GAE framework encodes high-dimensional HSI data into low-dimensional latent space where the diffusion model works, thereby alleviating the difficulty of training the diffusion model while maintaining band correlation and considerably reducing inference time. Experimental results on both natural and remote sensing hyperspectral datasets demonstrate that the proposed method is superior to other state-of-the-art methods both visually and metrically.
Zhaoyang Wang 0003, Mingyang Zhang 0002, Hao Luo 0004, Maoguo Gong
AAAI3
2024 PointMC: Multi-instance Point Cloud Registration based on Maximal Cliques
abstract
Multi-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle registration scenarios where multiple instances overlap. To address these problems, we propose a maximal clique based multiinstance point cloud registration framework called PointMC. The key idea is to search for maximal cliques on the correspondence compatibility graph to estimate multiple transformations, and cluster these transformations into clusters corresponding to different instances to efficiently and accurately estimate all poses. PointMC leverages a correspondence embedding module that relies on local spatial consistency to effectively eliminate outliers, and the extracted discriminative features empower the network to circumvent missed pose detection in scenarios involving multiple overlapping instances. We conduct comprehensive experiments on both synthetic and real-world datasets, and the results show that the proposed PointMC yields remarkable performance improvements.
Yue Wu 0004, Xidao Hu, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Qiguang Miao, Wenping Ma 0001
ICML7
2024 Unsupervised Domain Adaptation for Cross-Scene Hyperspectral Image Classification Based on Decoupled Contrastive Learning
abstract
Recent studies have highlighted the effectiveness of deep domain adaptation (DA) techniques in addressing cross-scene hyperspectral image (HSI) classification challenges. However, most of the existing DA methods often prioritize aligning data distributions while overlooking the intrinsic separability between source and target domain data. In this paper, we propose a decoupled contrastive learning based unsupervised domain adaptation (DCLUDA) method for HSI classification. Unlike conventional adversarial DA methods, our method introduces a unique DA loss specifically designed to minimize class confusion in the target domain. This not only simplifies model training but also enhances class discriminability. Moreover, we employ a decoupled contrastive learning strategy on both domains to enhance data separability within each domain. Finally, we propose a sample selection strategy based on confident learning to select high-confidence samples from the target domain for fine-tuning the DA model. Experiments on two cross-scene HSI classification tasks shown that our proposed DCLUDA outperforms several existing DA methods.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Xiangming Jiang, Yu Zhou 0051, Dan Feng 0002
IJCNN2
2024 Data Customization-Based Multiobjective Optimization Pruning Framework for Remote Sensing Scene Classification
abstract
Pruning techniques have been utilized widely for convolutional neural networks (CNNs) to reduce the computation resources in remote sensing scene image classification. However, conventional pruning techniques are weight-based, which can not balance the pruning ratio and representation ability appropriately. In this paper, we propose a Data Customization-based Multiobjective Optimization Pruning (DCMOP) framework for the pruning in remote sensing scene image classification, which can not only trade-off between pruning ratio and capability for CNNs, but also speed up the evolutionary process for the pruning. We adopt the multiobjective evolutionary algorithms (MOEAs) to search for a trade-off between the pruning ratio and capability for CNNs. However, a big concern of pruning for networks via MOEAs is that the evaluation of sub-networks is time-costing. This originates that the slimmed sub-networks require a lot of retraining operation, which will burden the hardware. In order to alleviate this limitation, we design a Data Customization-based Proxy Mechanism (DCPM) to reduce the size of the input dataset in terms of the structure of the slimmed sub-network to accelerate significantly the evolutionary process for the pruning. According to this, our proposed DCMOP achieves the pruning with higher efficiency and performance by cooperating with MOEAs and DCPM. Experimental results based on four datasets of AID, NWPURESISC45, PatternNet, and WHU-RS19 show that the proposed DCMOP can achieve a balance between model performance and pruning rate, while obviously reducing the time cost of the pruning.
Zhuping Hu, Maoguo Gong, Yiheng Lu, Jianzhao Li, Yue Zhao 0024, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 MSANet: Multiscale Self-Attention Aggregation Network for Few-Shot Aerial Imagery Segmentation
abstract
Few-shot aerial imagery segmentation refers to the task of segmenting specific objects in scenes that have not been encountered during training with a small amount of annotated data for reference. However, most existing few-shot segmentation algorithms are primarily designed for natural images, and there is still a lack of exploration in the context of remote sensing aerial imagery. In this article, we propose a novel multiscale self-attention aggregation network (MS2A2Net), dubbed MS2A2Net, to address the challenge of few-shot aerial image segmentation in terms of scarce data and network architecture. Specifically, we first incorporate the designed asymmetric momentum contrastive learning (AMCL) into the pre-training stage, to improve the representation capability of the backbone without the expensive labeled data. Then the frozen encoder is transferred to the downstream few-shot segmentation task as the feature embedding. In terms of network architecture, we design self-attention aggregation in multiscale feature fusion, to construct the dual correlation of foreground and background between support and query features at the pixel level. Besides, the coordinate attention is designed to rearrange the distribution of feature importance in both horizontal and vertical spatial order perspectives, which facilitates adaptive fusion with the multiscale features. To verify the availability of the proposed MS2A2Net, we also reconstructed two novel datasets dedicated to few-shot aerial image segmentation, called DLRSD-$4^{i}$and iSAID-$4^{i}$. The experimental results show that our approach MS2A2Net is superior in three few-shot benchmark aerial imagery segmentation datasets, which achieves competitive segmentation performance. Extensive ablation experiments also reflect the effectiveness and scalability of the proposed components and overall network architecture.
Jianzhao Li, Maoguo Gong, Mingyang Zhang 0002, Yourun Zhang, Shanfeng Wang, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.4
2024 A Hybrid Multitask Learning Network for Hyperspectral Image Classification With Few Labels
abstract
Recently, the field of hyperspectral image (HSI) classification has witnessed advancements with the emergence of deep learning models. Promising approaches, such as self-supervised strategies and domain adaptation, have effectively tackled the overfitting challenges posed by limited labeled samples in HSI classification. To extract comprehensive semantic information from different types of auxiliary tasks, which view the problem from multiple perspectives, and efficiently integrate multiple tasks into a single network, this paper proposes a hybrid multi-task learning framework (HyMuT) by sharing representations across multiple tasks. Based on the similarity between the data and target classification task, we construct three auxiliary tasks that are similar, related and weakly correlated to the target task, while three corresponding multi-task learning methods are integrated. The framework establishes a backbone network with a hard parameter sharing mechanism, which handles the main task and a similar spatial mask classification task. Subsequently, a hierarchical transfer multi-task learning approach is introduced to transfer the knowledge of a spatial-spectral joint mask reconstruction task from the autoencoder to the backbone network. Furthermore, a new source domain HSI dataset is introduced as an auxiliary task weakly correlated. To solve the source domain classification task and assist the hard parameter sharing mechanism, a dual adversarial classifier based on adversarial learning is employed. This classifier effectively extracts domain and task invariance. Extensive experiments are conducted on four benchmark HSI datasets to evaluate the performance. The results demonstrate that HyMuT outperforms state-of-the-art methods. This code will be available from the website: https://github.com/HaoLiu-XDU/HyMuT.
Hao Li 0009, Mingyang Zhang 0002, Ziqi Di, Maoguo Gong, Tianqi Gao, A. K. Qin 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Adversarial Feature Equilibrium Network for Multimodal Change Detection in Heterogeneous Remote Sensing Images
abstract
Change detection (CD) methods have been crucial in exploring geo-environmental science. With the advancement of remote sensing (RS) technology, multimodal images acquired from different platforms and sensors are widely used for CD tasks. As an emerging task, multimodal CD (MCD) aims to achieve more comprehensive and precise detection of land cover changes through complementary information in multimodal images. However, there are significant differences between modalities, particularly in heterogeneous images. How to deal with modal differences while effectively integrating change information remains a challenge in MCD. In this article, we propose a novel adversarial feature equilibrium network (AFENet), which establishes an additional adversarial optimization to solve the equilibrium problem between modal differences and land cover changes. Our AFENet aligns the features and reduces the modal gap through a multiscale adversarial domain adaptation (MADA) approach. Meanwhile, a divergence-aware contrastive module (DCM) is designed as a regularization term for adversarial optimization. DCM affects the sensitivity of feature extractors by constraining the mutual information between changed and unchanged pixels. In this case, AFENet can maintain the consistency of feature representation while maximizing the discriminability of change targets. The features extracted from AFENet will then be integrated by our multistream feature fusion (MFF) module and utilized to generate change maps. The effectiveness of our approach is demonstrated on two scene-level multimodal RS datasets. Compared with existing methods, our AFENet achieves state-of-the-art (SOTA) performance on both datasets and outperforms the second-best$F1$score by 4.64% and 1.1%, respectively.
Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Tianqi Gao, Fenlong Jiang
IEEE Trans. Geosci. Remote. Sens.4
2024 Gradient-Guided Multiscale Focal Attention Network for Remote Sensing Scene Classification
abstract
Remote sensing scene classification (RSSC) aims to understand and analyze the semantic information at the scene level with complex geographical properties. Despite the profound success of advanced deep models in automatically capturing hierarchical embedding representations and the gradual dominant trend in RSSC, it still remains a great challenge to precisely focus on targets at variable scales that are considered highly relevant to the corresponding scene and separated from the background. Motivated by this recognition, in this article, we present the gradient-guided multiscale focal attention network (GMFANet) for RSSC to adaptively localize the representative multiscale semantic representation for complex scenes. In particular, a lightweight parameterized hierarchical multiscale attention (HMA) mechanism is proposed, which constitutes the main aim of adaptively enhancing physical detail and high-level semantic information at different layers, rather than regarding each scale set with equivalent insight, while eliminating redundant information inherent in conventional attention mechanisms. Subsequently, a gradient-guided spatial focused attention (GSFA) module is specifically designed to accurately localize critical regions at multiple scales, with the dynamic combination of gradient-activated reference attention map and prediction attention map from supervised information-based learning. In addition, a curriculum-driven dynamic attention fusion (CDAF) strategy is tailored to fuse the spatial attention above from easy to hard for avoiding from poor local optimum and decreasing the early learning ambiguity. Our extensive comparative experiments and ablation analyses implemented on real-world public RSSC datasets indicate that our approach achieves the state-of-the-art performance exactly. The code is available athttps://github.com/bling2beyond/GMFANet.
Yue Zhao 0024, Maoguo Gong, A. K. Qin 0001, Mingyang Zhang 0002, Zhuping Hu, Tianqi Gao, Yan Pu
IEEE Trans. Geosci. Remote. Sens.4
2024 Spectral Knowledge Transfer for Remote Sensing Change Detection
abstract
Change detection (CD) in multispectral remote sensing (RS) imagery suffers from low spectral resolution which can lead to degraded recognition of change information from land cover objects. Considering that natural hyperspectral imagery (HSI) is much higher in spectral resolution and more accessible, using it to enhance the spectral information of RS multispectral imagery for CD can improve performance. To achieve this, we propose a spectral knowledge transfer (SKT) framework to allow the creation of pseudo-hyperspectral RS images from the available RS multispectral ones without the need for the real pairs of RS multispectral and hyperspectral images, typically required by existing RS spectral enhancement methods. Specifically, an autoencoder is first trained based on the available pairs of natural HSI and its multispectral counterparts and then calibrated via the available RS multispectral images. The finally obtained decoder module is used to generate the pseudo-hyperspectral image from an input RS multispectral image. We further propose a multispectrum collaborative CD (MCCD) framework that leverages both the real multispectral images and the pseudo hyperspectral images generated from them in a collaborative way to achieve performance improvement. Extensive experiments on two large-scale RS CD datasets and eight existing deep learning-based CD methods demonstrate the stronger efficacy of the proposed method.
Hanhong Zheng, Mingyang Zhang 0002, Maoguo Gong, A. K. Qin 0001, Tongfei Liu, Fenlong Jiang
IEEE Trans. Geosci. Remote. Sens.3
2024 RORNet: Partial-to-Partial Registration Network With Reliable Overlapping Representations
abstract
Three-dimensional point cloud registration is an important field in computer vision. Recently, due to the increasingly complex scenes and incomplete observations, many partial-overlap registration methods based on overlap estimation have been proposed. These methods heavily rely on the extracted overlapping regions with their performances greatly degraded when the overlapping region extraction underperforms. To solve this problem, we propose a partial-to-partial registration network (RORNet) to find reliable overlapping representations from the partially overlapping point clouds and use these representations for registration. The idea is to select a small number of key points called reliable overlapping representations from the estimated overlapping points, reducing the side effect of overlap estimation errors on registration. Although it may filter out some inliers, the inclusion of outliers has a much bigger influence than the omission of inliers on the registration task. The RORNet is composed of overlapping points' estimation module and representations' generation module. Different from the previous methods of direct registration after extraction of overlapping areas, RORNet adds the step of extracting reliable representations before registration, where the proposed similarity matrix downsampling method is used to filter out the points with low similarity and retain reliable representations, and thus reduce the side effects of overlap estimation errors on the registration. Besides, compared with previous similarity-based and score-based overlap estimation methods, we use the dual-branch structure to combine the benefits of both, which is less sensitive to noise. We perform overlap estimation experiments and registration experiments on the ModelNet40 dataset, outdoor large scene dataset KITTI, and natural data Stanford Bunny dataset. The experimental results demonstrate that our method is superior to other partial registration methods. Our code is available at https://github.com/superYuezhang/RORNet.
Yue Wu 0004, Yue Zhang 0040, Wenping Ma 0001, Maoguo Gong, Xiaolong Fan, Mingyang Zhang 0002, A. K. Qin 0001, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.6
2024 Semisupervised Change Detection Based on Bihierarchical Feature Aggregation and Extraction Network
abstract
With the rapid development of remote sensing (RS) technology, high-resolution RS image change detection (CD) has been widely used in many applications. Pixel-based CD techniques are maneuverable and widely used, but vulnerable to noise interference. Object-based CD techniques can effectively utilize the abundant spectrum, texture, shape, and spatial information but easy-to-ignore details of RS images. How to combine the advantages of pixel-based methods and object-based methods remains a challenging problem. Besides, although supervised methods have the capability to learn from data, the true labels representing changed information of RS images are often hard to obtain. To address these issues, this article proposes a novel semisupervised CD framework for high-resolution RS images, which employs small amounts of true labeled data and a lot of unlabeled data to train the CD network. A bihierarchical feature aggregation and extraction network (BFAEN) is designed to achieve the pixelwise together with objectwise feature concatenation feature representation for the comprehensive utilization of the two-level features. In order to alleviate the coarseness and insufficiency of labeled samples, a confident learning algorithm is used to eliminate noisy labels and a novel loss function is designed for training the model using true- and pseudo-labels in a semisupervised fashion. Experimental results on real datasets demonstrate the effectiveness and superiority of the proposed method.
Mingyang Zhang 0002, Tianqi Gao, Maoguo Gong, Shengqi Zhu 0001, Yue Wu 0004, Hao Li 0009
IEEE Trans. Neural Networks Learn. Syst.1
2024 Self-Supervised Monocular Depth Estimation With Self-Perceptual Anomaly Handling
abstract
It is attractive to extract plausible 3-D information from a single 2-D image, and self-supervised learning has shown impressive potential in this field. However, when only monocular videos are available as training data, moving objects at similar speeds to the camera can disturb the reprojection process during training. Existing methods filter out some moving pixels by comparing pixelwise photometric error, but the illumination inconsistency between frames leads to incomplete filtering. In addition, existing methods calculate photometric error within local windows, which leads to the fact that even if an anomalous pixel is masked out, it can still implicitly disturb the reprojection process, as long as it is in the local neighborhood of a nonanomalous pixel. Moreover, the ill-posed nature of monocular depth estimation makes the same scene correspond to multiple plausible depth maps, which damages the robustness of the model. In order to alleviate the above problems, we propose: 1) a self-reprojection mask to further filter out moving objects while avoiding illumination inconsistency; 2) a self-statistical mask method to prevent the filtered anomalous pixels from implicitly disturbing the reprojection; and 3) a self-distillation augmentation consistency loss to reduce the impact of ill-posed nature of monocular depth estimation. Our method shows superior performance on the KITTI dataset, especially when evaluating only the depth of potential moving objects.
Yourun Zhang, Maoguo Gong, Mingyang Zhang 0002, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.3
2023 Superpixel-based multiobjective change detection based on self-adaptive neighborhood-based binary differential evolution
Tianqi Gao, Hao Li 0009, Maoguo Gong, Mingyang Zhang 0002, Wenyuan Qiao
Expert Syst. Appl.4
2023 Context-content collaborative network for building extraction from high-resolution imagery
Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Qingfu Zhang 0001, Di Lu 0004, Hanhong Zheng, Fenlong Jiang
Knowl. Based Syst.3
2023 Autonomous perception and adaptive standardization for few-shot learning
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
Knowl. Based Syst.5
2023 Features kept generative adversarial network data augmentation strategy for hyperspectral image classification
Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yue Wu 0004, Hao Li 0009
Pattern Recognit.1
2023 Self-structured pyramid network with parallel spatial-channel attention for change detection in VHR remote sensed imagery
Mingyang Zhang 0002, Hanhong Zheng, Maoguo Gong, Yue Wu 0004, Hao Li 0009, Xiangming Jiang
Pattern Recognit.1
2023 A Bilevel Gene-Based Multiobjective Memetic Algorithm for Passive Localization System Deployment Optimization
abstract
The passive localization system (PLS) is fundamental to many wireless applications. The deployment of the monitoring stations plays a key role in the performance of the PLSes. However, the workflow of the emerging cutting-edge PLSes is becoming more flexible in the complicated environment, which makes it hard to optimize the deployment. To fulfill the requirement of the real-world applications, we propose a multiobjective PLS deployment optimization model, including a surrogate geometric dilution of precision (S-GDOP) model and a system coverage indicator to meet the demand for the detection performance of the known and unknown targets. The proposed S-GDOP is separable and open to various performance-related factors in this article. Motivated by the various cooperation mechanisms and the empirical deployment patterns, we propose a bilevel gene-based multiobjective memetic algorithm within the decomposition framework to solve this problem. By maintaining an adaptive multicomponent gene population (MCGP) and a local pivot (LP)-based local search, the population evolves on two precise and consecutive gene levels, which effectively utilizes the problem and evolution-related heuristic information. The proposed algorithm outperforms another four popular algorithms in 83.3% bilateral comparisons and obtains more implicit deployment patterns, clearer deployment structures, and better converged Pareto fronts.
Zhao Wang 0011, Maoguo Gong, Fei Xie 0007, Mingyang Zhang 0002
IEEE Trans. Evol. Comput.5
2023 An M-Nary SAR Image Change Detection Based on GAN Architecture Search
abstract
Change detection (CD) in synthetic aperture radar (SAR) images aims to detect changed areas by considering the changes in backscattering coefficients. However, the changes can be further divided into positive and negative changes in terms of the increase or decrease of backscattering coefficient, so the CD task can be divided into binary and ternary according to the number of existent categories. This paper introduces an M-nary (binary or ternary) SAR change detection procedure based on the generative adversarial network (GAN) and neural architecture search (NAS) strategy to detect which changes exist in the SAR image-pair and design specialized classifiers for both binary and ternary CD. First, a difference image generation approach based on the salient changed region extraction and neighborhood information is designed for a robust difference representation on the M-nary CD. Due to the further subdivision of changes, the insufficiency of labeled data presents the M-nary change detection with a dilemma. Concerning the lack of labeled information, this paper presents a labeled sample generation strategy based on the GAN architecture search to supplement sample data. Since GAN training is inherently unstable, NAS provides an effective means of searching GAN architecture automatically and ameliorates the reliability of generated samples. During the architecture search procedure, a double-phase evolutionary search strategy is introduced to further improve the stability of GAN training. The experimental results with theoretical analysis prove the validity, robustness, and potential of our method in synthetic as well as real SAR datasets.
Maoguo Gong, Tianqi Gao, Mingyang Zhang 0002, Wei Li 0032, Zhibin Wang 0004, Dezhong Li
IEEE Trans. Geosci. Remote. Sens.3
2023 Self-Supervised Global-Local Contrastive Learning for Fine-Grained Change Detection in VHR Images
abstract
Self-supervised contrastive learning (CL) can learn high-quality feature representations that are beneficial to downstream tasks without labeled data. However, most CL methods are for image-level tasks. For the fine-grained change detection (FCD) tasks, such as change or change trend detection of some specific ground objects, it is usually necessary to perform pixel-level discriminative analysis. Therefore, feature representations learned by image-level CL may have limited effects on FCD. To address this problem, we propose a self-supervised global–local contrastive learning (GLCL) framework, which extends the instance discrimination task to the pixel level. GLCL follows the current mainstream CL paradigm and consists of four parts, including data augmentation to generate different views of the input, an encoder network for feature extraction, a global CL head, and a local CL head to perform image-level and pixel-level instance discrimination tasks, respectively. Through GLCL, features belonging to different perspectives of the same instance will be pulled closer, while features of different instances will be alienated, which can enhance the discriminativeness of feature representations from both global and local perspectives, thereby facilitating downstream FCD tasks. In addition, GLCL makes a targeted structural adaptation to FCD, i.e., the encoder network is undertaken by the common backbone networks of FCD, which can accelerate the deployment on downstream FCD tasks. Experimental results on several real datasets show that compared with other parameter initialization methods, the FCD models pretrained by GLCL can obtain better detection performance.
Fenlong Jiang, Maoguo Gong, Hanhong Zheng, Tongfei Liu, Mingyang Zhang 0002, Jia Liu 0020
IEEE Trans. Geosci. Remote. Sens.5
2023 Multiform Ensemble Self-Supervised Learning for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning is an effective way to solve model collapse for few-shot remote sensing scene classification (FSRSSC). However, most self-supervised contrastive learning auxiliary tasks perform poorly on the high interclass similarity problem in FSRSSC. Furthermore, it is time-consuming and computationally expensive to obtain the best combination among numerous self-supervised auxiliary tasks. In practical applications, we may encounter difficulties in remote sensing data acquisition and labeling, while most FSRSSC studies only focus on the former. To alleviate the above problems, we propose a multiform ensemble self-supervised learning (MES2L) framework for FSRSSC in this article. Based on the transfer learning-based few-shot scheme, we design a novel global–local contrastive learning auxiliary task to solve the low interclass separability problem. The self-attention mechanism is designed in the local contrast features to investigate the intrinsic associations between different remote sensing scene objectives. We also present a multiform ensemble enhancement (MEE) training method. Ensemble enhancement involves the concatenation of features extracted from different backbones trained by a combination of multiform self-supervised auxiliary tasks. MEE can not only be regarded as a more straightforward alternative to knowledge distillation but also can achieve an effective compromise between expensive computational cost and classification accuracy. In addition, we provide two scene classification schemes of inductive and transductive settings, corresponding to solving the difficulties of remote sensing data acquisition and labeling. The proposed network achieves state-of-the-art results on three benchmark FSRSSC datasets. The potential of the MES2L framework is also demonstrated in combination with classical metalearning-based and metric learning-based few-shot algorithms.
Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Mingyang Zhang 0002, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.5
2023 Energy-Based CNN Pruning for Remote Sensing Scene Classification
abstract
Convolutional neural networks (CNNs) have been adopted to classify the remote sensing scene image. However, the application of these complicated networks on the satellite platform is difficult because of the limited computation resources. Therefore, we propose an energy-based filter pruning framework (EFPF) to reduce the size of the original model. The energy can be obtained through the eigenvalues of each weight tensor by singular value decomposition (SVD). Specifically, we calculate the energy of each layer by the ratio of eigenvalues that are lower than a specified truncation parameter and then remove filters from the original layer in light of the degree of energy. The EFPF is reliable because SVD techniques can capture the covariance among all filters from the original weight tensor, and therefore, the energy from the eigenvalues can reflect the redundancy of the filters. (i.e., if the distribution of eigenvalues is sharp, then the energy among filters will be lower, and the redundancy will be higher.) Surprisingly, the EFPF can reduce the FLOPs and parameters, as well as improve the top1 accuracy obviously when the VGG-16 and ResNet-50 are adopted to classify the AID, NWPU45, PatternNet, and WHU19 datasets. Additionally, the EFPF can achieve similar pruning results when the original model is fully-trained (converge) and under-trained (In-converge), which means we can save the training computation resources for the original model.
Yiheng Lu, Maoguo Gong, Zhuping Hu, Wei Zhao 0019, Ziyu Guan, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.6
2023 Cross-Domain Self-Taught Network for Few-Shot Hyperspectral Image Classification
abstract
In recent years, deep learning models, which possess powerful feature extraction abilities, have achieved remarkable success in the classification of hyperspectral images (HSIs). Nevertheless, a common challenge faced by most deep learning models, including few-shot learning models, is the scarcity of valid labeled samples. To address this issue, we propose a cross-domain self-taught network (CDSTN) for few-shot hyperspectral image classification. The proposed CDSTN merges domain adaptation and semi-supervised self-taught strategy to implement the few-shot learning, which utilizes adequate labeled and unlabeled samples from source as well as target domain respectively. For the feature information extraction of HSI, we propose a deep spatial-spectral feature embedded extractor composed of four residual blocks and a channel attention module. Additionally, a set of domain classifiers are introduced behind each residual block for the purpose of domain alignment by extracting more domain information at different depths of the network. Finally, plenty of unlabeled samples are assigned with pseudo labels through the trained network, and a pseudo label refinement module is designed to select the most confident pseudo label sample for each class to further enrich the labeled database of target domain. Experiments conducted on four widely used benchmark HSI data sets demonstrate that CDSTN can obtain superior and stable performance with limited labeled samples compared with some state of the arts.
Mingyang Zhang 0002, Hao Liu 0123, Maoguo Gong, Hao Li 0009, Yue Wu 0004, Xiangming Jiang
IEEE Trans. Geosci. Remote. Sens.1
2022 Evolutionary Multitasking CNN Architecture Search for Hyperspectral Image Classification
abstract
In recent years, convolutional neural networks (CNNs) have shown excellent effectiveness on hyperspectral image classification (HSI) tasks. However, it is a challenge to design a suitable CNN architecture to obtain great performance according to different tasks. Different from the traditional manual design, in this paper, an evolutionary multitasking CNN architecture search framework for HSI classification is proposed to search the optimal architectures and accomplish classification of different tasks simultaneously. Through encoding the CNN architectures, the proposed algorithm is able to achieve global search in the same search space and select well-adapted individuals for evolution. In the evolutionary multitasking environment, information can be transferred between and within tasks, which can accelerate the convergence and explore good architectures through beneficial transfer. In the experiments, the effectiveness of the proposed method is demonstrated by the comparison with different methods on two common data sets.
Yiting Liu 0004, Hao Li 0009, Maoguo Gong, Jieyi Liu, Yue Wu 0004, Mingyang Zhang 0002, Jiao Shi
IJCNN6
2022 An anti-jamming method in multistatic radar system based on convolutional neural network
abstract
Abstract For the existing jamming discrimination methods on the multistatic radar system, the single feature of target echo space correlation is utilised as the metric, which leads to the lack of comprehensive feature extraction and universal discrimination algorithm. In this study, a discrimination method in a multistatic radar system based on the convolutional neural network is proposed. This proposal combines the advantages of multiple‐radar systems cooperative detection technology with the convolutional neural network, and effectively applies to the field of anti‐deception jamming, which takes full advantage of unknown information of echo data to obtain multi‐dimensional, comprehensive, complete and deep feature differences besides correlation, so as to achieve a better jamming discrimination effect. The simulation results show that the proposed method can extract the multidimensional and separable essential features of echoes, and all these features have a strong degree of differentiation between targets and jamming, which effectively reduce the influence of noise and pulse number. At the same time, the influence of radar distribution on jamming discrimination under non‐ideal conditions is relieved, when the correlation coefficient of the true target reaches 0.4, the discrimination probability remains above 85%, which broadens the boundary conditions of the application process.
Jieyi Liu, Maoguo Gong, Mingyang Zhang 0002, Hao Li 0009
IET Signal Process.3
2022 HFA-Net: High frequency attention siamese network for building change detection in VHR remote sensing images
Hanhong Zheng, Maoguo Gong, Tongfei Liu, Fenlong Jiang, Tao Zhan 0005, Di Lu 0004, Mingyang Zhang 0002
Pattern Recognit.7
2022 Symmetric All Convolutional Neural-Network-Based Unsupervised Feature Extraction for Hyperspectral Images Classification
abstract
Recently, deep-learning-based feature extraction (FE) methods have shown great potential in hyperspectral image (HSI) processing. Unfortunately, it also brings a challenge that the training of the deep learning networks always requires large amounts of labeled samples, which is hardly available for HSI data. To address this issue, in this article, a novel unsupervised deep-learning-based FE method is proposed, which is trained in an end-to-end style. The proposed framework consists of an encoder subnetwork and a decoder subnetwork. The structure of the two subnetworks is symmetric for obtaining better downsampling and upsampling representation. Considering both spectral and spatial information, 3-D all convolution nets and deconvolution nets are used to structure the encoder subnetwork and decoder subnetwork, respectively. However, 3-D convolution and deconvolution kernels bring more parameters, which can deteriorate the quality of the obtained features. To alleviate this problem, a novel cost function with a sparse regular term is designed to obtain more robust feature representation. Experimental results on publicly available datasets indicate that the proposed method can obtain robust and effective features for subsequent classification tasks.
Mingyang Zhang 0002, Maoguo Gong, Haibo He, Shengqi Zhu 0001
IEEE Trans. Cybern.1
2022 A Spectral and Spatial Attention Network for Change Detection in Hyperspectral Images
abstract
Hyperspectral images (HSIs) contain rich spectral signatures that reveal more image details and, thus, enable the detection of less noticeable changes on the ground. However, HSI-based change detection (CD) is susceptible to a large amount of irrelevant or noisy spectral and spatial information due to massive spectral bands. To address these issues, we propose a novel spectral and spatial attention network (S2AN) for HSI-based CD, which is capable to suppress CD-irrelevant spectral and spatial information via adaptive spectral and spatial attention mechanisms. S2AN takes as input the image patch from the difference map between two HSIs and outputs the status of change for the patch. Specifically, S2AN is composed of several repeated attention blocks, each of which contains the spectral attention (SpeA) module for directly calculating the attention score for each input channel, the Gaussian spatial attention (GSpaA) module that first constructs an adaptive Gaussian distribution and then samples it to derive the attention scores for each spatial position, and the convolutional feature extraction (CFE) module for extracting features from the attention-weighted input. It is worth mentioning that, in addition to the advantage of the attention, GSpaA also reduces the sensitivity of patch size for patch-based methods. To effectively train S2AN when facing insufficient labeled data, a semisupervised strategy that combines supervised and unsupervised methods to augment labeled training data is proposed. Experiments on several HSI datasets in comparison to existing methods show the superiority of S2AN.
Maoguo Gong, Fenlong Jiang, A. K. Qin 0001, Tongfei Liu, Tao Zhan 0005, Di Lu 0004, Hanhong Zheng, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.8
2022 Two-Path Aggregation Attention Network With Quad-Patch Data Augmentation for Few-Shot Scene Classification
abstract
The few-shot scene classification is dedicated to identifying unseen remote sensing classes when only a very small number of labeled samples are available for reference. Most of the existing few-shot scene classification methods are based on meta-learning and employ the episodic learning for training, which lacks the consideration for the utilization of data efficiency. In this paper, instead of designing sophisticated meta-learning based algorithms, we are committed to training a feature extractor with good generalization performance and strong feature extraction capability. Specifically, we propose a novel two-path aggregation attention network with quad-patch data augmentation, called DANet, to solve the problem of few-shot scene classification from both data and architecture aspects. In terms of data, we design a new data augmentation strategy named quad-patch augmentation. We utilize the characteristics of remote sensing images to chunk and reassemble any existing data, thereby generating pseudo-new data to enrich the training set. In terms of architecture, we present a two-path aggregation attention module that makes it easier for the model to focus on the key clues in a targeted manner. The comparative experiments in natural image datasets and remote sensing image datasets demonstrate the effectiveness of our two innovations. In addition, DANet achieves competitive or state-of-the-art (SOTA) results on three benchmark scene classification datasets.
Maoguo Gong, Jianzhao Li, Yourun Zhang, Yue Wu 0004, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.5
2022 A Vertex-Directed Evolutionary Algorithm for Multiobjective Endmember Estimation
abstract
Hyperspectral unmixing including endmember extration and abundance estimation has been investigated successively in recent years due to increasingly hyperspectral processing requirements. As one type of decision after solution paradigm, multiobjective endmember estimation method is able to obtain a set of Pareto optimal solutions, thus providing a wealth of information to determine the most representative endmembers. In addition, multiobjective optimization methods also have the characteristics of flexible modeling, excellent global convergence and commendable adaptability, etc. However, the evolutionary algorithms designed for this kind of method generally use little spatial-spectral information of the hyperspectral image. In this paper, we delve into the memetic strategy by exploiting the topological structure of hyperspectral data in the high-dimensional space to establish a vertex-directed multiobjective endmember estimation method, termed VD-MoEE. According to the frequently used linear mixture model, endmembers of hyperspectral images are generally distributed at the vertices of hyperspectral data manifold. Therefore, we design a vertex-directed local search operator to guide the search direction of the individuals in evolutionary algorithms. Experimental results on synthetic as well as real data sets demonstrated that the proposed VD-MOEE is able to achieve appealing performance in terms of the solution selecting and accuracy in comparison with several classic and state-of-the-art endmember estimation methods.
Xiangming Jiang, Maoguo Gong, Tao Zhan 0005, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.5
2022 Building Change Detection for VHR Remote Sensing Images via Local-Global Pyramid Network and Cross-Task Transfer Learning Strategy
abstract
Building change detection (BCD) for very-high-spatial-resolution (VHR) remote sensing images is very important and challenging in the field of remote sensing, as the building is one of the most significant and valuable man-made ground targets. This article proposes a local–global pyramid network (LGPNet) that combines a local feature pyramid module (LFPM) and a global spatial pyramid module (GSPM) for various building feature extraction. The LFPM is constructed using the convolutional kernel with three different pyramid scales, and then, the local pyramid features are obtained by adding features of each scale. In the GSPM, the global spatial pyramid features are extracted by adaptive average pooling to acquire global contextual information from different fields of view on deep features. The LFPM and the GSPM work in a parallel and complementary manner to capture discriminative features of various buildings. In addition to the LFPM and the GSPM, the proposed LGPNet also employs two general attention mechanisms, i.e., the position attention module and the channel attention module, which can select and emphasize adaptively some building features with high semantic responses. Besides, in order to mitigate the influence of other ground targets to a certain extent, a cross-task transfer learning strategy is introduced to make the LGPNet focus on the building, which significantly improves the performance of our method. Extensive experiments on two public available BCD datasets show that the proposed LGPNet can achieve significant improvement compared with eight other state-of-the-art methods. The source code and the pretrained model will be released athttps://github.com/TongfeiLiu/LGPNet.
Tongfei Liu, Maoguo Gong, Di Lu 0004, Qingfu Zhang 0001, Hanhong Zheng, Fenlong Jiang, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.7
2022 Self-Supervised Monocular Depth Estimation With Multiscale Perception
abstract
Extracting 3D information from a single optical image is very attractive. Recently emerging self-supervised methods can learn depth representations without using ground truth depth maps as training data by transforming the depth prediction task into an image synthesis task. However, existing methods rely on a differentiable bilinear sampler for image synthesis, which results in each pixel in a synthetic image being derived from only four pixels in the source image and causes each pixel in the depth map to perceive only a few pixels in the source image. In addition, when calculating the photometric error between a synthetic image and its corresponding target image, existing methods only consider the photometric error within a small neighborhood of each single pixel and therefore ignore correlations between larger areas, which causes the model to tend to fall into the local optima for small patches. In order to extend the perceptual area of the depth map over the source image, we propose a novel multi-scale method that downsamples the predicted depth map and performs image synthesis at different resolutions, which enables each pixel in the depth map to perceive more pixels in the source image and improves the performance of the model. As for the locality of photometric error, we propose a structural similarity (SSIM) pyramid loss to allow the model to sense the difference between images in multiple areas of different sizes. Experimental results show that our method achieves superior performance on both outdoor and indoor benchmarks.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Mingyang Zhang 0002, Fenlong Jiang, Hongyu Zhao 0007
IEEE Trans. Image Process.4
2020 Multi-objective optimization for location-based and preferences-aware recommendation
Shanfeng Wang, Maoguo Gong, Yue Wu 0004, Mingyang Zhang 0002
Inf. Sci.4
2020 Multiobjective Endmember Extraction Based on Bilinear Mixture Model
abstract
Hyperspectral imagery is always composed of mixed pixels because of the limited spatial resolution of a sensor and the macroscopic/microscopic mixture of distinct substances. The linear mixing model (LMM) is proven to be simple and effective in extensive literature when the macroscopic mixture dominates the mixing process. But when the photons undergo multiple reflections before reaching the sensor, the LMM becomes invalid. In this circumstance, the bilinear mixture model (Bi-LMM), which considers secondary reflections with a bilinear term, is a viable alternative. However, the bilinear term in most existing Bi-LMMs is constructed based on the pre-estimated endmembers, and thus, most Bi-LMMs focus mainly on the abundance estimation. This may lead to inaccurate estimation of endmembers and abundances for a given hyperspectral image. In this article, we propose a multiobjective endmember extraction (Bi-MoEE) method within the bilinear mixture paradigm, which considers each secondary reflection as a virtual endmember. Then, Bi-MoEE selects real and virtual endmembers from an extended spectral library consisting of a standard spectral library and their virtual products. By imposing some intuitive constraints, the solution space is greatly reduced, and the multipoint crossover and restricted bit-flip mutation operators are specially designed. Finally, Bi-MoEE can efficiently obtain a set of tradeoff solutions by minimizing the unmixing residuals and the number of selected endmembers, and automatically determine the optimal solution with multiobjective decision-making techniques. Compared with some advanced endmember extraction methods, the proposed Bi-MoEE does not need to know the number of real endmembers. In addition, the time efficiency of Bi-MoEE is mainly related to the image size and the algorithmic parameters, and has little to do with the size of spectral library, thus facilitating the practical implementation of Bi-MoEE with regard to the oversized spectral library. The experiments on synthetic and real data sets demonstrated the excellent performance of Bi-MoEE.
Xiangming Jiang, Maoguo Gong, Tao Zhan 0005, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2020 Unsupervised Scale-Driven Change Detection With Deep Spatial-Spectral Features for VHR Images
abstract
The rapid development of remote sensing technology has enabled the acquisition of very high spatial resolution (VHR) multitemporal images in Earth observation. However, how to effectively exploit these existing data to accurately monitor land surface changes is still a challenging task. In this article, we propose an unsupervised scale-driven change detection (CD) framework for VHR images by jointly analyzing the spatial-spectral change information, which combines the advantages of deep feature learning and multiscale decision fusion. First, a well pretrained deep fully convolutional network (FCN) is used to automatically extract the deep spatial context information from the acquired images. Then, the uncertainty analysis incorporating the deep spatial feature and the image spectral feature is implemented to generate a pseudobinary change map. On this basis, it is easy to choose suitable samples to train an excellent support vector machine (SVM) classifier, thus detecting changes occurred on the ground. In addition, the multiscale superpixel segmentation technique is introduced to make full use of the spatial structural information, which takes an image-object as the basic analysis unit. Finally, a robust binary change map with high detection precision can be achieved by merging the CD results obtained at different scales. The impressive experimental results on four real data sets demonstrate the effectiveness and flexibility of the proposed framework.
Tao Zhan 0005, Maoguo Gong, Xiangming Jiang, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2019 Evolutionary Multiobjective Change Detection via Self-paced Learning and Fuzzy Clustering
abstract
Fuzzy clustering algorithm based on multiobjective optimization can achieve accurate and comprehensive clustering results. However, the estimation of objective values for this multiobjective optimization problem (MOP) might be expensive. Offspring's selection driven by simple evaluation is time consuming. Therefore, we integrate regression techniques to determine the superiority of the offspring solutions in the evolution process. However, it suffers from an issue that it is hard to collect reliable samples to train such a robust regression model. In this paper, an evolutionary multiobjective fuzzy clustering method via self-paced learning is proposed for change detection. In the proposed method, the self-paced learning process is implemented to collect reliable training samples for training a robust regression model, which can help to select promising offspring solutions from the candidate solutions for MOP. Experiments on three remote sensing image datasets demonstrate that the proposed method can significantly outperform those state-of-art methods for change detection in terms of accuracy and robustness.
Yingying Duan, Jingjing Ma 0001, Hao Li 0009, Mingyang Zhang 0002, Zedong Tang, Maoguo Gong
CEC4
2019 Unsupervised Feature Extraction in Hyperspectral Images Based on Wasserstein Generative Adversarial Network
abstract
Feature extraction (FE) is a crucial research area in hyperspectral image (HSI) processing. Recently, due to the powerful ability of deep learning (DL) to extract spatial and spectral features, DL-based FE methods have shown great potentials for HSI processing. However, most of the DL-based FE methods are supervised, and the training of them suffers from the absence of labeled samples in HSIs severely. The training issue of supervised DL-based FE methods limits their application on HSI processing. To address this issue, in this paper, a novel modified generative adversarial network (GAN) is proposed to train a DL-based feature extractor without supervision. The designed GAN consists of two components, which are a generator and a discriminator. The generator can focus on the learning of real probability distributions of data sets and the discriminator can extract spatial-spectral features with superior invariance effectively. In order to learn upsampling and downsampling strategies adaptively during FE, the proposed generator and discriminator are designed based on a fully deconvolutional subnetwork and a fully convolutional subnetwork, respectively. Moreover, a novel min-max cost function is designed for training the proposed GAN in an end-to-end fashion without supervision, by utilizing the zero-sum game relationship between the generator and discriminator. Besides, the proposed modified GAN replaces the original Jensen-Shannon divergence with the Wasserstein distance, aiming to mitigate the unstability and difficulty of the training of GAN frameworks. Experimental results on three real data sets validate the effectiveness of the proposed method.
Mingyang Zhang 0002, Maoguo Gong, Yishun Mao, Jun Li 0009, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.1
2018 A Two-Phase Multiobjective Sparse Unmixing Approach for Hyperspectral Data
abstract
With the sparse unmixing becoming increasingly popular recently, some advanced regularization algorithms have been proposed for settling this problem. However, they are limited by their “decision ahead of solution” attribute, i.e., the regularization parameters must be preset before the solution is obtained. In this paper, the sparse unmixing problem is first formulated as a two-phase multiobjective problem. The first phase simultaneously minimizes the unmixing residuals and the number of estimated endmembers for automatically finding the real active endmembers from the spectral library. A decomposition-based endmember selection algorithm considering the gene exchange in the population is specially designed for better and quicker search of the decision space. This algorithm can obtain a set of nondominated solutions for better decision of the active endmembers, which are important for the subsequent calculation of the abundance matrix. The second phase concurrently minimizes the unmixing residuals and the total variation term for estimating a preferable abundance matrix. A local search strategy based on the multiplicative update rule is designed in the evolution process for better approximation of the Pareto front. The experimental results on the synthetic as well as the real data reveal that the proposed framework has a better performance in finding the real active endmembers and estimating their corresponding abundances than some advanced regularization algorithms.
Xiangming Jiang, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.4
2017 Multi-objective endmember extraction for hyperspectral images
abstract
Endmember extraction is a critical step of spectral unmixing. In this paper, a novel endmember extraction algorithm based on evolutionary multi-objective optimization is proposed for hyperspectral remote sensing images. In the proposed method, endmember extraction is modeled as a multi-objective optimization problem. Then the root mean square error between the original image and its remixed image and the number of endmembers are chosen as two conflicting objective functions, which are simultaneously optimized by particle swarm optimization algorithm to find the best tradeoff solutions. In order to promote diversity and speed up the convergence of the algorithm, a new particle status updating strategy and a novel method for selecting leaders are designed. The experimental results on both simulated and real hyperspectral remote sensing images confirm the performance of the proposed approach over some existing methods.
Hao Li 0009, Jingjing Ma 0001, Jia Liu 0020, Maoguo Gong, Mingyang Zhang 0002
CEC5
2017 An improved multiobjective evolutionary approach for community detection in multilayer networks
abstract
The detection of shared community structure in multilayer network is an interesting and important issue that has attracted many researches. Traditional methods for community detection of single layer networks are not suitable for that of multilayer networks. In a previous work, the authors modeled the community discovery problem in multilayer network as a multiobjective one and devised a genetic algorithm to carry out it. In this paper, based on their model, we propose an improved multiobjective evolutionary approach MOEA-MultiNet for community detection in multilayer networks. The proposed MOEA-MultiNet is based on the framework of NSGA-II which employs the string-based representation scheme and synthesizes the genetic operation and local search to perform individual refinement. Experimental results on two real-world networks both demonstrate the ability and efficiency of the proposed MOEA-MultiNet in detecting community structure in multilayer networks.
Shanfeng Wang, Maoguo Gong, Mingyang Zhang 0002
CEC4
2017 Evolutionary multi-task learning for modular extremal learning machine
abstract
Evolutionary multi-tasking is a novel concept where algorithms utilize the implicit parallelism of population-based search to solve several tasks efficiently. In last decades, multi-task learning, which harnesses the underlying similarity of the learning tasks, has proved efficient in many applications. Extreme learning machine is a distinctive learning algorithm for feed-forward neural networks. Because of its similarity and low computational complexity comparing with the convenient neural network training algorithms, it has been used in many cases of data analyses. In this paper, a modular training technique by employing evolutionary multi-task paradigm is used to evolve the modular topologies of extreme learning machine. Though, extreme learning machine is much faster than the convenient gradient-based method, it needs more hidden neurons due to the random determination of input weights. In proposed method, we combine the evolutionary extreme learning machine and multi-task modular training. Each task is defined by an evolutionary extreme learning machine with different number of hidden neurons. This method produces a modular extreme learning machine which needs less number of hidden units and could be effective even if some hidden neurons and connections are removed. Experiment results show effectiveness and generalization of the proposed method for benchmark classification problems.
Zedong Tang, Maoguo Gong, Mingyang Zhang 0002
CEC3
2017 Memetic algorithm based feature selection for hyperspectral images classification
abstract
Band selection is a crucial preprocessing step for hyperspectral image classification, which is a classic feature selection method. Feature selection is designed to select feature subsets to represent the whole feature space. For feature selection, two crucial issues need to be handled: preserving information and redundancy reducing. In this paper, a novel feature selection method for hyperspectral image classification is proposed, which is based on a newly designed memetic algorithm. In the proposed method, a suitable objective function is designed, which can measure the contained crucial information and redundancy information in the selected feature subsets. To optimize this objective function efficiently, a novel memetic algorithm is designed. The genetic operator and local search strategy are newly designed according to the characteristic of hyperspectral images. Experiments are implemented on three real data sets compared with some state of arts. The experimental results show that the proposed method can obtain stable and superior feature subsets for classification.
Mingyang Zhang 0002, Jingjing Ma 0001, Maoguo Gong, Hao Li 0009, Jia Liu 0020
CEC1
2017 Unsupervised Hyperspectral Band Selection by Fuzzy Clustering With Particle Swarm Optimization
abstract
Due to the lack of label information and the intrinsic complexity of hyperspectral images (HSIs), unsupervised band selection is always one of the most challenging tasks in HSI processing. Fuzzy clustering is a promising technique for unsupervised band selection, which can partition unlabeled data into groups effectively. However, due to the limits of its optimization process, standard fuzzy clustering is sensitive to initialization and easy to be trapped in a local optimum. To address the limits, a novel unsupervised band selection method is proposed, combining fuzzy clustering with particle swarm optimization (PSO). A newly designed PSO algorithm is introduced to improve the performance of fuzzy clustering band selection. Moreover, a new strategy is designed to select representative cluster centers according to the characteristics of HSIs. The experimental results indicate that the proposed method has the ability to select high-quality band subsets with good and robust performance on HSI classification.
Mingyang Zhang 0002, Jingjing Ma 0001, Maoguo Gong
IEEE Geosci. Remote. Sens. Lett.1
2017 Deep learning and mapping based ternary change detection for information unbalanced images
Linzhi Su, Maoguo Gong, Puzhao Zhang, Mingyang Zhang 0002, Jia Liu 0020, Hailun Yang
Pattern Recognit.4
2016 Unsupervised Band Selection Based on Evolutionary Multiobjective Optimization for Hyperspectral Images
abstract
Band selection is an important preprocessing step for hyperspectral image processing. Many valid criteria have been proposed for band selection, and these criteria model band selection as a single-objective optimization problem. In this paper, a novel multiobjective model is first built for band selection. In this model, two objective functions with a conflicting relationship are designed. One objective function is set as information entropy to represent the information contained in the selected band subsets, and the other one is set as the number of selected bands. Then, based on this model, a new unsupervised band selection method called multiobjective optimization band selection (MOBS) is proposed. In the MOBS method, these two objective functions are optimized simultaneously by a multiobjective evolutionary algorithm to find the best tradeoff solutions. The proposed method shows two unique characters. It can obtain a series of band subsets with different numbers of bands in a single run to offer more options for decision makers. Moreover, these band subsets with different numbers of bands can communicate with each other and have a coevolutionary relationship, which means that they can be optimized in a cooperative way. Since it is unsupervised, the proposed algorithm is compared with some related and recent unsupervised methods for hyperspectral image band selection to evaluate the quality of the obtained band subsets. Experimental results show that the proposed method can generate a set of band subsets with different numbers of bands in a single run and that these band subsets have a stable good performance on classification for different data sets.
Maoguo Gong, Mingyang Zhang 0002, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.2
2015 Unsupervised Hyperspectral Image Band Selection via Column Subset Selection
abstract
In this letter, we proposed a novel band selection algorithm for hyperspectral images (HSIs) based on column subset selection. The main idea of the proposed algorithm comes from the column subset selection problem in numerical linear algebra. It selects a group of bands, which maximizes the volume of the selected subset of columns. Since the high dimensionality decreases the contrast between bands, we use Manhattan distance to obtain a higher selection quality. Experimental results on real HSIs show that the proposed algorithm obtains competitively good results, in terms of classification accuracy, and is robust to noisy bands.
Maoguo Gong, Mingyang Zhang 0002, Yongqiang Chan
IEEE Geosci. Remote. Sens. Lett.3