Yu Zhou 0051

dblp:36/2728-51 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
21since 2021 · last 2026
0009-0006-9950-4728ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Adaptive Personalized Federated Meta-Learning Recommendation
abstract
Federated recommendation systems aim to provide high-quality recommendations while protecting user privacy. However, existing federated recommendation algorithms typically use a unified item embedding framework, where all clients share the same item representations. Such approaches often fail to capture users’ personalized perceptions of the same item and struggle to reflect subtle changes in user preferences, limiting the effectiveness of personalized recommendations. Moreover, statistical heterogeneity among clients poses additional challenges for model optimization. To address these issues, we propose APFMRec, a personalized federated recommendation framework that enhances both personalization and global optimization. APFMRec incorporates an adaptive item embedding module to dynamically adjust item representations based on individual user preferences. In addition, a meta-learning update module is designed to mitigate statistical heterogeneity and improve collaborative optimization across clients. Extensive experiments on five real-world datasets demonstrate the effectiveness of APFMRec. The proposed framework consistently outperforms existing federated recommendation methods, achieving up to 8.1% improvement in HR@10 and 7.8% improvement in NDCG@10.
Shanfeng Wang, Shanyang Gao, Lanyu Yao, Maoguo Gong, Ke Pan 0001, Yu Zhou 0051
IEEE Trans. Comput. Soc. Syst.6
2026 Sparse Unmixing Guided Adversarial Attack for Hyperspectral Image Classification
abstract
In recent years, adversarial attacks in hyperspectral image (HSI) classification have garnered increasing attention. However, existing attack methods primarily manipulate individual pixel spectral to mislead deep neural networks (DNNs) into misclassification, overlooking the physical consistency of hyperspectral data. This oversight results in adversarial samples that lack physical interpretability and suffer from low attack efficiency. To alleviate these issues, this paper proposes a sparse unmixing guided adversarial attack framework (SUGAA) to efficiently generate hyperspectral adversarial samples that satisfy physical consistency. The proposed framework first employs sparse unmixing to extract the abundance matrix of HSI, introducing adversarial perturbations to the abundance matrix to generate physically consistent adversarial samples. Additionally, SUGAA leverages the compositional similarity of materials within intra-class HSI pixels to design a class-specific perturbation generation strategy, enhancing the applicability of adversarial perturbations across pixels of the same class. To further improve optimization effectiveness, SUGAA incorporates a class-specific perturbation optimization algorithm based on momentum iterative gradients to avoid local optima, ensuring stable and efficient perturbation generation. Experimental results on real HSI datasets demonstrate that SUGAA not only generates adversarial samples with high attack performance and physical consistency but also exhibits robustness to common preprocessing transformations.
Hao Li 0009, Kelin Dang, Maoguo Gong, A. K. Qin 0001, Yu Zhou 0051, Yue Wu 0004, Lining Xing 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Uncertainty-Aware Local Bayesian Framework for Hyperspectral Image Classification With Noisy Labels
abstract
Various deep learning-based methods have greatly improved hyperspectral image (HSI) classification performance, but these models are sensitive to noisy training labels. Human annotation on remote sensing images inevitably introduced label noise, which degrades the model prediction confidence. Understanding the spatial characteristics and distribution of such annotation errors is crucial for both diagnosing dataset annotation failures and guiding effective robust learning strategies. Current noisy label learning methods pay limited attention to visualizing noise label distributions, and these approaches often exhibit poor compatibility with noise-free models. Leveraging the relationship between the prediction uncertainty and label noise, we propose a Local Bayesian Framework (LBF) for noisy HSI classification and noise labels awareness. LBF adapts standard CNN, GCN, or Transformer backbones via local Bayesian adaptation (LBA) to evaluate prediction uncertainty and employs an uncertainty-monitoring optimization strategy (U-MOS) for training. Without major architectural changes, LBF delivers accurate uncertainty maps that highlight noisy regions, suppresses overfitting to corrupted labels, and consistently improves classification robustness across four benchmark HSI datasets.
Mingyang Zhang 0002, Ziqi Di, Hao Liu 0123, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Circuits Syst. Video Technol.5
2026 Multigranularity Adversarial Attacks on Large Language Models Using Genetic Programming
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks, but they remain vulnerable to adversarial attacks and pose significant security concerns. Existing attack methods often treat adversarial prompts as flat sequences, neglecting the rich hierarchical structure of natural language, which could limit their effectiveness. Advancing the methodologies for adversarial attacks is crucial for rigorously assessing the security of LLMs and identifying subtle vulnerabilities. This paper introduces AdvGP, a novel framework that leverages genetic programming (GP) to generate adversarial prompts for LLMs. AdvGP exploits the inherent structural similarities between GP trees and natural language syntax to optimize the structure of harmful prompts. The framework incorporates a multi-granularity hierarchical attack strategy, specialized genetic operators that leverage an assisting LLM for depth-aware crossover and multi-level mutation, and a comprehensive fitness function integrating semantic consistency and attack effectiveness. The proposed method achieves competitive attack performance on multiple LLMs, consistently generating harmful outputs despite higher perplexity than some baselines. Ablation studies confirm the significant contributions of both LLM-aided and depth-aware mechanisms to AdvGP’s effectiveness. Furthermore, transferability analysis reveals that the generated prompts are able to bypass the defenses of various state-of-the-art LLMs, such as ChatGPT and Gemini.
Wencheng Han, Hao Li 0009, Maoguo Gong, Yu Zhou 0051, Yue Wu 0004, A. K. Qin 0001, Lining Xing 0001
IEEE Trans. Evol. Comput.4
2026 Privacy-Enhanced Offline Data-Driven Evolutionary Optimization Based on Cloud Server
abstract
Data-driven evolutionary algorithms (DDEAs) have achieved significant success in numerous real-world optimization problems, where exact objective functions and constraint functions do not exist, and they mainly rely on available data. However, the existing DDEAs primarily focus on improving performance through data and surrogate, without considering that the users may lack the specialized domain knowledge and sufficient computing resources required for DDEAs. To address the aforementioned issues, this paper proposes a novel paradigm called Evolutionary Learning and Optimization as a Service (ELOaaS) and investigates the potential collusion attacks between machine learning modules and evolutionary computing modules on cloud server, which may lead to privacy leakage. Consequently, a privacy-enhanced DDEA (PEDDEA) is proposed as an instantiation algorithm of ELOaaS, which is designed to tackle offline data-driven evolutionary optimization within the ELOaaS paradigm. In the proposed PEDDEA, a subspace learning-based privacy protection strategy is designed to defense the collusion attacks. Additionally, a model management strategy based on Kendall tau metric is introduced to construct high-quality surrogate ensembles. PEDDEA enables users to outsource private offline data to cloud servers, thereby approaching the optimal solution while ensuring privacy protection. Comprehensive experiments are conducted on benchmark problems and safety evaluation problems of autonomous vehicles. According to the experimental results, the proposed algorithm has significant performance advantages over existing offline DDEAs while ensuring privacy protection.
Hao Li 0009, Zhibin Xu, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001, Yu Zhou 0051
IEEE Trans. Evol. Comput.7
2025 STEAM: Style Transfer Enabled Adversarial Attack With Attention Mechanism on Remote Sensing Image Scene Classification
abstract
Research on adversarial attacks in remote sensing tasks have predominantly focused on designing perturbations or patches, presenting challenges in balancing attack success rate and adversarial stealthiness. Instead of focusing on designing adversarial examples under adversarial perturbation constraints to ensure stealthiness, this paper proposes a Style Transfer Enabled Adversarial Attack with Attention Mechanism (STEAM), which leverages style transfer to generate adversarial examples with high visual fidelity. Specifically, STEAM transfers distinctive styles from critical regions of source samples to attackable areas in target samples, effectively incorporating natural textures from source samples. To further refine this process, an attention mechanism is introduced to selectively extract style features from key regions of the source samples, mitigating redundancy from global style information. Additionally, selective style transfer process also includes the consideration of semantic features across different regions in target samples, ensuring a more effective attack area selection. As a result, STEAM achieves high attack success rate by utilizing proper style selected from groups of style samples, and preserving high visual fidelity through selectively transfer the natural style feature into specific attackable region in target samples. Experimental results on the UCM and WHU-RS19 datasets demonstrate that STEAM not only enhances the visual fidelity of adversarial examples but also improves the attack success rate. Furthermore, experiments against state-of-the-art adversarial defense methods highlight the adversarial attack effectiveness and robostness of STEAM compared to other adversarial attack methods.
Tianshi Luo, Hao Li 0009, Maoguo Gong, Yu Zhou 0051, A. K. Qin 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Spatial-Spectral Aggregation Transformer With Diffusion Prior for Hyperspectral Image Super-Resolution
abstract
Constrained by imaging systems, hyperspectral images (HSIs) always have a low spatial resolution. Deep learning-based HSI super-resolution methods have achieved impressive results through learning the nonlinear mapping between low-resolution (LR) and high-resolution (HR) images. However, most of them take the LR image or its upsampled version through bicubic interpolation as input, leading to low-quality features and limited details captured by the network. As a powerful generative model, diffusion model has the ability to learn both contextual semantics and textual details from distinct timesteps, enabling the effective exploration of spatial-spectral distributions in high-dimensional data. In this paper, we propose a novel method that extracts high-quality prior information from original images to assist in super-resolution through pretraining a diffusion model. Specifically, we first train a diffusion model using original HSI patches in a self-supervised manner and then obtain prior features from the pretrained denoising U-Net decoder. To efficiently incorporate the prior features into the super-resolution model, we propose an adaptive fusion module based on spatial and spectral attention mechanisms, which enhances features in both dimensions while preserving the original characteristics. Additionally, to leverage the complementarity of spatial and spectral information, we design a spatial-spectral aggregation Transformer module that incorporates an adaptive interaction module to facilitate information exchange across different dimensions, thereby enhancing the representation capability. Extensive experiments on three public hyperspectral datasets demonstrate that the proposed method achieves excellent super-resolution performance and outperforms the state-of-the-art methods in terms of quantitative quality and visual results.
Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yu Zhou 0051, Fenlong Jiang, Yue Wu 0004
IEEE Trans. Circuits Syst. Video Technol.6
2025 A General Uncertainty-Guided Bayesian Adaptation Framework for Building Change Detection
abstract
Existing deep learning-based building change detection (BCD) methods are often hindered by sample imbalance and imagery noise, which leads to inaccurate predictions, particularly for building edges and minority changed class regions. To overcome these limitations, we propose a novel Uncertainty-guided Bayesian Adaptation (UBA) framework, designed as a plug-and-play module to enhance existing BCD methods. The UBA framework consists of two core components. First, a Local Bayesian Adaptation strategy (LBs) pragmatically adapts the output layer of any BCD network, enabling efficient prediction uncertainty estimation and decomposition. We demonstrate that the decomposed aleatoric and epistemic uncertainty terms semantically highlight building edges and minority changed class regions, respectively. Based on this insight, we propose an Uncertainty-Weighted Optimization Mechanism (U-Wom) that leverages these uncertainty maps to dynamically re-weight the loss function. This mechanism guides the model to focus its learning on these challenging, fine-grained regions. Extensive experiments on several widely-used BCD datasets show that the UBA framework consistently and significantly improves the performance of various state-of-the-art methods.
Ziqi Di, Mingyang Zhang 0002, Fenlong Jiang, Yu Zhou 0051, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.4
2025 Change Masked Modality Alignment Network for Multimodal Change Detection
abstract
Using multimodal remote sensing images for change detection (CD) can significantly improve the feasibility and reliability in challenging environments. However, the differences in imaging mechanisms make multimodal images highly heterogeneous. A key challenge for multimodal CD (MCD) is that the heterogeneity of the modalities and changes in ground objects are intertwined during processing. To address this issue, this article proposes a change masked modality alignment network (CMMAN), which uses a multitask framework consisting of one CD branch and two image modal transformation (IMT) branches. Specifically, to ensure a unified feature space, bi-temporal multimodal images are first input into the same Swin-Transformer-based encoder. The extracted features are then fed simultaneously into the CD branch and separately into the two IMT branches. In the CD branch, the decoder is also designed based on the Swin-Transformer, and a weakly modality-correlated feature enhancement (WMCFE) module is introduced to mitigate the interference of modality heterogeneity on CD. For the two IMT branches, both employ a generative adversarial network (GAN) to transform between modalities, and the distributions of features from different modalities are aligned through simultaneous optimization. Uniquely, the change probability map predicted by the CD branch is utilized to mask the change regions in IMT, further decoupling ground object changes and modal heterogeneity. Experimental results on multiple public datasets demonstrate that the proposed CMMAN significantly improves MCD performance and shows good compatibility and portability with various common backbone networks.
Fenlong Jiang, Husheng Wu, Dan Feng 0002, Yu Zhou 0051, Mingyang Zhang 0002, Maoguo Gong, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.5
2025 D3PM: Dual-Stream Denoising Diffusion Probabilistic Model for Change Detection in Multimodal Remote Sensing Images
abstract
Detecting land cover changes from multi-temporal and multi-modal remote sensing images acquired by different sensors at the same location is a complex yet highly valuable task. Recently, diffusion models, exemplified by the Denoising Diffusion Probabilistic Model (DDPM), have garnered significant attention for their remarkable performance and straightforward architecture. These models excel in image generation, distribution modeling, and feature extraction, making them highly promising for advancing Multimodal Change Detection (MCD). In this paper, we propose a Dual-stream Denoising Diffusion Probabilistic Model (D3PM) to address the challenges of MCD. Specifically, D3PM leverages DDPM to design two distinct processing streams, one for each image modality. The first stream employs an unconditional DDPM, whose denoising encoder-decoder network can achieve robust feature extraction. The second stream employs a conditional DDPM to facilitate modal translation, enabling the extracted features to align with the characteristics of the other modality, thereby improving cross-modality comparability. To further enhance performance, we constructed a CD task branch based on the decoder features of the two DDPMs across multiple denoising time steps. Additionally, we designed a collaborative learning optimization strategy with asynchronous time steps, fostering cross-task knowledge sharing and mutual enhancement while preserving the integrity of individual task learning. Experimental results on multiple public datasets demonstrate the effectiveness and superiority of the proposed D3PM, which achieves efficient modal transformation and alignment, mitigates modal heterogeneity interference, and significantly improves detection performance.
Fenlong Jiang, Xinlong Huo, Mingyang Zhang 0002, Maoguo Gong, Yan Pu, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.6
2025 Adaptive Center-Focused Hybrid Attention Network for Change Detection in Hyperspectral Images
abstract
Hyperspectral images (HSIs) capture extensive spatial and spectral information, facilitating detailed change detection (CD) of complex land covers. However, the high correlation among spectral data can lead to information redundancy, increasing processing dimensions and introducing irrelevant or detrimental data to CD. To address these challenges, we propose an adaptive center-focused hybrid attention network (ACFHAN) for CD in HSIs. This network adaptively emphasizes the spatial regions and spectral channels most pertinent to CD while suppressing irrelevant information. The architecture establishes an end-to-end mapping from the two HSIs to the change results, featuring multiple center-focused hybrid attention blocks (CFHABs). Each CFHAB integrates two different attention modules, including an adaptive spatial–spectral hybrid self-attention (S2HSA) module that dynamically adjusts spatial–spectral feature weights and a center-focused attention (CFA) module that enhances the area most relevant to the center pixel to be classified. Additionally, to tackle the challenges of expensive labeling, we further designed a multiscale superpixel-based data augmentation method which combines traditional unsupervised and supervised methods to provide sufficient low-cost but high-confidence labeled data for CD. Experimental results across various HSI CD datasets validate the effectiveness of our proposed method.
Fenlong Jiang, Shining Zhang, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.5
2025 BARNet: Boundary-Aware Refinement Network for Weakly Supervised Change Detection
abstract
Remote sensing change detection plays a critical role in urban land management and disaster assessment. However, most existing methods rely on expensive and time-consuming pixel-level labels, limiting their practical applicability. Weakly supervised change detection methods, such as only using image-level labels, offer the potential to reduce annotation costs while maintaining robust detection performance. However, this coarse supervisory information often makes it difficult to accurately capture fine-grained details, resulting in poor pixel-level detection accuracy. To overcome these challenges, we propose a novel Boundary-Aware Refinement Network (BARNet) for weakly supervised change detection, which utilizes a two-stage framework that first generates pixel-level pseudo labels via image-level CD activation maps, then subsequently trains a pixel-level CD network using these generated pseudo labels. Specifically, the first stage adopted a teacher-student distillation image-level CD network, which integrated a multi-scale boundary feature attention module, along with activation ambiguity loss and contrastive learning loss as feature separation constraints, to generate high-quality pseudo labels. In the second stage, these pseudo labels are used to provide deep supervision a pixel-level CD network, where the boundary-aware decoupling module further refines boundary information, leading to more precise segmentation of change areas. Extensive experiments on three public datasets demonstrate that BARNet not only achieves state-of-the-art performance in the weakly supervised change detection domain but also shows competitive performance with existing fully supervised methods, significantly reducing annotation costs while maintaining detection accuracy. With its strong performance, BARNet demonstrates great potential for practical applications in scenarios with limited supervision.
Fenlong Jiang, Zikang Zhong, Mingyang Zhang 0002, Maoguo Gong, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan
IEEE Trans. Geosci. Remote. Sens.5
2025 Physical Adversarial Background Patch Against Aerial Object Detection Based on Pareto Efficiency
abstract
For adversarial attacks on aerial image object detection, some physical background attack methods have been proposed and demonstrated excellent performance. However, most of the methods restrict the protected object to a fixed area, and changes in the position and size of the object can affect the effectiveness of the attack. In addition, the size of the background is only related to the size of the object, without considering any reduction in the background area. To alleviate these issues, this paper proposes a novel adversarial background patch generation strategy, which generates an adversarial background patch where aircrafts parked at any position within it remain undetected. Besides we also aim to make the size of the background as small as possible to improve the concealment of the attack and reduce the economic cost. Specifically, a series of adversarial images are generated by placing the aircraft at random angles and positions on the adversarial background patch. Then, a novel background patch optimization strategy is proposed, which enhances the occlusion robustness of the adversarial background patch by weighting the adversarial loss of each image in the batch based on adversarial difficulty. In addition, this paper proposes a novel area loss for achieving optimal area size, which is used to reduce the cost of producing the patch. Finally, a multi-objective optimization method based on Pareto efficiency is introduced for balancing two conflicting losses, the adversarial loss and the area loss. Experimental results show that the adversarial background patches generated by this method have excellent attack effects in both digital and physical attacks, exhibiting strong occlusion robustness. In addition of that, the adversarial background patches generated by this method achieve effective attacks in a smaller area and reduce the physical implementation cost.
Hao Li 0009, Jiachang Li, Maoguo Gong, Haiyue Yu 0001, Kelin Dang, Yu Zhou 0051, A. K. Qin 0001, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.6
2025 Toward Federated Customized Neural Architecture Search for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scenarios usually involve sensitive geographic information on national security and regional development. In the commonly used centralized machine-learning paradigm, data dispersed in various locations are concentrated and processed on a single server, which is prone to privacy leakage and data security concerns. Besides, it is difficult to solve the high heterogeneity of RS images by simply applying federated learning (FL) algorithms to scene classification. In this article, we formulate a federated remote sensing scene classification (FedSC) framework, and design a customized neural architecture search (CNAS) to achieve both global generality for multiparty collaborative distributed training and local specificity for personalized RS scene customization. The proposed FedSC is generalizable to be implemented in any manually designed networks, network pruning strategies, or NAS methods related to remote sensing scene classification (RSSC). While the designed CNAS not only achieves collaborative distributed training in protecting participant data privacy to obtain a generalized global model, but also provides a customized local model for each participant that is more in line with the characteristics of private RS scenarios. Overall, the proposed FedSC$_{\textrm {CNAS}}$provides a novel federated collaborative training paradigm for RSSC in terms of data privacy, data heterogeneity, and personalized customization. Extensive analytical and comparative experiments on three benchmark RSSC datasets validate the versatility and effectiveness of our methods, and the proposed FedSC$_{\textrm {CNAS}}$exhibits superior competitiveness compared to state-of-the-art methods.
Jianzhao Li, Shanfeng Wang, Maoguo Gong, Zhuping Hu, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.8
2025 Meta-Collaborative Learning for Arbitrarily Scaled Hyperspectral Image Super-Resolution
abstract
Deep learning-based methods for hyperspectral image super-resolution (SR) have achieved significant success in recent years. These methods typically consist of feature extraction module (FEM) and upsampling module. However, due to structural limitations of the upsampling module, most current methods focus on training separate models for different scale factors, which ignores the exploration of potential feature interdependence among different scale factors. In response to these challenges, we introduce a novel framework, called “meta-collaborative learning for arbitrarily scaled hyperspectral image super-resolution” (MCArb). Specifically, MCArb integrates a collaborative learning framework with a meta-learning-based 3-D upsampling module (3DMetaUM) and a scale-aware feature adaptation module (SAFAM). It enables training multiple SR tasks at different scale factors within a single network at the same time. This strategy is able not only to process arbitrary-scale-factor SR for hyperspectral images but also to harness the latent feature interdependence among different scales. In this study, we applied the MCArb framework to transform three deep learning-based hyperspectral image SR networks to MCArb methods, resulting in significant performance enhancements across five hyperspectral datasets. These improvements showcase the proposed MCArb framework’s ability to enhance feature extraction efficiency and capitalize on latent interscale correlations. This code is available athttps://github.com/ShuangWu-XDU/MCArb_HSI_SR.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Yu Zhou 0051, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.6
2025 Evolutionary Multiobjective Cross-Spectral Adversarial Attacks With Synergistic Patches
abstract
DNN have demonstrated vulnerability to adversarial attacks in object detection tasks. While significant progress has been made in single-spectrum attacks, cross-spectral adversarial attacks remain challenging due to the complex tradeoffs between visible and infrared domains. To address this, an evolutionary multiobjective cross-spectral attack (MoXAttack) framework, for developing adversarial patches in closed-box cross-spectral scenarios is proposed. MoXAttack incorporates a multipopulation constraint-handling technique, which uses both penalty functions and feasibility rules to guide the search process. Spectrum-aware genetic operators are introduced to enhance solution diversity and feasibility. The framework automatically optimizes the smooth to cross-spectral shared patch shape using curvature energy. In addition, MoXAttack utilizes SVD for visible spectrum texture perturbations and adjustable thermal shielding material thickness for infrared spectrum control. Experiments on the LLVIP dataset demonstrate that MoXAttack achieves competitive performance across multiple object detection models. Ablation studies reveal the positive impact of improved components on attack effectiveness. The multipatch strategy improves attack success rates by at least 17%, while optimized patch shapes outperform conventional geometric shapes by at least 25% in terms of mAP drop. In the physical world test, the proposed method shows stability in different viewing angles.
Wencheng Han, Hao Li 0009, Maoguo Gong, Yue Wu 0004, A. K. Qin 0001, Lining Xing 0001, Yu Zhou 0051
IEEE Trans. Syst. Man Cybern. Syst.7
2024 Unsupervised Domain Adaptation for Cross-Scene Hyperspectral Image Classification Based on Decoupled Contrastive Learning
abstract
Recent studies have highlighted the effectiveness of deep domain adaptation (DA) techniques in addressing cross-scene hyperspectral image (HSI) classification challenges. However, most of the existing DA methods often prioritize aligning data distributions while overlooking the intrinsic separability between source and target domain data. In this paper, we propose a decoupled contrastive learning based unsupervised domain adaptation (DCLUDA) method for HSI classification. Unlike conventional adversarial DA methods, our method introduces a unique DA loss specifically designed to minimize class confusion in the target domain. This not only simplifies model training but also enhances class discriminability. Moreover, we employ a decoupled contrastive learning strategy on both domains to enhance data separability within each domain. Finally, we propose a sample selection strategy based on confident learning to select high-confidence samples from the target domain for fine-tuning the DA model. Experiments on two cross-scene HSI classification tasks shown that our proposed DCLUDA outperforms several existing DA methods.
Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Xiangming Jiang, Yu Zhou 0051, Dan Feng 0002
IJCNN6
2024 ShiftAttack: Toward Attacking the Localization Ability of Object Detector
abstract
State-of-the-art (SOTA) adversarial attacks expose vulnerabilities in object detectors, often resulting in erroneous predictions. However, existing adversarial attacks neglect the stealth and flexibility of adversarial examples, which are crucial for conducting contextually consistent and inconspicuous attacks. To address these issues, leveraging the observed phenomenon of predicted box offsets in real-world object detection scenarios, this paper presents a novel adversarial attack framework called ShiftAttack. It leverages the concept of dense detection in prevalent object detectors, by boosting the confidence of low Intersection over Union (IoU) predictions within the positive samples (the set of predicted boxes responsible for localizing the same target), which leads to the erroneous exclusion of true positive predictions during the post-processing stage. Such a paradigm is highly stealthy as the shifted predictions seem like natural detector mistakes rather than obvious manipulations. To enhance the flexibility of ShiftAttack this paper proposes a generative approach called ShiftAttack Generator (SAG), which can not only shift predicted boxes for any target in arbitrary directions and distances but also facilitate adaptive feature exchange between pre- and post-shift regions to optimize the attack. Additionally, the proposed SAG incorporates the Dynamic Hinge Loss (DHL) to ensure the imperceptibility of perturbations, effectively mitigating the Patch-Pattern associated with the use of$\mathcal {L}_{2}$norm. Extensive experiments confirm that SAG surpasses other SOTA adversarial attacks in effectiveness, speed and stealthiness.
Hao Li 0009, Maoguo Gong, Shiguo Chen, A. K. Qin 0001, Zhenxing Niu, Yue Wu 0004, Yu Zhou 0051
IEEE Trans. Circuits Syst. Video Technol.8
2024 Toward Multiparty Personalized Collaborative Learning in Remote Sensing
abstract
The powerful deep learning models in remote sensing are inseparable from the support of massive data. However, the privacy and sensitivity of remote sensing data (RSD) restrict the possibility of each party to collaboratively train and share a large general model. Although multi-party learning (MPL) is a feasible solution, it is difficult for the existing MPL methods to uniformly process different remote sensing tasks (RSTs), and the data held by each party is non-independent and identically distributed, heterogeneous and multi-sources. Therefore, it is urgent to explore a solution for the personalized processing of different RSTs. In this paper, we formulate a novel multi-party personalized collaborative learning (MPCL) framework in terms of models and tasks. Specifically, in each iteration of the communication round, we aim to decouple personalized model optimization from global model learning. Different participants are allowed to explore their personalized local models at a certain distance from the global aggregation models according to the characteristics of their local data. In terms of task personalization, MPCL provides different personalized global models to handle the corresponding RSTs. For participants with different RSTs, it can be implemented in the multi-task collaborative training strategy to explore the connection between different tasks. To demonstrate the feasibility of MPCL, we take remote sensing image classification as a case study and provide a detailed feasibility scheme. We constructed four benchmark datasets compliant with MPL and personalized MPL, including single-source and multi-source about SAR, hyperspectral and optical RSD. The experimental results demonstrate that our MPCL is superior in these four RSD, which ranked first in the competition with the classic or state-of-the-art MPL and personalized MPL algorithms. In addition, the scalability of MPCL is also verified on image segmentation RSTs of building and road extraction.
Jianzhao Li, Maoguo Gong, Zaitian Liu, Shanfeng Wang, Yourun Zhang, Yu Zhou 0051, Yuan Gao 0019
IEEE Trans. Geosci. Remote. Sens.6
2024 Collaborative Self-Supervised Evolution for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning, which leverages unlabeled data to learn useful feature representations by constructing auxiliary tasks, has been widely explored in few-shot scene classification to improve the feature representation and generalization capabilities of deep models in scarce data. However, most of the current related work adopts specific self-supervised auxiliary tasks (SSATs) for combinatorial improvement, and does not explore the intrinsic connection between different pretext tasks. In practice, the linkage of SSATs is complex, and the optimization of task-sharing parameters by minimizing linear combinations of losses can be conflicting. In addition, although a single combination of SSAT can improve certain performance on the baseline, it is not the personalized optimal solution on various remote sensing datasets with diverse properties. In this article, we propose a collaborative self-supervised evolution (so-called CSENet) framework for few-shot remote sensing scene classification to automatically search for appropriate weights in balancing the task conflicts. In contrast to most existing methods, which consider all SSATs to be equally efficacious or fixed-weighted for the few-shot main task, CSENet achieves autonomous co-evolutionary optimization by encoding arbitrary self-supervised weights. Specifically, the complex self-supervised combinations for different remote sensing data are transformed into an evolutionary optimization problem, where chromosomes with weighting variables obtain the optimal combination with genetic operators. Based on the transfer learning few-shot training paradigm, CSENet first efficiently searches for optimal self-supervised combinations with potential by the proposed automatic collaborative evolution strategy and automatically adjusts the weights without manual settings. Importantly, CSENet provides both inductive and transductive inference, and supports the embedding of arbitrary SSATs. The effectiveness of the proposed framework is demonstrated by state-of-the-art (SOTA) results on three benchmark datasets.
Yiting Liu 0004, Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Zedong Tang, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.8
2024 Personalized Multiparty Few-Shot Learning for Remote Sensing Scene Classification
abstract
The existing few-shot scene classification (FSSC) algorithms have achieved satisfactory results, but they are limited by the paradigm of centralized machine learning, i.e., private remote sensing data need to be centralized on a certain server for training. However, remote sensing images generally contain sensitive information such as national security and company privacy, so it is realistically difficult to collect remote sensing data from all the parties. Therefore, there is a pressing requirement in FSSC for a novel paradigm to achieve multi-party collaborative learning without compromising remote sensing data privacy. In this paper, we formulate a novel personalized multi-party few-shot learning (PMPFSL) paradigm for remote sensing scene classification. In PMPFSL, different participants can achieve multi-party collaborative learning without sacrificing the privacy of their local data, and their respective local models are able to recognize the unseen remote sensing scene categories with a small number of labeled samples. Importantly, the proposed PMPFSL is applicable to various multi-party learning algorithms and few-shot scene classification networks. Moreover, to address the problems of local model overfitting and poor discriminability of few-shot metrics, we propose the personalized adaptive distillation (PAD) scheme and multi-scale feature matching network (MSFMNet) on PMPFSL, respectively. Specifically, each participant obtains the MSFMNet with initialization parameters, and implements a certain number of local training on their respective private machines. Global aggregation is subsequently achieved by uploading only the local models to the central server. In a new round of local training, the participants realize personalized data adaptation to the global model based on the PAD. Overall, the proposed PMPFSL customizes a personalized few-shot model for each participant that is more tailored to their respective remote sensing scenarios. The experimental results demonstrate that our PMPFSL is superior in three benchmark FSSC datasets. We also extensively studied and analyzed the contributions of PAD and MSFMNet in the proposed PMPFSL framework.
Shanfeng Wang, Jianzhao Li, Zaitian Liu, Maoguo Gong, Yourun Zhang, Yue Zhao 0024, Boya Deng, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.8