Xiaowei Zhou 0003

dblp:30/1273-3 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-5871-2762ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Part-aware cross-integration transformer with proximity guided regularization for domain generalizable animal re-identification
Zeyuan Sun, Junyu Dong, Xiaowei Zhou 0003, Huiyu Zhou 0001, Hao Fan 0004
Expert Syst. Appl.3
2025 MSFMamba: Multiscale Feature Fusion State Space Model for Multisource Remote Sensing Image Classification
abstract
In the field of multisource remote sensing image classification, remarkable progress has been made by using the convolutional neural network (CNN) and Transformer. While CNNs are constrained by their local receptive fields, Transformers mitigate this issue with their global attention mechanism. However, Transformers come with the tradeoff of higher computational complexity. Recently, Mamba-based methods built upon the state space model (SSM) have shown great potential for long-range dependence modeling with linear complexity, but they have rarely been explored for multisource remote sensing image classification tasks. To address this issue, we propose the Multi-Scale Feature Fusion Mamba (MSFMamba) network, a novel framework designed for the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR)/synthetic aperture radar (SAR) data. The MSFMamba network is composed of three key components: the Multi-Scale Spatial Mamba (MSpa-Mamba) block, the Spectral Mamba (Spe-Mamba) block, and the fusion Mamba (Fus-Mamba) block. The MSpa-Mamba block employs a multiscale strategy to reduce computational cost and alleviate feature redundancy in multiple scanning routes, ensuring efficient spatial feature modeling. The Spe-Mamba block focuses on spectral feature extraction, addressing the unique challenges of HSI data representation. Finally, the Fus-Mamba block bridges the heterogeneous gap between HSI and LiDAR/SAR data by extending the original Mamba architecture to accommodate dual inputs, enhancing cross-modal feature interactions and enabling seamless data fusion. Together, these components enable MSFMamba to effectively tackle the challenges of multisource data classification, delivering improved performance with optimized computational efficiency. Comprehensive experiments on four real-world multisource remote sensing datasets (Berlin, Augsburg, Houston2018, and Houston2013) demonstrate the superiority of MSFMamba outperforms several state-of-the-art methods and achieves overall accuracies of 76.92%, 91.38%, 92.38%, and 92.86%, respectively. The source codes of MSFMamba will be publicly available athttps://github.com/oucailab/MSFMamba.
Feng Gao 0005, Xuepeng Jin, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Prototype-Based Information Compensation Network for Multisource Remote Sensing Data Classification
abstract
Multi-source remote sensing data joint classification aims to provide accuracy and reliability of land cover classification by leveraging the complementary information from multiple data sources. Existing methods confront two challenges: inter-frequency multi-source feature coupling and inconsistency of complementary information exploration. To solve these issues, we present a Prototype-based Information Compensation Network (PICNet) for land cover classification based on HSI and SAR/LiDAR data. Specifically, we first design a frequency interaction module to enhance the inter-frequency coupling in multi-source feature extraction. The multi-source features are first decoupled into high- and low-frequency components. Then, these features are recoupled to achieve efficient inter-frequency communication. Afterward, we design a prototype-based information compensation module to model the global multi-source complementary information. Two sets of learnable modality prototypes are introduced to represent the global modality information of multi-source data. Subsequently, cross-modal feature integration and alignment are achieved through cross-attention computation between the modality-specific prototype vectors and the raw feature representations. Extensive experiments on three public datasets demonstrate the significant superiority of our PICNet over state-of-the-art methods. The codes are available at https://github.com/oucailab/PICNet.
Feng Gao 0005, Chuanzheng Gong, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 LSKSANet: A Novel Architecture for Remote Sensing Image Semantic Segmentation Leveraging Large Selective Kernel and Sparse Attention Mechanism
abstract
In this paper, we proposed large selective kernel and sparse attention network (LSKSANet) for remote sensing image semantic segmentation. The LSKSANet is a lightweight network that effectively combines convolution with sparse attention mechanisms. Specifically, we design large selective kernel module to decomposing the large kernel into a series of depth-wise convolutions with progressively increasing dilation rates, thereby expanding the receptive field without significantly increasing the computational burden. In addition, we introduce the sparse attention to keep the most useful selfattention values for better feature aggregation. Experimental results on the Vaihingen and Postdam datasets demonstrate the superior performance of the proposed LSKSANet over state-of-the-art methods.
Miao Fu, Feng Gao 0005, Ruzhuang Hua, Yanhai Gan, Xiaowei Zhou 0003
IGARSS5
2024 On Adversarial Training with Incorrect Labels
Benjamin Zi Hao Zhao, Junda Lu 0001, Xiaowei Zhou 0003, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar
WISE (4)3
2024 Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is critical for the effective analysis and interpretation of hyperspectral data. However, simultaneously modeling global and local features is rarely explored to enhance HSI denoising. In this letter, we propose a hybrid convolution and attention network (HCANet), which leverages both the strengths of convolution neural networks (CNNs) and Transformers. To enhance the modeling of both global and local features, we have devised a convolution and attention fusion module aimed at capturing long-range dependencies and neighborhood spectral correlations. Furthermore, to improve multi-scale information aggregation, we design a multi-scale feed-forward network to enhance denoising performance by extracting features at different scales. Experimental results on mainstream HSI datasets demonstrate the rationality and effectiveness of the proposed HCANet. The proposed model is effective in removing various types of complex noise. Our codes are available at https://github.com/summitgao/HCANet.
Shuai Hu, Feng Gao 0005, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Wavelet-Based Bi-Dimensional Aggregation Network for SAR Image Change Detection
abstract
Synthetic aperture radar (SAR) image change detection is critical in remote sensing image analysis. Recently, the attention mechanism has been widely used in change detection tasks. However, existing attention mechanisms often use downsampling operations such as average pooling on the key and value components to enhance computational efficiency. These irreversible operations result in the loss of high-frequency components and other important information. To address this limitation, we develop wavelet-based bi-dimensional aggregation network (WBANet) for SAR image change detection. We design a wavelet-based self-attention block that includes discrete wavelet transform (DWT) and inverse DWT (IDWT) operations on key and value components. Hence, the feature undergoes downsampling without any loss of information, while simultaneously enhancing local contextual awareness through an expanded receptive field. In addition, we have incorporated a bi-dimensional aggregation module (BAM) that boosts the nonlinear representation capability by merging spatial and channel information via broadcast mechanism. Experimental results on three SAR datasets demonstrate that our WBANet significantly outperforms contemporary state-of-the-art methods. Specifically, our WBANet achieves 98.33%, 96.65%, and 96.62% of percentage of correct classification (PCC) on the respective datasets, highlighting its superior performance. Source codes are available athttps://github.com/summitgao/WBANet.
Jiangwei Xie, Feng Gao 0005, Xiaowei Zhou 0003, Junyu Dong
IEEE Geosci. Remote. Sens. Lett.3
2023 LADDER: Latent boundary-guided adversarial training
abstract
Abstract Deep Neural Networks (DNNs) have recently achieved great success in many classification tasks. Unfortunately, they are vulnerable to adversarial attacks that generate adversarial examples with a small perturbation to fool DNN models, especially in model sharing scenarios. Adversarial training is proved to be the most effective strategy that injects adversarial examples into model training to improve the robustness of DNN models against adversarial attacks. However, adversarial training based on the existing adversarial examples fails to generalize well to standard, unperturbed test data. To achieve a better trade-off between standard accuracy and adversarial robustness, we propose a novel adversarial training framework called LAtent bounDary-guided aDvErsarial tRaining (LADDER) that adversarially trains DNN models on latent boundary-guided adversarial examples. As opposed to most of the existing methods that generate adversarial examples in the input space, LADDER generates a myriad of high-quality adversarial examples through adding perturbations to latent features. The perturbations are made along the normal of the decision boundary constructed by an SVM with an attention mechanism. We analyze the merits of our generated boundary-guided adversarial examples from a boundary field perspective and visualization view. Extensive experiments and detailed analysis on MNIST, SVHN, CelebA, and CIFAR-10 validate the effectiveness of LADDER in achieving a better trade-off between standard accuracy and adversarial robustness as compared with vanilla DNNs and competitive baselines.
Xiaowei Zhou 0003, Ivor W. Tsang, Jie Yin 0001
Mach. Learn.1
2022 Edge but not Least: Cross-View Graph Pooling
Xiaowei Zhou 0003, Jie Yin 0001, Ivor W. Tsang
ECML/PKDD (2)1
2022 Detecting adversarial examples by additional evidence from noise domain
abstract
Abstract Deep neural networks are widely adopted powerful tools for perceptual tasks. However, recent research indicated that they are easily fooled by adversarial examples, which are produced by adding imperceptible adversarial perturbations to clean examples. Here the steganalysis rich model (SRM) is utilized to generate noise feature maps, and they are combined with RGB images to discover the difference between adversarial examples and clean examples. In particular, a two‐stream pseudo‐siamese network that fuses the subtle difference in RGB images with the noise inconsistency in noise features is proposed. The proposed method has strong detection capability and transferability, and can be combined with any model without modifying its architecture or training procedure. The extensive empirical experiments show that, compared with the state‐of‐the‐art detection methods, the proposed approach achieves excellent performance in distinguishing adversarial samples generated by popular attack methods on different real datasets. Moreover, this method has good generalization, it trained by a specific adversary can defend against other adversaries effectively.
Shui Yu 0001, Liwen Wu, Shaowen Yao 0001, Xiaowei Zhou 0003
IET Image Process.5
2021 Human-Understandable Decision Making for Visual Recognition
Xiaowei Zhou 0003, Jie Yin 0001, Ivor W. Tsang, Chen Wang 0008
PAKDD (3)1