Xudong Zhao 0003

dblp:02/735-3 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-6942-136XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Dual-Domain Fractional Fourier Transformer for Underwater Image Degradation Removal
abstract
In this letter, we propose DEFriT, a dual-domain framework for underwater image restoration targeting degradation caused by scattering and absorption. Supervised training is enabled by simulating underwater degradation from land images using a physics-based formation model. To address the spectrally non-stationary nature of underwater attenuation, we introduce the fractional Fourier transform (FrFT) to bridge spatial and spectral representations. A Fractional Transform Block (FrTB) with real-imaginary decomposition and Hybrid Time-Frequency Self-Attention (HTFSA) supports joint spatial-spectral modeling. Empirical analysis shows that degradation mainly affect the amplitude spectrum, with learned fractional orders concentrated in a narrow range ($p \in [0.48, 0.52]$). To the best of our knowledge, this is the first work to introduce FrFT into underwater image restoration and to systematically establish a fractional order prior. DEFriT achieves state-of-the-art performance on five benchmark datasets.
Chuangxi Chen, Xudong Zhao 0003, Yixiao Yang, Ran Tao 0003, Binghua Su
IEEE Signal Process. Lett.3
2025 Unsupervised Domain Adaptation With Hierarchical Masked Dual-Adversarial Network for End-to-End Classification of Multisource Remote Sensing Data
abstract
Although unsupervised domain adaptation (UDA) has been successfully applied for cross-scene classification of multisource remote sensing (MSRS) data, there are still some tough issues: 1) The vast majority of them are patch-based, requiring pixel by pixel processing at high complexity and ignoring the roles of unlabeled data between different domains. 2) Traditional masked autoencoder (MAE)-based methods lack effective multiscale analysis and require pre-training, ignoring the roles of low-level representations. As such, a hierarchical masked dual-adversarial DA network (HMDA-DANet) is proposed for cross-domain end-to-end classification of MSRS data. Firstly, a hierarchical asymmetric MAE (HAMAE) without pre-training is designed, containing a frequency dynamic large-scale convolutional (FDLConv) block to enhance important structural information in the frequency domain, and an intramodality enhancement and intermodality interaction (IAEIEI) block to embed some additional information beyond the domain distribution by expanding the cross-modal reconstruction space. Representative multimodal multiscale features can be extracted, while to some extent improving their generalization to the target domain. Then, a multimodal multiscale feature fusion (MMFF) block is built to model the spatial and scale dependencies for feature fusion and reduce the layer by layer transmission of redundancy or interference information. Finally, a dual-discriminator-based DA (DDA) block is designed for class-specific semantic feature and global structural alignments in both spatial and prediction spaces. It will enable HAMAE to model the cross-modal, cross-scale, and cross-domain associations, yielding more representative domain-invariant multimodal fusion features. Extensive experiments on five cross-domain MSRS datasets verify the superiority of the proposed HMDA-DANet over other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 Degradation-Noise-Aware Deep Unfolding Transformer for Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) play a pivotal role in fields, such as medical diagnosis and agriculture. However, it often contends with significant noise stemming from narrowband spectral filtering. Existing denoising techniques have their limitations: model-driven methods rely on manual priors and hyperparameters, while learning-based methods struggle to discern intrinsic noise patterns, as they require paired images with specific example noise for training, fail to capture critical noise distribution information, leading to unrobust denoising results. This work addresses the issue by presenting a degradation-noise-aware unfolding network (DNA-Net). Unlike training directly with the simulated noise, DNA-Net initially models general sparse and Gaussian noise through statistic distributions. It then explicitly represents image priors with a customized spectral transformer. The model is subsequently unfolded into an end-to-end (E2E) network, with hyperparameters adaptively estimated from noisy HSI and degradation models, effectively regulating each iteration. Furthermore, a novel U-shaped local-nonlocal–spectral transformer (U-LNSA) is introduced, simultaneously capturing spectral correlations, local features, and nonlocal dependencies. The integration of U-LNSA into DNA-Net establishes the first Transformer-based deep unfolding method for HSI denoising. Experimental results on synthetic and real noise validate DNA-Net’s superior performance over state-of-the-art (SOTA) methods. Moreover, the DNA-Net, trained exclusively on mixed Gaussian noise and impulse noise, demonstrates the ability to generalize to unseen noise present in real images. Code and models will be released at:https://github.com/NavyZeng/DNA-Net.
Haijin Zeng, Xudong Zhao 0003, Jiezhang Cao, Shaoguang Huang, Hiêp Quang Luong, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.3
2025 Chirplet Fourier Analysis Network for Cross-Scene Classification of Multisource Remote Sensing Data
abstract
The joint application of multisource remote sensing (MSRS) data, such as hyperspectral image (HSI) and light detection and ranging (LiDAR), offers significant potential for accurate land cover classification. However, the existing applications often struggle with domain shifts across scenes caused by sensor, illumination, and phase variations. Focusing on this domain adaptation problem, a Chirplet Fourier analysis network (ChirpFAN) is proposed for cross-scene classification of MSRS data in this paper. Firstly, a fractional spatial-frequency-phase feature extraction module including the fractional Fourier transform and a learnable phase-aware weighting block is proposed to capture multi-domain features. Secondly, a Chirplet swin transformer (ChirpST) block integrates a Chirplet Fourier analysis (ChirpFA) layer within a Swin transformer is designed to analyze multi-scale textural and oscillatory patterns. Finally, a modality-shared network including ChirpST blocks is designed for inter-modal fusion and alignment. Extensive experiments demonstrate that the ChirpFAN framework achieves state-of-the-art performance with 3% average improvements on three challenging cross-scene MSRS datasets. Code will be released on GitHub.
Xudong Zhao 0003, Qi Ming, Yixiao Yang, Wen-Shuai Hu, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 Cross-Scene Joint Classification of Multisource Data With Multilevel Domain Adaption Network
abstract
Domain adaption (DA) is a challenging task that integrates knowledge from source domain (SD) to perform data analysis for target domain. Most of the existing DA approaches only focus on single-source-single-target setting. In contrast, multisource (MS) data collaborative utilization has been extensively used in various applications, while how to integrate DA with MS collaboration still faces great challenges. In this article, we propose a multilevel DA network (MDA-NET) for promoting information collaboration and cross-scene (CS) classification based on hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this framework, modality-related adapters are built, and then a mutual-aid classifier is used to aggregate all the discriminative information captured from different modalities for boosting CS classification performance. Experimental results on two cross-domain datasets show that the proposed method consistently provides better performance than other state-of-the-art DA approaches.
Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Fractional Fourier Image Transformer for Multimodal Remote Sensing Data Classification
abstract
With the recent development of the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data, deep learning methods have achieved promising performance owing to their locally sematic feature extracting ability. Nonetheless, the limited receptive field restricted the convolutional neural networks (CNNs) to represent global contextual and sequential attributes, while visual image transformers (VITs) lose local semantic information. Focusing on these issues, we propose a fractional Fourier image transformer (FrIT) as a backbone network to extract both global and local contexts effectively. In the proposed FrIT framework, HSI and LiDAR data are first fused at the pixel level, and both multisource feature and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contextual and sequential feature extraction. Unlike the attention-based representations in classic VIT, FrIT is capable of speeding up the transformer architectures massively and learning valuable contextual information effectively and efficiently. More significantly, to reduce redundancy and loss of information from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. Five HSI and LiDAR scenes including one newly labeled benchmark are utilized for extensive experiments, showing improvement over both CNNs and VITs.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Lianfang Tian, Wilfried Philips
IEEE Trans. Neural Networks Learn. Syst.1
2023 Morphological Transformation and Spatial-Logical Aggregation for Tree Species Classification Using Hyperspectral Imagery
abstract
Hyperspectral image (HSI) consists of abundant spectral and spatial characteristics, which contribute to a more accurate identification of materials and land covers. However, most existing methods of hyperspectral image analysis primarily focus on spectral knowledge or coarse-grained spatial information while neglecting the fine-grained morphological structures. In the classification task of complex objects, spatial morphological differences can help to search for the boundary of fine-grained classes, e.g., forestry tree species. Focusing on subtle traits extraction, a spatial-logical aggregation network (SLA-NET) is proposed with morphological transformation for tree species classification. The morphological operators are effectively embedded with the trainable structuring elements, which contributes to distinctive morphological representations. We evaluate the classification performance of the proposed method on two tree species datasets, and the results demonstrate that the proposed SLA-NET significantly outperforms the other state-of-the-art classifiers.
Mengmeng Zhang 0005, Wei Li 0032, Xudong Zhao 0003, Huan Liu 0015, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Multisource Cross-Scene Classification Using Fractional Fusion and Spatial-Spectral Domain Adaptation
abstract
To solve the limitation of labeled samples in hyperspectral image (HSI) classification, cross-scene learning methods are developed recently. However, the disparity caused by environmental variation between HSI scenes is still a challenge. As a supplement, light detection and ranging (LiDAR) data provides elevation and spatial information regardless the variations. In this paper, we propose a multisource cross-scene classification method using fractional fusion and spatial-spectral domain adaptation to reduce disparity between scenes. The spatial information of HSI is preserved by fractional differential masks (FrDM) firstly. Then the LiDAR data is utilized for spectral alignment of HSI. The utilization of LiDAR data reduces the pixel-level disparity between scenes. At last, a spatial-spectral domain adaptation network is proposed for feature extraction and classification. Experimental results on HSI and LiDAR scenes show 5% improvements in overall accuracy compared with state-of-the-art methods.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips
IGARSS1
2022 Multisource Remote Sensing Data Classification Using Fractional Fourier Transformer
abstract
Focusing on joint classification of Hyperspectral image (HSI) and Light detection and ranging (LiDAR) data, a fractional Fourier image transformer (FrIT) is proposed as a backbone network in this paper. In the proposed FrIT, HSI and LiDAR data are firstly fused at pixel-level. Both multi-source and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contexts and sequential feature extraction. Unlike the attention-based representations in classic visual image transformer (VIT), FrIT is capable of speeding up the transformer architectures massively. To reduce the information loss from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. At last, to evaluate the performance of FrIT, a new HSI and LiDAR benchmark is provided for extensive experiments, on which the proposed FrIT gains an improvement of 3% over state-of-the-art methods.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips
IGARSS1
2022 Multi-Source Remote Sensing Data Cross Scene Classification Based on Multi-Graph Matching
abstract
Multi-source joint classification has been extensively investigated in single scenario setting; however, for cross scene (CS) classification, few studies have been conducted for evaluating the collaborative performance of multi-sources. In this paper, using hyperspectral image (HSI) and light detection and ranging (LiDAR) data, we propose a multi-source CS classification method, and build source-related alignment to reduce statistical shift. Both geometrical and statistical alignments are considered to learn common-subspaces of each source with preserving discrimination information. Finally, the aligned features from both sources are integrated for final classification. Experimental results demonstrate the superior of the proposed method over other state-of-the-art CS approaches.
Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005
IGARSS2
2022 Fractional Gabor Convolutional Network for Multisource Remote Sensing Data Classification
abstract
Remote sensing using multisensor platforms has been systematically applied for monitoring and optimizing human activities. Several advanced techniques have been developed to enhance and extract the spatially and spectrally semantic information in the hyperspectral image (HSI) and light detection and ranging (LiDAR) data processing and analysis. However, an abundance of redundant information and sometimes a lack of discriminative features reduce the efficiency and effectiveness of multisource classification methods. This article proposes a fractional Gabor convolutional network (FGCN), focusing on efficient feature fusion and comprehensive feature extraction. First, the proposed FGCN uses Octave convolution layers to perform multisource information fusion and preserve discriminative information. Second, fractional Gabor convolutional (FGC) layers are proposed to extract multiscale, multidirectional, and semantic change features. The completeness and discrimination of the multisource features using different FGC kernels are improved, which yield robust feature extraction against semantic changes. Finally, the fractional Gabor feature and spectral feature are combined with two weighting factors which can be learned during the network training. Experimental results and comparisons with state-of-the-art multisource classification methods indicate the effectiveness of the proposed FGCN. With the FGCN, we can obtain an 89.90% overall accuracy on the challenging Muufl Gulfport (MUUFL) data set, with an improvement of 3% over state-of-the-art methods.
Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Wilfried Philips, Wenzi Liao
IEEE Trans. Geosci. Remote. Sens.1
2020 Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Random Walk and Deep CNN Architecture
abstract
Earth observation using multisensor data is drawing increasing attention. Fusing remotely sensed hyperspectral imagery and light detection and ranging (LiDAR) data helps to increase application performance. In this article, joint classification of hyperspectral imagery and LiDAR data is investigated using an effective hierarchical random walk network (HRWN). In the proposed HRWN, a dual-tunnel convolutional neural network (CNN) architecture is first developed to capture spectral and spatial features. A pixelwise affinity branch is proposed to capture the relationships between classes with different elevation information from LiDAR data and confirm the spatial contrast of classification. Then in the designed hierarchical random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects the local similarity of pixel pairs, which enforce spatial consistency in the deeper layers of networks. Finally, a classification map is obtained by calculating the probability distribution. Experimental results validated with three real multisensor remote sensing data demonstrate that the proposed HRWN significantly outperforms other state-of-the-art methods. For example, the two branches CNN classifier achieves an accuracy of 88.91% on the University of Houston campus data set, while the proposed HRWN classifier obtains an accuracy of 93.61%, resulting in an improvement of approximately 5%.
Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001, Wenzi Liao, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.1
2019 Multisource Remote Sensing Data Classification Using Deep Hierarchical Random Walk Networks
abstract
Collaborative classification of hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data is investigated using effective hierarchical random walk networks, denoted as HRWN. The proposed HRWN jointly optimizes dual-tunnel CNN, pixelwise affinity and seeds map via a novel random walk layer, which enforces spatial consistency in the deepest layers of the network. In designed random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects local similarity of pixel pairs, which preserves boundary localization and spatial consistency well. Experimental results validated with two real multisource remote sensing data demonstrate that the proposed HRWN can significantly outperform other state-of-art methods.
Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032
ICASSP1
2019 Collaborative Classification of Hyperspectral and Lidar Data With Information Fusion and Deep Nets
abstract
Convolutional neural network (CNN) receives extensive attention in hyperspectral image classification. While hyper-spectral images contain abundant spectral information but lack spatial information, which usually contributes to poor classification results. In this paper, a novel classification framework called information fusion based CNN (IF-CNN) is proposed to compensate for the shortcomings of hyper-spectral images. The proposed method merges hyperspectral images with abundant spectral information and LiDAR images with rich spatial information as the input of classification framework. Furthermore, the framework consists of two convolutional neural networks: one-dimensional CNN for extracting spectral features, and two-dimensional CNN for extracting spatial correlation features. Experimental results demonstrate that the proposed method achieves excellent performance compared with some existing methods.
Chen Chen 0001, Xudong Zhao 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IGARSS2