Zhu Han 0002

dblp:83/514-2 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0002-8602-864XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 10 first-author · 16 since 2021
YearPublicationVenuePosition
2025 Subpixel Spectral Variability Network for Hyperspectral Image Classification
abstract
Deep learning-based frameworks have shown great potential in the field of hyperspectral image (HSI) classification owing to their superior modeling capabilities. However, the existence of mixed pixels and spectral heterogeneity limits the discriminant performance of the classifier, which makes it impossible to distinguish the mixed spectra effectively in actual scenarios. To address this gap, we propose a subpixel spectral variability network ($\text {S}^{2}\text {VNet}$) for HSI classification, which incorporates complete subpixel information and class features modeled by spectral variability and nonlinear mixture characteristics to enhance classification performance.$\text {S}^{2}\text {VNet}$is capable of extracting endmembers and abundances based on the nonlinear autoencoder (AE) framework and estimating variability parameters by simultaneously considering scaling factors and perturbation terms to ensure accurate endmember construction. The enhanced subpixel fusion module is further designed to automatically integrate three aspects of abundances, spectral cosine correlation information, and pixel-level class features to provide a robust joint representation for the classifier. Extensive experiments on four public HSI datasets demonstrate the superiority and generalization of the proposed method when benchmarked with state-of-the-art methods. The code will be available athttps://github.com/hanzhu97702/S2VNet.
Zhu Han 0002, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2025 Enhanced Deep Image Prior for Unsupervised Hyperspectral Image Super-Resolution
abstract
Depending on a large-scale paired dataset of low-resolution hyperspectral image (LrHSI), high-resolution multispectral image (HrMSI), and corresponding high-resolution hyperspectral image (HrHSI), the supervised paradigm has achieved impressive performance in the hyperspectral image super-resolution (HISR). However, the intrinsic data-intensive manner hinders its further application in real scenarios. Fortunately, deep image prior (DIP) allows us to achieve unsupervised super-resolution (SR) by solely utilizing degraded observations. However, its potential to accurately model complicated hyperspectral priors is still not fully exploited due to the following two factors: 1) existing methods tend to reconstruct the unknown HrHSI directly from a randomly generated noise, leaving it hard to leverage the scene-relevant information for prior learning and 2) the vanilla architecture is handcrafted for the generator network, which shows limitations in feature representation and thus fails to characterize the complicated image properties. To unleash the potential of DIP for the HISR task, we propose an enhanced DIP network, called EDIP-Net, by addressing the aforementioned impediments. Specifically, EDIP-Net is built with a two-stage four-component scheme, with a zero-shot learning (ZSL) stage for input image establishment and a deep image generation (DIG) stage for prior learning. First, we exploit the cross-scale spectral relationship inside the observations and thus design a degradation learning network to generate paired training samples from the observations themselves. As such, two image-coarse estimations are derived in a ZSL manner by learning an interactive spectral learning network. By replacing random noise with two estimations, we design a double U-shape architecture for the generator network to capture their hyperspectral prior, each independently generating one HrHSI candidate. Under this premise, we further propose a degradation-aware decision fusion strategy to integrate the optimal results in a pixel-to-pixel manner. Extensive experiments demonstrate our superiority in achieving high-quality SR performance. The code will be available athttps://github.com/JiaxinLiCAS.
Jiaxin Li 0002, Lianru Gao, Zhu Han 0002, Zhi Li 0083, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 GRetNet: Gaussian Retentive Network for Hyperspectral Image Classification
abstract
Vision transformer (ViT) is a prevalent technique for capturing long-distance dependencies and has shown impressive performance in the field of hyperspectral image (HSI) classification. However, the core component of ViT, namely, self-attention, faces challenges in balancing high-computational complexity and global modeling within entire input sequences. To alleviate this issue, a novel Gaussian retentive network, called GRetNet, is devised in this letter to enhance the comprehension of fine-grained spatial and spectral features while reducing computational costs. This method provides a powerful classification backbone and can adaptively generate priors to perceive more effective spatial information by introducing a spatial decay mask to assign different weights at various positions. Furthermore, the Gaussian multi-head attention (GMA) is designed to provide dynamic recalibration of feature significance based on statistical distribution and focuses on distinct spectral patterns across different heads, thereby rendering a more concise and robust modeling for HSI classification. Compared with the state-of-the-art classification algorithms, the proposed GRetNet method can yield better classification results and computational efficiency on four benchmark hyperspectral datasets, which verifies its effectiveness and superiority.
Zhu Han 0002, Shuyi Xu, Lianru Gao, Zhi Li 0083, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.1
2024 Dual-Branch Subpixel-Guided Network for Hyperspectral Image Classification
abstract
Deep learning (DL) has been widely applied to hyperspectral image (HSI) classification, owing to its promising feature learning and representation capabilities. However, limited by the spatial resolution of sensors, existing DL-based classification approaches mainly focus on pixel-level spectral and spatial information extraction through complex network architecture design while ignoring the existence of mixed pixels in actual scenarios. To tackle this difficulty, we propose a novel dual-branch subpixel-guided network for HSI classification, called DSNet, which automatically integrates subpixel information and convolutional class features by introducing a deep autoencoder unmixing architecture to enhance classification performance. DSNet is capable of fully considering physically nonlinear properties within subpixels and adaptively generating diagnostic abundances in an unsupervised manner to achieve more reliable decision boundaries for class label distributions. The subpixel fusion module is designed to ensure high-quality information fusion across pixel and subpixel features, further promoting stable joint classification. Experimental results on three benchmark datasets demonstrate the effectiveness and superiority of DSNet compared with state-of-the-art DL-based HSI classification approaches. The codes will be available athttps://github.com/hanzhu97702/DSNet, contributing to the remote sensing community.
Zhu Han 0002, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2024 Multisource Collaborative Domain Generalization for Cross-Scene Remote Sensing Image Classification
abstract
Cross-scene image classification aims to transfer prior knowledge of ground materials to annotate regions with different distributions and reduce hand-crafted cost in the field of remote sensing. However, existing approaches focus on single-source domain generalization to unseen target domains, and are easily confused by large real-world domain shifts due to the limited training information and insufficient diversity modeling capacity. To address this gap, we propose a novel multi-source collaborative domain generalization framework (MS-CDG) based on homogeneity and heterogeneity characteristics of multi-source remote sensing data, which considers data-aware adversarial augmentation and model-aware multi-level diversification simultaneously to enhance cross-scene generalization performance. The data-aware adversarial augmentation adopts an adversary neural network with semantic guide to generate MS samples by adaptively learning realistic channel and distribution changes across domains. In views of cross-domain and intra-domain modeling, the model-aware diversification transforms the shared spatial-channel features of MS data into the class-wise prototype and kernel mixture module, to address domain discrepancies and cluster different classes effectively. Finally, the joint classification of original and augmented MS samples is employed by introducing a distribution consistency alignment to increase model diversity and ensure better domain-invariant representation learning. Extensive experiments on three public MS remote sensing datasets demonstrate the superior performance of the proposed method when benchmarked with the state-of-the-art methods.
Zhu Han 0002, Ce Zhang 0005, Lianru Gao, Michael Kwok-Po Ng, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2023 Few-Shot SAR Target Recognition Through Meta-Adaptive Hyperparameters' Learning for Fast Adaptation
abstract
In synthetic aperture radar automatic target recognition (SAR-ATR), the limitations of imaging environment and observation conditions make it challenging to acquire a substantial amount of high-value targets, resulting in a severe shortage of datasets. This scarcity leads to poor performance and instability in few-shot SAR target recognition. To address these shortcomings, this paper proposes Mada-SGD, a novel inner-loop parameter update approach based on meta adaptive hyper-parameter learning. By considering the correlation information between multiple update steps, Mada-SGD learns the weight distribution information of initialization parameters across previous and current update steps, akin to a memory mechanism. This approach enhances feature extraction and representation ability for few-shot SAR targets. Additionally, an adaptive hyper-parameter update strategy is introduced to simultaneously learn the initialization, weight factor, update factor, and update direction in the meta-learner. This effectively resolves parameter updating issues in meta-learning models while improving fast adaptation for few-shot SAR targets. Experimental results on the specialized MSTAR-FSL dataset demonstrate that Mada-SGD outperforms the latest few-shot SAR target recognition model in terms of SAR target recognition performance, validating its advancement and superiority.
Jinping Sun, Dandan Gu, Zhu Han 0002, Wen Hong
IEEE Trans. Geosci. Remote. Sens.5
2022 Multimodal Hyperspectral Unmixing via Attention Networks
abstract
Owing to the powerful feature extraction and representation capabilities, deep learning (DL) has been successfully applied in hyperspectral unmixing (HU). However, only relying on hyperspectral data for unmixing fails to distinguish objects with similar spectral information, resulting in the degradation of unmixing performance. To this end, this paper presents a novel multimodal unmixing network, MUNet for short, by considering the height information of light detection and ranging (LiDAR) data in a squeeze-and-excitation (SE) attention fashion to guide the unmixing process toward a more accurate performance. MUNet is capable of efficiently embedding the height information obtained from LiDAR data into the autoencoder unmixing architecture through the attention mechanism, thereby fusing more spatial information to obtain ideal unmixing results. Experimental results conducted on the real multimodal dataset demonstrate the effectiveness and superiority of the proposed MUNet compared to several state-of-the-art deep unmixing approaches.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Jing Yao 0002, Bing Zhang 0001, Jocelyn Chanussot
IGARSS1
2022 Reinforcement Learning for Neural Architecture Search in Hyperspectral Unmixing
abstract
In this letter, a novel neural architecture search (NAS) method based on reinforcement learning, called RLNAS, is devised to realize the automatic architecture design in the field of hyperspectral unmixing (HU). This method first train the search network in the constructed self-supervised datasets based on hyperspectral images. The block-based searching and weight-sharing strategies are then introduced to reduce the computational cost in the training phase. The final optimal architecture is obtained by optimizing the multi-objective reward function to balance the trade-off between accuracy and computational efficiency. Compared with the state-of-the-art unmixing algorithms, the proposed RLNAS method can yield better unmixing results on synthetic and real hyperspectral datasets, which verifies its effectiveness and superiority. In addition, the proposed method offers promising potential of the NAS for HU.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Swalpa Kumar Roy, Bing Zhang 0001, Jocelyn Chanussot
IEEE Geosci. Remote. Sens. Lett.1
2022 Radar HRRP Target Recognition Method Based on Multi-Input Convolutional Gated Recurrent Unit With Cascaded Feature Fusion
abstract
Over the past decades, radar high-resolution range profile (HRRP) has been one of the research highlights in the field of radar automatic target recognition (RATR) due to its advantages of easy acquisition, small amount of data, and rich target structure information. However, most of existing methods only consider its amplitude (time domain) characteristics, thereby neglecting the temporal dependence and multi-domain features inside the HRRP sequence. To this end, we propose an end-to-end multi-input convolutional gated recurrent unit neural network, called MIConvGRU, for RATR by both exploiting the multi-domain and temporal information to improve the recognition performance of HRRP target. Initially, the data-preprocessing module is employed to extract the multi-domain features of the target, including time domain, frequency domain, and time-frequency domain features, in order to further enhance the target representation. In addition, a cascaded multi-input GRU structure is designed to acquire the multi-domain temporal dependence feature of HRRP sequence from low to high level. Finally, these temporal features are adaptively fused by a parameter learnable strategy. The experimental results show that the proposed MIConvGRU can effectively learn the multi-domain temporal dependence correlation features in HRRP sequences, improving the target recognition performance.
Jinping Sun, Zhu Han 0002, Wen Hong
IEEE Geosci. Remote. Sens. Lett.3
2022 AutoNAS: Automatic Neural Architecture Search for Hyperspectral Unmixing
abstract
Owing to the powerful and automatic representation capabilities, deep learning (DL) techniques have made significant breakthroughs and progress in hyperspectral unmixing (HU). Among the DL approaches, autoencoders (AEs) have become a widely-used and promising network architecture. However, these AE-based methods heavily rely on manual design and may not be a good fit for specific datasets. To unmix hyperspectral images more intelligently, we propose an automatic neural architecture search model for HU, AutoNAS for short, to determine the optimal network architecture by considering channel configurations and convolution kernels simultaneously. In AutoNAS, the self-supervised training mechanism based on hyperspectral images is first designed for generating the training samples of the supernet. Then, the affine parameter sharing strategy is adopted by applying different affine transformations on the supernet weights in the training phase, which enables finding the optimal channel configuration. Furthermore, on the basis of the obtained channel configuration, the evolutionary algorithm with additional computational constraints is introduced into networks to achieve flexible convolution kernel search by evaluating unmixing results of different architectures in the supernet. Extensive experiments conducted on four hyperspectral datasets demonstrate the effectiveness and superiority of the proposed AutoNAS in comparison with several state-of-the-art unmixing algorithms.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2022 CyCU-Net: Cycle-Consistency Unmixing Network by Learning Cascaded Autoencoders
abstract
In recent years, deep learning (DL) has attracted increasing attention in hyperspectral unmixing (HU) applications due to its powerful learning and data fitting ability. The autoencoder (AE) framework, as an unmixing baseline network, achieves good performance in HU by automatically learning low-dimensional embeddings and reconstructing data. Nevertheless, the conventional AE-based architecture, which focuses more on the pixel-level reconstruction loss, tends to lose some significant detailed information of certain materials (e.g., material-related properties) in the reconstruction process. Therefore, inspired by the perception mechanism, we propose a cycle-consistency unmixing network, called CyCU-Net, by learning two cascaded AEs in an end-to-end fashion, to enhance the unmixing performance more effectively. CyCU-Net is capable of reducing the detailed and material-related information loss in the process of reconstruction by relaxing the original pixel-level reconstruction assumption to cycle consistency dominated by the cascaded AEs. More specifically, cycle consistency can be achieved by a newly proposed self-perception loss, which consists of two spectral reconstruction terms and one abundance reconstruction term. By taking advantage of the self-perception loss in the network, the high-level semantic information can be well preserved in the unmixing process. Moreover, we investigate the performance gain of CyCU-Net with extensive ablation studies. Experimental results on one synthetic and three real hyperspectral data sets demonstrate the effectiveness and competitiveness of the proposed CyCU-Net in comparison with several state-of-the-art unmixing algorithms.
Lianru Gao, Zhu Han 0002, Danfeng Hong, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.2
2022 Multimodal Hyperspectral Unmixing: Insights From Attention Networks
abstract
Deep learning (DL) has aroused wide attention in hyperspectral unmixing (HU) owing to its powerful feature representation ability. As a representative of unsupervised DL approaches, autoencoder (AE) has been proven to be effective to better capture nonlinear components of hyperspectral images than the traditional model-driven linearized methods. However, only using hyperspectral images for unmixing fails to distinguish objects in complex scene, especially for different endmembers with similar materials. To overcome this limitation, we propose a novel multimodal unmixing network for hyperspectral images, called MUNet, by considering the height differences of light detection and ranging (LiDAR) data in a squeeze-and-excitation (SE)-driven attention fashion to guide the unmixing process, yielding performance improvement. MUNet is capable of fusing multimodal information and using the attention map derived by LiDAR to aid network that focuses on more discriminative and meaningful spatial information regarding scenes. Moreover, attribute profile (AP) is adopted to extract the geometrical structures of different objects to better model the spatial information of LiDAR. Experimental results on synthetic and real datasets demonstrate the effectiveness and superiority of the proposed method compared with several state-of-the-art unmixing algorithms. The codes will be available athttps://github.com/hanzhu97702/IEEE_TGRS_MUNet, contributing to the remote sensing community.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Jing Yao 0002, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2022 SpectralFormer: Rethinking Hyperspectral Image Classification With Transformers
abstract
Hyperspectral (HS) images are characterized by approximately contiguous spectral information, enabling the fine identification of materials by capturing subtle spectral discrepancies. Owing to their excellent locally contextual modeling ability, convolutional neural networks (CNNs) have been proven to be a powerful feature extractor in HS image classification. However, CNNs fail to mine and represent the sequence attributes of spectral signatures well due to the limitations of their inherent network backbone. To solve this issue, we rethink HS image classification from a sequential perspective with transformers, and propose a novel backbone network called \ul{SpectralFormer}. Beyond band-wise representations in classic transformers, SpectralFormer is capable of learning spectrally local sequence information from neighboring bands of HS images, yielding group-wise spectral embeddings. More significantly, to reduce the possibility of losing valuable information in the layer-wise propagation process, we devise a cross-layer skip connection to convey memory-like components from shallow to deep layers by adaptively learning to fuse "soft" residuals across layers. It is worth noting that the proposed SpectralFormer is a highly flexible backbone network, which can be applicable to both pixel- and patch-wise inputs. We evaluate the classification performance of the proposed SpectralFormer on three HS datasets by conducting extensive experiments, showing the superiority over classic transformers and achieving a significant improvement in comparison with state-of-the-art backbone networks. The codes of this work will be available at https://github.com/danfenghong/IEEE_TGRS_SpectralFormer for the sake of reproducibility.
Danfeng Hong, Zhu Han 0002, Jing Yao 0002, Lianru Gao, Bing Zhang 0001, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.2
2022 SAR Automatic Target Recognition Method Based on Multi-Stream Complex-Valued Networks
abstract
In synthetic aperture radar automatic target recognition (SAR-ATR), target information is usually propagated and reserved in complex-valued form, namely magnitude information and phase information. However, most of the existing SAR target recognition methods only focus on real-valued (magnitude information) calculations and ignore the phase information of targets, yielding poor recognition performance. To overcome this limitation, this paper proposes a multi-stream feature fusion SAR target recognition method based on complex-valued operations, called MS-CVNets, to utilize the phase information of the target effectively. First of all, a series of complex-valued operation blocks are constructed to satisfy the network training in the complex field, such as complex convolution, complex batch normalization, complex activation, complex pooling, complex full connection, etc. Besides, a multi-stream structure is employed by applying different convolution kernels to extract multi-scale information of targets, further enhancing the representation ability of the model. Experimental results on the MSTAR dataset illustrate that, compared with current state-of-the-art real-valued based models, MS-CVNets can achieve better recognition results under both standard operating conditions (SOC) and extended operating conditions (EOC), validating the effectiveness and superiority of the proposed method.
Jinping Sun, Zhu Han 0002, Wen Hong
IEEE Trans. Geosci. Remote. Sens.3
2021 EvoNAS: Evolvable Neural Architecture Search for Hyperspectral Unmixing
abstract
Owing to the powerful ability in learning low-dimensional representations and reconstruction, autoencoders (AEs) have been successfully applied in hyperspectral unmixing (HU). However, AE-based unmixing architectures, to a great extent, need to be carefully designed in a manual fashion, leading to the bulk of costs in manpower and time. To unmix hyperspectral images more intelligently, we propose an AI-powered evolvable neural architecture search method for HU, EvoNAS for short, to optimally determine the network architecture by the means of the evolutionary algorithm instead of gradient-based or reinforcement learning-based rewards. In EvoNAS, a supernet with all candidate architectures is first trained to learn the unmixing mapping in a self-supervised manner. The optimal network is then constructed by evaluating unmixing results of different architectures in the supernet. EvoNAS is capable of saving tremendous computational cost, since it inherits the weights of the pre-trained supernet and avoids training from scratch during the search phase. Experimental results conducted on two real hyperspectral datasets verify the effectiveness and superiority of the EvoNAS and show the huge potential of the NAS for HU.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Jocelyn Chanussot, Bing Zhang 0001
IGARSS1
2021 Deep Half-Siamese Networks for Hyperspectral Unmixing
abstract
Over the past decades, numerous methods have been proposed to solve the linear or nonlinear mixing problems in hyperspectral unmixing (HU). The existence of spectral variabilities and nonlinearity limits, to a great extent, the unmixing ability of most traditional approaches, particularly in complex scenes. In recent years, deep learning (DL) has been garnering increasing attention in nonlinear HU owing to its powerful learning and fitting ability. However, the DL-based methods tend to generate trivial unmixing results due to the lack of considering physically meaningful endmember information. To this end, we propose a novel siamese network, called the deep half-siamese network (Deep HSNet), for HU by fully considering diverse endmember properties extracted using different endmember extraction algorithms. Moreover, the proposed Deep HSNet, beyond the previous autoencoder-like architecture, adopts another subnetwork to learn the endmember information effectively to guide the unmixing process in a reasonable and accurate way. The experimental results conducted on the synthetic and real hyperspectral data sets validate the effectiveness and superiority of the Deep HSNet over several state-of-the-art unmixing algorithms.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot
IEEE Geosci. Remote. Sens. Lett.1