EDBT 2026 Demo / reviewers in the wild / expert
Zhi Li 0083
dblp:43/3166-83
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0004-4835-2037ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhanced Deep Image Prior for Unsupervised Hyperspectral Image Super-ResolutionabstractDepending on a large-scale paired dataset of low-resolution hyperspectral image (LrHSI), high-resolution multispectral image (HrMSI), and corresponding high-resolution hyperspectral image (HrHSI), the supervised paradigm has achieved impressive performance in the hyperspectral image super-resolution (HISR). However, the intrinsic data-intensive manner hinders its further application in real scenarios. Fortunately, deep image prior (DIP) allows us to achieve unsupervised super-resolution (SR) by solely utilizing degraded observations. However, its potential to accurately model complicated hyperspectral priors is still not fully exploited due to the following two factors: 1) existing methods tend to reconstruct the unknown HrHSI directly from a randomly generated noise, leaving it hard to leverage the scene-relevant information for prior learning and 2) the vanilla architecture is handcrafted for the generator network, which shows limitations in feature representation and thus fails to characterize the complicated image properties. To unleash the potential of DIP for the HISR task, we propose an enhanced DIP network, called EDIP-Net, by addressing the aforementioned impediments. Specifically, EDIP-Net is built with a two-stage four-component scheme, with a zero-shot learning (ZSL) stage for input image establishment and a deep image generation (DIG) stage for prior learning. First, we exploit the cross-scale spectral relationship inside the observations and thus design a degradation learning network to generate paired training samples from the observations themselves. As such, two image-coarse estimations are derived in a ZSL manner by learning an interactive spectral learning network. By replacing random noise with two estimations, we design a double U-shape architecture for the generator network to capture their hyperspectral prior, each independently generating one HrHSI candidate. Under this premise, we further propose a degradation-aware decision fusion strategy to integrate the optimal results in a pixel-to-pixel manner. Extensive experiments demonstrate our superiority in achieving high-quality SR performance. The code will be available athttps://github.com/JiaxinLiCAS. Jiaxin Li 0002, Lianru Gao, Zhu Han 0002, Zhi Li 0083, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Feature Reconstruction Guided Fusion Network for Hyperspectral and LiDAR ClassificationabstractDeep learning has become increasingly popular in hyperspectral image (HSI) and light detection and ranging (LiDAR) data classification, thanks to its powerful feature learning and representation capabilities. However, HSI often contains substantial redundant information, which can hinder efficient data utilization. Furthermore, the significant disparity in information content between HSI and LiDAR data poses a major challenge in representing and aligning semantic information across these two modalities. To address these challenges, we propose a fusion network structure guided by feature reconstruction embedding. This approach employs feature decomposition to reconstruct HSI features and incorporates weight embedding to seamlessly integrate the reconstructed information into classification features. Furthermore, we introduce a cross-modal attention fusion module designed to merge extracted HSI and LiDAR features. This module fully exploits the complementary nature of these two type of feature, facilitating effective information exchange and semantic alignment across multimodal data. We evaluated our method on three widely used HSI and LiDAR datasets: Houston 2013, Augsburg and MUUFL. Experimental results demonstrate that our proposed FRGFNet significantly outperforms traditional probabilistic methods and state-of-the-art deep learning networks, showcasing its effectiveness in multi-source data fusion. Zhi Li 0083, Lianru Gao, Nannan Zi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Hyperspectral Image Classification via Inverse Mahalanobis Attention NetworkabstractDue to its powerful feature extraction and representation capabilities, deep learning has been successfully applied in the field of hyperspectral image classification. In patchbased hyperspectral image classification, where the central pixel represents the true category, extracting spectral features similar to the central pixel spectrum is crucial. To this end, we develop a new framework, called the inverse Mahalanobis attention network (IMAN), to address the spectral similarity feature extraction problem. The proposed framework develops a self-attention mechanism module based on the Mahalanobis distance to better learn the correlation between feature vectors of pixels, thereby efficiently suppressing noise generated by different land cover categories within patches. The feature extraction capability is further enhanced by integrating a dual-stream network structure that separates spatial and spectral information in hyperspectral images. Experiments conducted on on real hyperspectral datasets demonstrate the effectiveness and superiority of the proposed method compared to several state-of-the-art hyperspectral image classification methods. Zhi Li 0083, Longfei Ren, Lianru Gao |
IGARSS | 1 |
| 2024 | GRetNet: Gaussian Retentive Network for Hyperspectral Image ClassificationabstractVision transformer (ViT) is a prevalent technique for capturing long-distance dependencies and has shown impressive performance in the field of hyperspectral image (HSI) classification. However, the core component of ViT, namely, self-attention, faces challenges in balancing high-computational complexity and global modeling within entire input sequences. To alleviate this issue, a novel Gaussian retentive network, called GRetNet, is devised in this letter to enhance the comprehension of fine-grained spatial and spectral features while reducing computational costs. This method provides a powerful classification backbone and can adaptively generate priors to perceive more effective spatial information by introducing a spatial decay mask to assign different weights at various positions. Furthermore, the Gaussian multi-head attention (GMA) is designed to provide dynamic recalibration of feature significance based on statistical distribution and focuses on distinct spectral patterns across different heads, thereby rendering a more concise and robust modeling for HSI classification. Compared with the state-of-the-art classification algorithms, the proposed GRetNet method can yield better classification results and computational efficiency on four benchmark hyperspectral datasets, which verifies its effectiveness and superiority. Zhu Han 0002, Shuyi Xu, Lianru Gao, Zhi Li 0083, Bing Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Cross-Semantic Heterogeneous Modeling Network for Hyperspectral Image ClassificationabstractThe adequate and finer spectral information in hyperspectral images (HSIs) are benefit for various downstream applications like smart agriculture and environmental monitoring. In HSI classification, dual-stream convolutional networks have gained much attention and have been widely used. In patch-based hyperspectral classification tasks, however, merely using center-labeled patches could lead to an increased unlabeled noise in the data. Moreover, in the application of dual-stream network structures, heterogeneity existed in both the data and feature semantic levels to capture more representative features. To tackle these challenges, we have devised a framework called cross-semantic heterogeneous modeling network (CreatingNet), which aligns more closely with the design principles of dual-stream networks by adjusting the input size. This framework introduces a distance metric attention mechanism (DMAM) based on spectral and spatial distances to strengthen the influence of the center pixel on the entire patch. Additionally, we present a fusion module named CrossViT, which combines features with diverse structures and characteristics, leveraging their complementarity. The proposed multiscale heterogeneous fusion module allows for more effective integration of spatial and spectral features in the images. Extensive experiments on four well-known HSI datasets (Indian Pines, Pavia University, Salinas, and Houston 2013) demonstrate the superior classification performance of the proposed CreatingNet to several state-of-the-art methods. The effectiveness of the proposed model is further validated through ablation studies. Zhi Li 0083, Jiaxin Li 0002, Lianru Gao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Model-Guided Coarse-to-Fine Fusion Network for Unsupervised Hyperspectral Image Super-ResolutionabstractFusing a low-resolution hyperspectral image (LrHSI) with an auxiliary high-resolution multispectral image (HrMSI) is a burgeoning technique to realize hyperspectral image super-resolution, in which learning-based methods have dominated the mainstream direction. However, the underutilization of degradation models and strong dependence on large-scale training triplets severely impedes their applicability and performance. Considering these issues, we reformulate the fusion task as a spectral mapping problem and hence propose an unsupervised model-guided coarse-to-fine fusion network. Specifically, degradation knowledge learning is first performed to fully excavate latent model information, which will serve as guidance for better mapping learning. Following that, a coarse-to-fine fusion network is constructed with a multi-scale attentional fusion module in the head and a coarse-to-fine structure in the tail. The former is deployed to achieve a more informative compression, and the latter is adopted to capture the spectral relationship, including a spectral degradation-guided subnetwork for group-by-group coarse reconstruction and a refinement subnetwork for inter-group correlation and dependencies. Finally, high-resolution HSI can be recovered via established spectral mapping. Extensive experiments on simulated and real datasets verify the superiority of our proposed method. The code is available at https://github.com/JiaxinLiCAS/UMC2FF_GRSL. Jiaxin Li 0002, Wengu Liu, Zhi Li 0083, Haoyang Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | X-Shaped Interactive Autoencoders With Cross-Modality Mutual Learning for Unsupervised Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution can compensate for the incompleteness of single-sensor imaging and provide desirable products with both high spatial and spectral resolution. Among them, unmixing-inspired networks have drawn considerable attention owing to their straightforward unsupervised paradigm. However, most do not fully capture and utilize the multi-modal information due to their limited representation ability of constructed networks, hence leaving large room for further improvement. To this end, we propose an X-shaped interactive autoencoders network with cross-modality mutual learning between hyperspectral and multispectral data, XINet for short, to cope with this problem. Generally, it employs a coupled structure equipped with two autoencoders, aiming at deriving latent abundances and corresponding endmembers from input correspondence. Inside the network, a novel X-shaped interactive architecture is designed by coupling two disjointed U-Nets together via a parameter-shared strategy, which not only enables sufficient information flow between two modalities but also leads to informative spatial-spectral features. Considering the complementarity across each modality, a cross-modality mutual learning module is constructed to further transfer knowledge from one modality to another, allowing for better utilization of multi-modal features. Moreover, a joint self-supervised loss is proposed to effectively optimize our proposed XINet, enabling an unsupervised manner without external triplets supervision. Extensive experiments, including super-resolved results in four datasets, robustness analysis, and extension to other applications, are conducted, and the superiority of our method is demonstrated. Jiaxin Li 0002, Zhi Li 0083, Lianru Gao, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |