EDBT 2026 Demo / reviewers in the wild / expert
Ge Zhang 0006
dblp:70/6366-6
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0003-1308-5149ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyperspectral Tracker With Constrained Object Adaptive Learning and Trajectory ConstructionabstractHyperspectral imaging offers significant potential for precise object tracking, yet the scarcity of dataset volumes specifically tailored for hyperspectral tracking algorithms hinders progress, particularly for deep models with complex structures. Additionally, current deep learning-based hyperspectral trackers typically enhance model accuracy via online or adversarial learning, adversely affecting tracking speed. To address these challenges, this paper introduces the Constrained Object Adaptive Learning hyperspectral Tracker (COALT), an effective parameter-efficient fine-tuning tracker tailored for hyperspectral tracking. COALT integrates Pixel-level Object Constrained Spectral Prompt (POCSP) and Temporal Sequence Trajectory Prompt (TSTP) through Adaptive Learning with Parameter-efficient Fine-tuning (ALPEFT), enabling a transformer-based tracker to capture detailed spectral features and relationships in hyperspectral image sequences through trainable rank decomposition matrices. Specifically, POCSP is designed to retain optimal spectral information with low internal correlation and high object representativeness, enabling rapid image reconstruction. Then, the most representative spectral template and search are fused into a single stream as spectral prompts for the Encoder and Decoder layers. Concurrently, the previous coordinates within the same sequence are tokenized and utilized as temporal prompts by TSTP in the decoder layers. The model is trained with ALPEFT to optimize spectral information learning, which substantially reduces the number of training parameters, alleviating overfitting issues arising from limited data. Meanwhile, the proposed tracker not only retains the ability of pre-trained model to estimate object trajectories in an autoregressive manner but also effectively utilizes spectral information and enhances target location perception during the fine-tuning process. Extensive experiments and evaluations are conducted on two public hyperspectral tracking datasets. The results demonstrate that the proposed COALT tracker achieves satisfactory performance with leading processing speed. The code will be available at https://github.com/PING-CHUANG/COALT. Ye Wang 0020, Mingyang Ma 0004, Ge Zhang 0006, Tao Gao 0001, Shaohui Mei |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Spectral Variability-Aware Cascaded Autoencoder for Hyperspectral UnmixingabstractSpectral variability inevitably presents in hyperspectral images (HSIs), resulting in significant unmixing errors when using the conventional linear mixture model (LMM). Though several variants of LMM have been proposed to encounter such spectral variability, they cannot well model the complex characteristics of spectral variability, and the performance of these variants strongly depends on the prior knowledge of the scene. In this article, spectral variability within an image is classified into class-dependent variability and class-independent one, which can be tackled by a novel fully linear mixture model (FLMM) introducing a class-dependent multiplicative scaling term, a class-dependent additive perturbation term, and a class-independent variability term into the conventional LMM. Moreover, a spectral variability-aware cascaded autoencoder (SVACA) is designed to realize the automatic learning and representation of unmixing targets and spectral variability in different hyperspectral scenarios, which consists of a class-independent variability autoencoder and a cascaded class-dependent variability autoencoder. Such a network is able to handle different spectral variability autonomously without any scene prior by parallel inference structure. Experimental results over synthetic and real hyperspectral datasets demonstrate that the proposed SVACA network not only outperforms several state-of-the-art unmixing networks but also presents a stronger capability to handle spectral variability within HSIs. Ge Zhang 0006, Shaohui Mei, Huiyang Han, Yan Feng 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Improving Depth Completion via Depth Feature UpsamplingabstractThe encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods, but its working mechanism is ambiguous. In this paper, we vi-sualize the internal feature maps to analyze how the net-work densifies the input sparse depth. We find that the en-coder feature of ED-Net focus on the areas with input depth points around. To obtain a dense feature and thus esti-mate complete depth, the decoder feature tends to comple-ment and enhance the encoder feature by skip-connection to make the fused encoder-decoder feature dense, resulting in the decoder feature also exhibits sparse. However, ED-Net obtains the sparse decoder feature from the dense fused feature at the previous stage, where the “dense-i-sparse‘’ process destroys the completeness of features and loses in-formation. To address this issue, we present a depth feature upsampling network (DFU) that explicitly utilizes these dense features to guide the upsampling of a low-resolution (LR) depth feature to a high-resolution (HR) one. The completeness of features is maintained throughout the up-sampling process, thus avoiding information loss. Fur-thermore, we propose a confidence-aware guidance module (CGM), which is confidence-aware and performs guidance with adaptive receptive fields (GARF), to fully exploit the potential of these dense features as guidance. Experimental results show that our DFU, a plug-and-play module, can significantly improve the performance of existing ED-Net based methods with limited computational overheads, and new SOTA results are achieved. Besides, the generalization capability on sparser depth is also enhanced. Project page: https://npucvr.github.iolDFU. Ge Zhang 0006, Shaoqian Wang, Bo Li 0090, Qi Liu 0054, Le Hui, Yuchao Dai |
CVPR | 2 |
| 2024 | Bridging CNN and Transformer With Cross-Attention Fusion Network for Hyperspectral Image ClassificationabstractFeature representation is crucial for hyperspectral image (HSI) classification. However, existing convolutional neural network (CNN)-based methods are limited by the convolution kernel and only focus on local features, which causes it to ignore the global properties of HSIs. Transformer-based networks can make up for the limitations of CNNs because they emphasize the global features of HSIs. How to combine the advantages of these two networks in feature extraction is of great importance in improving classification accuracy. Therefore, a cross-attention fusion network bridging CNN and Transformer (CAF-Former) is proposed, which can fully utilize the advantages of CNN in local features and Transformer’s long time-dependent feature learning for hyperspectral classification. In order to fully explore the local and global information within an HSI, a Dynamic-CNN branch is proposed to effectively encode local features of pixels, while a Gaussian Transformer branch is constructed to accurately model the global features and long-range dependencies. Moreover, in order to fully interact with local and global features, a cross-attention fusion (CAF) module is proposed as a bridge to fuse the features extracted by the two branches. Experiments over several benchmark datasets demonstrate that the proposed CAF-Former significantly outperforms both CNN-based and Transformer-based state-of-the-art networks for HSI classification. Fulin Xu, Shaohui Mei, Ge Zhang 0006, Nan Wang 0026, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionabstractExisting deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish depth completion more efficiently, we propose a novel lightweight deep network framework, the Long-short Range Recurrent Updating (LRRU) network. Without learning complex feature representations, LRRU first roughly fills the sparse input to obtain an initial dense depth map, and then iteratively updates it through learned spatially-variant kernels. Our iterative update process is content-adaptive and highly flexible, where the kernel weights are learned by jointly considering the guidance RGB images and the depth map to be updated, and large-to-small kernel scopes are dynamically adjusted to capture long-to-short range dependencies. Our initial depth map has coarse but complete scene depth information, which helps relieve the burden of directly regressing the dense depth from sparse ones, while our proposed method can effectively refine it to an accurate depth map with less learnable parameters and inference time. Experimental results demonstrate that our proposed LRRU variants achieve state-of-the-art performance across different parameter regimes. In particular, the LRRU-Base model outperforms competing approaches on the NYUv2 dataset, and ranks 1st on the KITTI depth completion benchmark at the time of submission. Project page: https://npucvr.github.io/LRRU/. Bo Li 0090, Ge Zhang 0006, Qi Liu 0054, Tao Gao 0001, Yuchao Dai |
ICCV | 3 |
| 2023 | Lightweight Multiresolution Feature Fusion Network for Spectral Super-ResolutionabstractSpectral super-resolution (SR), which reconstructs high spatial-resolution hyperspectral images (HSIs) from RGB inputs, has been demonstrated to be one of the effective computational imaging techniques to acquire HSIs. Though deep neural networks have shown their superiority in such a complex mapping problem, existing networks generally involve a very complex structure with huge amounts of parameters, resulting in giant memory occupation. In this article, a lightweight multiresolution feature fusion network (MRFN) is proposed, which adopts a multiresolution feature extraction and fusion framework to fully explore RGB inputs in different scales of resolution. Specifically, a lightweight feature extraction module (LFEM), which adopts cheap convolution and attention mechanisms, is constructed to explore different scales of features under a lightweight structure. Moreover, a hybrid loss function is proposed by encountering not only pixel-value level reconstruction error but also spectral continuity and fidelity. Experiments over three benchmark datasets, i.e., CAVE, Interdisciplinary Computational Vision Laboratory (ICVL), and NTIRE2022 datasets, have demonstrated that the proposed MRFN can reconstruct HSIs from RGB inputs in higher quality with fewer parameters and computational floating-point operations (FLOPs) compared with several state-of-the-art networks. Shaohui Mei, Ge Zhang 0006, Nan Wang 0026, Mingyang Ma 0004, Yifan Zhang 0006, Yan Feng 0005 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Reconstruction-Assisted and Distance-Optimized Adversarial Training: A Defense Framework for Remote Sensing Scene ClassificationabstractDespite deep neural networks (DNNs) have been widely applied in remote sensing (RS) scene classification and achieved satisfying performance, the vulnerability of DNNs towards adversarial examples significantly degrades their performance. Moreover, the relatively limited labeled samples of RS scene classification make DNNs more likely to overfit, leading to weak generalizability and noise sensitivity. This may result in DNNs being more vulnerable to adversarial examples. Consequently, the defense of adversarial examples is of crucial importance to improve both the generalizability and robustness of DNNs in the RS scene classification task. However, few studies have been conducted on defense for RS scene classification, especially ignoring the intrinsic characteristics of RS images. In this paper, an effective defense framework for RS scene classification, named reconstruction-assisted and distance-optimized adversarial training (RDAT), is proposed to defend adversarial examples. In order to solve the problems caused by high interclass similarity, a distance-optimized (DO) strategy is designed for adversarial training to strengthen the learning of underfitting content, increase the interclass distance, and improve the robustness of the networks. Furthermore, in order to generate high quality samples for adversarial training, a reconstruction-assisted (RA) block is proposed to eliminate adversarial perturbations in adversarial examples. Specifically, in this block, by swin transformer (SwinT) block and multi-scale convolution (MSC) block, SwinT-MSC-UNet (SMUNet) is constructed to fully extract global and multi-scale local features to adapt to the characteristics of RS images with large variance of ground object scales. Extensive experiments on the benchmark datasets, i.e., UC Merced (UCM) and Aerial Image Dataset (AID), have demonstrate that the proposed RDAT can effectively resist multiple adversarial attacks and yield superior results than other defense methods for RS scene classification. Yuru Su, Ge Zhang 0006, Shaohui Mei, Jiawei Lian, Ye Wang 0020, Shuai Wan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multiscale and Cross-Level Attention Learning for Hyperspectral Image ClassificationabstractTransformer-based networks, which can well model the global characteristics of inputted data using the attention mechanism, have been widely applied to hyperspectral image (HSI) classification and achieved promising results. However, the existing networks fail to explore complex local land cover structures in different scales of shapes in hyperspectral remote sensing images. Therefore, a novel network named multiscale and cross-level attention learning (MCAL) network is proposed to fully explore both the global and local multiscale features of pixels for classification. To encounter local spatial context of pixels in the transformer, a multiscale feature extraction (MSFE) module is constructed and implemented into the transformer-based networks. Moreover, a cross-level feature fusion (CLFF) module is proposed to adaptively fuse features from the hierarchical structure of MSFEs using the attention mechanism. Finally, the spectral attention module (SAM) is implemented prior to the hierarchical structure of MSFEs, by which both the spatial context and spectral information are jointly emphasized for hyperspectral classification. Experiments over several benchmark datasets demonstrate that the proposed MCAL obviously outperforms both the convolutional neural network (CNN)-based and transformer-based state-of-the-art networks for hyperspectral classification. Fulin Xu, Ge Zhang 0006, Hui Wang 0017, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Gaussian Information Entropy based band Reduction for Unsupervised Hyperspectral Video TrackingabstractHyperspectral videos, which provide extra spectral characteristics besides spatial and temporal information, can improve the performance of object tracking using spectral signatures. However, there is a lack of labeled hyperspectral videos to support deep learning based model design. On the contrary, object tracking in the color space has been well developed in the past decade with many benchmark tracking models, e.g., SiamBAN. Therefore, how to transfer models designed in the color space to the hyperspectral space is of great importance. In this paper, hyperspectral videos are reduced into 3 bands using a band reduction algorithm, by which the existing well-trained trackers can be directly used. Specifically, Gaussian Information Entropy (GIE) is used to transform a hyperspectral video into a 3-band pseudo-color video, by which hyperspectral object tracking is conducted in an unsupervised mode. Experimental results demonstrate that object trackers designed in the color space can be transferred to hyperspectral videos using band reduction algorithms and the GIE based reduction is more effective than several well-known band reduction algorithms when using SiamBAN. Yuru Su, Shaohui Mei, Ge Zhang 0006, Ye Wang 0020, Mingyi He, Qian Du 0001 |
IGARSS | 3 |
| 2022 | Extended Collaborative Representation-Based Hyperspectral Imagery ClassificationabstractCollaborative representation (CR) has been demonstrated to be very effective for hyperspectral image classification. However, insufficient diversity of training samples often results in limited classification accuracy under small-training-sample conditions, especially when diverse spectral variation is presented in testing samples. In order to alleviate such a problem, a spectral variation augmented-based linear mixed model (SV-LMM) is proposed, in which the spectral variation is extracted by conducting singular value decomposition (SVD) over training samples. Such spectral variation is further utilized to extend the CR for hyperspectral classification. Experiments over two benchmark datasets, i.e., the Pavia Center dataset and the University of Houston dataset, demonstrate that the proposed extended CR-based classifier (ECRC) clearly improves the performance of conventional CRC for hyperspectral classification and outperforms several state-of-the-art algorithms. Bobo Xie, Shaohui Mei, Ge Zhang 0006, Yifan Zhang 0006, Yan Feng 0005, Qian Du 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Spectral Variability Augmented Two-Stream Network for Hyperspectral Sparse UnmixingabstractDeep learning-based methods have drawn great attention in hyperspectral unmixing and obtained promising performance due to their powerful learning capability. However, few existing networks explicitly deal with the spectral variability inevitably present in hyperspectral images, limiting their fitting performance. In this letter, a spectral variability augmented two-stream network (SVATN) is designed to explicitly address the problem of spectral variability in a deep convolutional network for sparse unmixing. Specifically, the proposed SVATN maps a random input to coefficients of spectral variability in addition to abundances of endmembers, in which spectral variability is accommodated by the linear mixture model as an augmented item. Moreover, a spatial-spectral correlation-based variability extraction method (SSCVE) is proposed to construct a spectral variability library, which serves as priors in the loss function to optimize the proposed SVATN. Experiments over synthetic and real data sets demonstrate the superiority of the proposed SVATN over several state-of-the-art methods. The code of our proposed method is released at: https://github.com/MeiShaohui/SVATN. Ge Zhang 0006, Shaohui Mei, Bobo Xie, Yan Feng 0005, Qian Du 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Spectral Variation Augmented Representation for Hyperspectral Imagery Classification With Few Labeled SamplesabstractDue to variation of imaging conditions, spectra of the same type of ground objects usually exhibit certain discrepancy, leading to intra-class spectral distance increase and inter-class distance decrease. As a result, classification accuracy is greatly affected, especially in cases with few labeled samples. For representation based classifiers, the spectral variability within limited training samples is far from sufficient to represent diverse variations within testing ones. To handle this problem, a spectral variation augmented representation for hyperspectral imagery classification (SVARC) with few labeled samples is proposed in this article. Firstly, a novel class-independent and class-dependent components based linear representation model (CICD-LRM) is proposed to emphasize the representation of spectral variation. Secondly, depending on spatial and spectral correlation, the CICD-LRM guided global and local spectral variation extraction schemes are designed, and a fused spectral variation dictionary is constructed by concatenation. Finally, a classifier for hyperspectral images based on the CICD-LRM and spectral variation dictionary is proposed, and specifically three different spectral variation reconstruction strategies are designed. Similar to most of the representation based classifiers, residual-driven decision is also employed in the proposed classifier. Comparative experiments are conducted with eight classical and state-of-the-art methods using two benchmark datasets. The experimental results demonstrate that the proposed SVARC method significantly outperforms the compared ones in cases with few labeled samples. Bobo Xie, Yifan Zhang 0006, Shaohui Mei, Ge Zhang 0006, Yan Feng 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Spectral Variability Augmented Sparse Unmixing of Hyperspectral ImagesabstractSpectral unmixing expresses the mixed pixels existing in hyperspectral images as the product of endmembers and their corresponding fractional abundances, which has been widely used in hyperspectral imagery analysis. However, the endmember spectra even for pixels from the same material of an image may include variability due to the influence of lighting conditions and inherent properties of materials within different pixels. Though thein situspectral library has been used to accommodate such variability by using multiplein situspectra to represent each kind of material, the performance improvement may be restricted due to the limited number of endmembers for each material. Therefore, in this article, spectral variability is directly extracted from anin situendmember library and considered to be transferable among different endmembers for the first time. Furthermore, such a spectral variability is further used to augment sparse unmixing by synchronously performing endmember-based reconstruction and spectral variability-augmented reconstruction in the sparse unmixing model. By, respectively, imposing sparse and smoothness regularization over abundances and variability coefficients, a convex optimization-based spectral variability augmented sparse unmixing (SVASU) is finally proposed, and its convergence performance is also analyzed. Experiments conducted over synthetic and real-world datasets demonstrate that the proposed SVASU method not only significantly improves the unmixing performance of conventional spectral library-based unmixing but also outperforms several state-of-the-art sparse unmixing algorithms. Ge Zhang 0006, Shaohui Mei, Bobo Xie, Mingyang Ma 0004, Yifan Zhang 0006, Yan Feng 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Local Sparse Representation Based Spatial Preprocessing For Endmember ExtractionabstractHyperspectral unmixing has been widely used to decompose a mixed pixel into a collection of endmembers weighted by their corresponding fractional abundances, in which endmember extraction step is of crucial importance. Many classical endmember extraction algorithms mainly identify spectrally pure endmembers according to spectra of pixels, e.g., NFINDR and vertex component analysis (VCA), ignoring spatial distribution or structure information that has been demonstrated to be complemental for spectral information in hyperspectral image processing. In order to improve the performance of these classical endmember extraction algorithms, a novel spatial preprocessing method is proposed to explore spatial information prior to endmember extraction step. Specifically, pixels in hyperspectral images are modified using their sparse linear approximation by neighboring pixels, such that spectral variation within a local spatial neighbor-hood can be alleviated. Experimental results on both simulated and real data sets demonstrate that the proposed local sparse representation based spatial preprocessing algorithm is capable of producing better unmixing result compared to several state-of-the-art spatial preprocessing methods. Ge Zhang 0006, Shaohui Mei, Yan Feng 0005, Qian Du 0001 |
IGARSS | 1 |