Zhi He

dblp:56/5307 · DBLP profile ↗
← Back
38ranked-venue papers
21as first author
18since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 12 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 HSACT: A hierarchical semantic-aware CNN-Transformer for remote sensing image spectral super-resolution
Chengle Zhou, Zhi He, Liwei Zou, Yunfei Li 0006, Antonio Plaza
Neurocomputing2
2025 Hardness-Aware Prototypical Contrastive Learning for Hyperspectral Coastal Wetlands Classification
abstract
Self-supervised contrastive learning performs well in representing hyperspectral data under limited labels. However, the heterogeneity poses challenges in balancing a fixed temperature to prevent model collapse with accurately identifying hard negatives. In this paper, we propose a hardness-aware prototypical contrastive learning method (HAPC) for hyperspectral coastal wetlands classification. Firstly, we propose a spatial-spectral integration data augmentation method to enhance the construction of positive counterparts. Additionally, we introduce a hardness-aware temperature reweighting strategy to effectively identify challenging negatives and preserve manifold integrity of data. Experiments on two hyperspectral datasets captured by Zhuhai-1 satellite indicate that HAPC can improve feature representation of hyperspectral coastal wetlands. Our code and datasets is available at https://github.com/sakurashine/HAPC.
Jian Dong 0004, Zhi He, Chengle Zhou
IEEE Geosci. Remote. Sens. Lett.2
2025 Wavelet-Inspired Sparse Learning Network for Hyperspectral Image Change Detection
abstract
In this letter, a novel wavelet-inspired sparse learning network (WISLNet) is proposed for hyperspectral image change detection (HSI-CD), which introduces wavelet convolution transform into a data-driven low-rank and sparse representation (LRSR) model to build a deep unfolding network in the frequency domain. The WISLNet mainly involves the following key steps. First, the difference image (DI) of the bi-temporal hyperspectral images (Bi-HSIs) is obtained through pixel-by-pixel and band-by-band subtraction operations. Then, the LRSR model is designed as a deep unfolding network with modules for low-rank learning, sparse learning, and DI reconstruction to capture deep semantics related to change identification in the DI. Meanwhile, wavelet convolution is embedded in the above modules to progressively mine the low-rank and sparse features of the DI in the frequency domain. Next, an iterative optimization scheme based on low-rank and sparse features is used to capture the change semantic details between Bi-HSIs. Finally, a wavelet convolution filter is designed and applied to the sparse components to reflect the image change information. Experiments on River and Farmland Bi-HSIs demonstrated that the proposed WISLNet method is able to achieve superior CD results compared to well-known and state-of-the-art deep networks. The code is available at: https://github.com/chengle-zhou/WISLNet.
Chengle Zhou, Zhi He, Jian Dong 0004, Liwei Zou
IEEE Geosci. Remote. Sens. Lett.2
2025 Hierarchical and Bidirectional Contrastive Learning for Hyperspectral Image Classification
abstract
Representing hyperspectral images (HSI) is a complex and challenging task, primarily due to spectral uncertainty. Learnable prototypical contrastive learning is specialized in discriminative instance representation. However, it requires a much low temperature to prevent model collapse, which can hinder the encoder’s ability to capture category relationships. Furthermore, self-supervision at network terminal could obscure specific semantics in hyperfine spectra. In this paper, we propose a hierarchical and bidirectional learnable prototypical contrastive learning method (HiBiCo) for HSI representation. We build learnable prototype dictionaries at shallow and final layers of the network for deep contrastive supervision. Moreover, we introduce reverse contrastive learning under negative sample dominance to address excessively uniform representations caused by traditional positive-dominated contrastive loss in dual-dictionary hierarchical supervision. By allowing deep contrastive supervision and two-way information flow along with the general InfoNCE loss, our approach alleviates the uniform distribution, maintains the latent data manifold, and enables diverse and effective representation. Experiments with linear probing demonstrate the effectiveness of our HiBiCo framework in handling complex scenes, highlighting the potential of self-supervised pretraining for hyperspectral image representation. The code is available at https://github.com/sakurashine/HiBiCo.
Jian Dong 0004, Miaomiao Liang, Zhi He, Chengle Zhou
IEEE Trans. Geosci. Remote. Sens.3
2025 Low-Rank and Sparse Representation Meet Deep Unfolding: A New Interpretable Network for Hyperspectral Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is a technique that intelligently checks the changed details in bitemporal hyperspectral images (Bi-HSIs). Deep learning (DL), with the ability to model nonlinear changing features, has achieved promising results in HSI-CD, but the feature mining mechanism is unclear and the architecture design lacks transparency in such DL models. To alleviate this problem, this paper proposes a new low-rank and sparse representation-based deep unfolding network (LRSRNet) for HSI-CD. For feature mining mechanism, the LRSRNet adopts a low-rank and sparse subnetwork (LRSnet) and a change detection sub-network (CDnet). The former is responsible for extracting low-rank features with valuable information and suppressing sparse features containing interference information, while the latter aims to obtain change information from low-rank features. For architecture design, the LRSnet formulates the HSI as a low-rank estimation, sparse estimation, and hyperspectral reconstruction in a low-rank and sparse model, and iteratively optimizes and updates the above sub-problems through deep networks. A new CDnet is designed as a concise convolutional architecture to extract change information from representative Bi-HSIs features. Experiments on three real datasets demonstrate the performance superiority of the proposed LRSRNet method over nine model-driven, datadriven, and model-data-joint-driven HSI-CD algorithms in both qualitative and quantitative evaluations. The proposed LRSRNet is available online: https://github.com/chengle-zhou/LRSRNet.
Chengle Zhou, Zhi He, Jian Dong 0004, Yunfei Li 0006, Jinchang Ren, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2024 Biscale Convolutional Self-Attention Network for Hyperspectral Coastal Wetlands Classification
abstract
Coastal wetlands classification is a hot but challenging issue. Hyperspectral image (HSI) can provide abundant spectral information for coastal wetlands, and deep learning excels at extracting abstract features. However, effectively leveraging global and local features to enhance the accuracy of coastal wetlands classification remains a significant challenge. In this letter, we propose a biscale convolutional self-attention network (termed as HyperBCS) for hyperspectral coastal wetlands classification. HyperBCS consists of biscale adding module (BSAM) and convolutional self-attention module (CSM). On the one hand, BSAM uses two branches to extract features of different scales. On the other hand, CSM is a paralleled structure of convolution and self-attention, which can effectively extract both local and global features. Experiments on two hyperspectral datasets captured by Zhuhai-1 satellite indicate that HyperBCS can improve the accuracy of hyperspectral coastal wetlands classification, showcasing the highest accuracy (OA = 98.29% and 96.82%, Kappa = 0.976 and 0.958 in two datasets) compared with other methods. Our code and datasets are available athttps://github.com/JeasunLok/HyperBCS.
Junshen Luo, Zhi He, Haomei Lin, Heqian Wu
IEEE Geosci. Remote. Sens. Lett.2
2024 RGB-to-HSV: A Frequency-Spectrum Unfolding Network for Spectral Super-Resolution of RGB Videos
abstract
Hyperspectral videos (HSVs) play an important role in the monitoring domain, as they can provide more information than RGB videos about the movement of interesting objects from the perspective of material interpretation. However, the acquisition of HSV data is expensive and time-consuming, whereas RGB videos are readily available. In order to obtain HSV data from its corresponding RGB data, this paper proposes a lightweight frequency-spectrum unfolding network (FSUF-Net) for spectral super-resolution (SSR) of RGB video data. Specifically, the proposed FSUF-Net method belongs to a data-knowledge-driven joint paradigm, which is an interpretable SSR model instead of an end-to-end black-box architecture. The FSUF-Net consists of five main steps. First, the conversion representation of RGB video data to HSV data is derived into an initial recovery term, a data term, and a prior term according to a variable splitting method. Second, the spectral response function between hyperspectral images (HSIs) and RGB images is utilized to achieve the initial recovery term. Third, a convolutional neural network (CNN)-based frequency-domain subnetwork (called F-Net) is designed to solve the data subproblem for recovering the spatial detail information from the HSI, and a Transformer-based spectrum-domain subnetwork (called S-Net) is developed to solve the prior subproblem for reconstructing the spectral information of the HSI. Fourth, two network modules are employed to conduct parametric self-learning. Finally, the HSV data can be obtained in a fixed number of iterations, including alternately solving the above data subproblem and the prior subproblem. Experiments performed on several real datasets demonstrated that the FSUF-Net can effectively reconstruct HSV from RGB videos as compared to traditional and state-of-the-art SSR methods. The proposed method is available online: https://github.com/chengle-zhou/HSV-SSR_FSUF-Net.
Chengle Zhou, Zhi He, Anjun Lou, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2024 SEGAC: Sample Efficient Generalized Actor Critic for the Stochastic On-Time Arrival Problem
abstract
This paper studies the problem in transportation networks and introduces a novel reinforcement learning-based algorithm, namely. Different from almost all canonical sota solutions, which are usually computationally expensive and lack generalizability to unforeseen destination nodes, segac offers the following appealing characteristics. segac updates the ego vehicle’s navigation policy in a sample efficient manner, reduces the variance of both value network and policy network during training, and is automatically adaptive to new destinations. Furthermore, the pre-trained segac policy network enables its real-time decision-making ability within seconds, outperforming state-of-the-art sota algorithms in simulations across various transportation networks. We also successfully deploy segac to two real metropolitan transportation networks, namely Chengdu and Beijing, using real traffic data, with satisfying results.
Hongliang Guo 0003, Zhi He, Wenda Sheng, Zhiguang Cao, Yingjie Zhou 0001, Weinan Gao
IEEE Trans. Intell. Transp. Syst.2
2023 Convolutional Transformer-Inspired Autoencoder for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) plays a vital role in military and civilian applications. However, compared with target detection or classification tasks, HAD is more challenging due to insufficient anomaly information and the difficulty of extracting local and global discriminative features. In this letter, a Convolutional Transformer-inspired Autoencoder (CTA) is proposed for HAD. The CTA consists of a clustering-based module and an autoencoder-based module. First, note that the number of anomalies is small and distinct from their surroundings, a clustering-based module is proposed to detect the pseudo background and anomaly samples. Second, the autoencoder module is composed of an encoder and a decoder formed from several skip-connected convolutions and multi-head attention-based transformers. The CTA is trained not only to distinguish the anomalies from the background but also to reconstruct the input hyperspectral images. Benefiting from integrating the convolution and transformer, the CTA has local and global receptive fields. Moreover, both background and anomaly information explored by the clustering-based module can be adopted to improve the separability of anomalies. Experiments on two hyperspectral datasets demonstrate that the proposed CTA achieves superior detection performance to its counterparts. The code is available at https://github.com/hzhdhz/CTA.
Zhi He, Dan He 0003, Anjun Lou, Guanglin Lai
IEEE Geosci. Remote. Sens. Lett.1
2023 Blind Superresolution of Satellite Videos by Ghost Module-Based Convolutional Networks
abstract
Deep learning (DL)-based video satellite superresolution (SR) methods have recently yielded superior performance over traditional model-based methods by using an end-to-end manner. Existing DL-based methods usually assume that the blur kernels are known and, thus, do not model the blur kernels during restoration. However, this assumption is rarely held for real satellite videos and leads to oversmoothed results. In this article, we propose a Ghost module-based convolution network model for blind SR of satellite videos. The proposed Ghost module-based video SR (GVSR) method, which assumes that the blur kernel is unknown, consists of two main modules, i.e., the preliminary image generation module and the SR results’ reconstruction module. First, the motion information from adjacent video frames and the wrapped images are explored by an optical flow estimation network, the blur kernel is flexibly obtained by a blur kernel estimation network, and the preliminary high-resolution image is generated by feeding both blur kernel and wrapped images. Second, a reconstruction network consisting of three paths with attention-based Ghost (AG) bottlenecks is designed to remove artifacts in the preliminary image and obtain the final high-quality SR results. Experiments conducted on Jilin-1 and OVS-1 satellite videos demonstrate that the qualitative and quantitative performance of our proposed method is superior to current state-of-the-art methods.
Zhi He, Dan He 0003, Rongning Qu
IEEE Trans. Geosci. Remote. Sens.1
2022 Unsupervised Video Satellite Super-Resolution by Using Only a Single Video
abstract
Recent studies have shown that deep-learning (DL)-based methods lead to improved performance in video satellite super-resolution (SR). However, the vast majority of prior work is supervised, which is restricted to artificially generated training data (e.g., predetermined bicubic downsampling). Unfortunately, in the real world, the low-resolution (LR) satellite video frames rarely obey these restrictions. To solve this problem, we resort to unsupervised learning and propose a video satellite SR method by using only a single video. The single video SR (SingleVSR) method takes advantage of the power of DL without relying on prior high-resolution (HR) and LR pairs. In the training phase, the LR frames are alternately processed by both downsampling network (i.e., NetLR) and upsampling network (i.e., NetHR). The losses obtained by LR frames and network outputs are used to optimize both NetLRand NetHR. In the testing phase, the trained NetHRis applied to generate the SR results of LR frames. In contrast to the existing video satellite SR methods, our SingleVSR does not require any assumption on degradation or any additional training data except for the single video to be tested. Experiments performed on Jilin-1 and OVS-1 satellite videos demonstrate the superiority of the proposed method.
Zhi He, Jiani Xu
IEEE Geosci. Remote. Sens. Lett.1
2022 Power Transformations and Feature Alignment Guided Network for SAR Ship Detection
abstract
Due to the capacity of full-time and full-weather working, synthetic aperture radar (SAR) images have been frequently applied to ship detection. However, the interference of speckle noise and shores poses enormous challenges to the accuracy of detection. Extracting multi-scale features is regarded as a good way to detect ships of different sizes, but features at different scales are not strictly aligned, which may further affect the detection results. Therefore, this letter proposes an anchor-free method, namely power transformations and feature alignment guided network (Pow-FAN), to solve the above problems. In Pow-FAN, we first utilize a power-based convolution block called PCB to extract features, which can suppress speckle noise and shores and enhance the ship targets. Furthermore, a novel feature alignment block named FAB is put forward to avoid the dislocation problems when integrating features of different scales. Compared with other state-of-the-art methods, Pow-FAN can achieve competitive detection results on the SAR ship detection dataset (SSDD) and the High-Resolution SAR Images Dataset (HRSID). The ablation study and discussion on the backbone and different situations further demonstrate the superiority of our network structure.
Zhi He, Anjun Lou
IEEE Geosci. Remote. Sens. Lett.2
2022 Multiframe Video Satellite Image Super-Resolution via Attention-Based Residual Learning
abstract
Video satellite can generate video image sequences with rich dynamic information, thus providing a new way for monitoring moving objects. However, to maintain high temporal resolution, video satellite images usually sacrifice their spatial resolution. Therefore, super-resolution (SR) plays a vital role in improving the quality of video satellite images. In this article, we propose a multiframe video SR neural network (MVSRnet) for video satellite image SR reconstruction. The proposed MVSRnet consists of three main subnetworks: an optical flow estimation subnetwork (OFEnet), an upscaling subnetwork (Upnet) and an attention-based residual learning subnetwork (ARLnet). The OFEnet aims to estimate low-resolution (LR) optical flow of multiple image frames. Upnet is then constructed to enhance the resolution of both input frames and the estimated LR optical flows. Motion compensation is subsequently performed according to the high-resolution (HR) optical flows. Finally, the compensated HR cube is fed to the ARLnet to generate SR results. Different from existing video satellite image SR methods, the proposed MVSRnet is a multiframe-based method with an attention mechanism, which can merge the motion information among adjacent frames and highlight the importance of extracted features. Experiments conducted on Jilin-1 and OVS-1 video satellite images demonstrate that the proposed MVSRnet significantly outperforms some state-of-the-art SR methods.
Zhi He, Jun Li 0009, Lin Liu 0005, Dan He 0003
IEEE Trans. Geosci. Remote. Sens.1
2022 Convolutional Two-Stream Generative Adversarial Network-Based Hyperspectral Feature Extraction
abstract
Hyperspectral image processing is faced with difficulties considering its redundant features and complex information. Studies on hyperspectral feature extraction in the deep learning domain have become increasingly popular. The mainstream techniques fully consider the spatial information in local neighborhoods when extracting spectral features by constructing deep neural networks. Deep generative models simulate the intrinsic structure of samples by adequately training, showing their potential values for signal processing. In this article, a convolutional two-stream network (cs2GAN-FE) based on the improved Wasserstein generative adversarial network (WGAN) is proposed for unsupervised hyperspectral spatial–spectral feature extraction. The improved WGAN is composed of one generator and one discriminator; the former perceives real data distributions, and the latter determines the attribution of generated data. The designed two-stream strategy is not a simple extension of a one-stream strategy and considers both the static spectral–spatial information and the dynamic spectral reflectance variation in multiple bands. Intrinsic spatial–spectral features are extracted by the trained discriminator considering sample distributions and feature relationships. The loss function is also improved for the unique structure of cs2GAN-FE. Various state-of-the-art techniques are chosen for comparison. Experimental results show the feasibility and potential of this network. Besides, experiments with the random split and the disjointed split both show that the proposed method can outperform other comparison techniques.
Wenbo Yu 0001, Miao Zhang 0001, Zhi He, Yi Shen 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Navigation With Time Limits in Transportation Networks: A Fourth Moment Approach
abstract
This paper investigates the stochastic on-time arrival (SOTA) problem in transportation networks. We propose a fourth moment approach (FMA), which calculates the tight lower bound of a given routing policy’s on-time-arrival probability, through estimating the first four moments of the policy’s travel time. Then, we employ the generalized policy iteration (GPI) scheme to gradually improve the policy towards the optimal one. Different from state-of-the-art algorithms for the SOTA problem, which require the full travel time distribution and usually incur high computational cost due to the convolution integration operation, FMA only requires the moments of travel-time statistics, which are easily estimated from the statistics perspective. Moreover, the algorithm’s computational complexity analysis indicates the relatively light computational load requirement of FMA. Experimental results in a range of transportation networks show FMA’s superior performance over state of the arts.
Hongliang Guo 0003, Zhi He, Chen Gao 0009, Daniela Rus
IEEE Trans. Intell. Transp. Syst.2
2022 Traffic Information Mining From Social Media Based on the MC-LSTM-Conv Model
abstract
Social media (e.g., Sina Weibo) have the advantage of reflecting traffic information, including the reasons for jams, illegal behaviors, and emergency recourses on roads. However, there remains a challenging issue regarding how to sufficiently mine traffic information. In this paper, we propose a deep learning-based method that uses social media data for traffic jam management. The core ideas of the proposed method are twofold. First, a multichannel network with a Long Short-Term Memory layer (LSTM-layer) and a Convolution layer (Conv-layer) (termed as MC-LSTM-Conv) is proposed. This model consists of two information channels for extracting abstract features from input text. Each channel includes two Conv-layers, and an LSTM-layer is added to one of the four Conv-layers. The MC-LSTM-Conv model is used to extract check-in microblogs reflecting traffic jams from mass Sina Weibo data. Second, a series of matching rules are constructed based on the keywords that are related to traffic-jam scenes. These rules further classify the microblogs extracted by the first step into four classes, and each of the classes reflects a specific road condition (i.e., traffic accidents or large-scale activities, road construction, traffic lights, and the low efficiency of government agencies). Experiments on Sina Weibo data demonstrate that the proposed multichannel network has superior performance in extracting microblogs about traffic jams. The keyword fuzzy matching method can fetch detailed information about traffic jams efficiently.
Zhi He
IEEE Trans. Intell. Transp. Syst.2
2021 Bilinear Squeeze-and-Excitation Network for Fine-Grained Classification of Tree Species
abstract
Tree species classification is beneficial to multiple applications but is difficult because the categories should be discriminated by subtle differences, and the acquisition of class labels is both expensive and time-consuming. In this letter, a bilinear squeeze-and-excitation network (BiSENet) is proposed for fine-grained classification of tree species. First, objects of the remote sensing data are constructed based on superpixel segmentation. Second, a deep neural network (i.e., BiSENet) is constructed and trained to distinguish different tree species. The proposed BiSENet is inspired by the fine-grained image classification, which is more subtle than traditional classification since it classifies the images within a subordinate category. Moreover, the AdaBound optimization method is adopted to obtain the optimal parameters of BiSENet. Experiments on the Haizhu Lake data acquired by the Jilin-1 satellite demonstrate that the proposed method exhibits superior quantitative and qualitative performance than existing state-of-the-art methods.
Zhi He, Dan He 0003
IEEE Geosci. Remote. Sens. Lett.1
2021 A Unified Network for Arbitrary Scale Super-Resolution of Video Satellite Images
abstract
Super-resolution (SR) has attracted increasing attention as it can improve the quality of video satellite images. Most previous studies only consider several integer magnification factors and focus on obtaining a specific SR model for each scale factor. However, in the real world, it is a common requirement to zoom the videos arbitrarily by rolling the mouse wheel. In this article, we propose a unified network for arbitrary scale SR (ASSR) of video satellite images. The proposed ASSR consists of two modules, i.e., feature learning module and arbitrary upscale module. The feature learning module accepts multiple low-resolution (LR) frames and extracts useful features of those frames by using many 3-D residual blocks. The arbitrary upscale module takes the extracted features as input and enhances the spatial resolution by subpixel convolution and bicubic-based adjustment. Different from existing video satellite image SR methods, ASSR can continuously zoom LR video satellite images with arbitrary integer and noninteger scale factors in a single model. Experiments have been conducted on real video satellite images acquired by Jilin-1 and OVS-1. Quantitative and qualitative results have demonstrated that ASSR has superior reconstruction performance compared with the state-of-the-art SR methods.
Zhi He, Dan He 0003
IEEE Trans. Geosci. Remote. Sens.1
2020 PySRResNet: Super Resolution for Video Satellite Imagery via Pyramid Residual Network
abstract
Video satellite is of great significance in change detection and military reconnaissance due to its high temporal resolution. However, restricted by the hardware conditions, spatial resolution must be sacrificed if temporal resolution is to be guaranteed. Because of this, how to reconstruct super resolution (SR) video satellite data is particularly important. Based on the proposed SR Residual Network (SRResNet), we proposed a Pyramid Residual Network (PySRResNet) model, using a pyramid structure to obtain features of different scales, including 1, 1/2 and 1/4, and concatenate them together to provide more detailed information for SR reconstruction. In addition, we reduced the number of blocks and removed the batch normalization layer to achieve good performance. Training with “Jilin-1” video satellite images, our PySRResNet can get superior grades than other comparing models both in PSNR and SSIM, which demonstrates the effectiveness of PySRResNet in video satellite imagery SR reconstruction.
Zhi He, Jiemin Wu
IGARSS2
2020 Object-Oriented Mangrove Species Classification Using Hyperspectral Data and 3-D Siamese Residual Network
abstract
Mangrove species classification is of particular importance for coastal conservation and restoration. However, it is challenging to distinguish species-level differences with limited training data. In this letter, we propose an object-oriented classification method for mangrove forests by using the hyperspectral image (HSI) and the 3-D Siamese residual network. First, superpixel segmentation is utilized to obtain objects with various shapes and scales. Second, 3-D patches of each object are extracted from the original HSI, and those patches containing training samples are adopted to pairwise train the network. The 3-D spatial pyramid pooling (3-D-SPP) is added in the network to extract features in multiple scales. Finally, the abstract features of test samples are learned by the trained network, and the labels are determined by the nearest neighbor classifier within the metric space. Experiments on real mangrove hyperspectral data demonstrate the effectiveness of the proposed method in species classification of mangroves.
Zhi He, Qian Shi 0001, Kai Liu 0003, Jingjing Cao, Wen Zhan, Beifen Cao
IEEE Geosci. Remote. Sens. Lett.1
2019 Video Satellite Imagery Super-Resolution via a Deep Residual Network
abstract
Recently, as a new remote sensing system, video satellite develops rapidly for long-time observation. Thanks to its high temporal resolution, video satellite has been extensively used for environmental detection, especially for dynamic target monitoring. However, limited by the imaging device, it sacrifices some of its spatial resolution. Therefore, the super-resolution (SR) technology applied to these images is crucial. Based on deep residual learning, which has obtained a great success in the single-image SR, we propose a SR network structure which consists of two main steps. First, we use multi-scale feature extraction to exploit more contextual information on video satellite imagery, which is aimed at inferring high frequency components. Then, we utilize a series of residual blocks to learn the mapping between low resolution and high resolution images in a deeper and more stable network. In our experiment, the SR reconstruction results on Jinlin-1 satellite images greatly indicate the effectiveness of our method and the potential of the residual network for video satellite imagery SR.
Jiemin Wu, Zhi He, Li Zhuo 0002
IGARSS2
2018 Monte Carlo Non-Local Means Method for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising has become an important research topic in the research community due to its significance improvements in many applications (e.g. classification). In this paper, we introduce a Monte Carlo non-local means (MCNLM) method for noise reduction of the HSI. Each band of the HSI is processed by the MCNLM, which is a randomized algorithm suitable for large-scale patch-based image (e.g. HSI) filtering. More specifically, the MCNLM is achieved by randomly choosing a fraction of the similarity weights to obtain an approximated result. Compared to the classical non-local means (NLM), the MCNLM consumes less time while achieves comparable performance. Experimental results on the real hyperspectral data set demonstrate the promising performance of the MCNLM for HSI denoising.
Chuyin Deng, Liyan Li, Zhi He, Jun Li 0009, Yuanhui Zhu
IGARSS3
2018 Wide Contextual Residual Network with Active Learning for Remote Sensing Image Classification
abstract
In this paper, we propose a wide contextual residual network (WCRN) with active learning (AL) for remote sensing image (RSI) classification. Although ResNets have achieved great success in various applications (e.g. RSI classification), its performance is limited by the requirement of abundant labeled samples. As it is very difficult and expensive to obtain class labels in real world, we integrate the proposed WCRN with AL to improve its generalization by using the most informative training samples. Specifically, we first design a wide contextual residual network for RSI classification. We then integrate it with AL to achieve good machine generalization with limited number of training sampling. Experimental results on the University of Pavia and Flevoland datasets demonstrate that the proposed WCRN with AL can significantly reduce the needs of samples.
Shengjie Liu 0001, Haowen Luo, Ying Tu, Zhi He, Jun Li 0009
IGARSS4
2018 Low-rank tensor learning for classification of hyperspectral image with limited labeled samples
Zhi He, Jie Hu 0012
Signal Process.1
2018 Kernel Low-Rank Multitask Learning in Variational Mode Decomposition Domain for Multi-/Hyperspectral Classification
abstract
Multitask learning (MTL) has recently yielded impressive results for classification of remotely sensed data due to its ability to incorporate shared information across multiple tasks. However, it remains a challenging issue to achieve robust classification results in the case that the data are from nonlinear subspaces. In this paper, we propose a kernel low-rank MTL (KL-MTL) method to handle multiple features from the 2-D variational mode decomposition (2-D-VMD) domain for multi-/hyperspectral classification. On the one hand, a nonrecursive 2-D-VMD method is applied to extract various features [i.e., intrinsic mode functions (IMFs)] of the original data concurrently. Compared with the existing 2-D empirical mode decomposition, 2-D-VMD has much stronger mathematical foundation and does not need any recursive sifting process. On the other hand, KL-MTL is proposed for classification by taking the extracted IMFs as features of multiple tasks. In KL-MTL, the low-rank representation formulated by nuclear norm can capture global structure of multiple tasks, while the kernel tricks are utilized for nonlinear extension of the low-rank MTL. Moreover, the optimization problem in KL-MTL is solved by the inexact augmented Lagrangian method. Compared with several state-of-the-art feature extraction and classification methods, the experimental results using both multi-/hyperspectral images demonstrate that the proposed method has satisfactory classification performance.
Zhi He, Jun Li 0009, Kai Liu 0003, Lin Liu 0005, Haiyan Tao
IEEE Trans. Geosci. Remote. Sens.1
2017 Hyperspectral classification based on kernel low-rank multitask learning
abstract
In this paper, we propose a kernel low-rank multitask learning (KL-MTL) method to handle multiple features from the variational mode decomposition (VMD) domain for hyperspectral (HSI) classification. Core ideas of the proposed method are twofold: 1) a non-recursive VMD method is applied to extract various features (i.e. intrinsic mode functions (IMFs)) of the original data concurrently; 2) KL-MTL is proposed for classification by taking the extracted IMFs as multiple tasks. In KL-MTL, the low-rank representation formulated by nuclear norm can capture global structure of multiple tasks while the kernel tricks are utilized for nonlinear extension of the low-rank multitask learning (MTL). Experimental results using the real hyperspectral data demonstrate that the proposed methods have satisfactory classification performance.
Zhi He, Jun Li 0009, Lin Liu 0005
IGARSS1
2017 Social Media: New Perspectives to Improve Remote Sensing for Emergency Response
abstract
Remote sensing is a powerful technology for Earth observation (EO), and it plays an essential role in many applications, including environmental monitoring, precision agriculture, resource managing, urban characterization, disaster and emergency response, etc. However, due to limitations in the spectral, spatial, and temporal resolution of EO sensors, there are many situations in which remote sensing data cannot be fully exploited, particularly in the context of emergency response (i.e., applications in which real/near-real-time response is needed). Recently, with the rapid development and availability of social media data, new opportunities have become available to complement and fill the gaps in remote sensing data for emergency response. In this paper, we provide an overview on the integration of social media and remote sensing in time-critical applications. First, we revisit the most recent advances in the integration of social media and remote sensing data. Then, we describe several practical case studies and examples addressing the use of social media data to improve remote sensing data and/or techniques for emergency response.
Jun Li 0009, Zhi He, Javier Plaza, Shutao Li 0001, Jinfen Chen, Henglin Wu
Proc. IEEE2
2017 Three-dimensional empirical mode decomposition (TEMD): A fast approach motivated by separable filters
Zhi He, Jun Li 0009, Lin Liu 0005, Yi Shen 0001
Signal Process.1
2016 Multi-way projections-based reconstruction for hyperspectral image denoising
abstract
In this paper, we propose a multi-way projections-based reconstruction method for noise reduction of hyperspectral image (HSI). Core ideas of the proposed method are twofold: 1) the original HSI is partitioned into many small three-dimensional (3D) patches. Each of the patch is taken as a third-order tensor, on which compressive multi-way measurements are performed; 2) denoised patches are produced by the approximate low multilinear-rank reconstructions, and the final denoised HSI can be obtained by putting the denoised patches back to where they are in the original HSI. Experiments conducted on the real hyperspectral data set demonstrate the promising performance of the proposed method.
Zhi He, Jun Li 0009, Lin Liu 0005
IGARSS1
2016 Learning group-based sparse and low-rank representation for hyperspectral image classification
Zhi He, Lin Liu 0005, Suhong Zhou, Yi Shen 0001
Pattern Recognit.1
2016 Low-rank group inspired dictionary learning for hyperspectral image classification
Zhi He, Lin Liu 0005, Ruru Deng, Yi Shen 0001
Signal Process.1
2016 Fast Three-Dimensional Empirical Mode Decomposition of Hyperspectral Images for Class-Oriented Multitask Learning
abstract
In this paper, we propose a fast 3-D empirical mode decomposition (fTEMD) method for hyperspectral images (HSIs) to achieve class-oriented multitask learning (cMTL). The major steps of the proposed method are twofold: 1) fTEMD and 2) cMTL. On the one hand, the traditional empirical mode decomposition is extended to its 3-D version, which naturally treats the HSI as a cube and effectively decomposes the HSI into several 3-D intrinsic mode functions (TIMFs). To accelerate the fTEMD, 3-D Delaunay triangulation is adopted to determine the distances of extrema, whereas separable filters are implemented to generate the envelopes. On the other hand, cMTL is performed on the TIMFs by taking those TIMFs as features of different tasks. The proposed cMTL learns the representation coefficients by taking advantage of the class labels and fully exploiting the information contained in each TIMF. Experiments conducted on three benchmark data sets demonstrate the effectiveness of the proposed method.
Zhi He, Jun Li 0009, Lin Liu 0005, Kai Liu 0003, Li Zhuo 0002
IEEE Trans. Geosci. Remote. Sens.1
2015 Multiple data-dependent kernel for classification of hyperspectral images
Zhi He, Junbao Li
Expert Syst. Appl.1
2015 Regularized multivariable grey model for stable grey coefficients estimation
Zhi He, Yi Shen 0001, Junbao Li, Yan Wang 0047
Expert Syst. Appl.1
2014 Kernel Sparse Multitask Learning for Hyperspectral Image Classification With Empirical Mode Decomposition and Morphological Wavelet-Based Features
abstract
Recently, many researchers have attempted to exploit spectral–spatial features and sparsity-based hyperspectral image classifiers for higher classification accuracy. However, challenges remain for efficient spectral–spatial feature generation and combination in the sparsity-based classifiers. This paper utilizes the empirical mode decomposition (EMD) and morphological wavelet transform (MWT) to gain spectral–spatial features, which can be significantly integrated by the sparse multitask learning (MTL). In the feature extraction step, the sum of the intrinsic mode functions extracted by an optimized EMD is taken as spectral features, whereas the spatial features are formed by the low-frequency components of one-level MWT. In the classification step, a kernel-based sparse MTL solved by the accelerated proximal gradient is applied to analyze both the spectral and spatial features simultaneously. Experiments are conducted on two benchmark data sets with different spectral and spatial resolutions. It is found that the proposed methods provide more accurate classification results compared to the state-of-the-art techniques with various ratio of training samples.
Zhi He, Qiang Wang 0001, Yi Shen 0001, Mingjian Sun
IEEE Trans. Geosci. Remote. Sens.1
2013 Discrete multivariate gray model based boundary extension for bi-dimensional empirical mode decomposition
Zhi He, Qiang Wang 0001, Yi Shen 0001, Yan Wang 0047
Signal Process.1
2012 Boundary extension for Hilbert-Huang transform inspired by gray prediction model
Zhi He, Yi Shen 0001, Qiang Wang 0001
Signal Process.1
2006 OMVD: An Optimization of MVD
Zhi He, Shengfeng Tian, Houkuan Huang
ADMA1