Cheng Shi 0002

dblp:86/1102-2 · DBLP profile ↗
← Back
35ranked-venue papers
17as first author
22since 2021 · last 2026
0000-0001-8530-2005ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2026 Frequency-aware cosine similarity alignment network in remote sensing semantic segmentation
Fu-Lin He, Zhiyong Lv, Cheng Shi 0002, Jón Atli Benediktsson
Expert Syst. Appl.3
2026 Forward consistency learning with gated context aggregation for video anomaly detection
Jiahao Lyu 0001, Minghua Zhao, Xuewen Huang, Yifei Chen 0006, Shuangli Du, Jing Hu 0005, Cheng Shi 0002, Zhiyong Lv
Knowl. Based Syst.7
2026 MoBA: Motion memory-augmented deblurring autoencoder for video anomaly detection
Jiahao Lyu 0001, Minghua Zhao, Jing Hu 0005, Xuewen Huang, Shuangli Du, Cheng Shi 0002, Zhiyong Lv
Knowl. Based Syst.6
2026 Open-set domain adaptation via unknown sample exploration for hyperspectral image classification
Cheng Shi 0002, Qiguang Miao, Zhiyong Lv
Pattern Recognit.2
2026 Bidirectional skip-frame prediction for video anomaly detection with intra-domain disparity-driven attention
Jiahao Lyu 0001, Minghua Zhao, Jing Hu 0005, Runtao Xi, Xuewen Huang, Shuangli Du, Cheng Shi 0002
Pattern Recognit.7
2025 A Method for Removing Reflections from Water Surface Images Based on Pre-trained Image Restoration
abstract
Reflections on the water surface hinder the extraction of valuable information from water surface images. To remove reflections from water surface images, we construct a synthetic dataset and propose a multi-task network for water surface reflection detection and removal. Specifically, we first use a U-Net-based reflection detection module to generate a reflection mask, followed by a GAN-based network to remove the reflection. To extract multi-level features from the images, we design a color feature extraction network and a detail feature extraction network. Finally, to enhance the model's ability to remove large-area reflections, we pre-train the reflection removal network on an image restoration dataset. Experimental results on the proposed synthetic dataset and real water surface reflection images from the Internet show that our method significantly outperforms other methods in water surface reflection detection and removal.
Minghua Zhao, Rui Zhi, Shuangli Du, Jing Hu 0005, Cheng Shi 0002
ICASSP5
2025 Flexible visually secure image encryption with meta-learning compression and chaotic systems
Wei Chen 0155, Yichuan Wang 0003, Cheng Shi 0002, Guanglei Sheng, Yu Liu 0148, Xinhong Hei 0001
Neural Networks3
2025 Learning hyperspectral noisy label with global and local hypergraph laplacian energy
Cheng Shi 0002, Linfeng Lu, Minghua Zhao, Xinhong Hei 0001, Chi-Man Pun, Qiguang Miao
Pattern Recognit.1
2025 Adaptive Multitype Contrastive Views Generation for Remote Sensing Image Semantic Segmentation
abstract
Self-supervised contrastive learning is a powerful pre-training framework for learning the invariant features from the different views of remote sensing images, therefore, the performance of contrastive learning heavily depends on the generation of views. Current view generation is primarily accomplished through different transformations, and the types and parameters of the transformations are require hand-crafted. Hence, the diversity and discriminability of generated views cannot be guaranteed. To address this, we propose a multi-type views optimization method to optimize these transformations. We formulate contrastive learning as a min-max optimization problem, and transformation parameters are optimized by maximizing the contrastive loss. The optimized transformations encourage the negative sample pairs to be close and the positive sample pairs to be far apart. Different from the current adversarial view generation methods, our method can optimize both photometric transformations and geometric transformations. For remote sensing images, the geometric transformation is more critical for view generation, while the existing view optimization methods fail to achieve this. We consider the hue, saturation, brightness, contrast, and geometric rotation transformations in contrastive learning, and evaluate the optimized views on the downstream remote sensing images semantic segmentation task. Extensive experiments are carried on the three remote sensing image segmentation datasets, including ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and LoveDA dataset. Results show that the learned views obtain highly advantages compared to the hand-crafted views and other optimized views. The code associated with this paper has been released and can be accessed at https://github.com/AAAA-CS/AMView.
Cheng Shi 0002, Peiwen Han, Minghua Zhao, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.1
2025 Adaptive Pitfall: Exploring the Effectiveness of Adaptation in Skeleton-Based Action Recognition
abstract
Graph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition by exploiting the adjacency topology of body representation. However, the adaptive strategy adopted by the previous methods to construct the adjacency matrix is not balanced between the performance and the computational cost. We assume this concept ofAdaptive Trap, which can be replaced by multiple autonomous submodules, thereby simultaneously enhancing the dynamic joint representation and effectively reducing network resources. To effectuate the substitution of the adaptive model, we unveil two distinct strategies, both yielding comparable effects. (1) Optimization.Individuality and Commonality GCNs (IC-GCNs)is proposed to specifically optimize the construction method of the associativity adjacency matrix for adaptive processing. The uniqueness and co-occurrence between different joint points and frames in the skeleton topology are effectively captured through methodologies like preferential fusion of physical information, extreme compression of multi-dimensional channels, and simplification of self-attention mechanism. (2) Replacement.Auto-Learning GCNs (AL-GCNs)is proposed to boldly remove popular adaptive modules and cleverly utilize human key points as motion compensation to provide dynamic correlation support. AL-GCNs construct a fully learnable group adjacency matrix in both spatial and temporal dimensions, resulting in an elegant and efficient GCN-based model. In addition, three effective tricks for skeleton-based action recognition (Skip-Block, Bayesian Weight Selection Algorithm, and Simplified Dimensional Attention) are exposed and analyzed in this paper. Finally, we employ the variable channel and grouping method to explore the hardware resource bound of the two proposed models. IC-GCN and AL-GCN exhibit impressive performance across NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UAV-Human datasets, with an exceptional parameter-cost ratio.
Qiguang Miao, Wentian Xin, Ruyi Liu 0001, Cheng Shi 0002, Chi-Man Pun
IEEE Trans. Multim.6
2024 Attack-invariant attention feature for adversarial defense in hyperspectral image classification
Cheng Shi 0002, Minghua Zhao, Chi-Man Pun, Qiguang Miao
Pattern Recognit.1
2024 Novel Distribution Distance Based on Inconsistent Adaptive Region for Change Detection Using Hyperspectral Remote Sensing Images
abstract
Change detection with remote sensing images (RSIs) plays an important role in the community of remote sensing applications. However, when change detection is conducted with hyperspectral remote sensing images (HRSIs), how to measure the change magnitude between bitemporal HRSIs becomes challenging due to the high dimension of HRSIs. In this article, a novel Distribution Distance based on Inconsistent Adaptive Region (D2IAR) change detection approach is proposed to measure the change magnitude between bitemporal HRSIs for improving the performance of change detection with HRSIs. First, a band selection algorithm called optimal neighborhood reconstruction is employed to reduce the dimensions of HRSIs. Then, an adaptive region around each pixel is generated to explore the contextual feature around each pixel, and kernel density estimation is suggested to estimate the spectral distribution of the pixels within an adaptive region. A distribution distance is defined based on the adaptive region to measure the change magnitude between bitemporal HRSIs. Finally, the change magnitude between pairwise adaptive regions is measured by the proposed distance between the pairwise distributions. Experimental results based on four datasets and comparisons with eight methods indicated the feasibility and superiorities of the proposed D2IAR-based change detection approach with HRSIs. The improvement rates are approximately 0.13%-24.04% for overall accuracy. The code and datasets can be available at: https://github.com/ImgSciGroup/2024-HSICD.
Zhiyong Lv, Zhengjie Lei, Linfu Xie, Nicola Falco, Cheng Shi 0002, Zhenzhen You
IEEE Trans. Geosci. Remote. Sens.5
2023 Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action Recognition
abstract
Vision Transformer, which performs well in various vision tasks, encounters a bottleneck in skeleton-based action recognition and falls short of advanced GCN-based methods. The root cause is that the current skeleton transformer depends on the self-attention mechanism of the complete channel of the global joint, ignoring the highly discriminative differential correlation within the channel, so it is challenging to learn the expression of the multivariate topology dynamically. To tackle this, we present Skeleton MixFormer, an innovative spatio-temporal architecture to effectively represent the physical correlations and temporal interactivity of the compact skeleton data. Two essential components make up the proposed framework: 1) Spatial MixFormer. The channel-grouping and mix-attention are utilized to calculate the dynamic multivariate topological relationships. Compared with the full-channel self-attention method, Spatial MixFormer better highlights the channel groups' discriminative differences and the joint adjacency's interpretable learning. 2) Temporal MixFormer, which consists of Multiscale Convolution, Temporal Transformer and Sequential Holding Module. The multivariate temporal models ensure the richness of global difference expression and realize the discrimination of crucial intervals in the sequence, thereby enabling more effective learning of long and short-term dependencies in actions. Our Skeleton MixFormer demonstrates state-of-the-art (SOTA) performance across seven different settings on four standard datasets, namely NTU-60, NTU-120, NW-UCLA, and UAV-Human. Related code will be available on https://github.com/ElricXin/Skeleton-MixFormer.
Wentian Xin, Qiguang Miao, Ruyi Liu 0001, Chi-Man Pun, Cheng Shi 0002
ACM Multimedia6
2023 Auto-Learning-GCN: An Ingenious Framework for Skeleton-Based Action Recognition
Wentian Xin, Ruyi Liu 0001, Qiguang Miao, Cheng Shi 0002, Chi-Man Pun
PRCV (1)5
2023 Novel Piecewise Distance Based on Adaptive Region Key-Points Extraction for LCCD With VHR Remote-Sensing Images
abstract
Land cover change detection (LCCD) with very high-resolution remote-sensing images (VHR_RSIs) is important in observing surface change on Earth. However, pseudo changes usually reduces the accuracy of the detection map. In this paper, novel piecewise distance based on adaptive region key-points extraction called sparse key-point distance (SKPD) is developed to measure the change magnitude between the bitemporal VHR_RSIs for LCCD. The proposed approach consists of three steps. First, an adaptive region generation algorithm is promoted for exploring spatial-contextual information. Then, the adaptive region around each pixel is sparsely represented with the box-whisker plot theory and the adaptive region is converted into a sparse key point vector. Finally, a piecewise distance is defined to measure the change magnitude between the bi-temporal images. While the entire VHR_RSIs are scanned and the proposed SKPD method proceeds on a pixel by pixel basis, a change magnitude image (CMI) can be generated and a binary threshold method can be applied on the CMI to obtain a change detection map. Experimental results based on four pairs of real VHR_RSIs and four state-of-the-art methods effectively demonstrated the superiority of the proposed approach for achieving LCCD with VHR_RSIs, such as the improvements for the four datasets are 5.25%, 14.76%, 18.13%, and 22.24%, respectively in terms of overall accuracy.
Zhiyong Lv, Pingdong Zhong, Zhenzhen You, Jón Atli Benediktsson, Cheng Shi 0002
IEEE Trans. Geosci. Remote. Sens.6
2023 Universal Object-Level Adversarial Attack in Hyperspectral Image Classification
abstract
The vulnerability of deep neural networks has garnered significant attention. Various advanced adversarial attack methods have been proposed. However, these methods exhibit higher attack performance on three-band natural images while struggling to handle high-dimensional attacks in terms of attack transferability and robustness. Hyperspectral images, unlike natural images, possess high-dimensional and redundant spectral information. On one hand, different classification models focus on distinct discriminative spectral bands, leading to poor transferability. On the other hand, most existing attack methods are implemented at the pixel-level, making them less resilient to image processing-based defenses. In this paper, we address the improvement of transferability and robustness in high-dimensional attacks and introduce a universal object-level adversarial attack method in hyperspectral image classification. We found that perturbations with higher similarity in a local region can decrease the sensitivity of adversarial attacks to various discriminative spectral patterns and enhance resistance to image processing-based defenses. Consequently, we construct spatial and spectral oversegmented templates by utilizing the local smooth properties of hyperspectral images, aiming to promote similarity among perturbations within a local region. Extensive experiments conducted on two real hyperspectral image datasets validate that our method enhances the attack transferability and robustness of several existing attack methods. By incorporating the object-level adversarial attack with baseline fast gradient sign method (FGSM), momentum iterative FGSM (MI-FGSM), and variance tuning MI-FGSM (VMI-FGSM), the average transferability success rate of the proposed method has increased by 7.38% on the PaviaU dataset and 9.30% on the HoustonU 2018 dataset than the baselines, respectively. Meanwhile, the proposed method outperforms the baselines by an average of 6.19% on the PaviaU dataset and 10.05% on the HoustonU 2018 dataset in attacking image processing-based defense models. The code is available at https://github.com/AAAA-CS/SS_FGSM_HyperspectralAdversarialAttack.
Cheng Shi 0002, Mengxin Zhang, Zhiyong Lv, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.1
2022 Simple Multiscale UNet for Change Detection With Heterogeneous Remote Sensing Images
abstract
Change detection with heterogeneous remote sensing images (HRSIs) is attractive for observing the Earth’s surface when homogeneous images are unavailable. However, HRSIs cannot be compared directly because the imaging mechanisms for bitemporal HRSIs are different, and detecting change with HRSIs is challenging. In this letter, a simple yet effective deep learning approach based on the classical UNet is proposed. First, a pair of image patches are concatenated together to learn a shared abstract feature in both image patch domains. Then, a multiscale convolution module is embedded in a UNet backbone to cover the various sizes and shapes of ground targets in an image scene. Finally, a combined loss function, which incorporates the focal and dice losses with an adjustable parameter, was incorporated to alleviate the effect of the imbalanced quantity of positive and negative samples in the training progress. By comparisons with five state-of-the-art methods in three pairs of real HRSIs, the experimental results achieved by our proposed approach have the best overall accuracy (OA), average accuracy (AA), recall (RC), and F-Score that are more than 95%, 79%, 60%, and 61%, respectively. The quantitative results and visual performance indicated the feasibility and superiority of the proposed approach for detecting land cover change with HRSIs.
Zhiyong Lv, Jón Atli Benediktsson, Minghua Zhao, Cheng Shi 0002
IEEE Geosci. Remote. Sens. Lett.6
2022 Hyperspectral Image Classification With Adversarial Attack
abstract
The performance of a neural network is highly dependent on the labeled samples. However, the labeled samples are primarily clean, which prevents the network from capturing the features of the samples near the decision boundary. For hyperspectral images (HSIs), high spectral dimensions and same-spectra foreign matter lead to more boundary samples in the data. In this letter, we investigate an adversarial attack algorithm against these problems for HSIs. A modified DeepFool algorithm is implemented to generate boundary adversarial samples with minimal disturbance, and the generated boundary adversarial samples are simply added to the training set to improve the accuracy of the boundary samples in the data. Furthermore, we iteratively complete network training and boundary adversarial sample generation so that the decision boundary can be adjusted according to the real-time classification situation. Extensive experiments are carried out on the two HSI datasets, and the results demonstrate that the modified DeepFool algorithm can improve the accuracy of the decision boundary. Our findings also show that adversarial attacks are sensitive to high-dimensional and multiple-category data and are worthy of further study.
Cheng Shi 0002, Yenan Dang, Zhiyong Lv, Minghua Zhao
IEEE Geosci. Remote. Sens. Lett.1
2022 Improved Generative Adversarial Networks for VHR Remote Sensing Image Classification
abstract
With increasing spatial resolution of remote sensing images, accurate classification of land classes depends more on the number of labeled samples. However, the acquisition of labeled samples is difficult and time-consuming. Hence, generative adversarial networks (GANs) have become a new method for collecting training samples for very-high-resolution (VHR) remote sensing image classification. A traditional GAN generates new samples with the same distribution as the labeled samples. However, the generated samples have features close to their class center, and the network cannot obtain effective discriminative ability for the samples close to the decision boundary. This letter presents an improved GAN (IGAN) for VHR remote sensing image classification. In the proposed framework, the generator aims to generate synthetic samples close to the classification boundary, and the discriminator aims to constrain the labels of the synthetic samples. The obtained synthetic samples can effectively improve the classification accuracy of the classification boundary. Experiments are conducted on two VHR remote sensing images, and the results show that the proposed method performs better than several state-of-the-art methods.
Cheng Shi 0002, Zhiyong Lv, Huifang Shen
IEEE Geosci. Remote. Sens. Lett.1
2022 Explainable scale distillation for hyperspectral image classification
Cheng Shi 0002, Zhiyong Lv, Minghua Zhao
Pattern Recognit.1
2022 Multifeature Collaborative Adversarial Attack in Multimodal Remote Sensing Image Classification
abstract
Deep neural networks have strong feature learning ability, but their vulnerability cannot be ignored. Current research shows that deep learning models are threatened by adversarial examples in remote sensing (RS) classification tasks, and their robustness drops sharply in the face of adversarial attacks. Therefore, many adversarial attack methods have been studied to predict the risks faced by a network. However, the existing adversarial attack methods mainly focus on single-modal image classification networks, and the rapid growth of RS data makes multimodal RS image classification a research hotspot. Generating multimodal adversarial examples needs to consider a high attack success rate, subtle perturbation, and collaborative attack ability between different modalities. In this article, we investigate the vulnerability of multimodal RS classification networks and propose a multifeature collaborative adversarial network (MFCANet) for generating multimodal adversarial examples. Two modality-specific generators are designed to generate the multimodal collaborative perturbations with strong attack ability, and two modality-specific discriminators make the generated multimodal adversarial examples closer to the real instances. In addition, a modality-specific generative loss and a modality-specific discriminative loss are proposed, and an alternating optimization strategy is designed for training the proposed MFCANet. Extensive experiments are carried out on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen 2D dataset and ISPRS Potsdam 2D dataset. The results show that the attack performance of the proposed method is stronger than that of the fast gradient sign method (FGSM), project gradient descent (PGD), and Carlini and Wagner (C&W) attack methods.
Cheng Shi 0002, Yenan Dang, Minghua Zhao, Zhiyong Lv, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.1
2021 Local Histogram-Based Analysis for Detecting Land Cover Change Using VHR Remote Sensing Images
abstract
The majority of the change detection (CD) methods consider spatial information by using a regular window or strict mathematical model. Moreover, these methods use the spectra directly to measure the change magnitude between bitemporal images. To solve this problem, local histogram-based analysis (LHBA) is proposed for detecting a land cover change in this letter. This new approach aims to inhibit the pseudo change by defining the local histogram trend (LHT) in an adaptive manner instead of using spectral values to measure change magnitude directly. In the proposed approach, the spatial information around each pixel is first exploited by defining an adaptive local histogram. The LHT distance between the pairwise local histograms is then developed to measure the change magnitude between the pairwise pixels of bitemporal images. Finally, the change magnitude image is generated, and a binary CD is achieved by a threshold method. Experiments based on two pairs of very high-resolution remote sensing images, which refer to land use change and landslides events, demonstrate the advantages and performance of the proposed approach.
Zhiyong Lv, Tongfei Liu, Cheng Shi 0002, Jón Atli Benediktsson
IEEE Geosci. Remote. Sens. Lett.3
2020 Towards Ghost-Free Shadow Removal via Dual Hierarchical Aggregation Network and Shadow Matting GAN
abstract
Shadow removal is an essential task for scene understanding. Many studies consider only matching the image contents, which often causes two types of ghosts: color in-consistencies in shadow regions or artifacts on shadow boundaries (as shown in Figure. 1). In this paper, we tackle these issues in two ways. First, to carefully learn the border artifacts-free image, we propose a novel network structure named the dual hierarchically aggregation network (DHAN). It contains a series of growth dilated convolutions as the backbone without any down-samplings, and we hierarchically aggregate multi-context features for attention and prediction, respectively. Second, we argue that training on a limited dataset restricts the textural understanding of the network, which leads to the shadow region color in-consistencies. Currently, the largest dataset contains 2k+ shadow/shadow-free image pairs. However, it has only 0.1k+ unique scenes since many samples share exactly the same background with different shadow positions. Thus, we design a shadow matting generative adversarial network (SMGAN) to synthesize realistic shadow mattings from a given shadow mask and shadow-free image. With the help of novel masks or scenes, we enhance the current datasets using synthesized shadow images. Experiments show that our DHAN can erase the shadows and produce high-quality ghost-free images. After training on the synthesized and real datasets, our network outperforms other state-of-the-art methods by a large margin. The code is available: http://github.com/vinthony/ghost-free-shadow-removal/
Xiaodong Cun, Chi-Man Pun, Cheng Shi 0002
AAAI3
2020 BioExpDNN: Bioinformatic Explainable Deep Neural Network
abstract
In recent years, machine learning is applied in the bioinformatics and medical fields to analyze relationships among biological features and behaviors. However, it is difficult to discover the significant features of large-scale datasets. A novel feature extraction method called bioinformatic explainable deep neural network (BioExpDNN) is proposed to filter the critical features with strong influences on the dataset and to explain the interaction of features. In the practical experiments, this study adopted three biomedical science datasets from the UCI (University of California, Irvine) Machine Learning Repository: (1). Cryotherapy Data Set (CDS) contains 6 attributes and 2 classes (i.e., recovery and non-recovery); (2). Cervical Cancer Behavior Risk Data Set (CCBRDS) consists of 18 attributes and 2 classes (i.e., cervical cancer patient and healthy body); (3). Heart Failure Clinical Records Data Set (HFCRDS) includes 12 clinical attributes and 2 classes (i.e., death and life). In comparison results, extracted features were considered as inputs of a classifier based on deep neural network for classification. The classification accuracy was selected as an evaluation factor to evaluate the performance of feature extraction methods. The experimental results showed that the classification accuracies of CDS, CCBRDS, and HFCRDS were 92.59%, 100%, and 78.9%, respectively.
Cheng Shi 0002, Chi-Hua Chen 0002
BIBM2
2020 Multiscale Superpixel-Based Hyperspectral Image Classification Using Recurrent Neural Networks With Stacked Autoencoders
abstract
This paper develops a novel hyperspectral image (HSI) classification framework by exploiting the spectral-spatial features of multiscale superpixels via recurrent neural networks with stacked autoencoders. The superpixels can be used to segment an HSI into shape-adaptive regions, and multiscale superpixels can capture the object information more accurately. Therefore, the superpixel-based classification methods have been studied by many researchers. In this paper, we propose a multiscale superpixel-based classification method. In contrast to current research, the proposed method not only captures the features of each scale but also considers the correlation among different scales via recurrent neural networks. In this way, the spectral-spatial information within a superpixel is more efficiently exploited. In this paper, we first segment the HSI from coarse to fine scales using the superpixels. Then, the spatial features within each superpixel and among superpixels are sufficiently exploited by the local and nonlocal similarity measure. Finally, recurrent neural networks with stacked autoencoders are proposed to learn the high-level multiscale spectral-spatial features. Experiments are conducted on real HSI datasets. The results demonstrate the superiority of the proposed method over several well-known methods in both visual appearance and classification accuracy.
Cheng Shi 0002, Chi-Man Pun
IEEE Trans. Multim.1
2019 Adaptive multi-scale deep neural networks with perceptual loss for panchromatic and multispectral images classification
Cheng Shi 0002, Chi-Man Pun
Inf. Sci.1
2018 Perceptual Loss for Superpixel-Level Multispectral and Panchromatic Image Classification
abstract
Convolutional neural networks (CNNs) have proven to be an effective way for deep feature extraction. However, multispectral and panchromatic images are susceptible to illumination unevenness and noise, and the default cross entropy loss function consider only the local information, resulting in misclassification. In this paper, we propose a novel super-pixel-level deep neural networks for multispectral and panchromatic images classification, and define a novel percep-tualloss function via non-local spectral and structure similarity to suppress the interference of unbalanced light and noise. We also propose the corresponding iteration optimization algorithm in this paper. Experimental results show that the proposed method performs better than the state-of-the-art methods.
Cheng Shi 0002, Chi-Man Pun
ICASSP1
2018 Multi-scale hierarchical recurrent neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Neurocomputing1
2018 Superpixel-based 3D deep neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Pattern Recognit.1
2018 Adaptive Hierarchical Multinomial Latent Model With Hybrid Kernel Function for SAR Image Semantic Segmentation
abstract
Synthetic aperture radar (SAR) images have been one of the important tools to support earth observations and topographic measurements. It means that SAR images are essentially rich in structures. However, the single spatial relationship is difficult to deal with the heterogeneous structures of the SAR images. In this paper, we propose an adaptive hierarchical multinomial latent model with hybrid kernel function for SAR image semantic segmentation. In the proposed approach, we design a hybrid kernel function combing Gaussian radial basis function (GRBF) and ridgelet kernel function to adaptively describe the spatial relationships between the central pixel and the surrounding pixels. Then, based on the hybrid kernel function, adaptive methods are proposed for semantic segmentation. Specifically, an SAR image is divided into different characteristics subspaces, homogeneous, structural, and aggregated subspaces, by SAR hierarchical semantic model. For the homogeneous subspace, GRBF is used to describe the isotropic spatial relationships. Then, multilayer multinomial latent model with GRBF is used for segmentation to improve the labeling consistency and reduce the wrong segmentation. For the structural subspace, the ridgelet kernel function is used to describe the anisotropic spatial relationships. Then, we adopt the single-layer multinomial latent model with ridgelet kernel function for segmentation to preserve the details (such as edge, lines, and small objects). For aggregated subspace, bag-of-words model is used to extract the features of the aggregated portions, and then affinity propagation cluster is used for segmentation. Finally, the segmentation results of different subspaces are integrated together to obtain the final segmentation result. Comprehensive experiments on both synthetic and real SAR images demonstrate that the segmentation results by our proposed approach achieve the semantic consistency, labeling consistency, and detail preservation simultaneously.
Yiping Duan, Fang Liu 0001, Licheng Jiao, Xiaoming Tao 0001, Jie Wu 0016, Cheng Shi 0002, Martin O. Wimmers
IEEE Trans. Geosci. Remote. Sens.6
2017 3D multi-resolution wavelet convolutional neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Inf. Sci.1
2015 Pan-sharpening via regional division and NSST
Cheng Shi 0002, Fang Liu 0001, Qiguang Miao
Multim. Tools Appl.1
2015 Learning Interpolation via Regional Map for Pan-Sharpening
abstract
Although the bandwidth of the high-resolution panchromatic (HR PAN) image is wide, it is narrow in each band of the low-resolution multispectral (LR MS) image. Hence, the spatial resolution of the HR PAN image is much higher than that of the LR MS image. However, HR PAN image only has a single band. The purpose of the Pan-sharpening algorithm is to make the Pan-sharpened image with both high spatial resolution and good spectral information. In this paper, a novel learning interpolation method for Pan-sharpening is proposed by expanding the sketch information in the HR PAN image. The sketch information contains the edges and lines features of the image, and each segment of the sketch information has its own direction. According to the primal sketch graph of the HR PAN image, a regional map is obtained by a designed geometrical template. Since the size of the HR PAN image is different from that of the LR MS image, the LR MS image is interpolated into an interpolated multispectral (IMS) image by the nearest interpolation method. In addition, the IMS image can be mapped into the structure and the nonstructure regions by this regional map. The nonstructure regions are divided into the smooth and the texture regions by a variance value. For the structure and texture regions, the interpolated pixels in the IMS image are relearned and readjusted by the proposed structure and texture learning interpolation method, respectively. Experimental results show that the proposed Pan-sharpening method can provide superior performance in both visual effect and quality metrics, particularly for the images with a large spectral difference.
Cheng Shi 0002, Fang Liu 0001, Lingling Li 0002, Licheng Jiao, Yiping Duan, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2013 A novel algorithm of remote sensing image fusion based on Shearlets and PCNN
Cheng Shi 0002, Qiguang Miao, Pengfei Xu 0003
Neurocomputing1
2012 An edge detection algorithm based on the multi-direction shear transform
Pengfei Xu 0003, Qiguang Miao, Cheng Shi 0002, Weisheng Li 0001
J. Vis. Commun. Image Represent.3