Zhen Ye 0007

dblp:46/5245-7 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0001-5410-863XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 9 first-author · 14 since 2021
YearPublicationVenuePosition
2025 Attention Multiscale Network for Semantic Segmentation of Multimodal Remote Sensing Images
abstract
Due to recent advancements in deep learning, techniques for urban structure extraction and semantic segmentation of multimodal remote sensing images have significant improvements. However, the challenge arises from the variable color intensity and complex texture of urban structures in optical images, particularly in buildings and roads. Fortunately, the light detection and ranging (LiDAR) images promote the task of developing an optimal multimodal fusion network that effectively leverages information from different modalities. In this article, we propose an attention multiscale network (AMSNet) for binary semantic segmentation tasks focused on building extraction, as well as multiclass semantic segmentation tasks, by integrating optical and LiDAR remote sensing images. AMSNet introduces two feature fusion modules—spatial scale adaptive fusion (S2AF) and semantic guided fusion (SGF). S2AF facilitates feature fusion between optical and LiDAR images within the same layer. This module contains a spatial scale selection strategy and an adaptive weight learning strategy, which enables the network to adaptively extract and intentionally select multiscale features from multimodal data. SGF addresses the semantic gap between different layered block features through semantic feature guidance strategy while achieving feature fusion. Furthermore, we introduce robust feature learning (RFL) to ensure the network robustness in rotation and variation in objects, making it resilient to images captured from different viewpoints and sensors. RFL incorporates point-to-point similarity learning strategy and multiscale feature reuse strategy. Experimental results on publicly available datasets demonstrate that AMSNet outperforms other state-of-the-art models. Extensive ablation studies further confirm the significance of all key components in the proposed approach. The source code of this method is available athttps://github.com/B-LG-J/AMSNet.git.
Zhen Ye 0007, Yuan Li 0037, Zhen Li 0063, Huan Liu 0015, Yuxiang Zhang 0005, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2024 FRCFNet: Feature Reassembly and Context Information Fusion Network for Road Extraction
abstract
Existing road extraction methods based on very high resolution (VHR) satellite imagery suffer from insufficient multidimensional feature expression and difficulty capturing global context. We propose a grouping multidimensional feature reassembly (GMFR) module, performing channel, height, and width reassembly of multiscale features between network layers via gating to focus on valid information. Given the distinct geometric structure of roads, we propose a novel module, multidirectional context information fusion (MCIF), utilizing four strip convolutions to capture the long-distance context in various directions within VHR images. It aggregates global information through two pooling branches. Based on these, we designed a road extraction network, FRCFNet, with an encoder–decoder structure and skip connections. The proposed network efficiently fuses multiscale features while capturing global context from various directions and reducing complexity. Experimental results show that the proposed method achieves 68.97% and$80.23\%~F1$-score on CHN6-CUG and DeepGlobe datasets, respectively, outperforming other comparison methods. The code will be posted athttps://github.com/CHD-IPAC/FRCFNet.
Haijuan Wang, Danni Xue, Moslema Chowdhuray Momi, Zhen Ye 0007, Siwen Quan
IEEE Geosci. Remote. Sens. Lett.5
2024 ELLK-Net: An Efficient Lightweight Large Kernel Network for SAR Ship Detection
abstract
ELLK-Net, an efficient, lightweight network with a large kernel, is proposed for synthetic aperture radar (SAR) ship detection. It addresses background variations, different ship scales, and noise interference challenges. ELLK-Net uses an anchor-free detector framework and sequentially decomposes large kernel convolutions to capture comprehensive global information and long-range dependencies. It adaptively selects convolution kernels on the basis of target characteristics, enhancing multiscale feature expression. A novel large kernel multiscale attention (LKMA) module is introduced to enhance interlayer feature fusion and semantic alignment, mitigating the impacts of overlapping ships and scattering noise. Structural reparameterization techniques optimize inference speed across devices without compromising accuracy. The experimental results on the SAR ship detection dataset (SSDD) and high-resolution SAR image dataset (HRSID) datasets demonstrate that ELLK-Net achieves impressive AP50 values of 95.6% and 90.6% for horizontal box detection and 89.7% and 79.7% for rotating box detection, respectively. The reparameterized detector exhibits a significant 48.7% FPS improvement on the Nvidia Jetson NX platform, indicating its suitability for edge computing deployment. The code is available athttps://github.com/CHD-IPAC/ELLK-Net.
Moslema Chowdhuray Momi, Siwen Quan, Zhen Ye 0007
IEEE Trans. Geosci. Remote. Sens.6
2024 Self-Supervised Learning With Multiscale Densely Connected Network for Hyperspectral Image Classification
abstract
In recent years, deep learning-based methods have exhibited remarkable performance in the field of hyperspectral image (HSI) classification. However, conventional supervised methods heavily rely on a substantial number of labeled samples. Self-supervised learning, as a prominent unsupervised representation learning technique, offers the potential to extract valuable information from unlabeled data. In this article, we introduce a novel unsupervised approach called self-supervised learning with the multiscale densely connected network (SS-MSDCNet) to make full use of unlabeled samples for HSI classification. First, a two-stream structure was designed to generate more positive pairs, which enables the contrast self-supervised training to learn more useful information from unlabeled data. Subsequently, a data augmentation technique based on spectral splitting was proposed to coordinate the two-stream structure of SS-MSDCNet, enhancing spectral information expression. The backbone of the proposed approach is the multiscale densely connected network (MSDCNet), which obtains input HSIs with various spatial scales by removing peripheral pixels from the original input and subsequently extracts multiscale spatial-spectral features using 3-D densely connected modules and 3-D spatial attention modules. The 3-D densely connected module effectively harnesses multiscale features extracted by various convolutional layers, while the 3-D spatial attention module enhances the network’s focus on features conducive to accurate classification. To validate the efficacy of our approach, we conducted extensive experiments using four distinct HSI datasets. The results unequivocally demonstrate that SS-MSDCNet outperforms several well-established supervised and unsupervised classification methods. Furthermore, we designed a transfer experiment to confirm SS-MSDCNet’s robust generalization capabilities. The code is available athttps://github.com/mrblank99/SS-MSDCNet.
Zhen Ye 0007, Zhan Cao, Huan Liu 0015, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2024 Cross-Domain Few-Shot Learning Based on Graph Convolution Contrast for Hyperspectral Image Classification
abstract
Training a deep-learning classifier notoriously requires hundreds of labeled samples at least. Many practical hyperspectral image (HSI) scenarios suffer from a substantial cost associated with obtaining a number of labeled samples. Few-shot learning (FSL), which can realize accurate classification with prior knowledge and limited supervisory experience, has demonstrated superior performance in the HSI classification. However, previous few-shot classification algorithms assume that the training and testing data are distributed in the same domains, which is a stringent assumption in realistic applications. To alleviate this limitation, we propose a cross-domain FSL based on graph convolution contrast (GCC-FSL). The proposed method leverages cross-domain learning to acquire transferable knowledge from the source domain for classifying samples in the target domain. Specifically, a positive and negative pairs module is designed for constructing positive and negative pairs by matching the class prototypes of the target domain with those of the source domain, which aligns the data distribution of the source and target domains. In addition, a graph convolution contrast (GCC) module is proposed for extracting global graph-structure information of HSI to improve the ability of feature expression and constructing a graph-contrast loss to solve a domain-shift problem. Finally, a multiscale feature extraction network is designed to expand convolutional receptive fields through feature reuse and increase information interaction for fine-grained feature extraction. The experimental results demonstrate the improved performance for the proposed FSL framework relative to both state-of-the-art convolutional neural network (CNN)-based methods as well as other few-shot techniques. The source code of this method can be found athttps://github.com/JieW-ww/GCC-FSL.
Zhen Ye 0007, Jie Wang 0135, Tao Sun 0021, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2024 A Multiscale Incremental Learning Network for Remote Sensing Scene Classification
abstract
To infer unknown remote sensing scenarios, for remote sensing scene classification (RSSC) most existing deep neural networks (DNNs) are trained on closed datasets. When the acquisition speed and quantity of remote sensing images increases rapidly, these models cannot be used to classify new scenes. Currently, incremental learning as an effective solution for solving thecatastrophic forgettingissue, but ignoringthe stability-plasticity dilemma. In this paper, we propose a new incremental learning network, named efficient channel attention-based multiscale depthwise network (ECA-MSDWNet), in which efficient channel attention (ECA) improve the model’s ability to focus on critical information in complex context and multiscale depthwise convolution (MSDW Conv) extracts multiscale features in a fine-grained way. In addtion, in incremental learning process, we expand new modules based on a dynamic-structure method to fit the residuals between the labels and the outputs of the old model, enhancing the plasticity of the new model for new tasks while maintaining the performance of the old tasks. Finally, we compress the model to reduce redundant parameters and feature dimensions through an effective knowledge distillation strategy. Experiments on four open datasets demonstrate the effectiveness of our method. Our code is available at https://github.com/zhangyu-chd/ECA-MSDWNet.
Zhen Ye 0007, Yu Zhang 0200, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2023 A Novel Anchor-Free Detector Using Global Context-Guide Feature Balance Pyramid and United Attention for SAR Ship Detection
abstract
Most SAR ship detectors based on convolutional neural networks (CNNs) needed preset anchor boxes to object classification and bounding box coordinate regression. However, the sparsity and unbalanced distribution of ships in SAR images mean that most anchor boxes are redundant. Thus, the anchor settings directly affect the performance and generalization ability of the detector. In addition, a variety in ship scales and the substantial interference of inshore backgrounds bring significant challenges to the SAR ship detector’s performance improvement. In this letter, a novel anchor-free based detector, named FBUA-Net, is proposed. We adopt a keypoint-based strategy to predict bounding boxes to eliminate the influence of anchors. Besides, we propose a global context-guided feature balanced pyramid (GC-FBP), which balances the semantic information at different levels of the feature pyramid by aggregation and averaging and uses the global context module (GCM) to learn global contextual information to construct long-range dependencies between ship targets and the background. Considering the interference of scattering noise to the detector, a united attention module (UAM) is designed to reduce the interference of surrounding noise by focusing on the spatial shape and scale size of ship targets in both the spatial and scale domains. Experimental results on the SSDD and HRSID datasets show that our detector achieves state-of-the-art (SOTA) performance. The source code can be found at https://github.com/so-bright/FBUA-Net.
Zhen Ye 0007, Dongling Xue, Xiangyuan Lin, Meng Hui
IEEE Geosci. Remote. Sens. Lett.3
2023 Computationally Lightweight Hyperspectral Image Classification Using a Multiscale Depthwise Convolutional Network With Channel Attention
abstract
Convolutional networks have been widely used for the classification of hyperspectral images; however, such networks are notorious for their large number of trainable parameters and high computational complexity. Additionally, traditional convolution-based methods are typically implemented as a simple cascade of a number of convolutions using a single-scale convolution kernel. In contrast, a lightweight multiscale convolutional network is proposed, capitalizing on feature extraction at multiple scales in parallel branches followed by feature fusion. In this approach, 2D depthwise convolution is used instead of conventional convolution in order to reduce network complexity without sacrificing classification accuracy. Furthermore, multiscale channel attention is also employed to selectively exploit discriminative capability across various channels. To do so, multiple 1D convolutions with varying kernel sizes provide channel attention at multiple scales, again with the goal of minimizing network complexity. Experimental results reveal that the proposed network not only outperforms other competing lightweight classifiers in terms of classification accuracy but also exhibits a lower number of parameters as well as significantly less computational cost.
Zhen Ye 0007, Cuiling Li, Qingxin Liu, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.1
2023 Few-Shot Learning Using Residual Channel Attention and Prototype Domain Adaptation for Hyperspectral Image Classification
abstract
While deep learning has been widely employed for the classification of hyperspectral imagery, many scenarios arise in practice in which too few labeled samples exist to effectively train the networks. Few-shot learning has been recently used to deploy classifiers trained on source-domain datasets comprising a large number of labeled samples to datasets from a target domain with only few labeled samples. However, most techniques in this vein effectively assume that the source and target domains possess the same data distribution, whereas the distributions between the two domains often differ widely in practice. Adversarial domain adaption driven by prototype classifiers deployed independently in the source and target domains is proposed to handle such differing source and target distributions, while an attention-based feature extractor with residual skip connections is developed in order to weight spectral bands according to their importance to the hyperspectral classification task. Experimental results demonstrate improved performance for the proposed few-shot-learning framework relative to both fully-supervised classifiers as well as other few-shot techniques.
Zhen Ye 0007, Tao Sun 0021, Zhan Cao, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.1
2023 Local-Global Active Learning Based on a Graph Convolutional Network for Semi-Supervised Classification of Hyperspectral Imagery
abstract
Deep learning is being increasingly employed for hyperspectral classification, although such use is often predicated on the availability of a sufficiently large set of labeled samples for training. To improve classification performance under a limited training-set size, a semi-supervised network with end-to-end local–global active learning (AL) based on graph convolutional networks (GCNs) is proposed. The proposed AL extracts both global as well as local graph-based features to gauge the discriminative information in unlabeled samples, while semi-supervised classification expands the training set of a fully supervised classifier by attaching pseudo-labels to high-confidence unlabeled samples. Experimental results demonstrate that the proposed network outperforms not only other approaches to semi-supervised classification but also several existing fully supervised methods. The source code of this method can be found athttps://github.com/XtaoS/semi-LG-AGCN.
Zhen Ye 0007, Tao Sun 0021, Shihao Shi, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.1
2023 Adaptive Domain-Adversarial Few-Shot Learning for Cross-Domain Hyperspectral Image Classification
abstract
The process of annotating hyperspectral image (HSI) data is characterized by its time-consuming and labor-intensive nature. To address this challenge, researchers often employ a meta-learning paradigm known as few-shot learning (FSL), which leverages source domains containing a substantial number of labeled samples to assist in the classification of target domains with limited labeled samples. Many existing FSL methods rely on a conditional domain-adversarial strategy to mitigate the domain shift between source and target domains. However, these methods overlook the fact that the degrees of conditional distribution discrepancies between the two domains can vary significantly across different classes, leading to suboptimal conditional distribution alignment. To address this problem, we propose a framework called Adaptive Domain-Adversarial Few-Shot Learning (ADAFSL). Overall, the proposed ADAFSL employs an adaptive strategy that assigns varying weights to the conditional adversarial losses for different classes based on their respective degrees of discrepancies, thereby achieving global conditional distribution alignment. Specifically, a local alignment score map is constructed by measuring the similarity between labeled and unlabeled samples using both Euclidean and class-covariance metrics. This map is then multiplied with the conditional adversarial loss map, thus allocating more emphasis to the classes exhibiting greater discrepancies between the two domains. Moreover, to enhance cross-domain FSL, we design a multi-scale spectral-spatial feature extraction (MSFE) module, which incorporates cascaded multi-scale dilated convolutions. Experimental results on four public HSI datasets demonstrate that the proposed ADAFSL outperforms other state-of-the-art methods. The source code of this method can be found at https://github.com/JieW-ww/ADAFSL.
Zhen Ye 0007, Jie Wang 0135, Huan Liu 0015, Yu Zhang 0200, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2022 A Lightweight and Multiscale Network for Remote Sensing Image Scene Classification
abstract
Remote sensing image (RSI) scene classification plays an active role in many application areas. Due to the excellent performance of the convolutional neural networks (CNNs), which have widely applied in RSI scene classification in recent years. However, most existing methods improve the classification accuracy by improving the model parameters or fusing the features of CNNs. This will make the whole model very complicated and unable to extract multiscale features at a more granular level. This letter proposes a novel and lightweight multiscale depthwise network (MSDWNet) with efficient spatial pyramid attention (ESPA), namely ESPA-MSDWNet, with low model parameters and high accuracy in solving this problem. The ESPA-MSDWNet uses MobileNet V2 as a backbone. We represent multiscale features at a more granular level and expand the receptive fields by multiscale depthwise convolution (MSDW Conv). We also propose the ESPA module to extract dependencies between channels. The ablation experiment verifies the effectiveness of our proposed MSDW Conv and ESPA module. Experimental results on three public RSI datasets show that ESPA-MSDWNet has advantages in classification accuracy and execution efficiency over current state-of-the-art (SOTA) methods.
Qingxin Liu, Cuiling Li, Chunlin Zhu, Zhen Ye 0007
IEEE Geosci. Remote. Sens. Lett.5
2022 MsanlfNet: Semantic Segmentation Network With Multiscale Attention and Nonlocal Filters for High-Resolution Remote Sensing Images
abstract
With the development of deep learning, remote sensing image semantic segmentation has produced significant advances. The majority of existing methods use fully convolutional network (FCN) that lacks fine-grained multi-scale representation and fails to extract global context information. Thus, we improve FCN by adding two modules—multi-scale attention (MSA) and non-local filter (NLF). The MSA module enhances the network’s fine-grained multi-scale representation capability and allows modeling the inter-dependencies of feature maps among different channels. The NLF module can capture global context information by sequential using fast Fourier transform, parameter learnable filters and inverse fast Fourier transform. By using MSA module for encoder and NLF module for decoder in the FCN framework, MsanlfNet can obtain both fine-grained multi-scale spatial feature and global context information, thus achieving a balance between performance and computational effort. Experimental results on the remote sensing semantic segmentation public data sets demonstrate that our method can achieve better performance. The code is available at https://github.com/xyuanLin/MsanlfNet.
Xiangyuan Lin, Zhen Ye 0007, Dongling Xue, Meng Hui
IEEE Geosci. Remote. Sens. Lett.3
2022 Remote Sensing Image Scene Classification Using Multiscale Feature Fusion Covariance Network With Octave Convolution
abstract
In remote sensing scene classification (RSSC), features can be extracted with different spatial frequencies where high-frequency features usually represent detailed information and low-frequency features usually represent global structures. However, it is challenging to extract meaningful semantic information for RSSC tasks by just utilizing high- or low-frequency features. The spatial composition of remote sensing images (RSIs) is more complex than that of natural images, and the scales of objects vary significantly. In this article, a multiscale feature fusion covariance network (MF2CNet) with octave convolution (Oct Conv) is proposed, which can extract multifrequency and multiscale features from RSIs. First, the multifrequency feature extraction (MFE) module is used to obtain fine-grained frequency features by Oct Conv. Then, the features of different layers in MF2CNet are fused by the multiscale feature fusion (MF2) module. Finally, instead of using global average pooling (GAP), global covariance pooling (GCP) extracts high-order information from RSIs to capture richer statistics of deep features. In the proposed MF2CNet, the obtained multifrequency and multiscale features can effectively improve the performance of CNNs. Experimental results on four public RSI datasets show that MF2CNet has advantages in RSSC over current state-of-the-art methods. The source codes of this method can be found athttps://github.com/liuqingxin-chd/MF2CNet.
Qingxin Liu, Cuiling Li, Zhen Ye 0007, Meng Hui, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.4
2014 Classification Based on 3-D DWT and Decision Fusion for Hyperspectral Image Analysis
abstract
In this letter, a fusion-classification system is proposed to alleviate ill-conditioned distributions in hyperspectral image classification. A windowed 3-D discrete wavelet transform is first combined with a feature grouping-a wavelet-coefficient correlation matrix (WCM)-to extract and select spectral-spatial features from the hyperspectral image dataset. The adjacent wavelet-coefficient subspaces (from the WCM) are intelligently grouped such that correlated coefficients are assigned to the same group. Afterwards, a multiclassifier decision-fusion approach is employed for the final classification. The performance of the proposed classification system is assessed with various classifiers, including maximum-likelihood estimation, Gaussian mixture models, and support vector machines. Experimental results show that with the proposed fusion system, independent of the classifier adopted, the proposed classification system substantially outperforms the popular single-classifier classification paradigm under small-sample-size conditions and noisy environments.
Zhen Ye 0007, Saurabh Prasad, Wei Li 0032, James E. Fowler, Mingyi He
IEEE Geosci. Remote. Sens. Lett.1
2012 Locality-preserving discriminant analysis for hyperspectral image classification using local spatial information
abstract
Locality-preserving projection as well as local Fisher discriminant analysis is applied for dimensionality reduction of hyperspectral imagery based on both spatial and spectral information. These techniques preserve the local geometric structure of hyperspectral data into a low-dimensional subspace wherein a Gaussian-mixture-model classifier is then considered. In the proposed classification system, local spatial information—which is expected to be more multimodal than strictly spectral features—is used. Results with experimental hyperspectral data demonstrate that this system outperforms traditional classification approaches.
Wei Li 0032, Saurabh Prasad, Zhen Ye 0007, James E. Fowler, Minshan Cui
IGARSS3