VLDB 2026 Research / reviewers in the wild / expert
Yanfei Zhong
dblp:24/2247
· DBLP profile ↗
254ranked-venue papers
21as first author
141since 2021 · last 2026
0000-0001-9446-5850ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 213 · 18 first-author · 110 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 14 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaRS: A Multi-modality Very-high-resolution Remote Sensing Foundation Model with Cross-Granularity Meta-Modality LearningabstractThe multi-modality remote sensing foundation model (MM-RSFM) has made notable progress recently. However, most existing approaches remain limited to medium-resolution, single-modality, restricting their performance in fine-grained downstream applications such as disaster response and urban planning. In this work, MaRS is proposed, a multi-modality very-high-resolution (VHR) remote sensing foundation model designed for cross-modality granularity interpretation of complex scenes. To achieve this, a multi-modality VHR SAR-optical dataset, MaRS-16M, is constructed through large-scale collection and semi-automated processing, comprising over 16 million paired samples. Unlike previous work, MaRS tackles two fundamental challenges in VHR SAR-optical self-supervised learning (SSL) techniques. Cross-granularity contrastive learning (CGCL) is introduced to alleviate alignment inconsistencies caused by imaging differences, and meta-modality attention (MMA) is designed to unify heterogeneous physical characteristics across modalities. Compared to existing remote sensing foundation models (RSFMs) and general vision foundation models (VFMs), MaRS performs better as a pre-trained backbone across nine multi-modality VHR downstream tasks. Ruoyu Yang, Yinhe Liu, Heng Yan, Yiheng Zhou, Yihan Fu, Yanfei Zhong |
AAAI | 7 |
| 2026 | DisasterKD: Frequency-guided cross-decoder knowledge distillation for UAV real-time disaster damage assessment
Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
Pattern Recognit. | 5 |
| 2026 | Robust Fine-Grained Oriented Ship Detection for Remote Sensing Imagery via Controllable Generative PretrainingabstractFine-grained ship recognition in remote sensing imagery is essential for maritime applications. However, its development is hindered by two challenges: 1) the limited granularity of existing ship detection datasets, and 2) the disturbance of complex maritime conditions as well as the arbitrary ship orientations and distributions. To address the first issue, we annotated a large-scale fine-grained ship instance detection dataset (LAFI), comprising 48,717 ship instances worldwide with 49 categories. To tackle the challenges of marine disturbance and diverse ship status, we proposed a controllable generative knowledge-driven ship detection framework (COSD). It employs a controllable diffusion model guided by ship-marine textual prompt to generate millions of synthetic images that not only preserve ship structures but also cover diverse sea and weather conditions for robust pretraining. The pretraining stage then utilizes masked reconstruction to learn component-level cues under occlusion, clutter, fog, and illumination changes. Furthermore, a heterogeneous feature alignment decoder is designed to align multi-modal metrics of orientation and distribution features in the latent space, allowing for accurate representation of diverse ship status. Extensive experiments on two benchmark datasets showed that our method respectively increased 0.011 and 0.030 mean average precision (mAP@50) over SOTA methods, particularly in scenarios involving small, densely packed and arbitrary oriented ships. Da He, Xikun Hu, Ping Zhong 0001, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 10 |
| 2025 | HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryabstractAdvanced interpretation of hyperspectral remote sensing images benefits many precise Earth observation tasks. Recently, visual foundation models have promoted the remote sensing interpretation but concentrating on RGB and multi-spectral images. Due to the varied hyperspectral channels, existing foundation models would face image-by-image tuning situation, imposing great pressure on hardware and time resources. In this paper, we propose a tuning-free hyper-spectral foundation model called HyperFree, by adapting the existing visual prompt engineering. To process varied channel numbers, we design a learned weight dictionary covering full-spectrum from 0.4 ∼ 2.5 μm, supporting to build the embedding layer dynamically. To make the prompt design more tractable, HyperFree can generate multiple semantic-aware masks for one prompt by treating feature distance as semantic-similarity. After pre-training HyperFree on constructed large-scale high-resolution hyperspectral images, HyperFree (1 prompt) has shown comparable results with specialized models (5 shots) on 5 tasks and 11 datasets. Code and dataset are accessible at https://rsidea.whu.edu.cn/hyperfree.htm. Yingyi Liu, Xinyu Wang 0003, Yunning Peng, Shaoyu Wang 0003, Zhendong Sun, Tian Ke, Tangwei Lu, Anran Zhao, Yanfei Zhong |
CVPR | 12 |
| 2025 | DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City UnderstandingabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite comprises 14,871 high-resolution (1.0m) multi-temporal images spanning 42 major cities in the U.S. from 2005 to 2023, organized into two components: DVL-Bench and DVL-Instruct. The DVL-Bench includes six urban understanding tasks, from fundamental change detection (pixel-level) to quantitative analyses (regional-level) and comprehensive urban narratives (scene-level), capturing diverse urban dynamics including expansion/transformation patterns, disaster assessment, and environmental challenges. We evaluate 18 state-of-the-art MLLMs and reveal their limitations in long-term temporal understanding and quantitative analysis. These challenges motivate the creation of DVL-Instruct, a specialized instruction-tuning dataset designed to enhance models' capabilities in multi-temporal Earth observation. Building upon this dataset, we develop DVLChat, a baseline model capable of both image-level question-answering and pixel-level segmentation, facilitating a comprehensive understanding of city dynamics through language interactions. Project: https://github.com/weihao1115/dynamicvl. Weihao Xuan, Heli Qi, Zihang Chen 0001, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya |
NeurIPS | 6 |
| 2025 | A demographic optimization proximity model for air pollution exposure assessmentabstractAir pollution poses a significant threat to human health. Effective and efficient assessing individual exposure intensity will provide a crucial reference to mitigate potential air pollution risk. Previous proximity models in scenario simulation method neglect the variations among individual absorption efficiency. In this study, we proposed a novel demographic optimization proximity model (DOPM) to quantified sulfur dioxide (SO2) exposure risk under eight groups’ respiratory rates. In assessing exposure risk in Wuhan, one of the megacities in central China, we compared the capacity of DOPM and a classic proximity model on simulating exposure intensity, utilizing the near-Gaussian bi-square diffusion function and the filter optimal bandwidth. Subsequently, we utilized ordinary kriging (OK) interpolation to visualize the SO2 exposure risk map based on the simulation results of DOPM. We found that 9.5 km was the optimal bandwidth when near-Gaussian bi-square diffusion function utilized in proximity model. We also observed DOPM has stronger ability to simulate individual exposure intensity with a R-square of 0.44 when distinguished respiratory rate among receptors. These insights will aid public health researchers in assessing exposure risk across absorption efficiency and help local government officials devise more effective air pollution control strategies. Dingming Zhang, Xinyu Wang 0003, Wanqiang Yao, Yanfei Zhong |
Int. J. Geogr. Inf. Sci. | 4 |
| 2025 | MDTNet: Partial transformer with degradation-aware module for restoring old photos with multiple degradations
Liqin Cao, Xuan Zhang 0007, Ju Hua Liu, Yanfei Zhong |
Neurocomputing | 6 |
| 2025 | A Global-Local Collaborative and Decomposition-Based Multiobjective Evolutionary Optimization Method for UAV 3-D Path PlanningabstractIn the context of the widespread application of unmanned aerial vehicles (UAVs) across various industries, effective path planning in three-dimensional (3-D) environments has emerged as a crucial challenge in their deployment. In real-world applications, UAV path planning missions are usually converted into multi-objective tasks and solved using evolutionary computation, where the optimal flight path should consider both the overall flight route length and potential terrain threat. However, the existing methods usually treat complete paths as individuals, and this modeling approach lacks the evaluation of track points and is unable to fully reflect the quality of the path. In addition, as the quantity of track points increases, it is difficult for the traditional genetic crossover operator to quickly converge to the global optimum in complex high dimensional objective space. Thus, in this paper, we propose a UAV 3-D path planning method utilizing the global-local collaborative modeling approach with a decomposition-based method (P2GLCM). In the P2GLCM method, the global objective functions and the local objective functions are used to evaluate the path and track points, respectively, to achieve accurate modeling. In addition, to efficiently utilize the high-quality track points in the candidate paths, a dominance relationship approach is introduced to guide the generation of offsprings in a point-by-point manner, improving the search capability in complex objective space. The experimental results on 3-D environments with unified representation of voxels demonstrate that P2GLCM outperforms current methods in convergence and effectiveness. Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Internet Things J. | 4 |
| 2025 | Learning Global Context and Fine Structures for Enhanced Hyperspectral Subpixel MappingabstractSubpixel mapping (SPM) is a crucial technique in remote sensing imagery analysis, aimed at characterizing subpixel distribution within the mixed pixels. Traditional SPM methods and convolutional neural network (CNN)-based SPM methods primarily rely on local spatial autocorrelation, which limits their ability to capture long-range dependencies between distant locations or objects. To address this limitation, we propose a global-local spatial dependence integrator for the SPM method (GLSDSPM) that employs both CNN and the vision transformer as dual-path structures to model global context and local spatial dependencies efficiently. Besides, the previous SPM methods often struggle to accurately reconstruct high-quality spatial patterns for linear features, such as slender rivers and roads, due to insensitivity to textures and sharp, high-frequency details. To overcome this challenge, we integrate a linear pattern refinement module (LPRM) into GLSDSPM, which adaptively focuses on thin and long local structures to accurately capture high-frequency features and detailed information. Two experiments conducted on the Pavia and Houston hyperspectral images prove that the proposed method achieves superior performance, outperforming the state-of-the-art by 4.13% and 3.95% in overall accuracy (OA), respectively. Wen Zhou 0018, Ailong Ma, Da He, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Changen2: Multi-Temporal Remote Sensing Generative Change Foundation ModelabstractOur understanding of the temporal dynamics of the Earth's surface has been significantly advanced by deep vision models, which often require a massive amount of labeled multi-temporal images for training. However, collecting, preprocessing, and annotating multi-temporal remote sensing images at scale is non-trivial since it is expensive and knowledge-intensive. In this paper, we present scalable multi-temporal change data generators based on generative models, which are cheap and automatic, alleviating these data problems. Our main idea is to simulate a stochastic change process over time. We describe the stochastic change process as a probabilistic graphical model, namely the generative probabilistic change model (GPCM), which factorizes the complex simulation problem into two more tractable sub-problems, i.e., condition-level change event simulation and image-level semantic change synthesis. To solve these two problems, we present Changen2, a GPCM implemented with a resolution-scalable diffusion transformer which can generate time series of remote sensing images and corresponding semantic and change labels from labeled and even unlabeled single-temporal images. Changen2 is a "generative change foundation model" that can be trained at scale via self-supervision, and is capable of producing change supervisory signals from unlabeled single-temporal images. Unlike existing "foundation models", our generative change foundation model synthesizes change data to train task-specific foundation models for change detection. The resulting model possesses inherent zero-shot change detection capabilities and excellent transferability. Comprehensive experiments suggest Changen2 has superior spatiotemporal scalability in data generation, e.g., Changen2 model trained on 256 pixel single-temporal images can yield time series of any length and resolutions of 1,024 pixels. Changen2 pre-trained models exhibit superior zero-shot performance (narrowing the performance gap to 3% on LEVIR-CD and approximately 10% on both S2Looking and SECOND, compared to fully supervised counterpart) and transferability across multiple types of change tasks, including ordinary and off-nadir building change, land-use/land-cover change, and disaster assessment. Zhuo Zheng, Stefano Ermon, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Layout-Anchored Prioritizing Continual Learning for Continuous Building Footprint Extraction From High-Resolution Remote Sensing ImageryabstractContinuous building footprint extraction requires learning new building patterns from remote sensing imagery without forgetting old knowledge. It is inherently challenging due to the spatial layout heterogeneity, which leads to the problem of knowledge forgetting in two aspects: complex background could have distinct patterns (background diversity) and buildings could have similar patterns to the background (foreground-background similarity). To solve the issues, we propose a domain-incremental continual learning algorithm named layout-anchored prioritizing learning network (LAPNet), including a latent layout anchoring module and layout-aware prioritizing learning module. The latent layout anchoring aggregates background information into latent layout features and employs a herding strategy to select representative layout anchors iteratively. This module maintains a memory buffer to narrow the background differences by dynamically discarding unrepresentative experiences and storing layout-anchored experiences. Furthermore, layout-aware prioritizing learning uses these experiences to identify and emphasize the most valuable knowledge for maximizing interclass distance. This module leverages the layout variance metric to measure interclass discrepancies and employs prioritizing learning to reweight the optimization function based on this layout prior. We established a Global-CL dataset to validate the proposed LAPNet framework, containing six study areas across four continents with different remote sensing sensors. Experiments showed that LAPNet achieves state-of-the-art performance in continuous building footprint extraction by effectively correlating knowledge across various domains. The code is available at:https://github.com/Dingyuan-Chen/LAPNet. Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | BASHVS: A Multispectral and SAR Image Fusion Method Based on Bidirectional Aggregation of Saliency in Human Visual SystemabstractThe limitations of remote sensing sensor technology make it difficult to simultaneously capture earth observation information presented in different forms within a single remote sensing image. Acquiring more plentiful target information through fusion technology has therefore remained a research hotspot. The fusion of multispectral (MS) and synthetic aperture radar (SAR) imagery integrates spectral and backscatter information, thereby improving land cover (LC) classification effects. However, current pixel-level fusion methods often fail to adequately account for the model differences between SAR and MS images, leading to spectral-spatial inconsistencies and severe degradation from speckle noise. To address this problem, A fusion method is proposed based on Bidirectional Aggregation of Saliency in the Human Visual System (BASHVS). First, the SAR and MS images are decomposed into base and detail layers using a synchronized anisotropic diffusion algorithm. Subsequently, the detail layer is fused using a New Sum of Modified Anisotropic Laplacian (NSMAL) algorithm. Finally, for base layer fusion, pixel saliency and structural saliency are extracted bidirectionally. The BASHVS is compared with 16 existing fusion methods using 10 evaluation metrics. The results demonstrate that BASHVS achieves the best comprehensive performance and significantly improves the visual quality of the fused images. LC classification using BASHVS fused images shows an average increase of 1.050% in overall accuracy and 0.014 in Kappa coefficient compared to the original MS images, confirming its advantage for LC classification. The source code of BASHVS is shared at https://github.com/CHUANGL8346/BASHVS. Xunqiang Gong, Yichuang Luo, Yonglei Chang, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Confusion-Aware Contrastive-Based Semi-Supervised Semantic Change Detection Through Space CollaborationabstractSemantic change detection (SCD) involves detecting changed regions and classifying their corresponding semantic change categories in remote sensing images. However, in most SCD scenarios, only the land-cover category information of the changed regions is provided. The mainstream multi-task Siamese network optimization process fails to fully leverage the information from unchanged land-cover types, which limits its ability to recognize land-cover change types and ultimately limits SCD performance. To address the issue of suppressed change type recognition caused by sparse labels in the SCD task, this paper proposes a confusion-aware contrastive-based semi-supervised semantic change detection method through space collaboration (SC2A-SCD). The proposed SC2A-SCD framework first integrates bi-temporal land-cover classification (LCC)-predicted logits from both the classification and representation spaces using joint classification and representation space modeling (JSM), providing high-quality pseudo-labels for unchanged land-cover types and enhancing the model’s ability to recognize bi-temporal land-cover types. Meanwhile, confusion-aware contrastive learning (CACL) is conducted in the representation space, utilizing confusion sampling strategies to sample more confusable anchors and negatives that are prone to misclassification, thereby effectively improving the representation capability of bi-temporal land-cover types and enhancing the quality of the pseudo-labels obtained via JSM, ultimately boosting the ability of SC2A-SCD to recognize semantic changes. Compared to the state-of-the-art SCD methods, the comprehensive experimental results confirm that the proposed SC2A-SCD framework can effectively improve the recognition ability for change types, demonstrating its effectiveness in enhancing SCD performance. Yinhe Liu, Jue Wang 0011, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Disaster-Aware Path Planning Based on Reinforcement Learning for Postearthquake Emergency ResponseabstractAfter earthquake disasters, ensuring that emergency rescue operations reach affected areas as quickly as possible, with the aim of maximizing the rescue of lives and property and minimizing secondary disaster losses, is the primary task of post-earthquake emergency response. Remote sensing technology, characterized by its large-scale coverage, non-contact nature, and rapid response capabilities, provides valuable information for disaster area assessment and post-earthquake emergency response path planning. However, existing post-disaster emergency path planning studies often fail to utilize this rich information. Traditional path planning algorithms insufficiently consider the disaster situation, hindering the efficient utilization of rescue forces and resources. To address these challenges, this study proposes a disaster-aware path planning method based on reinforcement learning for post-earthquake emergency response (P2DARL). The P2DARL method utilizes real disaster information provided by high-resolution remote sensing images and models it in a reinforcement learning environment. This allows the intelligent agent to learn optimal post-earthquake emergency response path planning strategies through continuous trial-and-error interactions with the environment. This approach facilitates the optimal allocation of rescue resources and enhances rescue efficiency. Experiments using real earthquake disaster images demonstrate that the P2DARL method effectively plans paths that cover more affected centers in mid-short distance scenarios, maximizing rescue efficiency and significantly reducing casualties and economic losses resulting from earthquake disasters. Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Reconstruct Multiscale Features for Lightweight Small Object Detection in Remote Sensing ImagesabstractMost small objects are missed when object detection algorithms are transferred from natural images to remote sensing images. Constructing multi-scale features has been proven to be an effective approach for detecting small objects. However, existing methods for multi-scale features have two limitations: insufficient discriminative capability and sparse semantic-spatial information, which fail to fully leverage the potential of multi-scale features. To overcome these limitations, we propose the Multiscale Feature Reconstruction Network (MRN), which introduces three novel modules during feature extraction, fusion, and enhancement: the Composite Multi-scale Feature Extraction Module (CEM), Interlayer Feature Joint Module (IJM), and Spatial-Semantic Information Cross Module (SSM). First, CEM utilizes a multi-branch structure to aggregate scale information. Dilated convolution and asymmetric convolution are extensively used in the branches, which expand the receptive field and capture information of rectangular instances, respectively. Second, the IJM leverages a gating mechanism to achieve pixel-level feature enhancement for feature maps at different hierarchical depths. Finally, the SSM alleviates high-level semantic feature information imbalance through dual-branch information interaction. Furthermore, to utilize the limited computational resources, we propose a lightweight version called MRN_Lite. We evaluate MRN and MRN_Lite on three existing public datasets: AI-TOD, VEDAI, and VisDrone2019. Extensive experiments demonstrate the effectiveness of our method. In comparison experiments, both versions of the model outperform the state of the art (SOTA). And MRN_Lite has less than 50% of the FLOPs and parameters of MRN, which has comparable performance to the original version. Yuancheng Huang, Renwei Qin, Xiangtao Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Lightweight Multiscale and Multiattention Hyperspectral Image Classification Network Based on Multistage SearchabstractHyperspectral image (HSI) classification has become a core task in hyperspectral remote sensing interpretation, with deep learning dominating due to its ability to learn hierarchical features without manual engineering. As the model complexity has grown, manual design limitations have prompted a shift to automated approaches such as differentiable architecture search (DARTS), where the architectures are optimized for greater accuracy and efficiency. However, applying gradient-based neural architecture search (NAS) methods directly to hyperspectral classification presents several challenges. Regarding search space design, there is a lack of lightweight operators that can mitigate the spectral variability, spatial heterogeneity, and scale differences inherent in hyperspectral imagery. In terms of search strategy, the traditional DARTS approach directly derives the topology from operation weights, which can lead to suboptimal topological structures, and thus affects the performance of the network in HSI classification. In this article, to address these issues, we propose L3M, which is a lightweight multiscale and multiattention HSI classification network based on multistage search. The proposed approach introduces a novel lightweight operator to address the spectral variability, spatial heterogeneity, and scale differences in HSIs. The operation search and topology search are also decomposed into a multistage process to prevent a suboptimal network by searching for and determining the topological order of the candidate operations in a predefined operation space. L3M was validated on four public datasets, where the proposed model demonstrated a superior classification performance, compared to other lightweight models, while maintaining a low parameter count, low model complexity, and high inference speed. Kefan Li, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | FreMamba: A Frequency-Domain Mamba Model for Hyperspectral Image Change DetectionabstractThe task of hyperspectral image change detection (HSI-CD) is to identify the subtle category changes of land surfaces by utilizing the rich spectral information of bi-temporal hyperspectral images (HSIs). Recent advanced deep learning methods have improved the performance of HSI-CD. However, HSIs have the mixed-pixel phenomenon, which affects the ability of HSI-CD models to discriminate land-cover changes. In this paper, a frequency domain Mamba (FreMamba) model is proposed to precisely capture the details of the spatial and spectral changes in bi-temporal HSIs, with consideration of the abovementioned problem. The proposed FreMamba model utilizes the capability of the Mamba model to adaptively filter the critical change information in long-range feature sequences. This is combined with frequency domain information to enhance the low-frequency global information and high-frequency spatial detail representation, based on a dual-branch network structure. A low-frequency spectral-guided attention (LSGA) module is proposed for the low-frequency Mamba branch, where it is embedded within the Mamba block via residual connections. Global low-frequency features are aggregated by the LSGA module with noise suppression, while the spectral change discriminability of ground objects is adaptively enhanced via the subsequent Mamba blocks. In the high-frequency Mamba branch, a high-frequency spatial geometric enhancement (HSGE) module is proposed that preserves edge details by spatial frequency domain decomposition to extract high-frequency geometric features. Experiments on three public HSI-CD datasets demonstrate that the proposed FreMamba model can accurately capture detailed change information and can outperform the existing state-of-the-art HSI-CD methods. Pengyuan Lv, Feilong Shan, Yune Cao, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | QSCDNet: A Hybrid Quantum Spectral Change Detection Network for Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSI-CD) is an important remote sensing technique for identifying fine-grained land-cover change. Deep learning methods such as convolutional neural networks (CNNs) and transformers have achieved good performance in HSI-CD. However, due to the pseudo-changes caused by imaging conditions, the spectral change characteristics within each pixel often exhibit uncertainties. In this study, differing from the traditional deep learning methods, we aimed to relate the abovementioned spectral change uncertainty to the perspective of the quantum state and built a dual-branch hybrid quantum neural network for HSI-CD (QSCDNet). The quantum branch consists of several 2-D quantum spectral change convolutional blocks (QSCCBs). These blocks provide a wider variety of expression forms of the spectral change, independent of the pseudo-change impact, through a parameterized quantum circuit (PQC), to better extract the change information of the spectral features. The CNN branch is designed based on a channel attention network architecture to provide traditional network change features. The output features of the quantum network branch and the CNN branch are then fused based on a cross-domain change feature fusion module (CDFM). The impact of the number of QSCCBs was also analyzed to verify the effectiveness of increasing the depth of the quantum network structure. The proposed method was tested on three public HSI-CD datasets and compared with the state-of-the-art methods to validate its potential in the field of HSI-CD. Pengyuan Lv, Ye Gao 0006, Heng Hu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | SiamS²F: Satellite Video Single-Object Tracking Based on a Siamese Spectral-Spatial-Frame Correlation NetworkabstractIn recent years, satellite video object tracking has received widespread attention as a research hotspot, but is also faced with many challenges. In satellite video, moving objects usually consist of only a few pixels, which makes it more difficult for the tracker to distinguish the target from the background. Furthermore, occlusion factors such as clouds, trees, and bridges can bring challenges when relocating the target from adjacent frames. In this article, to solve the above-mentioned problems, we propose a Siamese spectral-spatial-frame (SiamS2F) correlation network, which combines deep spectral, spatial, and frame features to enhance the interaction between video frames, to better focus on the target continuity. First, the deep features of the video are extracted with a Siamese backbone, along with a channel-spatial attention module. Second, a spectral-spatial-frame (SSF) module is proposed, where the output features of the backbone are fused based on the graph attention layer, and then the intraframe and interframe information of the fused features is enhanced by a newly designed multiframe interactive attention (MFIA) mechanism. Third, to solve the problem of similar objects, an additional center loss function is proposed in the classification regression head (CRH), where the search area is limited to a local range by adding an inbox identifier, to reduce the impact of the similar objects around the target. The proposed method was tested on two challenging benchmark remote sensing datasets—SatSOT and SV248S—where SiamS2F outperformed the related state-of-the-art trackers. Pengyuan Lv, Xianyan Gao, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | NOT-156: Night Object Tracking Using Low-Light and Thermal Infrared: From Multimodal Common-Aperture Camera to Benchmark DatasetsabstractNight object tracking (NOT) is aimed at tracking objects under low-illumination conditions at night. Existing works concentrate on thermal infrared modality, while some RGB and thermal infrared (RGB-T) data also contain night scenes. However, night scenes in these datasets are mostly well-lit, making it challenging to fully cover low-illumination scenarios. In this article, we focus on the NOT task and build up a novel low-light visible and thermal infrared (LOL-T) multimodal benchmark dataset for NOT-156. To achieve night vision, we design a common-aperture LOL-T camera by integrating a highly dynamic low-light visible imaging sensor with a thermal infrared sensor in a common aperture optical system. The proposed dataset consists of 156 video sequences and a total of 170k annotated frames, including various low-illumination night scenes such as dark rooms, streets, corridors, and so on. Compared with existing datasets, NOT-156 has more comprehensive and distinctive attributes (thermal variation, noise, high illumination overexposure, etc.). Comprehensive experiments are carried out to evaluate the performance of the advanced visible, infrared, and visible-thermal trackers on the proposed NOT-156 dataset. The authors believe that NOT-156 has great potential in the application and development of night vision. The dataset will be made available athttp://rsidea.whu.edu.cn/NOT156_dataset.htm. Xinyu Wang 0003, Shenghua Fan, Xiaobing Dai, Yuting Wan, Zengliang Zhu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Learning Temporal Consistency for High Spatial Resolution Remote Sensing Imagery Semantic Change DetectionabstractSemantic change detection (SCD) is a crucial task in remote sensing imagery interpretation, which identifies where the changes are and what categories the objects before and after the changes belong to. Post-classification comparison (PCC), as the simplest method, suffers from severe false alarms. While multi-task methods face another issue, in that the changed areas in the binary change map have same object categories in the semantic maps, i.e., semantic change inconsistency. In this study, we leveraged the commutative property of temporal semantic labels in unchanged areas to generate temporal consistency embedding and propose temporal-semantic feature blending (TSFB), which is a feature interaction operation that can dynamically adjust the distance between bi-temporal features by controlling a blending weight. In order to realize adaptive temporal consistency learning, we propose the TEmporal-Semantic feature CalibratiOn (TESCO) module, which can estimate the optimal value of the blending weight for TSFB automatically and make the segmentation network learn temporal consistency from coarse to fine through recurrence and the end-to-end training process. The TESCO module is a plug-and-play module that can be combined with any SCD method. In this study, we added the TESCO module to both PCC and multi-task based SCD methods and conducted comprehensive experiments on three large-scale and different application scenario SCD datasets with different semantic change numbers and time series. The experimental results show that the TESCO module can effectively learn temporal consistency, resulting in a significant improvement in the performance of PCC methods and the semantic consistency of the multi-task based methods. The implementation of TESCO module will be available at https://github.com/Daisy-7/TESCO. Shiqi Tian, Ailong Ma, Zhuo Zheng, Xicheng Tan, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | MSAHiFiC: A Super-Prior Driven High-Fidelity Spectral Attention Network for Hyperspectral Image CompressionabstractTo address the limitations of existing deep learning-based hyperspectral image compression methods in accurately modeling the rate-distortion problem, we propose a coupled multi-scale attention spatial-spectral high-fidelity compression network (MSAHiFiC). MSAHiFiC employs a super-prior network to estimate bitrate and guide rate-distortion optimization, enhancing performance under constrained bitrate conditions. A multi-scale spectral attention module is introduced to capture spectral dependencies across varying inter-band distances and preserve key spectral features during downscaling. A spectral fidelity term is further incorporated into the loss function to improve reconstruction accuracy. Experiments on three benchmark hyperspectral datasets—HySpecNet-11k, XiongAn, and WHU-Hi—demonstrate that MSAHiFiC outperforms state-of-the-art methods by achieving 5% higher spectral fidelity and 6% improvement in reconstruction accuracy under a low bitrate of 0.5 bpp. Yuting Wan, Chao Chen 0029, Ailong Ma, Xunqiang Gong, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Low-Light and Infrared Multimodal Remote Sensing in Nighttime Rescue Mission: A Review of Anomaly Detection MethodsabstractWhen facing natural disasters like sudden floods at night, due to sudden nature disasters, high timeliness of rescue information and complexity and diversity of rescue environment, it is difficult for ground personnel to perform rescue operations in disaster area in time. Remote sensing UAV technology plays an escalating role in disaster relief due to fast response and high flexibility advantages. However, with low nighttime visibility level, complex post-disaster environment, and numerous obstructions, traditional UAV’s visible light remote sensing struggles to achieve accurate rescue detection at night. Therefore, the multimodal detection methods are investigated using low-light and infrared modalities, exploring the integration of data fusion and detection in night rescue applications, and examine the advantages and disadvantages of different anomaly detection methods here. This paper provides the following contributions: 1) a fully-annotated low-light infrared co-observation multimodal remote sensing image dataset for nighttime emergency rescue, termed MRSI-NERD; 2) a benchmark test for most state-of-the-art unsupervised anomaly detection methods to thoroughly explore their capability in extracting useful information from normal samples; and 3) a low-light infrared bimodal fusion method based on frequency domain feature decomposition, which enhances the performance of detectors. The performance of eight types of traditional or deep learning-based detection methods on the MRSI-NERD dataset is reported, including metrics such as ROC-AUC, FPR, TPR, etc. Additionally, a comprehensive analysis of the principles and performance of each category of detection methods is provided. Yuting Wan, Haoyu Yao, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Source-Free Multitarget Unsupervised Domain Adaptation for Cross-City Local Climate Zone ClassificationabstractLocal Climate Zones (LCZs) offer a standardized urban classification system critical for studying climate variations and the urban heat island effect. Large-scale LCZ mapping facilitates cross-city climate comparisons. Remote sensing (RS), particularly supervised learning, has become the primary LCZ classification method due to satellite imagery’s broad coverage and high resolution. However, current RS approaches face challenges, including high labeling demands and limited transferability. In this work, we aim to leverage limited labeled data from a source city (i.e., source domain) to improve LCZ classification performance for multiple unlabeled target cities (i.e., target domains). To address these challenges, this work proposes a novel Source-free Multi-target Unsupervised Domain Adaptation (SFMT-UDA) framework for cross-city LCZ classification. Our approach operates in two stages: first performing single-target adaptation between source and target domains using a self-supervised pseudo-labeling module enhanced by an LCZ class similarity matrix to reduce label noise, then conducting multi-target adaptation among target domains through a Multi-head Multi-target Domain Adaptation (MH-MTDA) network with an integrated style-transfer module for consistent self-training. Extensive experiments on the newly developed VHRLCZ dataset demonstrate that the proposed SFMT-UDA framework achieves superior performance compared to state-of-the-art methods, showing significant improvements in overall accuracy across multiple target domains. The related code and data are available at https://github.com/ctrlovefly/SFMT-UDA. Qianqian Wu 0004, Yinhe Liu, Yanfei Zhong, Kexin Lin, Xianping Ma, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | MF-Mamba: Multiscale Convolution and Mamba Fusion Model for Semantic Segmentation of Remote Sensing ImageryabstractSemantic segmentation of remote sensing imagery plays an important role in applications such as environmental monitoring and disaster response. However, challenges such as complex spatial patterns of variable target objects, significant scale variations, and high inter-class similarity challenge accurate segmentation. Most existing methods based on convolutional neural networks (CNNs) and Transformers face limitations in modeling multi-scale global-local dependencies or often incur high computational costs. Therefore, we propose a multi-scale convolution and mamba fusion model (MF-Mamba) that integrates a CNN encoder with a Mamba-based decoder. The decoder incorporates a Global-Local State Space (GLSS) module with eight-directional selective scanning mechanisms and multi-kernel parallel convolutions to capture the rich global-local context. To enhance multi-scale feature representation, we developed a channel-spatial attention and dense multi-scale feature fusion (CSDF) module, which combines channel-spatial attention and atrous convolutions for multi-scale feature fusion. Additionally, a multi-scale lateral connection is developed to align encoder features for efficient integration. Experiments on the data sets of ISPRS Vaihingen, ISPRS Potsdam, and the Wuhan Dense Labeling Dataset (WHDLD) demonstrate the superior performance of MF-Mamba compared to existing state-of-the-art methods. It achieves Mean F1 scores of 86.71%, 90.70%, and 77.07%, respectively. The code is available at https://github.com/Mango-Mars/MF-Mamba. Pu Xiao, Ji Zhao 0006, Tieqi Peng, Christian Geiß, Yanfei Zhong, Hannes Taubenböck |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | CFDNet: Coupling Computational Fluid Dynamics With Convolutional Neural Networks for Gas Detection Using Thermal Infrared Multispectral VideoabstractGas detection is critically important in both industrial production and environmental monitoring. Due to the absorption characteristics of gases in the infrared wavelength domain, thermal infrared multispectral video imaging provides a convenient data sensing method for rapid, large-scale gas detection. In the process of gas detection, the shape of the detected gas leakage region is often distorted by the irregular motion patterns of gases. Although some studies have considered temporal motion information, these spatial-temporal feature extractors were designed originally for regular and salient objects such as vehicles and pedestrians, but not for gas. Moreover, the problem of weak gas signal and lack of a large well-labeled dataset also hinders the development of gas detection with deep learning. Regarding the above issues, a novel gas detection network, Computational Fluid Dynamics neural Network (CFDNet), was proposed for infrared multispectral gas detection. Firstly, to better fit the motion patterns of gas and keep the shape of the detected leakage region, a spatial-temporal fluid motion feature extractor, Computational Fluid Dynamics (CFD) Basic Block, was proposed. CFD Basic Block diffuses and displaces high-dimensional features through a convolution based on the Navier-Stokes equations in computational fluid dynamics. Secondly, to enhance the signal of gaseous targets, a Local Entropy and combination Difference (LED) data feature enhancement module was designed based on the instrument characteristics and local entropy information. Finally, a simulated dataset and a transfer learning framework were built to train a deep learning gas detection model with great generalizability. Experiments show that the proposed method, CFDNet, achieves a better performance than existing approaches. On the real-world dataset used for testing, it reaches an IoU of 50.128%, a Kappa of 59.902%, and an F1 Score of 44.885%. CFDNet demonstrates an excellent performance on keeping the shape of the detected gas leakage region under irregular motion patterns, especially for gas plumes with small area-ratio and low signal-noise-ratio. Haiyang Xiong, Liqin Cao, Du Wang, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Progressive Symmetric Registration for Multimodal Remote Sensing ImageryabstractImage registration forms the foundation of collaborative processing in multimodal remote sensing imagery (MRSI). However, high-resolution MRSIs frequently display complex distortions due to imaging characteristics and terrain variations, with both global and local distortions present. Effectively addressing these complex distortions necessitates the identification of uniformly and densely distributed corresponding points across the entire image. Existing methods primarily focus on global affine distortions and often extract only sparse and unevenly distributed corresponding points, which makes the effective handling of these coexisting distortions a significant challenge. To address this problem, we propose a progressive symmetric registration learning network (PSRNet) for MRSIs. In PSRNet, multimodal remote sensing image registration (MRSIR) is redefined as a symmetric dense regression task, differing from the traditional pipeline that concentrates on unidirectional sparse transformation parameter prediction. Specifically, PSRNet consists of three primary components: 1) a multiscale feature projector (MFP), which employs a dual-branch structure with nonshared weights to achieve modality-specific representation of different modal images across multiple scales, 2) a progressive cross-modal transformer (PCMT) to further mine modality-invariant features and progressively predict symmetric deformation fields, and 3) a symmetric consistency loss (SCL) function capable of elegantly achieving high-precision reversible alignment of image pairs, encompassing endpoint error loss, bidirectional alignment loss, and smoothness loss. Experimental results demonstrate that PSRNet achieves more comprehensive and advanced registration performance on our self-constructed large-scale high-resolution MRSIR dataset, which includes complex global-local geometric distortions and significant nonlinear radiometric differences (NRD). Heng Yan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Self-Supervised Joint Representation Learning for Urban Land-Use Classification With Multisource Geographic DataabstractUrban land-use patterns furnish the functional structures of cities, playing an essential role in urban planning and management. Recently, advancements in remote sensing and Internet technologies have made high-resolution remote sensing (HRS) images and social sensing data more widely available. These multisource geographic data contain comprehensive urban scene information, providing new opportunities for urban land-use classification. However, as the social sensing data are sparsely distributed, learning representations of parcels for land-use classification with limited multisource sample pairs remains challenging. In this article, a self-supervised joint representation learning (SJRL) framework for urban land-use classification with multisource geographic data is proposed. Specifically, a semantic-aware self-supervised representation learning approach is designed to mine visual information from the HRS images. This approach introduces a semantic-aware sampling strategy to identify regions with significant visual information, thereby enhancing the efficiency of the representation learning. For the points of interest (POIs), which are a type of social sensing data, a context-aware self-supervised representation learning approach is proposed to obtain a pretrained POI model suitable for land-use classification. To capture the correlation of multisource data, a cross-modal alignment (CMA) module and a modality-enhanced module (MEM) are designed to align the common information across modalities and adaptively learn the complementary information between multisource data. The CMA and MEM are then embedded into the land-use classification model for accurate classification. Experiments conducted on multisource samples from 34 Chinese provincial cities and the urban regions of Beijing, Shanghai, Wuhan, and Chengdu in China verified the effectiveness and generalizability of the proposed SJRL framework. Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Rethinking Semantic Segmentation With Multi-Grained Logical PrototypeabstractThe last decade has witnessed significant advances in semantic segmentation brought about by deep learning. However, existing methods only fit the data-label correspondence in a data-driven manner and do not fully conform to the abstraction and structuralization characteristics of the human visual cognition process, which limits the upper bounds of their performance. To this end, a multi-grained logical prototype (MGLP) method is proposed to rethink semantic segmentation based on these two key characteristics. Its novel design can be summarized as follows. 1) For abstraction, prototypes of the same class at different grain levels are established: a label generation method is proposed to automatically generate a multi-grained label space, which can guide the learning of the multi-grained prototypes for each class. 2) For structuralization, the intrinsic logical structure across different semantic levels is explicitly modeled: the horizontal metric relationships are established via metric relation operations on prototypes at the same grain level, to improve the discriminability between classes while taking the vertical semantic hierarchy into account. Moveover, the vertical logical relationships are established as the sub-to-super positive and super-to-sub negative constraints, to strengthen the semantic dependencies among prototypes at different grain levels. 3)MGLP is plug-and-play and can be directly combined with existing segmentation methods. Extensive experimental results indicate that MGLP can significantly improve the segmentation performance of existing methods, which opens up a new avenue for future research. Anzhu Yu, Kuiliang Gao, Xiong You, Yanfei Zhong, Bing Liu 0018, Chunping Qiu |
IEEE Trans. Image Process. | 4 |
| 2024 | EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question AnsweringabstractEarth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we develop a multi-modal multi-task VQA dataset (EarthVQA) to advance relational reasoning-based judging, counting, and comprehensive analysis. The EarthVQA dataset contains 6000 images, corresponding semantic masks, and 208,593 QA pairs with urban and rural governance requirements embedded. As objects are the basis for complex relational reasoning, we propose a Semantic OBject Awareness framework (SOBA) to advance VQA in an object-centric way. To preserve refined spatial locations and semantics, SOBA leverages a segmentation network for object semantics generation. The object-guided attention aggregates object interior features via pseudo masks, and bidirectional cross-attention further models object external relations hierarchically. To optimize object counting, we propose a numerical difference loss that dynamically adds difference penalties, unifying the classification and regression tasks. Experimental results show that SOBA outperforms both advanced general and remote sensing methods. We believe this dataset and framework provide a strong benchmark for Earth vision's complex analysis. The project page is at https://Junjue-Wang.github.io/homepage/EarthVQA. Zhuo Zheng, Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
AAAI | 5 |
| 2024 | A Multi-Level Fine-Grained Crop Classification Method Based on Multi-Expert Knowledge DistillabstractCrop mapping is an important task for agriculture-related activities and economic development. Most present studies focus on the crop mapping of staple crops, i.e., soybean, maize, and wheat, few concentrate on the multi-level fine-grained crop classification, which requires knowing finer crop type, i.e., barley or rye, winter wheat or spring wheat for different application. Deep learning methods with the strong ability to extract features automatically have great potential in fine-grained crop mapping. However, the classification of multi-level finer crops is challenged by the extremely similar phenological characters. In this paper, a multi-level fine-grained crop classification method based on multi-expert knowledge distill is proposed to learn the phenological features with different distinction degrees. Specifically, it uses three expert models to distinguish crop types with obvious, similar, and confusing phenological features. Then through a self-paced learning module, the student model first learns the knowledge from three expert models in the early stage and then learns to excavate the phenological features actively during the learning process. The experiment was carried on in Nordrhein Westfalen, Germany based on the EUROCROPS and time-series Sentinel-2 dataset and achieved great performance compared with popular deep learning methods. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2024 | MapChange: Enhancing Semantic Change Detection with Temporal-Invariant Historical Maps Based on Deep Triplet NetworkabstractSemantic Change Detection (SCD) is recognized as both a crucial and challenging task in the field of image analysis. Traditional methods for SCD have predominantly relied on the comparison of image pairs. However, this approach is significantly hindered by substantial imaging differences, which arise due to variations in shooting times, atmospheric conditions, and angles. Such discrepancies lead to two primary issues: the under-detection of minor yet significant changes, and the generation of false alarms due to temporal variances. These factors often result in unchanged objects appearing markedly different in multi-temporal images. In response to these challenges, the MapChange framework has been developed. This framework introduces a novel paradigm that synergizes temporal-invariant historical map data with contemporary high-resolution images. By employing this combination, the temporal variance inherent in conventional image pair comparisons is effectively mitigated. The efficacy of the MapChange framework has been empirically validated through comprehensive testing on two public datasets. These tests have demonstrated the framework's marked superiority over existing state-of-the-art SCD methods. Yinhe Liu, Sunan Shi, Zhuo Zheng, Jue Wang 0011, Shiqi Tian, Yanfei Zhong |
IGARSS | 6 |
| 2024 | Remote Sensing Image Land Cover Classification with Label Noise Based on Deep Reinforcement LearningabstractThe advent of high-resolution satellite technology has significantly improved the accuracy of land cover classification, but it also presents new challenges. Deep learning methods, particularly Convolutional Neural Networks (CNNs) and Transformers, have shown promise in overcoming these challenges, yet they are highly reliant on extensive annotated datasets. An alternative, using lower resolution land cover products for label generation, introduces Label Noise, which can adversely affect the learning process and model performance. To address the lack of quality labels and the Label Noise problem, this study introduces the LabelAgent framework. This framework combines high-resolution, recent remote sensing data with historical, low-resolution product data to produce higher quality training labels. Furthermore, it employs a deep reinforcement learning-based approach to intelligently identify and mitigate noisy labels, thereby stabilizing the training process and enhancing the classification efficacy. The outcome of this research offers a novel solution to the challenges in remote sensing image classification, promising improved accuracy and generalization capability for land cover classification models. Yinhe Liu, Sunan Shi, Ruoyu Yang, Yanfei Zhong |
IGARSS | 5 |
| 2024 | Self-Supporting Adaptive Prototype Learning for Remote Sensing Few-Shot Semantic SegmentationabstractThe segmentation of remote sensing images with few shots is valuable both in theory and application. Many existing few-shot segmentation methods rely on prototype learning, where a single support prototype guides predictions for the query set. However, due to visual differences between the support and query sets, a single support prototype struggles to capture all semantic information effectively. This paper introduces an adaptive self-supporting prototype learning network to address these challenges. We propose adaptive hyper prototype representation (HPR), comprising hyper prototype clustering (HPC) and guided prototype matching (GPM). HPC, a parameter-free and adaptive method, extracts more representative prototypes by clustering similar feature vectors using superpixel features. GPM selects matched prototypes to offer more accurate guidance, ensuring a uniform representation of multiple prototypes and complex semantic information. We also present self-supporting matching (SSM) prototype learning, guiding query set segmentation by obtaining query set prototypes. SSM generates initial pseudo-labels for the query set based on support set prototypes and further guides the query set using its own features, avoiding visual differences between support and query sets. Our adaptive self-supporting prototype learning network significantly enhances prototype quality and outperforms on object-level remote sensing datasets. Weihao Shen, Yanfei Zhong, Ailong Ma |
IGARSS | 2 |
| 2024 | MVDC: A Multi-View Domain Confusion Strategy for Cross Domain Change DetectionabstractRemote sensing (RS) change detection (CD) refers to the difference of RS images at different times in the same geographic area. The method of using high-resolution RS images based on deep learning (DL) has the problem of poor domain adaptation (DA) under geographical isolation conditions, especially in the cropland scene with different phenology and planting type. Based on the universal Siamese network architecture, a multi-view domain confusion (MVDC) module is designed. The module performs discriminant operations on the generated source domain features and target, and guides the generator to generate more domain invariant features. The time consistency information is taken into account while the difference between geographical regions is reduced. The module can theoretically be plugged into any universal CD network. The experimental results show that the adaptive ability of multiple CD methods is improved after the addition of MVDC. Zhendong Sun, Yanfei Zhong, Xinyu Wang 0003 |
IGARSS | 2 |
| 2024 | A Saliency-Aware Deep Network for Narrow Road Extraction of High-Resolution Remote Sensing ImageryabstractRoad extraction from high-resolution remote sensing imagery is important and efficacious due to deep learning. However, most methods face challenges in capturing narrow roads, i.e. rural roads. In this paper, we propose a saliency-aware deep network (SAN) for narrow road extraction. Specifically, a multi-scale context module is employed to extract contextual features in different scales, and a global context module is followed to aggregate the above multi-scale context features, to improve the connectivity of narrow roads. In addition, motivated by visual saliency, a saliency-aware module is proposed to highlight roads during skip connections, to further separate narrow roads from complex backgrounds. In the experiments, SAN is compared with some state-of-the-art methods by using the DeepGlobe road dataset and achieves better performances, especially for narrow roads, which proves its superiority. Ningjing Wang, Xinyu Wang 0003, Wanqiang Yao, Yanfei Zhong |
IGARSS | 5 |
| 2024 | Large-Scale Tidal Wetland Classification Based on Label Augmentation and Error CorrectionabstractTidal wetlands, situated crucially at the land-sea intersection, play a pivotal role in sustainable development. Remote sensing has become a useful method for large-scale tidal wetland mapping. However, it encounters challenges due to tidal wetlands exhibit considerable similarity in spectral characteristics and show significant seasonal variations. The existing machine-learning-based methods are difficult to achieve high-efficiency mapping, hence restricting their utility in wetland monitoring. Conversely, deep learning methods possess a significant edge in managing intricate data, resulting in better classification accuracy. However, the requirement of pixel-level labels for deep learning has posed difficulties for extensive network training. In this article, a framework based on historical point samples is introduced for large-scale tidal wetland mapping. The proposed method uses existing point labels to generate per-pixel pseudo-labels, maximizing the utilization of historical data, and utilizing methods such as CutMix and bootstrapping loss to deal with label noise. Yinhe Liu, Yanfei Zhong |
IGARSS | 4 |
| 2024 | Amalgamating Convolutional and Graph Neural Networks for Fast Multimodal Remote Sensing Image RegistrationabstractMultimodal remote sensing image registration is crucial for the comprehensive utilization of multimodal data. Existing methods demonstrate limited robustness in addressing significant nonlinear radiometric differences and geometric distortions in multimodal remote sensing images. This paper presents ACGNet, a fast registration method that amalgamates convolutional and graph neural networks to address these challenges. ACGNet first uses a deep convolutional network with a dual-branch encoder-decoder architecture to simultaneously extract feature points and corresponding advanced descriptors. Subsequently, a graph neural network based on attention mechanisms is employed for feature matching and outlier rejection, aiming to obtain a large number of high-confidence correspondences, which are crucial for estimating transformation parameters. Experimental results demonstrate that the proposed method exhibits robustness to variations in scale and rotation, and can efficiently and stably perform multimodal remote sensing image registration, outperforming other methods in terms of speed and accuracy. Heng Yan, Ailong Ma, Yanfei Zhong |
IGARSS | 3 |
| 2024 | Spliting Road Network to Road Segment Instances for Vector Road MappingabstractExtracting road network structures with high accuracy from very high-resolution (VHR) remote sensing imagery is a challenging task in the field of computer vision. Previous end-to-end algorithms modeled the road graph as a general graph structure with defined vertices and edges, represented as G = (V, E), which is a default approach in graph generation tasks. However, the use of excessively small edge units conflicts with the network’s geometric structure and potential object features of roads, leading to issues such as false alarm roads and disconnected roads. In this study, we propose a road graph generation model based on road segments, represented as G = (I, S), leveraging the long-distance connectivity of minimal road topological units. This approach aims to balance the connectivity and precision of the generated road graph. Subsequently, we introduce a road graph generation framework named RoadSegment, based on the aforementioned modeling, to validate the feasibility of the proposed road segment modeling. RoadSegment utilizes an Element Detector to identify intersections and road segment objects. Additionally, we propose an Intersection Connection Strategy (ICS) to establish connectivity between road segments. Empirical experiments conducted on the SpaceNet3 Dataset demonstrate the superiority of road segment modeling. Ruoyu Yang, Yinhe Liu, Yanfei Zhong |
IGARSS | 3 |
| 2024 | Urban Land-Use Classification with Multi-Source Self-Supervised Representation Learning and Correlation ModelingabstractUrban land-use classification aims to identify the functions of land parcels within cities, which can provide a reference for urban planning and environmental monitoring. Existing urban land-use classification methods generally integrate multi-source data to obtain comprehensive information about land parcels. However, these methods ignore the parcels with incomplete modalities, and the correlations between multisource data have not been explored. To this end, an urban land-use classification framework with multi-source self-supervised representation learning and correlation modeling (LUSC) is proposed. Specifically, the LUSC introduces self-supervised approaches to mine the multi-source representations from high-resolution remote sensing (HRS) images and points-of-interest (POIs), which consider the characteristics of parcels. The multi-source information joint learning module (MIL) is designed to explore the correlations between multi-source data, ensuring consistency of multisource data. The comparative and ablation experiments on the land-use dataset covering 34 Chinese provincial cities demonstrate that the proposed LUSC framework outperforms the advanced baselines. Yanfei Zhong |
IGARSS | 3 |
| 2024 | Historical Product Driven Large-Scale High-Resolution Land Cover and Wetland ClassificationabstractUnderstanding land cover dynamics, especially in wetlands, is crucial for environmental monitoring and ecosystem management, given the threats posed by climate change and human activities. Despite the existence of global-scale land cover maps and detailed wetland thematic maps, a gap remains in their integration, limiting comprehensive ecosystem analysis. Typically, these methods utilize medium-resolution data, which fail to capture the finer details of land cover and wetland interfaces, highlighting the need for high-resolution mapping. As high-resolution mapping becomes more prevalent, the classification of land cover and wetland classes grows increasingly complex. Deep learning methods, particularly semantic segmentation, have emerged as solutions, yet they require extensive training labels, which are challenging and resource-intensive to acquire. This paper proposes a novel framework utilizing historical land cover products as training labels for high-resolution mapping, addressing the challenges of resolution mismatches and semantic errors. The developed cross-resolution mapping dataset, evaluated against a manually labeled test set, demonstrates the framework's effectiveness in providing detailed classification at a provincial scale. Siqi Zeng 0005, Xueying Huang, Yinhe Liu, Ailong Ma, Yanfei Zhong |
IGARSS | 7 |
| 2024 | Decoupling Features for Remote Sensing Missing Modality LearningabstractMissing modality learning aims to enhance the robustness of the model when certain modalities are missing during the testing phase, by transferring knowledge between different modalities during the training stage. Existing works directly reduce the distance between different modalities feature through auxiliary tasks. While simple, this approach shows limited capability in some complex scenarios. It disrupts the distribution of modality-specific features, leading to a decline in modality-discriminative capability. Therefore, we prosed the DFNet. DFNet is designed to decouple the feature into modality-invariant and modality-variant features, by whitening transformation loss. Only the content features of different modalities are brought closer through auxiliary tasks. Experiments are conducted on MSAW datasets, with results indicating that our model outperforms competing methods. Yiheng Zhou, Ailong Ma, Zihang Chen 0001, Yanfei Zhong |
IGARSS | 5 |
| 2024 | Segment Any ChangeabstractVisual foundation models have achieved remarkable results in zero-shot image classification and segmentation, but zero-shot change detection remains an open problem.
In this paper, we propose the segment any change models (AnyChange), a new type of change detection model that supports zero-shot prediction and generalization on unseen change types and data distributions.
AnyChange is built on the segment anything model (SAM) via our training-free adaptation method, bitemporal latent matching.
By revealing and exploiting intra-image and inter-image semantic similarities in SAM's latent space, bitemporal latent matching endows SAM with zero-shot change detection capabilities in a training-free way.
We also propose a point query mechanism to enable AnyChange's zero-shot object-centric change detection capability.
We perform extensive experiments to confirm the effectiveness of AnyChange for zero-shot change detection.
AnyChange sets a new record on the SECOND benchmark for unsupervised change detection, exceeding the previous SOTA by up to 4.4\% F$_1$ score, and achieving comparable accuracy with negligible manual annotations (1 pixel per image) for supervised change detection. Code is available at https://github.com/Z-Zheng/pytorch-change-models. Zhuo Zheng, Yanfei Zhong, Liangpei Zhang 0001, Stefano Ermon |
NeurIPS | 2 |
| 2024 | Single-Temporal Supervised Learning for Universal Remote Sensing Change Detection
Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Uncertainties Analysis and Improvement of Smoothness TES Algorithm Based on Simulation Hyper-Cam-LW Airborne ImageryabstractIn recent years, several commercial airborne thermal infrared (TIR) hyperspectral imagers have been available for remote sensing applications. However, unlike the sensors developed primarily for research and government use, commercial imagers have particular configurations and imaging features, leading to uncertainties of mainstream smoothness temperature and emissivity separation (TES) methods. This letter systematically investigates the feasibility and the performance of retrieving temperature and emissivity from the hyper-Cam-LW sensor using typical smoothness TES algorithms. In addition, the linear spectral emissivity constraint (LSEC) TES algorithm was improved by considering the correlation among atmospheric parameters. Analysis results illustrated that typical smoothness TES algorithms exhibit high sensitivity to noise and flight height. Moreover, the retrieval accuracy of different ground objects varies significantly for the same TES method. By contrast, the improved LSEC-TES algorithm is not susceptible to flight height and the type of ground objects, and its retrieval errors of LST and LSE are less than 0.2 K and 0.004, respectively. Finally, a helpful conclusion can be drawn that the correlation between atmospheric parameters must be considered for the TES of airborne TIR hyperspectral imagery (HSI), and noise must be suppressed to ensure smoothness TES is effective for airborne TIR HSI of commercial TIR hyperspectral imager. Lyuzhou Gao, Liqin Cao, Hongquan Sun, Zhonggen Wang, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | A Semi-Supervised Pyramid Cross-Temporal Attention Transformer for Change Detection in High-Resolution Remote Sensing ImagesabstractThe vision transformer (ViT) model has the advantage of being able to model the long-range dependencies in the imagery and has been studied for the task of remote sensing image change detection (CD). However, the performance of the existing transformer-based CD methods is not satisfactory in the case of limited labeled data. The original self-attention mechanism cannot effectively extract the change information, and the large number of parameters in the ViT model makes the model difficult to train. To solve the above-mentioned problems, a semi-supervised pyramid cross-temporal attention transformer for change detection (CT2RCDSS) is proposed in this letter. The CT2RCDSS method follows an encoder-decoder structure. The encoder utilizes a dual-branch structure, containing the combination of the proposed cross-temporal attention (PCTA) and pyramid self-attention (PSA) mechanisms, which is designed to consider the interaction of the features from different time phases and enhance the changes at different scales. In the decoder, a series of deconvolutional layers with skip connections are utilized, and a Softmax layer follows to acquire the final binary change map. In addition, a semi-supervised training strategy, which reduces the errors in the pseudo-labels generated from the models initialized with different parameters, is used to improve the model stability while using unlabeled data. The experiments showed that the proposed method can achieve a superior F1-score and intersection over union (IoU), which indicates the potential of the proposed method. Pengyuan Lv, Mengchen Li, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | QETR: A Query-Enhanced Transformer for Remote Sensing Image Object DetectionabstractRecently, transformer models have been introduced into the field of remote sensing image object detection, benefiting from their ability to model long-term information. However, the existing transformer-based object detection methods mainly consider the global interaction of local elements and have a limited ability to enhance the local information, which can bring some difficulties in distinguishing real objects and a complex background. In this letter, a query-enhanced transformer (QETR) model is proposed to solve the above problems. The proposed model consists of three main parts: an encoder, a decoder, and a detection head. A Swin transformer is used to extract deep features in the encoder. In the decoder, the object and anchor queries are initialized and the feature and position information of the objects is learned by the multi-head self-attention and cross-attention mechanisms, respectively. Furthermore, a query align module along with a scale controller are proposed to enhance the object information around the local queries by limiting the attention to a certain range without losing important information. Finally, the boundaries and types of the objects are acquired from the detection head based on bipartite matching. To verify the effectiveness of the proposed method, comparative experiments were carried out with other state-of-the-art methodologies on two public datasets: the High-Resolution Remote Sensing Detection (HRRSD) dataset and the object detection in optical remote sensing images (DIOR) dataset. The experimental results confirm the effectiveness and superiority of the QETR model, which achieved 71.5% and 91.1% mAP values on the DIOR and HRRSD datasets, respectively. Pengyuan Lv, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Explicable Fine-Grained Aircraft Recognition Via Deep Part Parsing Prior Framework for High-Resolution Remote Sensing ImageryabstractAircraft recognition is crucial in both civil and military fields, and high-spatial resolution remote sensing has emerged as a practical approach. However, existing data-driven methods fail to locate discriminative regions for effective feature extraction due to limited training data, leading to poor recognition performance. To address this issue, we propose a knowledge-driven deep learning method called the explicable aircraft recognition framework based on a part parsing prior (APPEAR). APPEAR explicitly models the aircraft's rigid structure as a pixel-level part parsing prior, dividing it into five parts: 1) the nose; 2) left wing; 3) right wing; 4) fuselage; and 5) tail. This fine-grained prior provides reliable part locations to delineate aircraft architecture and imposes spatial constraints among the parts, effectively reducing the search space for model optimization and identifying subtle interclass differences. A knowledge-driven aircraft part attention (KAPA) module uses this prior to achieving a geometric-invariant representation for identifying discriminative features. Part features are generated by part indexing in a specific order and sequentially embedded into a compact space to obtain a fixed-length representation for each part, invariant to aircraft orientation and scale. The part attention module then takes the embedded part features, adaptively reweights their importance to identify discriminative parts, and aggregates them for recognition. The proposed APPEAR framework is evaluated on two aircraft recognition datasets and achieves superior performance. Moreover, experiments with few-shot learning methods demonstrate the robustness of our framework in different tasks. Ablation analysis illustrates that the fuselage and wings of the aircraft are the most effective parts for recognition. Yanfei Zhong, Ailong Ma, Zhuo Zheng, Liangpei Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Unsupervised Adaptation Learning for Real Multiplatform Hyperspectral Image DenoisingabstractReal hyperspectral images (HSIs) are ineluctably contaminated by diverse types of noise, which severely limits the image usability. Recently, transfer learning has been introduced in hyperspectral denoising networks to improve model generalizability. However, the current frameworks often rely on image priors and struggle to retain the fidelity of background information. In this article, an unsupervised adaptation learning (UAL)-based hyperspectral denoising network (UALHDN) is proposed to address these issues. The core idea is first learning a general image prior for most HSIs, and then adapting it to a real HSI by learning the deep priors and maintaining background consistency, without introducing hand-crafted priors. Following this notion, a spatial-spectral residual denoiser, a global modeling discriminator, and a hyperspectral discrete representation learning scheme are introduced in the UALHDN framework, and are employed across two learning stages. First, the denoiser and the discriminator are pretrained using synthetic noisy-clean ground-based HSI pairs. Subsequently, the denoiser is further fine-tuned on the real multiplatform HSI according to a spatial-spectral consistency constraint and a background consistency loss in an unsupervised manner. A hyperspectral discrete representation learning scheme is also designed in the fine-tuning stage to extract semantic features and estimate noise-free components, exploring the deep priors specific for real HSIs. The applicability and generalizability of the proposed UALHDN framework were verified through the experiments on real HSIs from various platforms and sensors, including unmanned aerial vehicle-borne, airborne, spaceborne, and Martian datasets. The UAL denoising scheme shows a superior denoising ability when compared with the state-of-the-art hyperspectral denoisers. Zhaozhi Luo, Xinyu Wang 0003, Petri Pellikka, Janne Heiskanen, Yanfei Zhong |
IEEE Trans. Cybern. | 5 |
| 2024 | Multispectral and SAR Image Fusion for Multiscale Decomposition Based on Least Squares Optimization Rolling Guidance FilteringabstractMultispectral and SAR image fusion is one of the key technologies to improve image quality. The fusion method of multi-scale decomposition includes two aspects: the decomposition of image and the design of fusion rule. There are some problems in the traditional decomposition methods, such as gradient reversals, halos and other artifacts, and limited scale separation of space overlapping features. In addition, the quality of fusion images is greatly affected by the fusion rule design. Therefore, a novel method based on least squares optimization rolling guidance filtering for multi-scale decomposition is proposed. All gradients of rolling guidance filtering are optimized by least squares to suppress artifacts such as gradient inversion, and then combined with Gaussian filtering for image decomposition to eliminate interference texture and speckle noise while preserving edge details. At the same time, the decomposition of image is extended to multi-scale space to achieve scale separation of space overlapping features, which is convenient for multi-level fusion of image features. In the end, based on the scale of decomposition, the results fall into three layers, and coupled neural P system and other rules are designed for different layers of information fusion. The results indicate that this method has good visual effect and outperforms all the other comparison methods on nine evaluation indexes. Xunqiang Gong, Zhaoyang Hou, Yuting Wan, Yanfei Zhong, Kaiyun Lv |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Channel Adaptive Dual Siamese Network for Hyperspectral Object TrackingabstractHyperspectral object tracking (HOT) aims at tracking targets using the rich spectral information from hyperspectral video. Recently, dual Siamese network (DSN) has been proposed for HOT with advanced performances, via integrating a RGB Siamese branch with a hyperspectral Siamese branch to solve small sample challenge of hyperspectral modality. However, there are still challenges of DSN that reduce its practicality: a single DSN model is difficult to process hyperspectral videos with varied channels; the spatial features extracted by the pre-trained RGB branch plays a dominant role, while the hyperspectral features are not fully explored. To address the challenges, we propose a Channel AdapTive dual Siamese network, termed SiamCAT, for HOT with varied channels. Specifically, treating each frame of hyperspectral video as a grayscale image sequence varied with wavelengths, a channel adaptive module is introduced to encode the grayscale image sequence of different lengths into a uniform length, and so that SiamCAT can process hyperspectral video with varied channels. Meanwhile, a guided learning attention module is proposed to progressively learn spectral features of the tracked target highlighted by the spatial attention of the pre-trained RGB branch. Note that, to force spectral features play a leading role, instead of traditional features fusion, the spectral features extracted by the hyperspectral branch are utilized for confirming the target position. In the experiments, SiamCAT were verified by using the HOT competition dataset (i.e., 16-channel, 25-channel, and 15-channel hyperspectral videos with different wavelength ranges) and the WHU-Hi-H3dataset (25-channel hyperspectral videos), and achieved advanced performances. Xinyu Wang 0003, Zengliang Zhu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | PhenoCropNet: A Phenology-Aware-Based SAR Crop Mapping Network for Cloudy and Rainy AreasabstractCrop mapping in a cloudy area is always a challenge due to the lack of time-series clear optical satellite imagery. Making use of time-series synthetic aperture radar (SAR) imagery that is immune to cloud contamination is essential and promising for seamless and large-area crop mapping. However, existing deep learning (DL)-based crop classification methods give the extracted phenological features equal weights, without considering the different contributions of phenological features of the different crop growth periods. In this article, a phenology-based crop mapping network (PhenoCropNet) is proposed to extract the discriminative features from the two levels, including the key phenological dates in the phenological periods and key phenological periods in the whole growth stages. PhenoCropNet includes a phenological calendar information injection (PAI) module that divides the satellite imagery time series (SITS) into multiple sequences according to the phenological calendar information, and a hierarchical attention network structure that uses the two-level bidirectional gated recurrent unit-based self-attention (BiGRUA) modules to automatically extract the features containing the most important phenological information of key phenological dates and key phenological periods. The proposed PhenoCropNet was verified in Hubei province in China, around 185 933 km2, a typical cloudy area in China, for rapid winter crop mapping based on temporal Sentinel-1 SAR imagery. The mapping result shows that the$F1$-score of PhenoCropNet for winter crop mapping could achieve 0.90, showing great potential in large-scale and seamless crop mapping. The code is available on request:https://github.com/LL0912/PhenoCropNet. Xinyu Wang 0003, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | One-Step Detection Paradigm for Hyperspectral Anomaly Detection via Spectral Deviation Relationship LearningabstractHyperspectral anomaly detection (HAD) aims to find small targets deviating from surroundings in an unsupervised manner. Recently, various deep models have been applied to HAD, such as autoencoder series and generative adversarial networks (GAN) series, which mainly use a proxy task, i.e., iteratively reconstructing low-frequency components (backgrounds) to separate anomalies (two-step paradigm). However, in such an unsupervised manner, most deep HAD model is trained and tested on the same image. Since the learned low-frequency background varies from image to image and the trained model cannot be directly transferred to unseen images. In this paper, the one-step detection paradigm is first proposed, where the model is optimized directly for the HAD task and can be transferred to unseen datasets. The one-step paradigm is optimized to identify the spectral deviation relationship according to the anomaly definition. Compared to learning the specific background distribution in the two-step paradigm, the spectral deviation relationship is universal for different images and guarantees transferability. Further, we instantiated the one-step paradigm as an unsupervised transferred direct detection (TDD) model. To train the TDD model in an unsupervised manner, an anomaly sample simulation strategy is proposed to generate numerous pairs of anomaly samples. A global self-attention module and a local self-attention module are designed to help the model focus on the “spectrally deviating” relationship. The TDD model was validated on six public datasets. The results show that TDD is superior to the recent two-step methods in detection and transferability aspects. Xinyu Wang 0003, Shaoyu Wang 0003, Hengwei Zhao, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Segmenting Remote Sensing Anomalies at Instance Level via Anomaly Map-Guided AdaptationabstractEarth anomalies can locate valuable targets in an unsupervised manner for many defense and surveillance applications. Most models assign a continuous score at the pixel-level, resulting in object-agnostic results with higher false alarms than the instance level results. However, since the anomaly objects contain a variety of categories and have a large intraclass variance, the current state-of-the-art (SOTA) query-based models designed for certain categories perform unsatisfactorily when applied to the anomaly instances. The larger intraclass variance of anomalies makes the learning of general representation more difficult. To bridge this gap, we propose general adaptations guided by the pixel-level anomaly map for any query-based model, which adapts the model from learning certain category representation to learning anomaly-aware representation in different categories. The proposed adaptation first builds a separate branch to output the pixel-level anomaly map, where anomaly information is then extracted to guide the pixel embeddings and queries to focus on a variety of anomaly categories. Especially, the anomaly rank embeddings are devised to make the pixel embeddings aware of the anomaly rank order. The queries are dynamically selected from the anomaly candidates after aligning the anomaly map and pixel embeddings for better locating different anomalies. Finally, the selected queries dot-product the anomaly-aware pixel embeddings to output the anomaly instances. The proposed adaptations are simple, general, and additive, which bring the average improvements of +4.9 box AP and +5.1 mask AP in infrared, synthetic aperture radar (SAR), and hyperspectral modalities. Yanfei Zhong, Hengwei Zhao, Zhi Gao 0005, Xinyu Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Deep Temporal-Spectral-Spatial Anchor-Free Siamese Tracking Network for Hyperspectral Video Object TrackingabstractHigh spatial, high spectral, and high temporal ($\text {H}^{3}$) information of the objects of interest can be provided by hyperspectral video, which makes it possible to track objects in complex scenarios. However, during motion, changes in the target’s appearance, background, and spectral information can degrade the performance of existing hyperspectral trackers due to insufficient training data. Consequently, this results in weak generalization of these trackers. In this article, to solve the above problems, a deep temporal-spectral–spatial anchor-free Siamese tracking network for hyperspectral video object tracking, namely HA-Net, is proposed. In HA-Net, a Siamese spectral enhancement tracker module based on an RGB tracker (pseudo-color tracker) is designed, which uses the powerful feature expression capabilities of the deep network to learn more discriminative deep spectral features for identifying objects in complex scenarios. The pseudo-color tracker is introduced to solve the problem of model performance limitation due to insufficient training data. By introducing the temporal-spectral–spatial online discrimination learning module, the temporal-spectral–spatial information of the target can be dynamically modeled to adapt to new targets and the dynamic changes of targets. Benefiting from the double Siamese network architecture, the model can be effectively trained from scratch with less than 20 000 training samples. Online learning of temporal-spectral–spatial information for the target, particularly in cases of insufficient training data, can alleviate the issue of model degradation. This approach enhances the model’s robustness when tracking the target in complex scenes. In the 2021 IEEE WHISPERS Hyperspectral Object Tracking (HOT) Challenge, HA-Net obtained the best performance, with a distance precision (DP) score of 0.948 and an area under the curve (AUC) score of 0.688. The running speed is also 14 frames/s, which is superior to the existing hyperspectral object trackers for hyperspectral video. The source code is available athttps://github.com/zhenliuzhenqi/HOT. Zhenqi Liu, Yanfei Zhong, Guorui Ma, Xinyu Wang 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | PU-KBS: A Robust Positive and Unlabeled Learning Framework With Key Band Selection for One-Class Hyperspectral Image ClassificationabstractPositive and unlabeled (PU) learning is aimed at building a binary classifier to distinguish the target from the background using only the known positive samples, which is an advanced solution for the hyperspectral target detection (HTD) task. However, when PU learning (PUL) meets complex hyperspectral scenarios, there are two main challenges: 1) How to estimate the class prior accurately? The class prior, i.e., the target proportion, is an important prior for PUL to learn the discriminant boundary, but it is difficult to estimate in hyperspectral imagery, due to the interclass spectral similarity and 2) How to remove redundancy and improve the discriminative features of the target? The diagnostic spectral feature extraction is important for the weakly supervised PUL models as it can help with separating the target from the background. In this article, to tackle these challenges, a robust PUL framework with key band selection (PU-KBS) is proposed, which is modeled as an end-to-end and class prior free PUL framework, where the accurate class prior and the most discriminative key band subset are jointly initialized and iteratively updated until reaching the optimal result by evolutionary search. Meanwhile, a deep PUL detector is introduced for guiding the subsequent search direction and discriminative deep feature extraction. The proposed PU-KBS framework was verified using different hyperspectral datasets, where accurate class prior estimation, diagnostic spectral characteristics, and robust detection results could be obtained simultaneously by the PU-KBS framework. Furthermore, the improvement in band selection interpretability and detection performance was proven experimentally. Ziying Liu, Hengwei Zhao, Xinyu Wang 0003, Shaoyu Wang 0003, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Semi-Supervised Semantic and Spatial Change Detail Retention Network for Semantic Change Detection in Remote Sensing ImagesabstractThe task of semantic change detection (SCD) in remote sensing images (RSIs) is aimed at identifying the multiple “from-to” changes in land cover. However, the prevailing multitask SCD network structure does not directly model the semantic change information, leading to potential inaccuracies in extracting multiple detailed changes. In this article, a semi-supervised semantic and spatial change detail retention network (S3CDRNet) is proposed to precisely model the details of spatial changes and the distribution of multiple change categories for RSIs. The encoder of the proposed S3CDRNet is based on a Siamese structure, where a shallow convolutional neural network (CNN) followed by a transformer are used to extract the local-global change features. A precise semantic change perception module (PSCPM) based on large kernel convolution is then introduced to enhance the weak semantic changes. The decoder consists of several deconvolution layers to restore the original resolution, and a pixelwise semantic change map is acquired. To further distinguish the inherent imbalanced change types, an adaptive category-balanced semi-supervised learning (ACBSS) strategy is developed to better use the abundant class distribution information from the unlabeled image pairs. In the experiments conducted in this study, we made an in-depth study of multiple SCD based on three public datasets: the SECOND dataset, the Landsat-SCD dataset, and the Hi-UCD mini dataset. The results show the potential of the proposed method under scenarios with various change classes and an imbalanced change class problem, compared with the state-of-the-art SCD methods. Pengyuan Lv, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | RBP-MTL: Agricultural Parcel Vectorization via Region-Boundary-Parcel Decoupled Multitask LearningabstractAgricultural parcel vectorization is important for precision agriculture when analyzing crops at the parcel scale. However, it is a challenging task to vectorize parcels from satellite imagery, where under-segmentation and the discontinuous boundaries of parcels are common problems, due to the varied shapes and sizes of parcels caused by the terrain and the farming mode, and the blurred boundaries caused by the limited resolution and the shadows. In this paper, a region-boundary-parcel decoupled multi-task learning (RBP-MTL) framework is proposed for agricultural parcel vectorization, where the local spatial constraints between the regions, boundaries, and the objects of each parcel are jointly modeled via multi-task learning, to promote the object separability of agricultural parcels and the boundary connectivity. In addition, simple and efficient boundary-object interaction vectorization is introduced to further ensure that each parcel is independent with a complete and separable boundary. In the experiments, the proposed framework was verified using three datasets with different spatial resolutions, i.e., the public parcel extraction dataset from the iFLYTEK Challenge 2021 (~1 m), in addition to a single-temporal Gaofen-1 dataset from Hubei province in China (2 m) and a multi-temporal Sentinel-2 dataset from the Netherlands (10 m), which were annotated and built by the authors. The proposed RBP-MTL framework achieved a state-of-the-art accuracy, with effective extraction of complete parcel boundaries in the different scenes. Xinyu Wang 0003, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Adaptive Self-Supporting Prototype Learning for Remote Sensing Few-Shot Semantic SegmentationabstractThe semantic segmentation of remote sensing images with few shots has important theoretical and application value. Most of the existing few-shot semantic segmentation frameworks are based on prototype learning methods, in which a single support prototype is designed to guide the query set for prediction. However, the visual differences between the support set and the query set make it difficult for a single support prototype, generated from the support set, to comprehensively encapsulate the semantic information of all the query images. This article introduces an adaptive self-supporting prototype learning network designed for few-shot segmentation (FSS), in order to tackle the challenges mentioned earlier. We propose adaptive hyperprototype representation (HPR), which consists of hyperprototype clustering (HPC) and guided prototype matching (GPM), to generate and assign multiple representative prototypes to compensate for the limitations of a single prototype in representing the semantic information of the query images. Specifically, HPC is a parameter-free and adaptive approach, which can extract more representative prototypes by aggregating similar feature vectors utilizing superpixel feature clustering. Meanwhile, GPM can select matched prototypes to provide more accurate guidance, allowing for uniformly aligned representation of multiple prototypes and complex image semantic information. We also introduce self-supporting matching (SSM) prototype learning, which can accurately guide the query set segmentation by acquiring query set prototypes. SSM generates initial pseudo labels for the query set based on the support set prototypes, and further guides the query set using the pseudo labels, along with the query prototypes generated by its own features, thus effectively avoiding visual differences between the support set and query set. The proposed adaptive self-supporting prototype learning network substantially improves the prototype quality and achieves a superior performance on object-level remote sensing datasets. Weihao Shen, Ailong Ma, Zhuo Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Multiobjective Spatiotemporal Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) aims to reconstruct a subpixel-level class distribution map from the pixel-level abundance maps, which is an under-determined problem that has nonunique solutions. To address this, the spatiotemporal SPM uses the abundance, spatial, and temporal constraints to reduce the uncertainty of the mapping solutions, so the spatiotemporal SPM is essentially a constrained optimization problem. However, it is hard to find the optimal weighting parameters to combine the three joint constraints. In addition, the existing spatiotemporal SPM methods mainly use the temporal information either for the unchanged subpixels detection or for the subpixel classification, which is insufficient in the utilization of the temporal information. In this article, a novel spatiotemporal SPM algorithm based on multiobjective optimization (STSPM_MO) is proposed. STSPM_MO is composed of an unchanged subpixels detection stage and a multiobjective spatiotemporal mapping stage. In the former stage, the historical thematic map is used for identifying the unchanged subpixels. In the latter stage, the historical thematic map is further used for providing the temporal dependence, so that the temporal information can be more fully utilized. Moreover, to solve the constrained optimization problem of the spatiotemporal SPM, the abundance, spatial, and temporal constraints are modeled as three objectives and are dynamically fused through the subfitness-based multiobjective evolution, to generate the optimal subpixel classification map. Both synthetic and real-data experiments have been conducted, and the results show the proposed method, and its two variants are superior, stable, and effective. Mi Song, Yanfei Zhong, Ailong Ma, Da He, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Attention in Attention for Hyperspectral With High Spatial Resolution (H) Image ClassificationabstractAn inevitable trend of hyperspectral remote sensing has been toward hyperspectral with high spatial resolution (H2) images. However, the higher resolution also brings higher spatial/spectral heterogeneity of surface features, which increases the difficulty of fine classification. Fully using global spatial–spectral features and contextual information is an effective method to alleviate spatial/spectral heterogeneity. Recently, to extract global spatial–spectral features with long-range dependencies, the self-attention mechanism has been widely used in H2 image classification and has achieved excellent results. As is well known, the simultaneous use of spatial and spectral information has always been a key aspect of hyperspectral image (HSI) processing; however, the current spatial and spectral attention modules only focus on the spatial and spectral features separately. This prevents further improvement in network performance, especially when the sample size is small. Therefore, a spatial–spectral attention-in-attention network (S2AiANet) is proposed, which solves the problem of the current spatial–spectral attention maps only focusing on single features through the spatial–spectral attention-in-attention (S2AiA) module. In addition, a multiscale attention (MSA) module is proposed to enhance the network’s adaptability to various complex scenarios. The experiments on two H2 datasets and one classic HSI dataset demonstrate that S2AiANet can achieve a significant performance improvement compared with the state-of-the-art hyperspectral classifiers. Ge Tang, Xinyu Wang 0003, Hengwei Zhao, Guang Jin, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | LPCN: Lightweight Precise Classification Network for Hyperspectral Remote Sensing Imagery Based on Multiobjective OptimizationabstractHyperspectral remote sensing image (HSI) has the unique advantages of spectral continuity as well as synchronous acquisition of both image and spectra of objects, which can achieve precise classification. For HSI classification, the deep learning (DL) methods have been fully developed due to its powerful layer by layer nonlinear feature learning ability, which is currently the mainstream method. However, current classification networks often refer directly to the fixed framework of natural image processing, and usually only focus on accuracy as an optimization goal, resulting in poor HSI data adaptation capability and parameter redundancy. In addition, HSIs usually have dozens of types of ground objects, with similar and mixed spectra, and the phenomenon of overlapping clusters is aggravated, making it difficult to distinguish similar objects. In this paper, a lightweight precise classification network (LPCN) for HSI based on multi-objective optimization was proposed. In LPCN, traditional fixed architectures are avoided, hierarchical lightweight search spaces are designed, and spectral attention mechanisms for similar objects are incorporated to improve the separability. Moreover, the accuracy and parameter function of the classification network are independently modeled and optimized simultaneously for reducing the amount of network parameters. The effectiveness of LPCN is proved by experiments with three HSI datasets, with 16, 16, and 22 types of ground objects, respectively. Yuting Wan, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Contrastive Scene Change Representation Learning for High-Resolution Remote Sensing Scene Change DetectionabstractScene change detection (SCD) involves recognizing the high-level semantic change types for a bitemporal remote sensing image scene pair. Previous methods have implicitly considered joint bitemporal features as a change representation to perform SCD. However, they have not effectively enhanced the discriminative power of the change representation. Consequently, these methods struggle to recognize scene changes with both intraclass variation and interclass similarity. In this paper, scene change contrastive (SCC) learning based on contrastive learning is proposed to ensure that bitemporal features are discriminative change representations for SCD recognition. Contrastive learning can learn specific discriminative features by gathering predefined specific positives and separating negatives in the projection space. In the SCC learning, the change representations, which are represented by the joint bitemporal features, are mapped to real and pseudo multi-view change projections by the proposed multi-view change (MVC) projector and the pseudo change augmentation (PeCA) strategy. The change projections are then guided to be discriminative, exploiting both local spatial and global information. By doing so, the joint bitemporal scene features become more discriminative change representations, which enable accurate recognition of scene changes. The extensive experimental results obtained on two public datasets consistently demonstrate the effectiveness of the proposed method. The code is available at https://github.com/wdczs/SCC. Jue Wang 0011, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Semantic Change Detection Based on Supervised Contrastive Learning for High-Resolution Remote Sensing ImageryabstractSemantic change detection (SCD) for high-resolution remote sensing imagery involves simultaneously locating the changed regions and identifying the semantic change categories. Recently, a series of multitask Siamese networks have been proposed to model the SCD task by merging the binary change detection (BCD) results and the bitemporal land-cover classification (LCC) results. However, due to the large reflectance variability of the land cover in bitemporal high-resolution images, the land-cover feature clusters extracted by these methods are still considerably mixed, leading to the misidentification of change and their semantic changes types. In this article, to handle this problem, the SiamContrast method is proposed to learn temporally invariant discriminative land-cover features for SCD. As part of SiamContrast, a novel SCD contrastive loss (SCD-CL) is proposed to enhance the temporally invariant feature discrimination across bitemporal images. SCD-CL utilizes supervised contrastive learning and consists of two complementary components: mono-temporal contrastive loss (MCL) and cross-temporal contrastive loss (CCL). In particular, MCL contrasts the land-cover features within each temporal image, to enhance the mono-temporal feature discrimination. Meanwhile, CCL with a change-aware hard anchor sampling (CHAS) strategy contrasts the land-cover features across bitemporal images, to align the land cover features of the same category. To validate the effectiveness of SCD-CL, the SiamContrast method incorporates a difference feature pyramid (DFP) decoder, which leverages feature distance to model changes, allowing it to directly benefit from the discriminative features learned by SCD-CL. The comprehensive experimental results consistently confirm the effectiveness of the proposed SiamContrast in improving SCD performance, compared with the existing SCD methods. The code is available athttps://github.com/wdczs/SiamContrast. Jue Wang 0011, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | IS-RoadDet: Road Vector Graph Detection With Intersections and Road Segments From High-Resolution Remote Sensing ImageryabstractExtracting road vector graphs with high accuracy from high-resolution remote sensing imagery presents a significant challenge. The prior end-to-end algorithms have typically modeled the road graph as a general graph structure with vertices and edges, denoted as$G=(V,E)$, which is a standard approach in graph generation tasks. However, the traditional$G=(V,E)$graph representation with vertices and edges, utilizing very small edge units, can conflict with the road network’s geometric structure and the inherent features of road instances, leading to issues such as false positives and disconnected roads. In this article, the IS-RoadDet framework is proposed to generate a road vector graph with intersections (I) and road segments (S), denoted as$G=(I,S)$, which leverages the minimum road topology unit features of road to improve road topology. Compared to$G=(V,E)$, instead of detecting a large number of vertices to maintain the road topology connectivity,$G=(I,S)$uses road segments with minimum road units to avoid false positives. In IS-RoadDet, the intersection and road segment detector (ISDetector) is introduced to detect intersections and road segments as independent instance objects with joint learning, and an intersection connectivity strategy (ICS) is designed to establish connectivity between road segments with intersections. Empirical experiments conducted on the SpaceNet3 and Sat2Graph datasets substantiated the superior performance of the proposed road segment modeling method. The code will be made available at:https://github.com/WanderRainy/IS-Roadandhttps://rsidea.whu.edu.cn/is-road.htm. Ruoyu Yang, Yanfei Zhong, Yinhe Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Occlusion-Aware Road Extraction Network for High-Resolution Remote Sensing ImageryabstractRoad occlusion seriously affects the connectivity of extracted roads, and has a negative effect in practical applications. The dense road occlusion problem is caused by high-rise buildings and street trees, and is a more serious and unique problem than simple occlusion caused by low buildings and scattered trees. The existing methods mainly solve the road occlusion problem by enhancing the encoder ability to capture the long connectivity feature of roads. Unfortunately, the existing methods can only solve small and sparse road occlusion situations, and they cannot deal with the dense road occlusions caused by dense high-rise buildings or trees. In this article, to solve the dense road occlusion problem, the occlusion-aware road extraction network, namely OARENet, is proposed for road extraction from high-resolution remote sensing imagery. In OARENet, an occlusion-aware decoder (OADecoder) is designed by explicit modeling the texture feature for road regions with dense occlusions. The OADecoder is made up of a regular occlusion-aware (ROA) module and a stochastic occlusion-aware (SOA) module. The ROA module is implemented by adopting different dilation rates to fit the texture feature in the semantic feature maps. The SOA module is proposed by designing stochastic convolutions to adaptively fit the spatial details of road regions with dense occlusions. In order to evaluate the dense occlusion problem, a dense occlusion road dataset (JHWV) was built and annotated. The experimental results obtained on the DeepGlobe dataset, the newly built JHWV dataset, and large-scale urban images demonstrate the superiority of OARENet, especially when faced with a dense road occlusion situation. Code has been made available at: https://github.com/WanderRainy/OARENet. Ruoyu Yang, Yanfei Zhong, Yinhe Liu, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Online Background Discriminative Learning for Satellite Video Object TrackingabstractSatellite video object tracking is a critical task in satellite video analysis, focused on locating targets within the imagery. Current research in satellite video single-object tracking primarily centers around traditional correlation filter-based methods and deep learning techniques based on Siamese networks. However, these methods encounter challenges, such as template inconsistency due to target appearance variations and imprecise object tracking caused by target-background similarities. In this paper, we introduce a novel approach called SiamOBR (Satellite Video Object Tracking Siamese Network with Online Background Discriminative Learning and Bounding Box Optimization) to address these issues. Building upon the baseline twin network framework, our model leverages an online background discriminative learning module to initialize using input satellite video frames. This module improves the network’s ability to distinguish the target by dynamically updating the template, effectively addressing target deformations resulting from motion. To tackle inaccuracies in bounding box estimation, we introduce a bounding box optimization module that refines tracking results obtained from the baseline tracker, further enhancing tracking accuracy. During network training, we employ probability regression in place of confidence regression and utilize a cross-entropy loss function to rectify labeling errors for small targets where the labeled box’s center point may not align perfectly with the object. We conducted quantitative experiments on real-world satellite video datasets, demonstrating that the SiamOBR framework outperforms existing satellite video object tracking models, showcasing its effectiveness in this domain. Yanfei Zhong, Xueting Fang, Meng Shu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Cross-Station Continual Aurora Image ClassificationabstractThe existing deep learning based methods have shown great potential for the aurora image classification problem. However, there are many differences in the morphology and distribution patterns of aurora images from different observation stations, and the differences between Antarctic and Arctic aurora images are particularly obvious. Currently, there are 76 research stations in 31 countries in Antarctica and more than 100 land-based stations in the Arctic. In the face of the difference in morphology and distribution patterns between Antarctic and Arctic auroras, the current popular methods cannot maintain a consistent classification ability. At the same time, it is important to effectively use both historical and real-time information to enable continual learning of aurora classification models to take full advantage of the high temporal resolution of streaming aurora image data. In this paper, a cross-station continual (CSC) aurora image classification framework is proposed to tackle these problems. To simulate a cross-station aurora image data stream, aurora images from three observation stations located in the Antarctic and the Arctic were selected and split into mini-batches in chronological order to form the cross-station streaming (CSS) aurora image dataset. Based on the vision transformer model, the CSC framework sequentially learns the semantic representation of streaming aurora data by learning dynamic prompts in the prompt bank selected by the average cosine distance. For the cross-station aurora discrepancy phenomenon, a local-global enhancement (LGE) module is designed, by organically combining the local and global semantics of aurora images to reduce the microscopic intra-class similarity and macroscopic inter-class confusion. Extensive experiments conducted on the CSS dataset show that the proposed method can achieve efficient continual learning of streaming aurora data and a competitive classification accuracy under the condition of joint training of data from multiple observation stations. Yanfei Zhong, Jingjun Yi, Richen Ye, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Self-Supervised Spaceborne Multispectral and Hyperspectral Image Fusion Unrolling NetworkabstractDeep learning has emerged as the predominant approach for multispectral and hyperspectral image fusion. However, most fusion networks are typically trained and validated on hyperspectral and multispectral image pairs generated from the same hyperspectral images, with degradation simulations inconsistent with real situations and relatively limited volumes of images. When transferring a pretrained multispectral and hyperspectral image fusion model from ground or airborne images to spaceborne images, it encounters a larger dataset and more complex spatial-spectral degradation, leading to spectral distortions and spatial artifacts in the fused images. In this article, the challenges associated with the transfer are addressed through the introduction of a self-supervised multispectral and hyperspectral image fusion unrolling network for spaceborne imagery, termed as MH-FUNet. MH-FUNet adopts a self-supervised paradigm to learn a robust mapping from spaceborne data. It utilizes a deep unrolling network to iteratively refine fusion results from coarse to fine. To account for spatial scale differences between the self-supervised training and test datasets, a multiscale fusion strategy is introduced. This strategy is combined with spectral and spatial attention mechanisms to restore spatial and spectral details. Additionally, a gradient constraint unit is proposed to maintain spatial consistency when up-scaling low-resolution hyperspectral imagery. Performance evaluation of the proposed method is conducted against state-of-the-art fusion techniques on both simulated Chikusei dataset and the proposed real WHU-MHF dataset, which consists of simultaneously observed hyperspectral and multispectral image pairs. MH-FUNet outperforms existing methods across all datasets, demonstrating superior performance in spaceborne multispectral and hyperspectral image fusion experiments. Zengliang Zhu, Xinyu Wang 0003, Guanzhong Li, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Learning a Cross-Modality Anomaly Detector for Remote Sensing ImageryabstractRemote sensing anomaly detector can find the objects deviating from the background as potential targets for Earth monitoring. Given the diversity in earth anomaly types, designing a transferring model with cross-modality detection ability should be cost-effective and flexible to new earth observation sources and anomaly types. However, the current anomaly detectors aim to learn the certain background distribution, the trained model cannot be transferred to unseen images. Inspired by the fact that the deviation metric for score ranking is consistent and independent from the image distribution, this study exploits the learning target conversion from the varying background distribution to the consistent deviation metric. We theoretically prove that the large-margin condition in labeled samples ensures the transferring ability of learned deviation metric. To satisfy this condition, two large margin losses for pixel-level and feature-level deviation ranking are proposed respectively. Since the real anomalies are difficult to acquire, anomaly simulation strategies are designed to compute the model loss. With the large-margin learning for deviation metric, the trained model achieves cross-modality detection ability in five modalities-hyperspectral, visible light, synthetic aperture radar (SAR), infrared and low-light-in zero-shot manner. Xinyu Wang 0003, Hengwei Zhao, Yanfei Zhong |
IEEE Trans. Image Process. | 4 |
| 2024 | E2SCNet: Efficient Multiobjective Evolutionary Automatic Search for Remote Sensing Image Scene Classification Network ArchitectureabstractRemote sensing image scene classification methods based on deep learning have been widely studied and discussed. However, most of the network architectures are directly reliant on natural image processing methods and are fixed. A few studies have focused on automatic search mechanisms, but they cannot weigh the interpretation accuracy and the parameter quantity for practical application. As a result, automatic global search methods based on multiobjective evolutionary computation have more advantages. However, in the ranking process, the network individuals with large parameter quantities are easy to eliminate, but a higher accuracy may be obtained after full training. In addition, evolutionary neural architecture search methods often take several days. In this article, in order to solve the above concerns, we propose an efficient multiobjective evolutionary automatic search framework for remote sensing image scene classification deep learning network architectures (E2SCNet). In E2SCNet, eight kinds of lightweight operators are used to build a diversified search space, and the coding connection mode is flexible. In the search process, a large model retention mechanism is implemented through two-step multiobjective modeling and evolutionary search, where one step involves the "parameter quantity and accuracy," and the other step involves the "parameter quantity and accuracy growth quantity." Moreover, a super network is constructed to share the weight in the process of individual network evaluation and promote the search speed. The effectiveness of E2SCNet is proven by comparison with several networks designed by human experts and networks obtained by gradient and evolutionary computing-based search methods. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Anomaly Segmentation for High-Resolution Remote Sensing Images Based on Pixel DescriptorsabstractAnomaly segmentation in high spatial resolution (HSR) remote sensing imagery is aimed at segmenting anomaly patterns of the earth deviating from normal patterns, which plays an important role in various Earth vision applications. However, it is a challenging task due to the complex distribution and the irregular shapes of objects, and the lack of abnormal samples. To tackle these problems, an anomaly segmentation model based on pixel descriptors (ASD) is proposed for anomaly segmentation in HSR imagery. Specifically, deep one-class classification is introduced for anomaly segmentation in the feature space with discriminative pixel descriptors. The ASD model incorporates the data argument for generating virtual abnormal samples, which can force the pixel descriptors to be compact for normal data and meanwhile to be diverse to avoid the model collapse problems when only positive samples participated in the training. In addition, the ASD introduced a multi-level and multi-scale feature extraction strategy for learning the low-level and semantic information to make the pixel descriptors feature-rich. The proposed ASD model was validated using four HSR datasets and compared with the recent state-of-the-art models, showing its potential value in Earth vision applications. Xinyu Wang 0003, Hengwei Zhao, Shaoyu Wang 0003, Yanfei Zhong |
AAAI | 5 |
| 2023 | Seeing Beyond the Patch: Scale-Adaptive Semantic Segmentation of High-resolution Remote Sensing Imagery based on Reinforcement LearningabstractIn remote sensing imagery analysis, patch-based methods have limitations in capturing information beyond the sliding window. This shortcoming poses a significant challenge in processing complex and variable geo-objects, which results in semantic inconsistency in segmentation results. To address this challenge, we propose a dynamic scale perception framework, named GeoAgent, which adaptively captures appropriate scale context information outside the image patch based on the different geo-objects. In GeoAgent, each image patch’s states are represented by a global thumbnail and a location mask. The global thumbnail provides context beyond the patch, and the location mask guides the perceived spatial relationships. The scale-selection actions are performed through a Scale Control Agent (SCA). A feature indexing module is proposed to enhance the ability of the agent to distinguish the current image patch’s location. The action switches the patch scale and context branch of a dual-branch segmentation network that extracts and fuses the features of multi-scale patches. The GeoAgent adjusts the network parameters to perform the appropriate scale-selection action based on the reward received for the selected scale. The experimental results, using two publicly available datasets and our newly constructed dataset WUSU, demonstrate that GeoAgent outperforms previous segmentation methods, particularly for large-scale mapping applications. Yinhe Liu, Sunan Shi, Yanfei Zhong |
ICCV | 4 |
| 2023 | Class Prior-Free Positive-Unlabeled Learning with Taylor Variational Loss for Hyperspectral Remote Sensing ImageryabstractPositive-unlabeled learning (PU learning) in hyperspectral remote sensing imagery (HSI) is aimed at learning a binary classifier from positive and unlabeled data, which has broad prospects in various earth vision applications. However, when PU learning meets limited labeled HSI, the unlabeled data may dominate the optimization process, which makes the neural networks overfit the unlabeled data. In this paper, a Taylor variational loss is proposed for HSI PU learning, which reduces the weight of the gradient of the unlabeled data by Taylor series expansion to enable the network to find a balance between overfitting and underfitting. In addition, the self-calibrated optimization strategy is designed to stabilize the training process. Experiments on 7 benchmark datasets (21 tasks in total) validate the effectiveness of the proposed method. Code is at: https://github.com/Hengwei-Zhao96/T-HOneCls. Hengwei Zhao, Xinyu Wang 0003, Yanfei Zhong |
ICCV | 4 |
| 2023 | Scalable Multi-Temporal Remote Sensing Change Data Generation via Simulating Stochastic Change ProcessabstractUnderstanding the temporal dynamics of Earth’s surface is a mission of multi-temporal remote sensing image analysis, significantly promoted by deep vision models with its fuel—labeled multi-temporal images. However, collecting, preprocessing, and annotating multi-temporal remote sensing images at scale is non-trivial since it is expensive and knowledge-intensive. In this paper, we present a scalable multi-temporal remote sensing change data generator via generative modeling, which is cheap and automatic, alleviating these problems. Our main idea is to simulate a stochastic change process over time. We consider the stochastic change process as a probabilistic semantic state transition, namely generative probabilistic change model (GPCM), which decouples the complex simulation problem into two more trackable sub-problems, i.e., change event simulation and semantic change synthesis. To solve these two problems, we present the change generator (Changen), a GAN-based GPCM, enabling controllable object change data generation, including customizable object property, and change event. The extensive experiments suggest that our Changen has superior generation capability, and the change detectors with Changen pre-training exhibit excellent transferability to real-world change datasets. Zhuo Zheng, Shiqi Tian, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 5 |
| 2023 | High-Resolution Fine-Grained Wetland Mapping Based on Class-Balanced Deep Semantic Segmentation NetworksabstractWetlands are among the most valuable environmental resources and are crucial to achieving sustainable development strategies. However, the number, distribution, and types of wetlands are poorly understood. Widely used medium-resolution wetland products do not provide enough surface feature for fine-grained wetland mapping. To fill this gap, a high-resolution wetland remote sensing mapping dataset covering the contiguous United States was constructed. On this dataset, the performance of deep semantic segmentation networks with various backbones and architecture for wetland mapping was evaluated. Several loss functions were implemented to address the issue of class imbalance of this dataset, resulting in an improvement in classification accuracy. High spatial resolution images offer an abundance of surface texture, shape, structure, and neighborhood relationship data that can be applied to the classification of large-scale wetlands. The combination of high-resolution imagery and deep semantic segmentation models enables the automatic classification of wetlands to be refined. Yinhe Liu, Sunan Shi, Yanfei Zhong |
IGARSS | 4 |
| 2023 | A graph-based framework to integrate semantic object/land-use relationships for urban land-use mapping with case studies of Chinese citiesabstractUrban land-use types, such as residential and administration, can be inferred through semantic objects and their relationships. Point of interest (POI) data can serve as the semantic objects for urban land-use mapping. However, the previous POI-based approaches have rarely considered the relationships between the semantic objects in the urban land-use mapping, and three main challenges remain: 1) the lack of paired semantic object/land-use samples; 2) the lack of a unified model for semantic objects and the relationships between sematic objects and urban land use; and 3) the difficulty of automatically learning semantic object/land-use mapping relationships. In this paper, to address these issues, a graph-based urban land-use mapping framework integrating semantic object/land-use relationships (GOLR) is proposed. Based on open-source area of interest (AOI) and POI data, an urban object/land-use (UOLU) dataset covering 34 cities in China was built. To model the spatial and mapping relationships, the semantic objects and their relationships are used to jointly build an urban land-use graph. The mapping from semantic objects to urban land use can then be learned by the urban land-use graph isomorphic network (ULGIN) model. Finally, the GOLR framework was applied to obtain accurate land-use mapping results for multiple Chinese cities. Yanfei Zhong, Yinhe Liu, Zhendong Zheng |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | FarSeg++: Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, a fundamental Earth vision task, always suffers from scale variation, the larger intra-class variance of background, and foreground-background imbalance in high spatial resolution (HSR) remote sensing imagery. Generic semantic segmentation methods mainly focus on the scale variation in natural scenarios. However, the other two problems are insufficiently considered in large area Earth observation scenarios. In this paper, we propose a foreground-aware relation network (FarSeg++) from the perspectives of relation-based, optimization-based, and objectness-based foreground modeling, alleviating the above two problems. From the perspective of the relations, the foreground-scene relation module improves the discrimination of the foreground features via the foreground-correlated contexts associated with the object-scene relation. From the perspective of optimization, foreground-aware optimization is proposed to focus on foreground examples and hard examples of the background during training to achieve a balanced optimization. Besides, from the perspective of objectness, a foreground-aware decoder is proposed to improve the objectness representation, alleviating the objectness prediction problem that is the main bottleneck revealed by an empirical upper bound analysis. We also introduce a new large-scale high-resolution urban vehicle segmentation dataset to verify the effectiveness of the proposed method and push the development of objectness prediction further forward. The experimental results suggest that FarSeg++ is superior to the state-of-the-art generic semantic segmentation methods and can achieve a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | An Accurate UAV 3-D Path Planning Method for Disaster Emergency Response Based on an Improved Multiobjective Swarm Intelligence AlgorithmabstractPlanning a practical three-dimensional (3-D) flight path for unmanned aerial vehicles (UAVs) is a key challenge for the follow-up management and decision making in disaster emergency response. The ideal flight path is expected to balance the total flight path length and the terrain threat, to shorten the flight time and reduce the possibility of collision. However, in the traditional methods, the tradeoff between these concerns is difficult to achieve, and practical constraints are lacking in the optimized objective functions, which leads to inaccurate modeling. In addition, the traditional methods based on gradient optimization lack an accurate optimization capability in the complex multimodal objective space, resulting in a nonoptimal path. Thus, in this article, an accurate UAV 3-D path planning approach in accordance with an enhanced multiobjective swarm intelligence algorithm is proposed (APPMS). In the APPMS method, the path planning mission is converted into a multiobjective optimization task with multiple constraints, and the objectives based on the total flight path length and degree of terrain threat are simultaneously optimized. In addition, to obtain the optimal UAV 3-D flight path, an accurate swarm intelligence search approach based on improved ant colony optimization is introduced, which can improve the global and local search capabilities by using the preferred search direction and random neighborhood search mechanism. The effectiveness of the proposed APPMS method was demonstrated in three groups of simulated experiments with different degrees of terrain threat, and a real-data experiment with 3-D terrain data from an actual emergency situation. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | Accurate Multiobjective Low-Rank and Sparse Model for Hyperspectral Image Denoising MethodabstractDue to the unavoidable influence of sparse and Gaussian noise during the process of data acquisition, the quality of hyperspectral images (HSIs) is degraded and their applications are greatly limited. It is therefore necessary to restore clean HSIs. In the traditional methods, low-rank and sparse matrix decomposition methods are usually applied to restore the pure data matrix from the observed data matrix. However, due to the fact that the optimization of the${l}_{0}$-norm for the sparse modeling is a nonconvex and NP-hard problem, convex relaxation and regularization parameters are usually introduced. However, convex relaxation often leads to inaccurate sparse modeling results, and the sensitive regularization parameters can lead to unstable results. Thus, in this article, to address these issues, an accurate multiobjective low-rank and sparse denoising framework is proposed for HSIs to achieve accurate modeling. The${l}_{0}$-norm is directly modeled as the sparse noise and is optimized by an evolutionary algorithm, and the denoising problem is converted into a multiobjective optimization problem through simultaneously optimizing the low-rank term, the sparse term, and the data fidelity term, without sensitive regularization parameters. However, since the low-rank clean image and sparse noise of the HSI are encoded into a solution, the length of the solution is too long to be optimized. In this article, a subfitness strategy is constructed to achieve effective optimization by comparing the objective function values corresponding to each band for each solution. The experiments undertaken with simulated images in 11 noise cases and four real noisy images confirm the effectiveness of the proposed method. Yuting Wan, Ailong Ma, Wei He 0003, Yanfei Zhong |
IEEE Trans. Evol. Comput. | 4 |
| 2023 | Unrolling Nonnegative Matrix Factorization With Group Sparsity for Blind Hyperspectral UnmixingabstractDeep neural networks have shown huge potential in hyperspectral unmixing (HU). However, the large function space increases the difficulty of obtaining the optimal solution with limited unmixing data. The autoencoder-based blind unmixing methods are sensitive to the hyperparameters, and the optimal solution can be difficult to obtain. Algorithm unrolling, which integrates deep learning and iterative algorithms, can shrink the search space and improve the efficiency of obtaining optimal results. Based on this, a model-driven deep neural network named the group sparsity regularized unmixing unrolling (GSUU) network, which unrolls a regularized matrix factorization objective function for blind HU, is proposed in this paper. Based on the nonnegative matrix factorization (NMF) optimization rules, the GSUU network contains two sub-networks—the A-Block and the S-Block—for alternately and iteratively estimating the optimal endmember spectra and abundance maps. The GSUU method incorporates the spatial group sparsity prior of the abundances, i.e., the fact that spatially adjacent mixed pixels share similar sparse abundances, into a deep unrolling network. The experimental results obtained with both synthetic and real hyperspectral data illustrate that the proposed algorithm can obtain a superior accuracy, compared to the other state-of-the-art unmixing algorithms. Chunyang Cui, Xinyu Wang 0003, Shaoyu Wang 0003, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Realistic Mixing Miniature Scene Hyperspectral Unmixing: From Benchmark Datasets to Autonomous UnmixingabstractMixed pixels that contain more than one material type are common in mid/low spatial resolution remote sensing imagery. Hyperspectral unmixing is aimed at decomposing the mixed pixels into endmembers and abundances. However, there are few datasets that are suitable for quantitatively evaluating unmixing accuracies, and the ground-truth abundances of the existing datasets are often generated in an approximate way. To address the lack of real unmixing datasets for quantitative evaluation, we built the realistic mixing miniature scenes (RMMS) dataset, which can be used to quantitatively evaluate the unmixing accuracy of different algorithms. The RMMS dataset consists of a simple mixture scene with homogeneous flat materials and a complex mixture scene with 3-D structural features. The features of the RMMS dataset also take point, line, and polygon characteristics into consideration, and the spectral similarity of the materials increases the challenge of the spectral unmixing. In the RMMS dataset, due to the multiscale observation characteristics of the spatiotemporal scanning modality, it can avoid the registration error between RGB and hyperspectral data, and it can ensure that the endmembers are pure pixels. Most of the autonomous hyperspectral unmixing algorithms focus on solving some of the unmixing problems and have difficulty achieving fully autonomous hyperspectral unmixing (FAHU). In this article, to overcome this shortcoming, a fully autonomous hyperspectral unmixing method called FAHU is proposed to take advantage of the spatial information. Some of the state-of-the-art autonomous hyperspectral unmixing algorithms are used to evaluate the performance with the RMMS dataset, and the experimental results show the advantages and disadvantages of the different autonomous unmixing algorithms. Chunyang Cui, Yanfei Zhong, Xinyu Wang 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Deep Hierarchical Pyramid Network With High- Frequency -Aware Differential Architecture for Super-Resolution MappingabstractSuper-resolution mapping (SRM) is a way to solve the mixed-pixel problem in urban land use/land cover caused by the limited spatial-resolving ability of satellite sensors, through resolution enhancement of the classification map. Recently, deep learning-based super-resolution mapping (DLSM) networks have been boomed, which can automatically learn a mapping pattern from low-resolution (LR) image to high-resolution (HR) land cover distribution to alleviate mixed-pixel problem. However, the urban compositions like buildings, trees, and roads exhibit a multiscale distribution with different size or orientation, which makes the traditional single-scale DLSM failed for an appropriate recognition. In addition, the urban compositions also show significant spatial heterogeneity with irregular distribution and intricate morphological shape, which are difficult to learn by simple convolutional layer. Therefore, it is necessary to explore the cue of these distribution characteristic to constrain the learning behavior of the network for better detail restoration. In this article, a deep hierarchical pyramid sub-pixel mapping network (HiSMNet) with high-frequency-aware differential architecture is proposed, which establishes an HP architecture to achieve explicit multiscale supervision of the feature map and prompt the network to learn a multiscale representation. In addition, a differential architecture is designed to enforce the network to intensify the learning of the high-frequency details. The validation experiments demonstrate that HiSMNet achieves superior performances in detailed delineation and outperformed the state-of-the-art DLSM models by up to 10% in terms of overall accuracy. Da He, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multiobjective Memetic Spatiotemporal Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) technology is an effective way to account for the distribution of the component objects within the mixed pixels and can alleviate the mixed pixel problem to some degree. However, the traditional SPM methods rely on only a single coarse-resolution image, and the limited information source can lead to great uncertainty. The rapid development of Earth observation systems has resulted in many fine spatial resolution remote sensing images now being accessible. Spatiotemporal SPM approaches utilize the finer spatial distribution of a historical thematic map to provide temporal prior information for the mapping process. However, spatiotemporal SPM is essentially a constrained optimization problem that aims to predict the optimal class distribution map subject to the abundance, spatial, and temporal constraints. It is a challenging task to properly model and optimize the multiple constraints. In this article, a multiobjective memetic spatiotemporal SPM (MOMSPM) framework is proposed. This model transforms the data fidelity term, spatial prior term, and temporal prior term into a multiobjective optimization problem (MOP) to discard the sensitive regularization parameters. The multiobjective model realizes the fusion of abundance, spatial, and temporal information. To optimize the three objective functions simultaneously, MOMSPM provides a multiobjective memetic algorithm framework in which the global multiobjective search method [multiobjective evolutionary algorithm based on decomposition (MOEA/D)] combines two commonly employed single-objective local search operators (geospatial distribution preference (GSDP) local search and maximum a posteriori (MAP)-based local search). The GSDP and MAP operators are employed to refine the solution and achieve an improved outcome. The hybridized method provides powerful search ability and achieves a good balance between the three objective functions. Experiments on synthetic and real datasets prove that the proposed method is superior to the state-of-the-art spatiotemporal SPM methods. Ailong Ma, Wen Zhou 0018, Mi Song, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Domain Adaptive Land-Cover Classification via Local Consistency and Global DiversityabstractUnsupervised domain adaptive (UDA) land-cover classification has recently gained more and more attention. UDA aimed to learn a model from the annotated source data and the unlabeled target data that can perform well on the target domain. The existing UDA frameworks based on adversarial training and self-training methods have boosted this field a lot. However, these methods almost all originate from the computer vision field, and they ignore the very nature of high-resolution remote sensing (HRS) images. The core insight of this paper is that a good land-cover classification result always has strong local consistency and good global diversity, which makes it possible to construct a metric representing the properties of good land-cover mapping, to improve the existing UDA algorithms. Firstly, based on this finding, we prove that local consistency and global diversity can be measured by the Frobenius norm and nuclear norm, respectively. Secondly, we propose a novel local consistency and global diversity metric (LCGDM), which can be easily integrated into the existing UDA frameworks. Finally, the experiments conducted on the LoveDA data set prove the validity of the proposed metric, which can not only improve the overall land-cover mapping but also the category-wise prediction. Ailong Ma, Chenyu Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | SiamOHOT: A Lightweight Dual Siamese Network for Onboard Hyperspectral Object Tracking via Joint Spatial-Spectral Knowledge DistillationabstractHyperspectral object tracking is aimed at tracking targets by using both the spatial information and abundant spectral information, overcoming the drawbacks of traditional RGB tracking in complex scenarios, such as the low resolution or background clutter. However, the current hyperspectral object tracking methods usually have a high computational complexity, due to the huge data volume, making them difficult to apply to real-time applications on edge devices (e.g., robots, unmanned aerial vehicles, and satellites) with limited computational resources. In this paper, a lightweight dual Siamese network for onboard hyperspectral object tracking—termed SiamOHOT—is proposed for real-time and onboard tracking. Specifically, a joint spatial-spectral knowledge distillation method is proposed to teach a lightweight dual Siamese tracker to learn from a deep tracker— SiamHYPER—so that the number of parameters can be compressed to improve the computational efficiency. In addition, a deep learning inference optimizer is introduced to fuse the layers with similar functions and quantify the parameters of the network, to further promote the processing speed when deployed on an embedded platform. The proposed lightweight model was verified using the 2021 WHISPERS Hyperspectral Object Tracking Challenge dataset, and achieved a superior efficiency and accuracy. In addition, a prototype system was built integrating a snapshot hyperspectral imager, the SiamOHOT tracking algorithm, and an artificial intelligence edge device (NVIDIA Jetson Xavier NX), to realize real-time imaging and tracking. The inference speed of the optimized SiamOHOT network is nearly doubled when compared to the teacher model on the prototype system. Xinyu Wang 0003, Zhenqi Liu, Yuting Wan, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Adaptive Multistrategy Particle Swarm Optimization for Hyperspectral Remote Sensing Image Band SelectionabstractHyperspectral remote sensing band selection picks out characteristic feature combination to weaken the strong correlation caused by spectral continuity. However, it is difficult for traditional methods with fixed strategies to search the entire space and make adjustments for the optimization process. Thus, the solutions obtained can be mostly local optima. In this paper, a novel adaptive multi-strategy particle swarm optimization for hyperspectral image remote sensing band selection (AMSPSO_BS) is introduced to obtain a subset solution suitable for classification. The problem is modeled as an effective fitness function, and the quotient of the linear discriminant value and the mean mutual information (LD/MMI) is used to remove the redundancy between bands. The randomly generated solutions are then encoded to form a population, which rely on various particle update strategies (PUS) with different reference positions for updating. During the particle motion, the effect of each strategy on population evolution is considered comprehensively and reflected in the change of selection probability. And the motion parameters are dynamically adjusted to balance the global and local capabilities. Four hyperspectral remote sensing image datasets were utilized to conduct band selection experiments, to confirm the effectiveness of AMSPSO_BS. Yuting Wan, Chao Chen 0029, Ailong Ma, Liangpei Zhang 0001, Xunqiang Gong, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Change Detection Based on Supervised Contrastive Learning for High-Resolution Remote Sensing ImageryabstractChange detection (CD) is a challenging task on high-resolution bitemporal remote sensing images. Many recent studies of CD have focused on designing fully convolutional Siamese network architectures. However, most of these methods initialize their encoders by random values or an ImageNet pretrained model, without any prior for the CD task, thus limiting the performance of the CD model. In this article, the novel supervised contrastive pretraining and fine-tuning CD (SCPFCD) framework, which is made up of two cascaded stages, is presented to train a CD network based on a pretrained encoder. In the first supervised contrastive pretraining stage, the encoder of the Siamese network is asked to solve a joint pretext task introduced by the proposed CDContrast pretraining method on labeled CD data. The proposed CDContrast pretraining method includes land contrastive learning (LCL), which is based on supervised contrastive learning, and proxy CD learning. The LCL focuses on learning the spatial relationships among the land cover from bitemporal images by solving a land contrast task, while the proxy CD learning performs a proxy CD task on the top of the upsampling projector to avoid local optima for the LCL and learn features for the CD. Then, in the second fine-tuning stage, the whole Siamese network initialized with the pretrained encoder is fine-tuned to perform the CD task in an end-to-end manner. The proposed SCPFCD framework was verified with three CD datasets of high-resolution remote sensing images. The extensive experimental results consistently show that the proposed framework can effectively improve the ability to extract change information for Siamese networks. Jue Wang 0011, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | One-Class Risk Estimation for One-Class Hyperspectral Image ClassificationabstractHyperspectral imagery (HSI) one-class classification is aimed at identifying a single target class from the HSI by using only knowing positive data, which can significantly reduce the requirements for annotation. However, when one-class classification meets HSI, it is difficult for classifiers to find a balance between the overfitting and underfitting of positive data due to the problems of distribution overlap and distribution imbalance. Although deep learning-based methods are currently the mainstream to overcome distribution overlap in HSI multi-classificaiton, few researches focus on deep learning-based HSI one-class classification. In this paper, a weakly supervised deep HSI one-class classifier, namelyHOneClsis proposed, where a risk estimator—theOne-Class Risk Estimator—is particularly introduced to make the full convolutional neural network (FCN) with the ability of one class classification in the case of distribution imbalance. Extensive experiments (20 tasks in total) were conducted to demonstrate the superiority of the proposed classifier. Hengwei Zhao, Yanfei Zhong, Xinyu Wang 0003, Hong Shu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | AutoLC: Search Lightweight and Top-Performing Architecture for Remote Sensing Image Land-Cover ClassificationabstractLand-cover classification has long been a hot and difficult challenge in remote sensing community. With massive High-resolution Remote Sensing (HRS) images available, manually and automatically designed Convolutional Neural Networks (CNNs) have already shown their great latent capacity on HRS land-cover classification in recent years. Especially, the former can achieve better performance while the latter is able to generate lightweight architecture. Unfortunately, they both have shortcomings. On the one hand, because manual CNNs are almost proposed for natural image processing, it becomes very redundant and inefficient to process HRS images. On the other hand, nascent Neural Architecture Search (NAS) techniques for dense prediction tasks are mainly based on encoder-decoder architecture, and just focus on the automatic design of the encoder, which makes it still difficult to recover the refined mapping when confronting complicated HRS scenes.To overcome their defects and tackle the HRS land-cover classification problems better, we propose AutoLC which combines the advantages of two methods. First, we devise a hierarchical search space and gain the lightweight encoder underlying gradient-based search strategy. Second, we meticulously design a lightweight but top-performing decoder that is adaptive to the searched encoder of itself. Finally, experimental results on the LoveDA land-cover dataset demonstrate that our AutoLC method outperforms the state-of-art manual and automatic methods with much less computational consumption. Chenyu Zheng, Ailong Ma, Yanfei Zhong |
ICPR | 4 |
| 2022 | Graph Laplacian Regularized Spectral-Spatial-Sparse Unmixing for Hyperspectral ImageryabstractSparse unmixing aims at finding the optimal subset of endmembers in a spectral library to approximate the observed data, and has received increasing attention as it can circumvent the estimation of the endmember. In this paper, a graph Laplacian regularized spectral-spatial-sparse unmixing algorithm is proposed, namely, gLapS3U, incorporating the graph Laplacian regularization to consider the similarity between pixels of the whole image, and enforcing the spectral-spatial-sparse constraints to enhance the local spatial information as well as the sparsity of the abundance solution jointly. Experimental results on simulated and real data show the superiority of the proposed algorithm compared with state-of-the-art existing methods. Zhi Li 0080, Ruyi Feng, Yichang Shi, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng |
IGARSS | 5 |
| 2022 | High Resolution Remote Sensing Image Semantic Segmentation Based on Ultra-Lightweight Fully Convolution Neural NetworkabstractIn recent years, fully convolutional neural networks (FCNs) have been widely used in the field of remote sensing image semantic segmentation. However, these networks have huge amount of parameters and cost much computational efficiency. In this paper, an ultra-lightweight network (ULN) is proposed to overcome this problem. The proposed ULN model uses the encoder-decoder architecture to acquire the pixelwise result. In ULN, the efficient spatial pyramid network (ESPNet) is used to extract deep semantic features with fewer parameters. Considering the dilated convolutions will lose some semantic information in the encoding process, the feature enhancement block (FEB) is proposed. The recurrent criss-cross attention module is added at the end of skip connection to acquire the global contextual information. The proposed ULN is tested on the ISPRS Vaihingen dataset, the results show that our network achieves competitive results with fewer parameters(1.5M). Pengyuan Lv, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2022 | Remote Sensing Image Super-Resolution via Dilated Convolution Network with Gradient PriorabstractDue to the limitations of the imaging sensor, the spatial resolution of satellite imagery is often insufficient, namely, low resolution (LR). Therefore, super-resolution (SR) is proposed, which strives to improve image resolution, perfectly to compensate for the shortcomings of satellite sensor imaging. In this study, we develop a unique dilated convolution network with gradient prior (DCNG) for remote sensing SR, aiming to extract powerful low-level features with gradient prior and efficitive network and then reconstruct the high-level feature details. The DCNG is built of two components: the Multi-Scale Feature Extraction Network and the Feature Reconstruction Network. In the Multi-Scale Feature Extraction Network, the Double-Path Dilated Residual Block (DPDRB) is designed with the dilation convolution operation to obtain the multi-scale features and increase the receptive field, the Global Self-attention Module (GSA) to catch the long-range dependency among picture patches, and a Gradient Propagation Network (GPN) is proposed to extract high-level gradient information. In the Feature Reconstruction Network, the Pixel Shuffle is introduced to reconstruct the feature by combining characteristics of different frequency bands. Experiments using Massachusetts_Roads and 3K VEHICLE_SR data sets indicate that our DCNG surpasses state-of-the-art algorithms in terms of quantitative and qualitative evaluations. Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng |
IGARSS | 4 |
| 2022 | GRE and Beyond: A Global Road Extraction DatasetabstractAccurate and timely road mapping that describes the road network geometry and topology is the key element of intelligent transport systems and smart city management. However, current global road maps like OpenStreetMap (OSM) are typically outdated and spatially incomplete with uneven accuracies. Although the development of remote sensing satellite technology and the advance of computer vision technology have made it possible to quickly extract road networks from massive very-high-resolution (VHR) remote sensing imagery, existing road extraction methods are limited by the problem: lacking of an accurate and diverse training dataset for global-scale road extraction, and manually labelling millions of road samples for training a global model is labor intensive. To address this problem, we utilized VHR satellite imagery and open-source crowdsourcing geospatial big data to build a robust global-scale road training dataset, termed GlobalRoadNet, for global road extraction (GRE) and beyond. The proposed GlobalRoadNet contains 47210 samples from 121 capital cities of six continents in Europe, Africa, Asia, South America, Oceania, and North America. Experimental results show that GlobalRoadNet can significantly improve model performance, not only can be applied for road extraction, but also has the potential to update OSM road data. Yanfei Zhong, Zhuo Zheng |
IGARSS | 2 |
| 2022 | Review of Vision Transformer Models for Remote Sensing Image Scene ClassificationabstractAs an important semantic understanding method of remote sensing images, scene classification has received much attention in recent years. Convolutional neural network (CNN) is the representative deep learning method for scene classification which has powerful ability in feature extraction. However, the multilevel features in CNN are acquired by hierarchical convolutional layers which have difficulty in considering the interaction of different objects in the scene. Vision transformer (ViT) model provides a new way to understand the image by directly modeling the contextual information of local patches. This paper makes a review of recent progress of ViT models in the field of computer vision and remote sensing. The major contributions are as follows: 1) A brief review of the traditional scene classification methods is made; 2) ViT based models for scene classification are introduced and compared with CNN models; 3) Experiments of recent ViT models are performed and analyzed on UCM and NWPU datasets. Pengyuan Lv, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2022 | Semantic Change Detection Based on a New Chinese Satellite Dataset and a Deep Conditional Random Field FrameworkabstractIn this paper, a new semantic change detection (CD) dataset based on Chinese Gaofen-2 (GF-2) satellite images with high spatial resolution (HSR) namely Wuhan Urban Semantic Understanding (WUSU) dataset is built up and a CD framework combining binary and semantic CD tasks based on deep learning and a conditional random field model (SDCRF) is proposed. Existing CD datasets mostly focus on “change/no change”. Traditional CD methods pay attention only on either of the binary CD task or the semantic CD task. Although there are methods to handle both tasks simultaneously but they ignore the inconsistency between the two tasks. In the SDCRF framework, any state-of-the-art feature extraction model can be used to extract the class and change probabilities as the unary potential of a fully connected conditional random field (FC-CRF) model which is adopted as a post-processing to enhance the location information of deep networks and reduce outlier noise. Sunan Shi, Yanfei Zhong, Yinhe Liu, Jue Wang 0011, DeRen Li |
IGARSS | 2 |
| 2022 | Mae-Net: A Micro Network Architecture Evolutionary Search Method for Remote Sensing Image Scene ClassificationabstractDeep learning based remote sensing scene classification methods have become a research hotspot, but they can not fully mine the image information due to the architecture comes directly from natural image. The automatic search method-based network architecture has then attracted a lot of attention benefits by its ability to independently learn the network structure suitable for remote sensing data. However, in the process of search and sorting, slightly larger models with better performance after full training are often eliminated due to insufficient training. Moreover, the methods often spend a lot of time searching. In this paper, a micro network architecture evolutionary search method is proposed (MAE-Net), the contributions are reflected in the slightly larger model retention mechanism by two-layer multi-objective functions and the super network mechanism used to reduce search time through weight sharing. The effectiveness is proved by comparison with human expert and search based networks on NWPU45 dataset. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2022 | Capformer: Pure Transformer for Remote Sensing Image CaptionabstractAccurately describing high-spatial resolution remote sensing images requires the understanding the inner attributes of the objects and the outer relations between different objects. The existing image caption algorithms lack the ability of global representation, which are not fit for the summarization of complex scenes. To this end, we propose a pure transformer (CapFormer) architecture for remote sensing image caption. Specifically, a scalable vision transformer is adopted for image representation, where the global content can be captured with multi-head self-attention layers. A transformer decoder is designed to successively translate the image features into comprehensive sentences. The transformer decoder explicitly model the historical words and interact with the image features using cross-attention layers. The comprehensive and ablation experiments on RSICD dataset demonstrate that the CapFormer outperforms the state-of-the-art image caption methods. Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
IGARSS | 4 |
| 2022 | HFGAN: A Heterogeneous Fusion Generative Adversarial Network for Sar-to-Optical Image TranslationabstractDue to the influence of the imaging mechanism of SAR images, it is difficult to interpret ground information through SAR images without expert knowledge. On the contrary, optical images have rich spatial and color information, so it is necessary to conduct research on the translation of SAR to optical remote sensing images. In this end, we propose a heterogeneous fusion generative adversarial network (HFGAN) for SAR-to-optical image translation. There are two main improvements: (1) Complementary generation of global structure and texture information. A heterogeneous fusion generator and a multi-scale discriminator are proposed to ensure that the global and detailed features of the generated image are more accurate and rich. (2) Color fidelity. Chromatic aberration loss are introduced to reduce the color difference between the generated image and the real optical image. Through qualitative and quantitative experiments, it is proved that the proposed method not only obtains better visual effects, but also has certain progress in the evaluation metrics, which proves that the proposed method is superior to the previous advanced methods in SAR-to-optical image translation. Ailong Ma, Yanfei Zhong, Xiaodong Gong |
IGARSS | 3 |
| 2022 | TypeFormer: Multiscale Transformer With Type Controller for Remote Sensing Image CaptionabstractImage captioning in remote sensing can help us understandthe inner attributes of the objects and the outer relations between different objects. However, the existing image captioning algorithms lack the ability of global representation, and cannot obtain object relations over long distances. In addition, these algorithmics generate captions randomly without consideration of the specific demands. To this end, we propose a pure transformer architecture with caption type controller for remote sensing image captioning. Specifically, a multi-scale vision transformer is adopted for the image representation, where the global and detailed content can be captured with multi-head self-attention layers. A transformer decoder is then introduced to successively translate the image features into comprehensive sentences. The optional block called the caption type controller is designed to consider the types of captions through caption type matrix sets according to the demands, embedding the learnable sentence feature with the required type. The comparison and ablation experiments conducted on the Remote Sensing Image Captioning Dataset (RSICD) dataset demonstrate that the proposed framework outperforms the current state-of-the-art image captioning methods. The experiments conducted on the FloodNet caption dataset further illustrate that the proposed methods can effectively generate specific types of captions. Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Spectral-Spatial Fusion Sub-Pixel Mapping Based on Deep Neural NetworkabstractSub-pixel mapping (SPM) has been widely adopted to alleviate the mixed pixel problem in hyperspectral image, as an extension of spectral unmixing (SU), providing a way to observe the spatial location of the endmember within mixed pixel. However, most of the SPM methods are unmixing-then-mapping (UTM), i.e., SPM process relies on the abundance images generated from SU, in which process uncertainty inherently exists and would be propagated to SPM. Furthermore, the prior knowledge toward the sub-pixel scale distribution is mainly model-driven/handcrafted, which has limitation for geographical-realistic distribution representation. In this letter, we proposed spectral–spatial fusion SPM based on deep neural network (SSNET), to realize the integrative modeling of SU and SPM problem in a unified network fashion to avoid uncertainty accumulation in UTM process, and it can simultaneously generate SU result and SPM result. Besides, SSNET provides a supervised manner to learn prior knowledge with external exemplar pairs of low- and high-resolution images for a geographical-realistic distribution representation. The experiment with two hyperspectral images validated the superiority of the proposed SSNET. Da He, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Xiaoding Liu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Multiobjective Sine Cosine Algorithm for Remote Sensing Image Spatial-Spectral ClusteringabstractRemote sensing image data clustering is a tough task, which involves classifying the image without any prior information. Remote sensing image clustering, in essence, belongs to a complex optimization problem, due to the high dimensionality and complexity of remote sensing imagery. Therefore, it can be easily affected by the initial values and trapped in locally optimal solutions. Meanwhile, remote sensing images contain complex and diverse spatial-spectral information, which makes them difficult to model with only a single objective function. Although evolutionary multiobjective optimization methods have been presented for the clustering task, the tradeoff between the global and local search abilities is not well adjusted in the evolutionary process. In this article, in order to address these problems, a multiobjective sine cosine algorithm for remote sensing image data spatial-spectral clustering (MOSCA_SSC) is proposed. In the proposed method, the clustering task is converted into a multiobjective optimization problem, and the Xie-Beni (XB) index and Jeffries-Matusita (Jm) distance combined with the spatial information term (SI_Jm measure) are utilized as the objective functions. In addition, for the first time, the sine cosine algorithm (SCA), which can effectively adjust the local and global search capabilities, is introduced into the framework of multiobjective clustering for continuous optimization. Furthermore, the destination solution in the SCA is automatically selected and updated from the current Pareto front through employing the knee-point-based selection approach. The benefits of the proposed method were demonstrated by clustering experiments with ten UCI datasets and four real remote sensing image datasets. Yuting Wan, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Cybern. | 4 |
| 2022 | A Spectral-Spatial-Dependent Global Learning Framework for Insufficient and Imbalanced Hyperspectral Image ClassificationabstractDeep learning techniques have been widely applied to hyperspectral image (HSI) classification and have achieved great success. However, the deep neural network model has a large parameter space and requires a large number of labeled data. Deep learning methods for HSI classification usually follow a patchwise learning framework. Recently, a fast patch-free global learning (FPGA) architecture was proposed for HSI classification according to global spatial context information. However, FPGA has difficulty in extracting the most discriminative features when the sample data are imbalanced. In this article, a spectral-spatial-dependent global learning (SSDGL) framework based on the global convolutional long short-term memory (GCL) and global joint attention mechanism (GJAM) is proposed for insufficient and imbalanced HSI classification. In SSDGL, the hierarchically balanced (H-B) sampling strategy and the weighted softmax loss are proposed to address the imbalanced sample problem. To effectively distinguish similar spectral characteristics of land cover types, the GCL module is introduced to extract the long short-term dependency of spectral features. To learn the most discriminative feature representations, the GJAM module is proposed to extract attention areas. The experimental results obtained with three public HSI datasets show that the SSDGL has powerful performance in insufficient and imbalanced sample problems and is superior to other state-of-the-art methods. Qiqi Zhu, Weihuan Deng, Zhuo Zheng, Yanfei Zhong, Qingfeng Guan 0001, Weihua Lin, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Cybern. | 4 |
| 2022 | A Decomposition-Based Multiobjective Clonal Selection Algorithm for Hyperspectral Image Feature SelectionabstractFeature selection is an effective way to handle the strong correlation of hyperspectral image data by screening the significant features, and is generally accepted to be a multiobjective optimization problem. Nevertheless, due to the randomness of the strategies and the ambiguity of the optimization directions, the existing multiobjective evolutionary optimization based feature selection methods can suffer from inefficient search and loss of search space with promising solutions when faced with the high-dimensional and multi-peak search space. The multiobjective evolutionary algorithm based on decomposition (MOEA/D) employs a decomposition framework to provide exact guidance for the optimization directions. Unfortunately, random operators are still used, leading to inadequate local optimization. Thus, evolutionary strategies with search preference such as clonal selection may be necessary for local search. In this paper, a novel decomposition-based multiobjective clonal selection algorithm for feature selection (MOCSA/D_FS) is proposed to obtain a feature subset with a superior classification performance. In MOCSA/D_FS, the information entropy and the ratio of the relative scatter value and mutual information are utilized as two objective functions to evaluate the information amount and redundancy. A series of subproblems are then obtained by decomposing the multiobjective problem through weight vectors, with anl2-norm constraint used to balance the search space. Subsequently, a clonal selection method with search space preference performs a detailed local search on each subproblem, which can fully exploit the potential optimal space. The effectiveness and generalizability of the proposed method was confirmed by experiments on four hyperspectral remote sensing image datasets. Chao Chen 0029, Yuting Wan, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Cross-Modality Image Matching Network With Modality-Invariant Feature Representation for Airborne-Ground Thermal Infrared and Visible DatasetsabstractThermal infrared (TIR) remote-sensing imagery can allow objects to be imaged clearly at night through the long-wave infrared, so that the fusion of thermal infrared and visible (VIS) imagery is a way to improve the remote-sensing interpretation ability. However, due to the large radiation difference between the two kinds of images, it is very difficult to match them. One of the most important issues is the lack of comprehensive consideration of the modality-specific information and modality-shared information, which makes it difficult for the existing methods to obtain a modality-invariant feature representation. In this article, a cross-modality image matching network, which we refer to as CMM-Net, is proposed to realize thermal infrared and visible image matching by learning a modality-invariant feature representation. First, in order to extract the modality-specific features of the imagery, the framework constructs a shallow two-branch network to make full use of the modality-specific information, without sharing parameters. Second, in order to extract the high-level semantic information between the different modalities, modality-shared layers are embedded into the deep layers of the network. In addition, three novel loss functions are designed and combined to learn the modality-invariant feature representation, that is, the discriminative loss of the non-corresponding features in the same modality, the cross-modality loss of the corresponding features between different modalities, and the cross-modality triplet (CMT) loss. The multimodal matching experiments conducted with ground- and airborne-based thermal infrared images and visible images showed that the proposed method outperforms the existing image matching methods by about 2% and 6% for the ground and airborne images, respectively. Ailong Ma, Yuting Wan, Yanfei Zhong, Bin Luo 0005, Miaozhong Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MAP-Net: SAR and Optical Image Matching via Image-Based Convolutional Network With Attention Mechanism and Spatial Pyramid Aggregated PoolingabstractThe complementarity of synthetic aperture radar (SAR) and optical images allows remote sensing observations to “see” unprecedented discoveries. Image matching plays a fundamental role in the fusion and application of SAR and optical images. However, both the geometric imaging pattern and the physical radiation mechanism of these two sensors are significantly different, so that the images show complex geometric distortion and nonlinear radiation differences. This phenomenon brings great challenges to image matching, which neither the handcrafted descriptors nor the deep learning-based methods have adequately addressed. In this article, a novel image-based matching method for SAR to optical images via an image-based convolutional network with spatial pyramid aggregated pooling (SPAP) and an attention mechanism is proposed, namely MAP-Net. The original image is embedded through the convolutional neural network to generate the feature map. Through the information extraction and abstraction of the original imagery, the embedded features containing the high-level semantic information are more robust to the geometric distortion and radiation variation among the different modal images, which is beneficial to the matching of cross-modal images. The adoption of the SPAP module makes the network more capable of integrating global and local contextual information. The attention block weights the dense features generated from the network to extract the key features that are invariant, distinguishable, repeatable, and suitable for the image matching task. In the experiments, five sets of multisource and multiresolution SAR and optical images with wide and varied ground coverage were used to evaluate the accuracy of MAP-Net, compared to both handcrafted and deep learning-based methods. The experimental results show that the MAP-Net method is superior to the current state-of-the-art image matching methods for SAR to optical images. Ailong Ma, Liangpei Zhang 0001, Miaozhong Xu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | CCANet: Class-Constraint Coarse-to-Fine Attentional Deep Network for Subdecimeter Aerial Image Semantic SegmentationabstractSemantic segmentation is important for the understanding of subdecimeter aerial images. In recent years, deep convolutional neural networks (DCNNs) have been used widely for semantic segmentation in the field of remote sensing. However, because of the highly complex subdecimeter resolution of aerial images, inseparability often occurs among some geographic entities of interest in the spectral domain. In addition, the semantic segmentation methods based on DCNNs mostly obtain context information using extra information within the added receptive field. However, the context information obtained this way is not explicit. We propose a novel class-constraint coarse-to-fine attentional (CCA) deep network, which enables the formation of class information constraints to obtain explicit long-range context information. Further, the performance of subdecimeter aerial image semantic segmentation can be improved, particularly for fine-structured geographic entities. Based on coarse-to-fine technology, we obtained a coarse segmentation result and constructed an image class feature library. We propose the use of the attention mechanism to obtain strong class-constrained features. Consequently, pixels of different geographic entities can adaptively match the corresponding categories in the class feature library. Additionally, we employed a novel loss function, CCA-loss to realize end-to-end training. The experimental results obtained using two popular open benchmarks, International Society for Photogrammetry and Remote Sensing (ISPRS) 2-D semantic labeling Vaihingen data set and Institute of Electrical and Electronics Engineers (IEEE) Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest Zeebrugge data set, validated the effectiveness and superiority of our proposed model. The proposed method achieved state-of-the-art performance on the IEEE GRSS Data Fusion Contest Zeebrugge data set. Guohui Deng, Zhaocong Wu, Miaozhong Xu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Local Spatial Constraint and Total Variation for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection, which is aimed at locating anomaly, has received widespread attention. In this article, a new anomaly detector, named local spatial constraint and total variation (LSC-TV), is proposed for hyperspectral imagery. In anomaly detection methods based on low-rank representation, background pixels are usually considered to have a global low-dimensional structure. However, the complex background distribution in hyperspectral images (HSIs) means that this global low-dimensional structure rarely occurs. In LSC-TV, the effective local spatial information is extracted by superpixel segmentation, and the regularization based on the F-norm is used to force the background within the same superpixel to show uniform spectral features. Moreover, each pixel is given a penalty based on the degree of anomaly determined during model iteration, while the anomaly is not considered by the background constraint. In addition, the background pixels in the neighborhood often show a high correlation, whereas the anomaly does not possess this feature. Nonisotropic TV is introduced into the proposed LSC model using the correlation of first-order neighborhoods to make it easier for anomalies to be separated. The proposed LSC-TV method and current state-of-the-art methods are tested on a set of simulated data and four sets of real data. The experimental results demonstrate that the proposed method is superior to the comparative method in terms of both color map detection and quantitative evaluation. Ruyi Feng, Hao Li 0058, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SPNet: Spectral Patching End-to-End Classification Network for UAV-Borne Hyperspectral Imagery With High Spatial and Spectral ResolutionsabstractIn deep learning (DL)-based hyperspectral imagery classification, “spatial patching” is primarily used as a preprocessing for incorporating local spatial information. This operation can help to promote classification accuracy but it is facing new challenges in the unmanned aerial vehicle (UAV)-borne hyperspectral imagery with high spatial and spectral resolutions (H2imagery). The ground objects’ various spatial scales result in it being challenging to determine the optimal size for the spatial patches. In addition, due to the severe spectral variability and spatial heterogeneity of the H2imagery, “spatial patching” only exploits the local spatial information and results in serious salt-and-pepper (SP) noise and isolated areas in the classification maps. In this article, to address these issues, a novel spectral patching network (SPNet) with an end-to-end DL architecture is proposed for UAV-borne H2imagery classification. The “spectral patching” approach is proposed to preserve the global spatial information and almost all the spectral information of the original hyperspectral imagery. An end-to-end deep encoder–decoder network is then constructed based on the spectral patching mechanism, which introduces the deep residual network (ResNet) and atrous spatial pyramid pooling (ASPP) modules to extract multiscale high-level semantic information for the H2imagery classification. The experimental results obtained with the Wuhan UAV-borne H2imagery (WHU-Hi) UAV-borne hyperspectral data set demonstrate that SPNet can achieve state-of-the-art accuracy and visualization performance in the classification of H2imagery. Yanfei Zhong, Xinyu Wang 0003, Chang Luo, Ji Zhao 0006, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Unsupervised Deep Hyperspectral Video Target Tracking and High Spectral-Spatial-Temporal Resolution (H³) Benchmark DatasetabstractTarget tracking has received increased attention in the past few decades. However, most of the target tracking algorithms are based on RGB video data, and few are based on hyperspectral video data. With the development of the new “snapshot” hyperspectral sensors, hyperspectral videos can now be easily obtained. However, hyperspectral video target tracking datasets are still rare. In this article, a high spectral-spatial-temporal resolution hyperspectral video target tracking algorithm framework (H3Net) based on deep learning is proposed. The proposed framework consists of two main parts: 1) an unsupervised deep learning-based target tracking training framework for hyperspectral video; and 2) a dual-branch network structure based on a Siamese network. Using the dual-branch network, the H3Net framework can utilize both the spatial and spectral information. The combination of deep learning and a discriminative correlation filter (DCF) makes the features extracted by deep learning more suitable for the DCF. Compared with hyperspectral images, hyperspectral video data require more manpower to annotate, so we propose an unsupervised approach to train H3Net, without any annotation. To solve the problem of the lack of hyperspectral video datasets, we built a 25-band hyperspectral video dataset (the high spectral-spatial-temporal resolution hyperspectral video dataset: the WHU-Hi-H3dataset) for target tracking. The experimental results obtained with the WHU-Hi-H3dataset confirm the potential of unsupervised deep learning in hyperspectral video target tracking. Zhenqi Liu, Yanfei Zhong, Xinyu Wang 0003, Meng Shu, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Cascaded Multi-Task Road Extraction Network for Road Surface, Centerline, and Edge ExtractionabstractRoad extraction from very high-resolution (VHR) remote sensing imagery remains a huge challenge, due to the shadows and occlusions of trees and buildings. Such complex backgrounds result in deep networks often producing fragmented roads with poor connectivity. Road extraction has three typical tasks: road surface segmentation (SS), centerline extraction (CE), and edge detection (ED), which are conducted in a wide range of real applications. Also, the three tasks have a symbiotic relationship, i.e., the road SS determines the location of the centerline and edges, and the CE and ED can allow the generation of more continuous road surfaces. However, most of the previous works have completed these three tasks separately, without exploiting the symbiotic relationship between them to boost the road connectivity. In this article, in order to improve road connectivity, a cascaded multitask (CasMT) road extraction framework for simultaneously extracting the road surface, centerline, and edges is proposed. In the proposed framework, topology-aware learning is applied to capture the long-distance topological relationships, and hard example mining (HEM) loss is employed to focus more on hard samples, to further enhance the road completeness. Extensive experiments were conducted on the DeepGlobe road dataset and a large-scale road dataset (called the LSCC dataset) from the three Chinese cities of Beijing, Shanghai, and Wuhan. The experimental results obtained on the public DeepGlobe dataset demonstrate that the proposed CasMT framework can significantly outperform the current state-of-the-art method. Moreover, the generalization capability of the model was verified on the LSCC dataset, where the proposed CasMT framework achieved the best performance in the average path length similarity (APLS) road topology metric, which further confirms the superiority of the proposed framework. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SCViT: A Spatial-Channel Feature Preserving Vision Transformer for Remote Sensing Image Scene ClassificationabstractConvolutional neural network (CNN)-based methods are widely used in remote sensing image scene classification and can obtain excellent performances. However, the stacked receptive fields in the CNN-based methods have limitations in modeling the long-range dependencies of local features. The vision transformer (ViT) model provides a good solution as it directly considers the global interactions of local patches by the self-attention mechanism. However, the vanilla ViT model, which simply splits images into fixed-size patches treated as tokens, mainly considers the global information in the spatial domain. In this article, a spatial-channel feature preserving ViT (SCViT) model is proposed, which considers both the detailed geometric information of the high-spatial-resolution (HSR) imagery and the contribution of the different channels contained in the classification token. First, in the proposed method, tokens are generated by progressively aggregating the neighboring overlapping patches to extract the local structural features of the imagery. Second, a multihead self-attention (MSA) mechanism is used to model the global interactions of the tokens in the encoder. A lightweight channel attention (LCA) module is then introduced to consider the importance of the different channels in the classification token. Finally, a multilayer perceptron (MLP) is used to acquire the final results. Compared with the state-of-the-art scene classification methods, the experimental results confirm the potential of using ViT models in remote sensing image scene classification. Pengyuan Lv, Yanfei Zhong, Fang Du, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FactSeg: Foreground Activation-Driven Small Object Semantic Segmentation in Large-Scale Remote Sensing ImageryabstractThe small object semantic segmentation task is aimed at automatically extracting key objects from high-resolution remote sensing (HRS) imagery. Compared with the large-scale coverage areas for remote sensing imagery, the key objects, such as cars and ships, in HRS imagery often contain only a few pixels. In this article, to tackle this problem, the foreground activation (FA)-driven small object semantic segmentation (FactSeg) framework is proposed from perspectives of structure and optimization. In the structure design, FA object representation is proposed to enhance the awareness of the weak features in small objects. The FA object representation framework is made up of a dual-branch decoder and collaborative probability (CP) loss. In the dual-branch decoder, the FA branch is designed to activate the small object features (activation) and suppress the large-scale background, and the semantic refinement (SR) branch is designed to further distinguish small objects (refinement). The CP loss is proposed to effectively combine the activation and refinement outputs of the decoder under the CP hypothesis. During the collaboration, the weak features of the small objects are enhanced with the activation output, and the refined output can be viewed as the refinement of the binary outputs. In the optimization stage, small object mining (SOM)-based network optimization is applied to automatically select effective samples and refine the direction of the optimization while addressing the imbalanced sample problem between the small objects and the large-scale background. The experimental results obtained with two benchmark HRS imagery segmentation datasets demonstrate that the proposed framework outperforms the state-of-the-art semantic segmentation methods and achieves a good tradeoff between accuracy and efficiency. Code will be available at:http://rsidea.whu.edu.cn/FactSeg.htm Ailong Ma, Yanfei Zhong, Zhuo Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Supervised Progressive Growing Generative Adversarial Network for Remote Sensing Image Scene ClassificationabstractRemote sensing image scene classification is a challenging task. With the development of deep learning, methods based on convolutional neural networks (CNNs) have made great achievements in remote sensing image scene classification. Since the training of a CNN requires a large number of labeled samples, a generative adversarial network (GAN) for sample generation represents a new opportunity to solve the problem of the limited samples. However, most of the existing GAN-based sample generation methods can only generate unlabeled samples, instead of samples labeled with the corresponding scene category. In this article, to solve the problem, a supervised progressive growing generative adversarial network (SPG-GAN) is proposed for remote sensing image scene classification. The proposed method can generate labeled samples for the remote sensing image scene classification, significantly improving the classification accuracy in the case of limited samples. The SPG-GAN method has two main improvements. First, a conditional generative framework for labeled samples is proposed, in which the label information is added in the channel dimension as the input. By considering the constraints of the label information in the loss function, the network can be trained in the direction of a specific category. As a result, the network can generate remote sensing image scene classification samples with label categories. Second, a progressive growing sample generation method is introduced. In order to ensure that the generated samples have more spatial details, they are generated by progressively adding modules to the generator and discriminator, thereby ensuring that the generated sample is of better quality. After testing on two benchmark datasets and carrying out a large-scale experiment in the central area of the city of Wuhan in China, it was found that the proposed method can obtain a superior scene classification accuracy in the case of limited samples. Ailong Ma, Zhuo Zheng, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Land-Use/Land-Cover Change Detection Based on Class-Prior Object-Oriented Conditional Random Field Framework for High Spatial Resolution Remote Sensing ImageryabstractHigh spatial resolution (HSR) remote sensing images can reflect more subtle changes and more specific types of land use and land cover (LULC) due to the abundant spatial geometric information. In this article, a class-prior object-oriented conditional random field (COCRF) framework consisting of a binary change detection (CD) task and a multiclass CD task is proposed to fill the application gap. In the proposed framework, the class-prior knowledge is used to improve the construction of the unary potential in both the binary and multiclass CD tasks, to reduce the influence of spectral variability. The binary CD result provides a constraint to the multiclass CD result. As a result, both parts have effective interaction. The class posterior probability images of two dates can be obtained automatically with the class-prior knowledge by sample migration. Furthermore, an object constraint described by the class dispersion within the objects is added to improve the smoothness in local objects, while the pairwise potential improves the smoothness of the whole area by using the eight-neighborhood spectral information of the center pixel. By integrating the above approaches, the problems of error accumulation and the manual intervention required in the traditional multiclass CD methods can be relieved. An adaptive parameter estimation strategy is also adopted in the proposed framework, to save the time required for manual parameter setting. The proposed COCRF framework was validated on two HSR remote sensing image data sets, where it achieved a better performance than the other state-of-the-art CD methods. Sunan Shi, Yanfei Zhong, Ji Zhao 0006, Pengyuan Lv, Yinhe Liu, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Joint Spectral Unmixing and Subpixel Mapping Framework Based on Multiobjective OptimizationabstractConventional subpixel mapping (SPM) is performed based on the abundance maps obtained by spectral unmixing (SU), to interpret the mixed pixels and improve the mapping resolution for hyperspectral remote-sensing imagery. However, the SU and SPM tasks are separately conducted, so that the unmixing error is propagated to SPM, and the mapping result is strongly reliant on the quality of the abundance maps. In this article, a novel joint SPM and SU framework (MO_SUSM) based on multiobjective optimization is proposed to simultaneously perform unmixing and mapping. Specifically, the multiobjective joint optimization model with a data fidelity term and a Laplacian prior term is constructed for SU and SPM. For the data fidelity term, since the unmixing result can be recovered by downsampling the mapping result, the unmixing model is joined with the mapping model by the downsampling matrix, so that the reconstruction errors of the unmixing and mapping results can be minimized together. Meanwhile, the Laplacian prior term is used to maximize the spatial dependence of the mapping result and provide the spatial constraint for SU. In addition, the multiobjective optimization algorithm with local search is designed to search for the optimal unmixing and mapping results that can balance the objective terms. Since the two objective terms are dynamically integrated during optimization, there is no need to set sensitive weights for the objectives combination. Four experiments were conducted on hyperspectral images of various data sources, including ground, airborne, and satellite images. The unmixing results show that MO_SUSM can reduce the unmixing error and can improve the quality of the abundance maps. The mapping results show that MO_SUSM can alleviate the dependence of SPM on the abundance maps and can improve the mapping accuracy. Mi Song, Yanfei Zhong, Ailong Ma, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Three-Dimensional Change Detection in Urban Areas Based on Complementary Evidence FusionabstractWith the acceleration of urbanization, it is essential to carry out change detection (CD) and obtain surface change information in urban areas. In the early stages, the spectral information of remote sensing images was used as a change index to capture the spectral and texture changes of ground objects in a two-dimensional plane. However, due to the dense buildings in urban areas, shadows and occlusions can be easily formed, and spectral information is sensitive to the imaging environment, such as the illumination, atmospheric conditions, and imaging angles, so that the detection results based on remote sensing images can often be incomplete. Most changes include not only 2-D plane changes but also 3-D elevation changes. Compared with spectral information, the elevation is more stable and more resistant to interference. Therefore, the fusion of remote sensing image and digital surface model (DSM) data has the potential to be used to detect the changes in urban areas. In this article, we propose a complementary evidence fusion 3-D CD framework based on the Dempster–Shafer theory (CDST). In this framework, DSM and normalized difference vegetation index (NDVI) data are combined using a complementary evidence combination rule. The DSM data can effectively overcome the impact of shadows, and the NDVI data can capture the relevant changes of height-insensitive ground objects, such as vegetation and water. When mapping the basic probability assignment (BPA) of the difference image (DI), prior knowledge is used to ensure that the BPA is not affected by the data distribution. Since DSM and remote sensing image data are heterogeneous data, there is a high degree of conflict when representing the change information characteristics of specific areas. For example, the change between grassland and road is small in elevation but significant in the spectral details, and the traditional Dempster’s combination rule no longer applies. The proposed CDST framework uses a complementary evidence combination rule, which can effectively alleviate the conflicts between the evidence sources and improve the integrity of the detected changes. The experimental results obtained on real datasets confirm that the proposed method does indeed perform well. Shiqi Tian, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Self-Supervised Denoising Network for Satellite-Airborne-Ground Hyperspectral ImageryabstractHyperspectral images (HSIs) are inevitably corrupted with various types of noise, which seriously degrades the data quality and usability. Denoising is an essential preprocessing task of HSI processing. Recently, benefiting from the great learning ability of deep learning, convolutional neural network (CNN) denoisers have obtained state-of-the-art performances for Gaussian noise removal. However, one central problem remains largely unsolved: how to deal with the complicated noise in the real-world HSIs, especially when a paired training data set is unavailable. In this article, a self-supervised hyperspectral image denoising network (SHDN) is proposed, which consists of a noise estimator and a CNN denoiser. Rather than defining a complex noise model to generate training pairs on the clean HSIs, a self-supervised training scheme is first proposed by considering the noisy HSI itself as the training data. Through the noise estimator, the realistic noise samples can be extracted and combined with the clean bands to make up the training pairs. In addition, to jointly restore the target noisy band and to maintain the spectral consistency, a flexible multi-to-single band convolutional network is designed, where the noisy band and the neighboring bands are jointly aggregated via multiscale contextualized dilated blocks and the spectral–spatial convolutional unit. Experiments on HSIs from spaceborne, airborne, unmanned aerial vehicle (UAV)-borne, and ground-based data sets demonstrate the applicability and the generalization of SHDN in the real scenarios. Additionally, the usability of the noisy bands and the suitability of the SHDN framework in the subsequent applications are verified in the land-cover mapping experiments. Xinyu Wang 0003, Zhaozhi Luo, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Auto-AD: Autonomous Hyperspectral Anomaly Detection Network Based on Fully Convolutional AutoencoderabstractHyperspectral anomaly detection is aimed at detecting observations that differ from their surroundings, and is an active area of research in hyperspectral image processing. Recently, autoencoders (AEs) have been applied in hyperspectral anomaly detection; however, the existing AE-based methods are complicated and involve manual parameter setting and preprocessing and/or postprocessing procedures. In this article, an autonomous hyperspectral anomaly detection network (Auto-AD) is proposed, in which the background is reconstructed by the network and the anomalies appear as reconstruction errors. Specifically, through a fully convolutional AE with skip connections, the background can be reconstructed while the anomalies are difficult to reconstruct, since the anomalies are relatively small compared to the background and have a low probability of occurring in the image. To further suppress the anomaly reconstruction, an adaptive-weighted loss function is designed, where the weights of potential anomalous pixels with large reconstruction errors are reduced during training. As a result, the anomalies have a higher contrast with the background in the map of reconstruction errors. The experimental results obtained on a public airborne data set and two unmanned aerial vehicle-borne hyperspectral data sets confirm the effectiveness of the proposed Auto-AD method. Shaoyu Wang 0003, Xinyu Wang 0003, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Deep Low-Rank Prior for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection is aimed at detecting observations that differ from their surroundings. To achieve this goal, low-rank models and autoencoders (AEs) have attracted a lot of attention. Although the low-rank model is self-explainable, a low-rank prior may not completely match real data. In contrast, AEs can automatically learn the discriminative features between anomalies and background, whereas AEs are not self-explainable. In this article, a deep low-rank prior-based method (DeepLR) is proposed, which combines a model-driven low-rank prior and a data-driven AE. To be specific, the low-rank prior and a fully convolutional AE architecture are incorporated through modeling an energy minimization problem solved by an iterative optimization framework, in which low-rank background estimation and network training serve as two subproblems. The low-rank background is input into the network to calculate a low-rank regularized loss, constraining the training of the network. Finally, the background can be approximately reconstructed, while the anomalies are reconstructed with significant reconstruction errors; thus, the reconstruction errors indicate the anomalous degree. The experimental results obtained on several public datasets and two large unmanned aerial vehicle (UAV)-borne datasets confirm the merit and viability of the proposed method. Shaoyu Wang 0003, Xinyu Wang 0003, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Domain Adaptation via a Task-Specific Classifier Framework for Remote Sensing Cross-Scene ClassificationabstractThe scene classification of high spatial resolution (HSR) imagery involves labeling an HSR image with a specific high-level semantic class according to the composition of the semantic objects and their spatial relationships. As such, scene classification has attracted increased attention in recent years, and many different algorithms have now been proposed for the cross-scene classification task. However, the recently proposed scene classification methods based on deep convolutional neural networks (CNNs) still suffer from domain shift problems, because of the training data and validation data not following the assumption of independent and identical distributions. The employment of generative adversarial networks has been found to be an effective way to bridge the domain shift/gap. However, the existing cross-scene classification methods do not use the classification information in the target domain, and the domain classifier is task-independent for different scene classification tasks. In this article, to solve this problem, domain adaptation via a task-specific classifier (DATSNET) framework is proposed for HSR image scene classification. Task-specific classifiers and minimizing and maximizing, “ i.e., minimaxing,” of the classifier discrepancy are integrated in the DATSNET framework. The task-specific classifiers are proposed to align the distributions of the source domain features and target domain features by utilizing task-specific decision boundaries in the target domain. In order to align the two task-specific classifiers’ feature distributions, minimaxing the defined discrepancy between the different classifiers in an adversarial manner is proposed to obtain better task-specific classifier boundaries in the target domain and a better-aligned feature distribution in both domains. The experimental results obtained with different remote sensing cross-scene classification tasks demonstrate that the proposed method achieves a significantly improved performance compared with the other state-of-the-art remote sensing cross-scene classification algorithms. Zhendong Zheng, Yanfei Zhong, Ailong Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spatio-Temporal Dual-Branch Network With Predictive Feature Learning for Satellite Video Object SegmentationabstractSatellite video is an important new earth observation data source that can be used to acquire large-scale dynamic information. Satellite video object segmentation (SVOS) is aimed at separating the foreground and background of satellite video and is a fundamental processing task for satellite video. To date, the state-of-the-art research into SVOS has mainly focused on unsupervised target extraction methods through hand-crafted features and post-processing operations, which are prone to obtaining incomplete contours of the targets and result in foreground aperture problem. Furthermore, the small size of targets and appearance deformation make the SVOS more difficult. Therefore, in this article, a spatio-temporal dual-branch network is proposed with predictive feature learning for the SVOS task. The proposed model consists of a temporal coherence branch and a spatial segmentation branch. In the temporal coherence branch, the Wasserstein generative adversarial network (WGAN) architecture is utilized for the future frame prediction to exploit temporal information, which captures the dynamic appearance and motion cues from the unlabeled satellite video data through predictive feature learning module in an adversarial manner. As a result, the proposed method can obtain segmentation results with temporal consistency, while avoiding the generation of optical flow images. In the spatial segmentation branch, a fully convolutional network (FCN) is used to extract the high-level spatial information of the satellite video and achieve end-to-end SVOS, without any post-processing operations. In the network implementation, boundary loss is used to solve the highly unbalanced segmentation problem caused by small size of the targets. The two branches of the network are also mutually constrained, to improve the final object segmentation results. The visual and quantitative results of three experiments all demonstrate that the proposed method outperforms the other current SVOS models. Yanfei Zhong, Meng Shu, Zhenqi Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | S3TRM: Spectral-Spatial Unmixing of Hyperspectral Imagery Based on Sparse Topic Relaxation-Clustering ModelabstractHyperspectral unmixing (HU) has been one of the hot spots in hyperspectral remote sensing research and has great potential in many applications. In recent years, the employment of the probabilistic topic model to mine latent topics in hyperspectral images has been an effective way for spectral unmixing. However, these methods fail to fully exploit the potential of topic models in uncovering image semantics and need extra sparsity constraints, which greatly increases the complexity of the model. In addition, the spatial information which can provide the correlation of features in adjacent pixels is usually ignored in topic model-based HU. To solve these problems, a spectral-spatial unmixing framework of hyperspectral imagery based on a sparse topic relaxation-clustering model ($\mathrm{S}^{3}{\mathrm {TRM}}$) is proposed. In$\mathrm{S}^{3}{\mathrm {TRM}}$, the topic model combined with implied sparse prior constraints are introduced, and the sparse characteristics of$\mathrm{S}^{3}{\mathrm {TRM}}$are used to capture the semantic representation of the spectrum. With the proposed relaxation-clustering strategy, multiple possible spectral representations of features can be obtained, which further alleviates the influence caused by endmember variability. Group clustering is used to locate the endmember quickly and accurately. Moreover, superpixel segmentation is considered to supplement the spatial distribution information of features, thereby improving the fractional abundance. Experiments on a synthetic dataset and three well-known real hyperspectral datasets confirm the excellent performance of the proposed framework in both qualitative assessment and quantitative evaluation, compared with the other state-of-the-art methods. Qiqi Zhu, Jiale Chen 0002, Wen Zeng 0003, Yanfei Zhong, Qingfeng Guan 0001, Zhijiang Yang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Oil Spill Contextual and Boundary-Supervised Detection Network Based on Marine SAR ImagesabstractOil spills have caused serious harm to the marine environment. Remote sensing technology is one of the important tools for marine environment monitoring. Synthetic aperture radar (SAR) has become an important technology for detecting marine pollution. Identifying dark spots is essential for oil spill detection based on SAR images. Dark spots’ detection can be achieved using image segmentation techniques. However, natural phenomena, such as waves and currents, can also cause dark spots, resulting in consistently uneven intensity, high noise, and blurred boundaries in oil spill images. In addition, existing oil spill detection models often perform well for large targets but have poor detection accuracy for small targets. To solve the above problems, the oil spill contextual and boundary-supervised detection network (CBD-Net) is proposed to extract refined oil spill regions by fusing multiscale features. To improve the internal consistency of oil spill regions, the spatial and channel squeeze excitation (scSE) block is introduced. In CBD-Net, boundary details are enhanced with optimized edge supervision. In addition, a manually labeled dataset is proposed, Deep-SAR Oil Spill (SOS) dataset, aiming to solve the problem of insufficient existing oil spill detection dataset. Experimental results demonstrate that CBD-Net outperforms other comparative models and is able to extract robust and accurate oil spill regions from complex SAR images. The highest mIoU of 83.42% and the highest F1 score of 87.87% were achieved on the SOS dataset. The CBD-Net model proposed in this article can play a guiding role in the marine oil spill decision support system. Qiqi Zhu, Xiaorui Yan, Qingfeng Guan 0001, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | SiamHYPER: Learning a Hyperspectral Object Tracker From an RGB-Based TrackerabstractHyperspectral videos can provide the spatial, spectral, and motion information of targets, which makes it possible to track camouflaged targets that are similar to the background. However, hyperspectral object tracking is a challenging task, due to the huge hyperspectral video data dimension and the "data hungry" problem for the model training. Insufficient training data can seriously interfere with the accuracy and generalization of the tracking models. In this paper, a dual deep Siamese network framework for hyperspectral object tracking (SiamHYPER) is proposed for learning a hyperspectral tracker from a pretrained RGB tracker in the case of the "data hungry" problem. Specifically, in addition to a pretrained RGB-based Siamese tracker, a hyperspectral target-aware module is designed to mine the spectral information during the target prediction, and a spatial-spectral cross-attention module is introduced to further fuse the deep spatial and spectral features extracted from the RGB tracker and the hyperspectral target-aware module. Benefiting from the guidance training of the RGB tracker, a robust hyperspectral object tracker can be trained effectively with only a small number of hyperspectral video samples, to overcome the "data hungry" problem. In the experiments conducted in this study, the SiamHYPER framework was verified using SiamBAN and SiamRPN++, with 13 000 frames of hyperspectral videos for training, and achieved the best performance on the publicly available hyperspectral dataset released as part of the WHISPERS Hyperspectral Object Tracking Challenge. The area under the curve (AUC) of SiamHYPER was increased by nearly 8.9% and 7.2%, respectively, when compared with the current state-of-the-art RGB-based and hyperspectral trackers. In addition, the processing speed of SiamHYPER was 19 FPS, which is much higher than that of the current state-of-the-art hyperspectral trackers. The source code is available at zhenliuzhenqi/HOT: Hyperspectral object tracking (github.com). Zhenqi Liu, Xinyu Wang 0003, Yanfei Zhong, Meng Shu |
IEEE Trans. Image Process. | 3 |
| 2022 | Open-Source Data-Driven Cross-Domain Road Detection From Very High Resolution Remote Sensing ImageryabstractHigh-precision road detection from very high resolution (VHR) remote sensing images has broad application value. However, the most advanced deep learning based methods often fail to identify roads when there is a distribution discrepancy between the training samples and test samples, due to their limited generalization ability. In this paper, to address this problem, an open-source data-driven domain-specific representation (OSM-DOER) framework is proposed for cross-domain road detection. On the one hand, as the spatial structure information of the source and target domains is similar, but the texture information is different, the domain-specific representation (DOER) framework is proposed, which not only aligns the distributions of the spatial structure information, but also learns the domain-specific texture information. Furthermore, in order to enhance the representation of the target domain data distribution, open-source and freely available OpenStreetMap (OSM) road centerline data are utilized to generate target domain samples, which are then used in the network training as the supervised information for the target domain. Finally, to verify the superiority of the proposed OSM-DOER framework, we conducted extensive experiments with the public SpaceNet and DeepGlobe road datasets, and large-scale road datasets from Birmingham in the UK and Shanghai in China. The experimental results demonstrate that the proposed OSM-DOER framework shows obvious advantages over the mainstream road detection methods, and the use of OSM road centerline data has great potential for the road detection task. Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing ImageryabstractFor high spatial resolution (HSR) remote sensing images, bitemporal supervised learning always dominates change detection using many pairwise labeled bitemporal images. However, it is very expensive and time-consuming to pairwise label large-scale bitemporal HSR remote sensing images. In this paper, we propose single-temporal supervised learning (STAR) for change detection from a new perspective of exploiting object changes in unpaired images as supervisory signals. STAR enables us to train a high-accuracy change detector only using unpaired labeled images and generalize to real-world bitemporal images. To evaluate the effectiveness of STAR, we design a simple yet effective change detector called ChangeStar, which can reuse any deep semantic segmentation architecture by the ChangeMixin module. The comprehensive experimental results show that ChangeStar outperforms the baseline with a large margin under single-temporal super-vision and achieves superior performance under bitemporal supervision. Code is available at https://github.com/Z-Zheng/ChangeStar. Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 4 |
| 2021 | Weakly Supervised Convolutional Neural Networks for Hyperspectral UnmixingabstractHyperspectral unmixing is an essential task in hyperspectral imagery applications. Because of the strong feature extract ability and satisfying performance, deep learning methods have been used for hyperspectral unmixing. However, there are still several problems in existing deep learning based spectral unmixing methods. Supervised learning methods can only accomplish a single task and lack a large amount of data for supervised learning. While the unsupervised learning unmixing methods are easily misled by the traditional way of initialization. In this paper, a weakly supervised deep convolutional neural network is proposed for hyperspectral unmixing. The experimental results show that competitive results can also be obtained by pretraining with a small number of samples, and weakly supervised learning still has potential for hyperspectral unmixing. Jiayu Bai, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2021 | Large-Scale Urban Road Vectorization Mapping Via A Road Node Proposal Network for High-Resolution Remote Sensing ImageryabstractUrban road vectorization mapping can reflect the urban development of cities, which consists of two separate tasks: road extraction and road vectorization. Most of the current road vectorization mapping methods focus on road extraction yet ignoring the importance of road vectorization, facing the problem of road connectivity. In this work, to implement urban road vectorization mapping in a unified way, a novel road vectorization mapping framework is proposed. The proposed framework consists of a node proposal network (NPN) module and a node connectivity based road refinement module. In the NPN module, a node proposal head is adopted, which improves the connectivity of the road mask by providing supervision of the road nodes, which are actually part of the road mask. The road mask is then converted into a road vector map by vectorization. In the node connectivity based road refinement module, road nodes are inserted into the road vector map to improve the connectivity. The experimental results of two public datasets (SpaceNet 3 and DeepGlobe) confirm the advantages of the proposed framework. The experiments for two areas in Shanghai and Wuhan demonstrate the generalization ability of the proposed framework. Yanfei Zhong, Ailong Ma |
IGARSS | 2 |
| 2021 | Rethinking the High Frequency Components in Deep Sub-Pixel Mapping NetworkabstractDeep sub-pixel mapping network (DSMNet) is a state-of-the-art approach in the field of sub-pixel mapping (SPM, also called super resolution mapping), combining deep learning theory, to solve the mixed pixel problem, which is ubiquitous in remote sensing images due to the spatial-resolving limitation. However, traditional DSMNet usually do not consider the multi-scale distribution characteristics of the real geographical distribution exposed in urban landscape. Furthermore, the heterogeneous distribution characteristics (high-frequency components) are the most important for SPM, but are difficult to learn and usually ignored in the tradition network models. In this paper, the high-frequency component aware (HFCA) module was proposed, based on the hierarchical supervised deep sub-pixel mapping network (HiDSMNet). HiDSMNet establishes a hierarchical supervised architecture for explicit multi-scale supervision to prompt the network to learn a multi-scale representation. Besides, HFCA module is integrated to prompt the network to intensify the learning of the high-frequency representation. The experimental results with three public datasets validated the superiority of the proposed HiDSMNet. Da He, Yanfei Zhong, Qian Shi 0001, Xiaoping Liu 0001 |
IGARSS | 2 |
| 2021 | Deep One-Class Crop Extraction Framework for Multi-Modal Remote Sensing ImageryabstractLarge scale crop mapping is an important task in agricultural resource monitoring. To obtain a distribution map of crops, traditional methods usually require the well-designed manufacture features and the ground-truth labels of all land-cover types for training a multi-class classifier. However, the redundant labeling for each land-cover type is time-consuming and labor-intensive, and the feature design for different remote sensing data is complex and limited to human prior knowledge. In this paper, a deep one-class crop extraction framework is proposed to solve the problems mentioned above. Specifically, it uses the deep one-class crop extraction module to extract the feature automatically for any remote sensing imagery and the one-class crop extraction loss to address the lack of negative class in the deep one-class classification model. In addition, the proposed framework can be applied to multi-modal remote sensing data, i.e. hyperspectral, multispectral, and SAR images, which is verified in the experiments and the proposed framework can achieve the highest accuracy on each multi-modal data. Xinyu Wang 0003, Hengwei Zhao, Chang Luo, Yanfei Zhong |
IGARSS | 6 |
| 2021 | Low-Rank Representation Incorporating Local Spatial Constraint for Hyperspectral Anomaly DetectionabstractRecently, hyperspectral anomaly detection methods based on low-rank representation(LRR) have been widely studied. However, the assumption of global low dimension of background may ignore the local structure information of hyperspectral image. In this paper, a novel LRR incorporating local spatial constraint method is proposed for hyperspectral anomaly detection. Different from LRR detector, the proposed method considers the spatial information based on the supe pixel in the background part. The proposed method and current state-of-the-art methods are tested on two sets of real data. The experimental results demonstrate that the proposed method is superior to the comparative method in terms of both colour map detection and quantitative evaluation. Hao Li 0058, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Lifei Wei |
IGARSS | 4 |
| 2021 | A Nov Al Global-Local Adversarial Network for Unsupervised Cross-Domain Road DetectionabstractRoad detection based on convolutional neural networks (CNNs) has achieved remarkable performances for very high resolution (VHR) remote sensing images. However, it relies on a large number of labeled samples and the problem of limited generalization for unseen images still remains. The manual pixel-level labeling process is extremely time-consuming, and the performance of CNN s degrades significantly when there is a domain gap between the training and test images. To address this problem, a global-local adversarial network (GLANet) is proposed for unsupervised cross-domain road detection. On the one hand, considering the spatial information similarities between the source and target domains, feature space driven adversarial learning is applied to explore the shared features across domains. On the other hand, the complex background of VHR remote sensing images, such as the occlusions and shadows of trees and buildings, makes some roads easy to recognize, while others are not. However, the traditional global adversarial learning approach cannot guarantee local semantic consistency. Therefore, a local alignment operation, which adaptively adjusts the weight of the adversarial loss according to the road recognition difficulty, is introduced. The experiments conducted on public road datasets show that the proposed method can effectively improve the cross-domain road detection performance, which demonstrates its strong generalization ability. Yanfei Zhong |
IGARSS | 2 |
| 2021 | Sensor-Specific Adversarial Network for Transferable Land-Cover ClassificationabstractAs the multi-source high-spatial-resolution (HSR) images are being daily acquired from different sensors, it brings the challenge of transferring the recognition model from labeled images to new unlabelled images obtained from other sensors. Existing deep transfer learning methods encode the land-cover features in the same architecture, which ignores the sensor divergence. In this paper, we tackle this problem by proposing a sensor-specific adversarial network for HSR land-cover classification. Specifically, the sensor-specific normalization (SN) is designed for decoupling the sensor divergence in different normalization weights. Moreover, the transferable adversarial optimization is proposed for effectively optimizing the source-related, target-related, and discriminator weights. Considering the sensor-specific characteristics, our proposed method improves the transferability of deep learning models between airborne and spaceborne sensors. The mutual transferability experiments on a self-constructed cross-sensor land-cover dataset demonstrate that the proposed method outperforms the state-of-the-art deep transfer learning methods. Yanfei Zhong, Zhuo Zheng, Ailong Ma |
IGARSS | 2 |
| 2021 | Weakly Supervised Semantic Change Detection via Label Refinement FrameworkabstractSemantic change detection is a meaningful but challenging task in the remote sensing community. The currently dominant approaches are mainly based on deep learning. However, the lack of high-resolution annotations is the main bottleneck for semantic change detection at scale when using these state-of-the-art deep learning models. In this paper, the label refinement framework is proposed for weakly-supervised semantic change detection, which allows the deep network to learn from low-resolution labels and produce high-resolution semantic change maps, thus alleviating the data-hungry problem. This framework contains four parts: coarse label training’ pseudo-label refinement, multitask change detection and post-process. The experimental results on 2021 IEEE GRSS Data Fusion Contest Track MSD dataset confirmed the effectiveness of the proposed method. Additionally, our method wins 4th place in the 2021 IEEE GRSS Data Fusion Contest Track MSD (DFC21-MSD). Zhuo Zheng, Yinhe Liu, Shiqi Tian, Ailong Ma, Yanfei Zhong |
IGARSS | 6 |
| 2021 | Deep Subpixel Mapping Based on Semantic Information Modulated Network for Urban Land Use MappingabstractMixed pixel problem is omnipresent in remote sensing images for urban land use interpretation due to the hardware limitations. Subpixel mapping (SPM) is a usual way to solve this problem by improving the observation scale and realizing a finer spatial resolution land cover mapping. Recently, deep learning-based subpixel mapping network (DLSMNet) was proposed, benefited from its strong representation and learning ability, to restore a visually pleasing finer mapping. However, the spatial context features of artifacts are usually aggregated and progressively lost during the forward pass of the network without sufficient representation, which make it difficult to be learned and restored. In this article, a semantic information modulated (SIM) deep subpixel mapping network (SIMNet) is proposed, which uses low-resolution semantic images as prior, to reinforce the representation of spatial context features. In SIMNet, SIM module is proposed to parametrically incorporate the semantic prior into the state-of-the-art (SOTA) feed forward network architecture in an end-to-end training fashion. Furthermore, stacked SIM module with residual blocks (SIM_ResBlock) is adopted to pass the representation of spatial context feature to the deep layers, to get it fully learned during backpropagation. Experiments have been implemented on three public urban scenario data sets, and the SIMNet generates a clearer outline of artificial facilities with sufficient spatial context, and is distinctive for even individual building, which is challenging for other SOTA DLSMNet. The results demonstrate that the proposed SIMNet is a promising way for high-resolution urban land use mapping from easily available lower resolution remote sensing images.Mixed pixel problem is omnipresent in remote sensing images for urban land use interpretation due to the hardware limitations. Subpixel mapping (SPM) is a usual way to solve this problem by improving the observation scale and realizing a finer spatial resolution land cover mapping. Recently, deep learning-based subpixel mapping network (DLSMNet) was proposed, benefited from its strong representation and learning ability, to restore a visually pleasing finer mapping. However, the spatial context features of artifacts are usually aggregated and progressively lost during the forward pass of the network without sufficient representation, which make it difficult to be learned and restored. In this article, a semantic information modulated (SIM) deep subpixel mapping network (SIMNet) is proposed, which uses low-resolution semantic images as prior, to reinforce the representation of spatial context features. In SIMNet, SIM module is proposed to parametrically incorporate the semantic prior into the state-of-the-art (SOTA) feed forward network architecture in an end-to-end training fashion. Furthermore, stacked SIM module with residual blocks (SIM_ResBlock) is adopted to pass the representation of spatial context feature to the deep layers, to get it fully learned during backpropagation. Experiments have been implemented on three public urban scenario data sets, and the SIMNet generates a clearer outline of artificial facilities with sufficient spatial context, and is distinctive for even individual building, which is challenging for other SOTA DLSMNet. The results demonstrate that the proposed SIMNet is a promising way for high-resolution urban land use mapping from easily available lower resolution remote sensing images. Da He, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Xinchang Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Deep Convolutional Neural Network Framework for Subpixel MappingabstractSubpixel mapping (SPM) is an effective way to solve the mixed pixel problem, which is a ubiquitous phenomenon in remotely sensed imagery, by characterizing subpixel distribution within the mixed pixels. In fact, the majority of the classical and state-of-the-art SPM algorithms can be viewed as a convolution process, but these methods rely heavily on fixed and handcrafted kernels that are insufficient in characterizing a geographically realistic distribution image. In addition, the traditional SPM approach is based on the prerequisite of abundance images derived from spectral unmixing (SU), during which process uncertainty inherently exists and is propagated to the SPM. In this article, a kernel-learnable convolutional neural network (CNN) framework for subpixel mapping (SPMCNN-F) is proposed. In SPMCNN-F, the kernel is learnable during the training stage based on the given training sample pairs of low- and high-resolution patches for learning a geographically realistic prior, instead of fixed priors. The end-to-end mapping structure enables direct subpixel information extraction from the original coarse image, avoiding the uncertainty propagation from the SU. In the experiments undertaken in this study, two state-of-the-art super-resolution networks were selected as application demonstrations of the proposed SPMCNN-F method. In experiment part, three hyperspectral image data sets were adopted, two in a synthetic coarse image approach and one in a real coarse image approach, for the validation. Additionally, a new data set with pairs of Moderate-resolution Imaging Spectroradiometer (MODIS) and Landsat images were adopted in a real coarse image approach, for further validation of SPMCNN-F in large-scale area. The restored fine distribution images obtained in all the experiments showed a perceptually better reconstruction quality, both qualitatively and quantitatively, confirming the superiority of the proposed SPM framework. Da He, Yanfei Zhong, Xinyu Wang 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Superpixel-Based Reweighted Low-Rank and Total Variation Sparse Unmixing for Hyperspectral Remote Sensing ImageryabstractSparse unmixing, as a semisupervised unmixing method, has attracted extensive attention. The process of sparse unmixing involves treating the mixed pixels of hyperspectral imagery as a linear combination of a small number of spectral signatures (endmembers) in a standard spectral library, associated with fractional abundances. Over the past ten years, to achieve a better performance, sparse unmixing algorithms have begun to focus on the spatial information of hyperspectral images. However, less accurate spatial information greatly limits the performance of the spatial-regularization-based sparse unmixing algorithms. In this article, to overcome this limitation and obtain more reliable spatial information, a novel sparse unmixing algorithm named superpixel-based reweighted low-rank and total variation (SUSRLR-TV) is proposed to enhance the performance of the traditional spatial-regularization-based sparse unmixing approaches. In the proposed approach, superpixel segmentation is adopted to consider both the spatial proximity and the spectral similarity. In addition, a low-rank constraint is enforced on the objective function as pixels within each superpixel have the same endmembers and similar abundance values, and they naturally satisfy the low-rank constraint. Differing from the traditional nuclear norm, a reweighted nuclear norm is used to achieve a more efficient and accurate low-rank constraint. Meanwhile, low-rank consideration is also used to enhance the spatial continuity and suppress the effects of random noise. Furthermore, TV regularization is introduced to promote the smoothness of the abundance maps. Experiments on three simulated data sets, as well as a well-known real hyperspectral imagery data set, confirm the superior performance of the proposed method in both the qualitative assessment and the quantitative evaluation, compared with the state-of-the-art sparse unmixing methods. Hao Li 0058, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Multiscale U-Shaped CNN Building Instance Extraction Framework With Edge Constraint for High-Spatial-Resolution Remote Sensing ImageryabstractBuilding extraction based on high-resolution remote sensing imagery has been widely used in automatic surveying and mapping. However, few methods have been developed for building instance extraction, i.e., extracting each building's footprint separately, which is required in a number of applications, such as the smallest unit of a cadastral database. In building instance extraction, there are two challenges: 1) buildings with various scales exist in the imagery and 2) precise building footprints are difficult to extract due to the blurry boundaries. In this article, to solve these problems, a multiscale U-shaped convolutional neural network building instance extraction framework with edge constraint (EMU-CNN) for high-spatial-resolution remote sensing imagery is proposed. The proposed framework consists of three components: 1) a multiscale fusion U-shaped network (MFUN); 2) a region proposal network (RPN); and 3) an edge-constrained multitask network (ECMN). First, in the proposed method, the MFUN includes three parallel branches to learn multiple building features with different scales. The RPN then detects the positions of the building instances, even for buildings that are connected with each other. Moreover, according to the instance positions, the ECMN is proposed to extract a precise mask and suppress overfitting. The experiments conducted on a self-annotated data set and two public data sets (the ISPRS Vaihingen semantic labeling contest data set and the WHU aerial image data set) show that the EMU-CNN method can achieve excellent performance and shows great robustness at different scales. Yuanyuan Liu 0004, Ailong Ma, Yanfei Zhong, Fang Fang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Autonomous Endmember Detection via an Abundance Anomaly Guided Saliency Prior for Hyperspectral ImageryabstractDetermining the optimal number of endmember sources, which is also called “virtual dimensionality” (VD), is a priority for hyperspectral unmixing (HU). Although the VD estimation directly affects the HU results, it is usually solved independently of the HU process. In this article, a saliency-based autonomous endmember detection (SAED) algorithm is proposed to jointly estimate the VD in the process of endmember extraction (EE). In SAED, we first demonstrate that the abundance anomaly (AA) value is an important feature of undetected endmembers since pure pixels have larger AA values than “distractors” (i.e., mixed pixels and pure pixels of detected endmembers). Then, motivated by the fact that endmembers usually gather in certain local regions (superpixels) in the scene, due to spatial correlation, a superpixel prior is introduced in SAED to distinguish endmembers from noise. Specifically, the undetected endmembers are defined as visual stimuli in the AA subspace, the EE is formulated as a salient region detection problem, and the VD is automatically determined when there are no salient objects in the AA subspace. Since the spatial-contextual information of the endmembers is exploited during the saliency analysis, the proposed method is more robust than the spectral-only methods, which was verified using both real and synthetic hyperspectral images. Xinyu Wang 0003, Yanfei Zhong, Chunyang Cui, Liangpei Zhang 0001, Yanyan Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | RSNet: The Search for Remote Sensing Deep Neural Networks in Recognition TasksabstractDeep learning algorithms, especially convolutional neural networks (CNNs), have recently emerged as a dominant paradigm for high spatial resolution remote sensing (HRS) image recognition. A large amount of CNNs have already been successfully applied to various HRS recognition tasks, such as land-cover classification and scene classification. However, they are often modifications of the existing CNNs derived from natural image processing, in which the network architecture is inherited without consideration of the complexity and specificity of HRS images. In this article, the remote sensing deep neural network (RSNet) framework is proposed using an automatically search strategy to find the appropriate network architecture for HRS image recognition tasks. In RSNet, the hierarchical search space is first designed to include module- and transition-level spaces. The module-level space defines the basic structure block, where a series of lightweight operations as candidates, including depthwise separable convolutions, is proposed to ensure the efficiency. The transition-level space controls the spatial resolution transformations of the features. In the hierarchical search space, a gradient-based search strategy is used to find the appropriate architecture. In RSNet, the task-driven architecture training process can acquire the optimal model parameters of the switchable recognition module for HRS image recognition tasks. The experimental results obtained using four benchmark data sets for land-cover classification and scene classification tasks demonstrate that the searched RSNet can achieve a satisfactory accuracy with a high computational efficiency and, hence, provides an effective option for the processing of HRS imagery. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, as a particular semantic segmentation task, always faces with larger-scale variation, larger intra-class variance of background, and foreground-background imbalance in the high spatial resolution (HSR) remote sensing imagery. However, general semantic segmentation methods mainly focus on scale variation in the natural scene, with inadequate consideration of the other two problems that usually happen in the large area earth observation scene. In this paper, we argue that the problems lie on the lack of foreground modeling and propose a foreground-aware relation network (FarSeg) from the perspectives of relation-based and optimization-based foreground modeling, to alleviate the above two problems. From perspective of relation, FarSeg enhances the discrimination of foreground features via foreground-correlated contexts associated by learning foreground-scene relation. Meanwhile, from perspective of optimization, a foreground-aware optimization is proposed to focus on foreground examples and hard examples of background during training for a balanced optimization. The experimental results obtained using a large scale dataset suggest that the proposed method is superior to the state-of-the-art general semantic segmentation methods and achieves a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma |
CVPR | 2 |
| 2020 | Semi-Supervised Hyperspectral Unmixing with Very Deep Convolutional Neural NetworksabstractHyperspectral unmixing is an essential task in hyperspectral imagery applications. Deep learning methods have been taken into hyperspectral unmixing because of its great feature extraction ability and better performance. However, there are several problems in existing deep learning based spectral unmixing methods. The networks are not deep enough to exploit their feature extraction capabilities in these unsupervised autoencoders based methods, and their effects are not stable. The main reason may be the limited prior information limited the ability of conducting the supervised method. In this manuscript, a semi-supervised deep learning based unmixing method is proposed. Unlike the existing methods, our model uses deeper neural networks without pooling layers, and the endmember spectrum are selected supervised from the original data, which uses nature and nurture cooperatively. The experimental results show that the proposed method achieves better performance and produces more accurate abundance maps, as well as higher quantitative results, compared with the current state-of-the-art deep learning unmixing algorithms. Jiayu Bai, Ruyi Feng, Lizhe Wang 0001, Hao Li 0058, Fengpeng Li, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 6 |
| 2020 | Dense Greenhouse Extraction in High Spatial Resolution Remote Sensing ImageryabstractGreenhouses are densely distributed across the cultivated land in high-resolution remote sensing imagery, resulting in the problem of dense object extraction. On the one hand, objects tend to be wrongly merged into one object since objects are connected; on the other hand, the existed random sampling scheme does not make good use of the object distribution density to improve the training effect. To meet the demand for dense greenhouse extraction, this paper proposes a novel deep learning-based greenhouse extraction algorithm. To solve the problem that dense objects tend to be wrongly merged, this paper proposes a dual-task learning module, which uses the discriminant object boundary of the pixel-based branch to distinguish the objects of the object-based branch; to take advantage of the object distribution density for effective training, a high-density biased sampler is proposed. Moreover, this paper provides a dataset of manually labeled imagery to train and develop the proposed method. Results on greenhouse extraction in six regions in China achieve peak mIoU and mAP value, surpassing the state-of-the-art methods. Finally, a product of a greenhouse map in China is provided for analysis. Yanfei Zhong, Ailong Ma, Liqin Cao |
IGARSS | 2 |
| 2020 | Urban Scenes Change Detection Based on Multi-Scale Irregular Bag of Visual Features for High Spatial Resolution ImageryabstractRemote sensing scene change detection (SCD) is to detect whether and what changes have occurred in the semantic category of corresponding scenes for a long time at the semantic level. This can provide detailed land use/land cover change information for Urban planning and environmental monitoring. Previous studies take regular patches divided by uniform grid sampling as scene units. This may lead to mosaic phenomenon, and use fixed Window to extract features, ignoring the multi-scale features of ground objects, while extracting scene features. To solve the problems, the multi-scale irregular bag of visual features (MIBVF) framework is proposed for high spatial resolution (HSR) imagery SCD. In this paper, we integrate image classification of the physical characteristics from remote sensing data with the socio-economic attributes from open source geographic data. Road network data is used to preserve the geological significance and semantic integrity of urban scenes, and multi-scale window sampling is used to solve the problem of different object sizes. To confirm the feasibility of the proposed method, experiments with multi-temporal images of the Pudong area in Shanghai indicate that the proposed method achieves a clearly higher change detection accuracy than current state-of-the-art methods. Jiale Chen 0002, Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 3 |
| 2020 | A Novel Global-Aware Deep Network for Road Detection of Very High Resolution Remote Sensing ImageryabstractRoad detection from very high-resolution (VHR) remote sensing imagery has great importance in a broad array of applications. However, the most advanced deep learning-based methods often produce fragmented road segments, due to the complex backgrounds of images, such as the occlusions and shadows caused by the trees and buildings, or the surrounding objects with similar textures. In this paper, the characteristics of existing models were analyzed and an effective road recognition method was explored, we found that capturing long-range dependencies helps improve road recognition. Therefore, a novel global-aware deep network (GAN) for road detection is proposed, in which the spatial-aware module (SAM) was applied to capture spatial context dependencies and the channel-aware module (CAM) was applied to capture the interchannel dependencies. Through establishing the relationships between spatial contexts and between channels, the GAN could effectively alleviate the road recognition problems, and the advantages of the proposed approach were validated on the public DeepGlobe road dataset. The experimental result demonstrates the superiority of our method. Yanfei Zhong, Zhuo Zheng |
IGARSS | 2 |
| 2020 | Cropnet: Deep Spatial-Temporal-Spectral Feature Learning Network for Crop Classification from Time-Series Multi-Spectral ImagesabstractThe Deep Learning (DL) methods can automatically extract features without artificial prior, provides an effective solution for multi-temporal crop classification. Convolutional Neural Networks (CNNs) have superb spatial-spectral feature extraction capabilities, but often lack consideration of the temporal relationship of multi-temporal images, while Recurrent Neural Networks (RNNs) can better learn the sequential variation pattern. This paper designed a deep spatial-temporal-spectral feature learning network (CropNet) combining the advantages of a deep spatial-spectral feature learning module and a deep temporal-spectral feature learning module for better feature extraction in crop classification from time-series remote sensing images. From the results, the proposed method has better crop classification effects from time-series multispectral images in our experimental areas compared with some common traditional machine learning approaches and common DL methods. Chang Luo, Shiyao Meng, Xinyu Wang 0003, Yanfei Zhong |
IGARSS | 5 |
| 2020 | RSSM-Net: Remote Sensing Image Scene Classification Based on Multi-Objective Neural Architecture SearchabstractThe deep learning (DL)-based scene classification methods have been obtained the remarkable attention for the high spatial resolution remote sensing (HRS) imagery. However, from one aspect, the existing DL methods in HRS image scene classification are usually the variations of the natural image processing methods and often the inherent network structures; from another aspect, the strenuous and significant efforts have been devoted to the design of relevant network structures by human experts. In this paper, learning from the natural evolution, the deep neural network is expected to be globally evolved by the machine for automatically adapting the structure of the HRS imagery, a multi-objective neural architecture search based HRS image scene classification method is proposed (RSSM-Net). The two objectives of minimizing a classification error and the computational complexity have been simultaneously optimized through the evolutionary multi-objective method, the competitive neural architectures in a Pareto solution set are then obtained. The effectiveness is proved by the experiment of the UC Merced dataset with several networks designed by human experts. Yuting Wan, Yanfei Zhong, Ailong Ma, Ruyi Feng |
IGARSS | 2 |
| 2020 | Semi-Automatic Fully Sparse Semantic Modeling Framework for Hyperspectral UnmixingabstractIn order to improve the accuracy of surface classification and meet the needs of sub-pixel-level target detection, spectral unmixing has been one of the hot spots in hyperspectral remote sensing research. The employment of the probabilistic topic model to acquire latent topics of hyperspectral image has been an effective way for spectral unmixing. However, this approach fails to consider the sparsity of the semantic representation and high computational complexity. In addition, the number of endmembers cannot be determined automatically. To solve the problem, in this paper, the novel spectral unmixing method based on semi-automatic fully sparse semantic modeling framework (SFSSM) is proposed. In SFSSM, modestly few arithmetic operations are required to identify the pure spectral signatures (endmembers) and the fractional abundances of the endmembers. Meanwhile, the sparsity and representativeness of the topics generated by SFSSM guarantee that the endmembers can be obtained automatically in low time consumption. The experimental results obtained with two real image confirm that the proposed method significantly improves the performance when compared with the other methods. Qiqi Zhu, Wen Zeng 0003, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 4 |
| 2020 | A Modified D-Linknet with Transfer Learning for Road Extraction from High-Resolution Remote SensingabstractRoad extraction, which aims to label remote images with a specific road detection, is a fundamental task for understanding remote sensing imagery. Deep learning has strong characteristic learning ability, for example, the state-of-the-art D-Linknet is an effective way to capture the road information. However, existing regularization methods either do not match the performance for large batches, or still exhibit degradation in performance for smaller batches. Besides, the roads in different areas have various characteristics and lack a good transfer. To remedy these issues, we proposed a novel road extraction network which integrated the filter response normalization (FRN) layer with D-Linknet (FND-Linknet). The FRN layer is effective and robust for road extraction task, and can eliminate the dependency on other batch samples. In addition, the multisource road dataset is collected and annotated to improve features transfer. Experimental results on three datasets verify that the proposed FND-Linknet framework outperforms the state-of-the-art methods both in accuracy and connectivity. Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 3 |
| 2020 | Mapping Local Climate Zones with Circled Similarity Propagation Based Domain AdaptationabstractLocal climate zone (LCZ) has the potential to describe the urban form and function among global cities, which is vital for urban-related researches. However, lacking high quality labeled samples makes it difficult to map global LCZ, and how to transfer labeled samples from one city to another becomes an urgent issue. In this paper, a circled similarity propagation-based domain adaptation method named CPDA for LCZ classification is proposed. In CPDA, the existed labeled samples and unseen samples are called the source and target domain, respectively. The circled similarity matrix of features from source to target and back to the source domain is measured using cosine similarity. Furthermore, an adaptation loss is proposed to make the feature distribution of these two domains as similar as possible, which is beneficial for predicting the target areas with no need for labels. Experiments on the LCZ42 dataset demonstrated the effectiveness of the proposed method. Yanfei Zhong, Ailong Ma |
IGARSS | 2 |
| 2020 | Super Resolution Generative Adversarial Network Based Image Augmentation for Scene Classification of Remote Sensing ImagesabstractHigh spatial resolution remote sensing image (RSI) scene classification, aimed at automatically labelling images with the given semantic categories, has been a hot issue. As it's difficult for RSI to quickly obtain a large number of training samples from a specific area. Traditional scene classification researches were mainly using deep learning models to transfer natural images to RSI. Considering the differences between natural images and RSI, we trained several Super Resolution GAN models by using different resolution RSI data from Google earth image. This paper proposed a novel SRGAN-CNN framework. Through transferring the data with scene classification dataset to obtain high resolution fake RSI. The experimental results demonstrate that the proposed framework can enhance transfer effect and help improve the accuracy of scene classification using low resolution RSI. Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 3 |
| 2020 | Topic Model for Remote Sensing Data: A Comprehensive ReviewabstractFrom text analysis to image interpretation, the topic model (TM) always plays an important role. With its powerful semantic mining capabilities, it is able to capture the latent spectral and spatial information from remote sensing (RS) images. Recent years have witnessed widespread use of TM to solve the problems in RS image interpretation, i.e., semantic segmentation, target detection, and scene classification. However, there has not yet been a study expatiating and summarizing the current situation of RS applications with TM. This paper intends to systematically summarize the application of TM in RS images and to conduct several typical experiments for comparison. Specifically, the architecture of our work can be explained as follows: 1) the theory of TM; 2) the applications of RS based on TM; 3) experimental analysis of typical TM methods to provide reference for further understanding, and 4) summary and prospects for guiding further research into TM for RS data. Qiqi Zhu, Jiangqin Wan, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 3 |
| 2020 | Precise object detection using adversarially augmented local/global feature fusionabstractObject detection, which aims at recognizing or locating the objects of interest in remote sensing imagery with high spatial resolutions (HSR), plays a significant role in many real-world scenarios, e.g., environment monitoring, urban planning, civil infrastructure construction, disaster rescuing, and geographic image retrieval . As a long-lasting challenging problem in both machine learning and geoinformatics communities, many approaches have been proposed to tackle it. However, previous methods always overlook the abundant information embedded in the HSR remote sensing images. The effectiveness of these methods, e.g., accuracy of detection, is therefore limited to some extent. To overcome the mentioned challenge, in this paper, we propose a novel two-phase deep framework, dubbed GLGOD-Net, to effectively detect meaningful objects in HSR images . GLGOD-Net firstly attempts to learn the enhanced deep representations from super-resolution image data . Fully utilizing the augmented image representations, GLGOD-Net then learns the fused representations into which both local and global latent features are implanted. Such fused representations learned by GLGOD-Net can be used to precisely detect different objects in remote sensing images . The proposed framework has been extensively tested on a real-world HSR image dataset for object detection and has been compared with several strong baselines. The remarkable experimental results validate the effectiveness of GLGOD-Net. The success of GLGOD-Net not only advances the cutting-edge of image data analytics , but also promotes the corresponding applicability of deep learning in remote sensing imagery . Xiaobing Han, Tiantian He 0001, Yew-Soon Ong, Yanfei Zhong |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | Exploiting Deep Features for Remote Sensing Image Retrieval: A Systematic InvestigationabstractRemote sensing (RS) image retrieval is of great significant for geological information mining. Over the past two decades, a large amount of research on this task has been carried out, which mainly focuses on the following three core issues: feature extraction, similarity metric, and relevance feedback. Due to the complexity and multiformity of ground objects in high-resolution remote sensing (HRRS) images, there is still room for improvement in the current retrieval approaches. In this article, we analyze the three core issues of RS image retrieval and provide a comprehensive review on existing methods. Furthermore, for the goal to advance the state-of-the-art in HRRS image retrieval, we focus on the feature extraction issue and delve how to use powerful deep representations to address this task. We conduct systematic investigation on evaluating correlative factors that may affect the performance of deep features. By optimizing each factor, we acquire remarkable retrieval results on publicly available HRRS datasets. Finally, we explain the experimental phenomenon in detail and draw conclusions according to our analysis. Our work can serve as a guiding role for the research of content-based RS image retrieval. Xin-Yi Tong 0003, Gui-Song Xia, Yanfei Zhong, Mihai Datcu, Liangpei Zhang 0001 |
IEEE Trans. Big Data | 4 |
| 2020 | Spectral-Spatial-Temporal MAP-Based Sub-Pixel Mapping for Land-Cover Change DetectionabstractThe maximum a posteriori (MAP) estimation model-based sub-pixel mapping (SPM) method is an alternative way to solve the ill-posed SPM problem. The MAP estimation model has been proven to be an effective SPM approach and has been extensively developed over the past few years, as a result of its effective regularization capability that comes from the spatial regularization model. However, various spatial regularization models do not always truly reflect the detailed spatial distribution in a real situation, and the over-smoothing effect of the spatial regularization model always tends to efface the detailed structural information. In this article, under the scenario of time-series observation by remote sensing imagery, the joint spectral-spatial-temporal MAP-based (SST_MAP) model for SPM is proposed. In SST_MAP, a newly developed temporal regularization model is added to the MAP model, based on the prerequisite for a temporally close fine image covering the same study region. This available fine image can provide the specific spatial structures most closely conforming to the ground truth for a more precise constraint, thereby reducing the over-smoothing effect. Furthermore, the three dimensions are mutually balanced and mutually constrained, to reach an equilibrium point and achieve restoration of both smooth areas for the homogeneous land-cover classes and a detailed structure for the heterogeneous land-cover classes. Four experiments were designed to validate the proposed SST_MAP: three synthetic-image experiments and one real-image experiment. The restoration results confirm the superiority of the proposed SST_MAP model. Notably, under the background of time-series observation, SST_MAP provides an alternative way of land-cover change detection (LCCD), achieving both detailed spatial-scale and high-frequency temporal LCCD observation for the study case of urbanization analysis within the city of Wuhan in China. Da He, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Multiobjective Subpixel Mapping With Multiple Shifted Hyperspectral ImagesabstractSubpixel mapping (SPM) is a useful technique that can interpret the spatial distribution inside mixed pixels and produce a finer-resolution classification map for hyperspectral remote-sensing imagery. However, SPM is essentially an ill-posed problem that requires additional information to produce the unique solution. The limited information of a single image is insufficient to make the mapping problem well posed, whereas the complementary spatial information of multiple shifted images is able to reduce the uncertainty and generate an accurate map. The maximum a posteriori model is a feasible way to incorporate auxiliary information for SPM with multiple shifted images, but it introduces a sensitive regularization parameter, which is difficult to preset. Furthermore, the fixed parameter in the iterations influences the incorporation of the multiple images and the spatial prior. In this article, to address these issues, a multiobjective SPM framework for use with multiple shifted hyperspectral images (MOMSM) is proposed. In the proposed algorithm, a multiobjective model consisting of two objective functions, i.e., data fidelity and spatial prior terms, is constructed to transform the SPM into a multiobjective optimization problem, to get rid of the sensitive regularization parameter. To simultaneously optimize the two objective functions, a multiobjective memetic algorithm with a local search operator and an adaptive global replacement strategy is proposed. The multiple images and spatial information can be dynamically fused and the optimal mapping solution with a good balance between the two objectives can be finally obtained. Experiments conducted on both synthetic and real data sets confirm that the proposed method outperforms the other tested SPM algorithms. Mi Song, Yanfei Zhong, Ailong Ma, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Multiobjective Hyperspectral Feature Selection Based on Discrete Sine Cosine AlgorithmabstractFeature selection is an effective way to reduce the data dimensionality of hyperspectral imagery and obtain a better performance in the subsequent applications, such as classification. The ideal approach is to obtain the optimal tradeoff between two criteria for hyperspectral image feature selection: 1) information preservation and 2) redundancy reduction. However, constructing a hyperspectral feature selection model for the above two criteria is difficult due to the complexity of hyperspectral imagery. Although evolutionary multiobjective optimization methods have been recently presented to simultaneously optimize the above criteria, they cannot control the global exploration versus local exploitation capabilities in the search space for the hyperspectral feature selection problem. Thus, in this article, a novel discrete sine cosine algorithm (SCA)-based multiobjective feature selection (MOSCA_FS) approach is proposed for hyperspectral imagery. In the proposed method, a novel and effective framework of multiobjective hyperspectral feature selection is designed. In the framework, the ratio between the Jeffries-Matusita (JM) distance and mutual information (MI) is modeled to minimize the redundancy and maximize the relevance of the selected feature subset. In addition, another measurement - the variance (Var) - is applied for maximizing the information amount. Furthermore, to resolve the discrete hyperspectral feature selection problem, a novel discrete SCA is first proposed, which enhances the selection of the ideal feature subset. The effectiveness and universality of the proposed method was verified by experiments with ten University of California at Irvine (UCI) data sets, five hyperspectral image data sets, and one spectral data set of typical surface features. Yuting Wan, Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Multi-Objective Sparse Subspace Clustering for Hyperspectral ImageryabstractHyperspectral images (HSIs) are typical high-dimensional and complex data. As such, the clustering of HSIs is a challenging task. Out of the motivation to find the low-dimensional structure representation of the high-dimensional data, sparse subspace clustering (SSC) methods have been proposed in recent studies. Sparse representation is an important technique in SSC, which is aimed at obtaining the sparse coefficient matrix of the HSI data. Generally speaking, the acquisition of the sparse coefficient matrix is an ill-posed problem, and the existing methods introduce an extra condition as a regularization term to resolve it. However, the regularization parameter is determined manually, which is difficult and lacks self-adaptability. Hence, in this article, a multi-objective SSC method for hyperspectral imagery is proposed, which simultaneously optimizes the sparse term and the data fidelity term. In addition, the spatial structure information of the HSIs is often neglected in the processing model, and thus, a spatial prior term, as the third optimization objective function, is also tested in this article. As a result, there is no need to manually set a regularization parameter. Furthermore, by using the l0norm as the sparse term, this reduces the error caused by the convex relaxation of the other norms. In the proposed method, a multi-objective optimization model is first used to acquire the sparse coefficient matrix, in which a strategy for constructing the dictionary is proposed for more precise and efficient multi-objective optimization. In addition, a knee point-based selection method is utilized to automatically select the optimal sparse representation solution from the Pareto front. The adjacency matrix is then constructed according to the sparse coefficient matrix. Finally, a spectral clustering method is used to obtain clustering results. Experiments undertaken with four HSI data sets confirm the effectiveness of the proposed method. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Hyperspectral Anomaly Detection via Locally Enhanced Low-Rank PriorabstractAnomaly detection is an active area of research in hyperspectral information processing. Recently, low-rank representation has been applied in hyperspectral anomaly detection. However, the existing low-rank-based methods either involve a complicated dictionary construction process or the anomaly-background separation which is not sufficient. In this article, to solve these problems, a novel hyperspectral anomaly detection method based on a locally enhanced low-rank prior (LELRP-AD) is proposed. This article is inspired by the observation that, in local homogeneous regions, the background signals hold an enhanced low-rank property while the anomalies exhibit spatial sparsity. Based on this observation, the background pixels can be low-rank reconstructed by a set of basis background signals, whereas anomalies can be represented as sparse residuals. First, image segmentation is performed to enhance the homogeneity of the background, in which a Potts-based image segmentation algorithm is adopted with postprocessing, thus avoiding the need for a complicated spectral dictionary for the representation of the background. Furthermore, the original hyperspectral data matrix is augmented with extracted background endmembers for the low-rank and sparse matrix decomposition, to further achieve anomaly-background separation. The experimental results obtained on four real hyperspectral data sets demonstrate the merit and viability of the proposed method compared with the current state-of-the-art methods. Shaoyu Wang 0003, Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image ClassificationabstractDeep learning techniques have provided significant improvements in hyperspectral image (HSI) classification. The current deep learning-based HSI classifiers follow a patch-based learning framework by dividing the image into overlapping patches. As such, these methods are local learning methods, which have a high computational cost. In this article, a fast patch-free global learning (FPGA) framework is proposed for HSI classification. The proposed framework consists of three main parts: 1) a designed sampling strategy; 2) an encoder-decoder-based fully convolutional network (FCN); and 3) lateral connections between the encoder and decoder. In FPGA, an encoder-decoder-based FCN is utilized to consider the global spatial information by processing the whole image, which results in fast inference. However, it is difficult to directly utilize the encoder-decoder-based FCN for HSI classification as it always fails to converge due to the insufficiently diverse gradients caused by the limited training samples. To solve the divergence problem and maintain the FCNs abilities of fast inference and global spatial information mining, a global stochastic stratified (GS2) sampling strategy is first proposed by transforming all the training samples into a stochastic sequence of stratified samples. This strategy can obtain diverse gradients to guarantee the convergence of the FCN in the FPGA framework. For a better design of FCN architecture, FreeNet, which is a fully end-to-end network for HSI classification, is proposed to maximize the exploitation of the global spatial information and boost the performance via a spectral attention-based encoder and a lightweight decoder. A lateral connection module is also designed to connect the encoder and decoder, fusing the spatial details in the encoder and the semantic features in the decoder. The experimental results obtained using three public benchmark data sets suggest that the FPGA framework is superior to the patch-based framework in both speed and accuracy for HSI classification. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | COLOR: Cycling, Offline Learning, and Online Representation Framework for Airport and Airplane Detection Using GF-2 Satellite ImagesabstractMonitoring airports using remote sensing imagery require us to first detect the airports and then perform airplane detection. Detecting airports and airplanes with large-scale remote sensing imagery are significant and challenging tasks in the field of remote sensing. Although many detection algorithms have been developed for detecting airports and airplanes in remote sensing imagery, the efficiency of the processing does not meet the needs of real applications in large-scale remote sensing imagery. In recent years, deep learning techniques, such as deep convolutional neural networks (DCNNs), have achieved great progress in image recognition. However, training a DCNN needs a large number of training examples to accurately fit the data distribution. Annotating training examples in large-scale remote sensing imagery is time-consuming, which makes the pipeline inefficient. In this article, to overcome the above two weaknesses, we propose a novel cycling data-driven framework for efficient and robust airport localization and airplane detection. The proposed method consists of three modules: cycling by example refinement (C), offline learning (OL), and online representation (OR), namely cycling, offline learning, and online representation (COLOR). The OR module is a coarse-to-fine cascaded convolutional neural network, which is used to detect airports and airplanes. The example refinement (ER) module implements the cycling and makes use of the unlabeled remote sensing images and the corresponding predictions obtained by the OR module, to generate training examples. The OL module aims to use the training examples from the ER module to update the OR module, to further improve the performance. The whole workflow involves COLOR. The COLOR framework was used to detect airplanes and airports in 512 large-scale Gaofen-2 (GF-2) remote sensing images with 29$200\times27$ 620 pixels. The results showed that the proposed method obtained a mean average precision (mAP) of 88.32% for the airplane detection. In addition due to the proposed coarse-to-fine cascaded OR module the proposed method is much faster than the traditional approaches in real-world applications. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Novel Robust Feature Descriptor for Multi-Source Remote Sensing Image RegistrationabstractNon-linear radiation difference (NRD) will lead to the corresponding features cannot be mapped one by one, so the traditional image feature matching methods based on intensity or gradient fail to be directly applied to the multisource remote sensing image registration. In this paper, a new robust feature descriptor is proposed, which has the invariance of radiation, scale and rotation. The nonlinear diffusion function which is insensitive to the radiation difference is used to construct the scale space so that the descriptors can be used in images with different resolutions. A pixel-by-pixel local phase congruency algorithm is used to extract the corresponding points, and then the features are described by means of rotation invariance description. Feature matching is completed based on feature vector constructed by the descriptor, thus to realize image registration. In the experimental part, three kinds of multisource remote sensing images with large radiation differences were used to test the descriptors. The results showed that the proposed method effectively extracted the corresponding features, and achieved the best effect in the quantitative evaluation of image registration. Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2019 | Local Block Grouping with Napca Spatial Preprocessing for Hyperspectral Remote Sensing Imagery Sparse UnmixingabstractSpatial regularization sparse unmixing (SRSU) has been widely studied and proved to be far better than the traditional spectral unmixing methods. These spatial sparse unmixing algorithms have obtained many competitive results except for the negative influences of inaccurate estimated unmixing abundances or outliers in abundances. In this paper, to obtain a more accurate SRSU results, a local block grouping with noise-adjusted principal component analysis method is used to do spatial preprocessing in sparse unmixing process. Here, local blocks are treated as a series of vector variables, and these variables are selected by grouping the pixels with similar local spatial structures to the underlying one in the local window. Then noise-adjusted principal component analysis (NAPCA) is taken to transform the original datasets into PCA domain and maintain only the most significant principal component as well as wipe off the inaccurate estimated fractional abundances. Compared with total variation-based and nonlocal means-based SRSU algorithms, the proposed joint local block grouping with NAPCA sparse unmixing method can yield competitive results with state-of-the-art spatial sparse unmixing algorithms using both simulated dataset and real hyperspectral imagery. Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong |
IGARSS | 3 |
| 2019 | SPNet: A Spectral Patching Network for End-To-End Hyperspectral Image ClassificationabstractDeep learning (DL)-based hyperspectral classification primarily use "spatial patching" as preprocessing for incorporating local spatial information. This operation can help to promote classification accuracy but faces the following problems. First, it is difficult to determine the optimal size of spatial patches for different hyperspectral images (HSIs). Second, this operation only exploits spatial features locally but not globally. In this paper, we propose a novel spectral patching network (SPNet) with an end-to-end deep learning architecture for HSI classification. SPNet uses "spectral patching" and Atrous Spatial Pyramid Pooling (ASPP) module to fully preserve the local and global spatial contextual information of original HSIs. The experimental results with UAV-borne hyperspectral dataset demonstrate that the SPNet achieved state-of-the-art accuracy and visualization performance in. Xinyu Wang 0003, Yanfei Zhong, Ji Zhao 0006, Chang Luo, Lifei Wei |
IGARSS | 3 |
| 2019 | D-Resunet: Resunet and Dilated Convolution for High Resolution Satellite Imagery Road ExtractionabstractReliably extracting information from satellite imagery is a difficult problem with many practical applications. One specific case of this problem is the task of automatically detecting roads. Road extraction from satellite images has been a hot research topic in the past decade. In this paper, we propose a semantic segmentation neural network, named D-ResUnet, which adopts U-Net structure, residual learning, and dilated convolutions for road area extraction. The network is built with ResUnet architecture and has dilated convolution layers in its center part. ResUnet architecture combines the strengths of residual units and feature concatenate, which help to ease training of networks and facilitate information propagation. Dilation convolution is a powerful tool that can enlarge the receptive field of feature points without reducing the resolution of the feature maps. We test our network and compare it with U-Net and ResUnet based road extraction methods. The proposed approach outperforms all the comparing methods, which demonstrates its superiority over recently developed state of the arts. Zhiqun Liu, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liqin Cao |
IGARSS | 4 |
| 2019 | Multi-Scale Enhanced Deep Network for Road DetectionabstractRoad detection is a hot research topic in the very high resolution (VHR) remote sensing field and has been applied in various practical applications. Many deep-learning based methods have been used to detect roads and achieved good performance. In this paper, a multi-scale enhanced road detection framework (DenseUNet) is proposed which based on the densely connected convolutional networks (DenseNet) and U-Net. The U-Net has strong capabilities of preserving spatial details due to its skip connections, and the DenseNet can better optimize the deep network. Meanwhile, atrous spatial pyramid pooling (ASPP) is employed to effectively capture multi-scale features for road detection. Finally, a public road dataset was used to verify the proposed approach, compared with other state-of-the-art methods. The proposed method achieve the best performance, which illustrates its superiority. Yanfei Zhong, Ji Zhao 0006 |
IGARSS | 2 |
| 2019 | Sub-Pixel Mapping with Multiple Shifted Hyperspectral Images Based on Multiobjective Evolutionary AlgorithmabstractSub-pixel mapping (SPM) can interpret the sub-pixel spatial distribution of land-cover classes in hyperspectral image, which is an ill-posed problem due to the inadequate information of a single image. Auxiliary information provided by multiple shifted (MS) images can make SPM problem well-posed and improve mapping accuracy. The maximum a posteriori (MAP) technique can incorporate the auxiliary information of MS images, but it introduces a fixed weight parameter to fuse the auxiliary information and spatial prior information, heavily influencing the mapping result. This paper proposed a novel SPM method to model the auxiliary information and spatial prior information into two objective functions, which can be simultaneously optimized by the devised multiobjective evolutionary algorithm. Therefore, there is no need of weight parameter, and the two objective functions can be intelligently integrated during the evolution. Experimental results and parameter analysis have indicated the superiority of the proposed method. Mi Song, Yanfei Zhong, Ailong Ma, Qiqi Zhu, Liqin Cao, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2019 | Tailings Reservoir Disaster and Environmental Monitoring Using the UAV-ground Hyperspectral Joint Observation and Processing: A Case of Study in Xinjiang, the Belt and RoadabstractThe tailings reservoir is an inevitable part of the production of metal mines, and due to it is usually the accumulation of waste residue and waste water, the risk source of artificial debris flow with high potential energy has been formed and the environmental risk cannot be underestimated. Thus, it is an important disaster and environmental protection project for the mining enterprises. However, the existing methods cannot conduct a comprehensive disaster and environmental monitoring, considering the remote sensing technology is an effective method for the large-scale monitoring, thus, this global monitoring will be carried out through a novel UAV-ground hyper-spectral joint observation and processing, where the UAV hyper-spectral image, the ground hyper-spectral data of the water and waste residue, and water quality testing report will be used. In addition, the study area is in Xinjiang, the Belt and Road. Yuting Wan, Yanfei Zhong, Ailong Ma, Lifei Wei, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2019 | Hyperspectral Remote Sensing Image Band Selection Via Multi-Objective Sine Cosine AlgorithmabstractFor hyper-spectral image band selection, there are two main key concerns, which are the curial information preservation and the redundancy information reduction. Since the two objectives are contradictory, the single-objective based band selection methods are usually unsatisfactory, thus, a superior approach which can obtain a trade-off between them is needed. In order to address this problem, an evolutionary computation method called sine cosine algorithm which has the capabilities of the global exploration and the local exploitation is applied, and its multi-objective discrete version is designed for multi-objective band selection. In addition, in this paper, two novel measures are utilized for meeting the requirements. The effectiveness of the proposed method is confirmed by the experimental results obtained with two real hyper-spectral images. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2019 | S3NET: Towards Real-Time Hyperspectral Imagery ClassificationabstractFast hyperspectral image classification methods are required to real time processing on UAV and airborne platform. But current state of the art methods are under patch based framework that is slow due to duplicate computation, while fully convolutional neural network (FCN) is difficult to run inference in one shot due to patch based training strategy. In this paper, we firstly identify overparameterization and inconsistent class ratio per mini-batch as the central causes impeding training of FCN based methods so we propose a novel framework based on a new lightweight network and a stratified sample based training strategy to overcome above problems. To refine redundant spectrum information for fast computation, the spectral attention module is proposed as a soft band selection. The proposed training strategy enables the stability of training by keeping class ratio consistent in mini-batches. On the WHU-HongHu UAV dataset, our method achieves great performance by obtaining a OA of 98.94 and running at 144 fps. Additionally, our method can obtain a flexible trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong |
IGARSS | 2 |
| 2019 | Pop-Net: Encoder-Dual Decoder for Semantic Segmentation and Single-View Height EstimationabstractThe single-view semantic 3D challenge in 2019 Data Fusion Contest is to predict both semantic labels and normalized digital surface model (nDSM) for urban scenes from single-view satellite images. We propose a novel pyramid on pyramid network (Pop-Net) based on Encoder-Dual Decoder framework to end-to-end multi-task learning. The encoder is a deformable ResNet-101 backbone network. Two feature pyramid networks, as decoders, are responsible for semantic segmentation and height estimation, respectively. Semantic information is crucial to estimate height. Therefore, regression pyramid on the semantic pyramid is introduced to leverage semantic features to help height estimation. To deal with outliers in heights, we leverage anchor-based regression and smooth L1 loss for optimization to obtain more robust height estimation. Without bells and whistles, our single model entry achieves 77.78% mIoU and 53.40% mIoU-3 on test set, ranking 2nd in the Single-view Semantic 3D Challenge of the 2019 IEEE GRSS Data Fusion Contest. The code is available at https://github.com/Z-Zheng/PopNet. Zhuo Zheng, Yanfei Zhong |
IGARSS | 2 |
| 2019 | High-Resolution Remote Sensing Image Scene Understanding: A ReviewabstractHigh-resolution remote sensing (HRS) image analysis is a fundamental but challenging problem. To bridge the semantic gap, scene understanding has been proposed to achieve higher-level interpretation, through classifying the HRS scene through spatial relationship cognition and semantic induction between the land-cover objects. As a new research field, however, there has not yet been a study expatiating and summarizing the current situation of scene understanding. This paper first defines the concept of scene understanding for HRS imagery, which is different from natural image scene classification. The theory of scene understanding for HRS imagery is investigated, and is classified into four main categories: 1) scene classification based on semantic objects; 2) scene classification based on mid-level features; 3) scene classification based on deep learning; and 4) scene understanding applications based on geographic data mining. Qiqi Zhu, Xiongli Sun, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2019 | Directive local color transfer based on dynamic look-up tableabstractColor transfer in image processing usually suffers from misleading color mapping and loss of details. This paper presents a novel directive local color transfer method based on dynamic look-up table (D-DLT) to solve these problems in two steps. First, a directive mapping between the source and the reference image is established based on the salient detection and the color clusters to obtain directive color transfer intention. Then, dynamic look-up tables are created according to the color clusters to preserve the details, which can suppress pseudo contours and avoid detail loss. Subjective and objective assessments are presented to verify the feasibility and the availability of the proposed approach. Experimental results demonstrate that our proposed method has better performance on natural color images than classical color transfer algorithms. Furthermore, the reference image can be extended to color blocks instead of images. Zhijiang Li, Zhenshan Tan, Liqin Cao, Lei Jiao 0001, Yanfei Zhong |
Signal Process. Image Commun. | 6 |
| 2019 | Spatiotemporal Subpixel Geographical Evolution MappingabstractIn recent decades, spatiotemporal subpixel mapping (SSM) approaches have been extensively developed to deal with the mixed-pixel problem by incorporating fine spatial resolution images with the same field of view from different acquisition times. This is an alternative to the conventional subpixel mapping (SPM) method, which is based on only monotemporal images. SSM has become one of the state-of-the-art SPM approaches, and has been widely applied in urban management and ecological monitoring. However, in the traditional SSM methods, the spatial correlation within the multitemporal images is insufficiently exploited and is ignored in the spatiotemporal model construction. In addition, the contribution of the land covers' spatial distribution in the multitemporal images is incompletely considered, and the geographic variation during the time interval is ignored, which underutilizes the spatiotemporal information. In this paper, an SSM algorithm based on a geographically weighted regression (GWR) model and evolutionary algorithm theory, called spatiotemporal subpixel geographical evolution mapping (STGEM), is proposed for multitemporal remote sensing images. The proposed algorithm considers the spatiotemporal dependence not only between the current subpixel and the corresponding fine pixel, but also with the neighboring fine distribution patterns within the fine image. Moreover, the potential temporal information of the geospatial variation is fully realized by considering not only the time interval between the bitemporal images, but also the ratio of changed area between them, based on the GWR model. Two synthetic-image experiments with bitemporal Landsat 8 images and bitemporal QuickBird images were carried out to validate the proposed algorithm. Furthermore, a real-image experiment using a bitemporal pair of Gaofen-2 images and a Landsat 8 image was also undertaken. A comparison was made with several traditional SPM methods, as well as the state-of-the-art SSM approaches, and the experimental results confirmed the superiority of the proposed STGEM algorithm. The proposed STGEM achieves a fine spatial and temporal resolution thematic map, both qualitatively and quantitatively, and has great potential for fine-scale and frequent time-series observation and monitoring. Da He, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Multi-Scale and Multi-Task Deep Learning Framework for Automatic Road ExtractionabstractRoad detection and centerline extraction from very high-resolution (VHR) remote sensing imagery are of great significance in various practical applications. Road detection and centerline extraction operations depend on each other, to a certain extent. The road detection constrains the appearance of the centerline, and the centerline enhances the linear features of the road detection. However, most of the previous works have addressed these two tasks separately and have not considered the symbiotic relationship between them, making it difficult to obtain smooth and complete roads. In this paper, a novel multi-scale and multi-task deep learning framework for automatic road extraction (MSMT-RE) is proposed to build the relationship between them and simultaneously complete the road detection and centerline extraction tasks. U-Net is selected as the basic network for multi-task learning due to its strong ability to preserve spatial details. Multi-scale feature integration is also applied in the framework to increase the robustness of the feature extraction. Meanwhile, an adaptive loss function is introduced to solve the problems of roads taking up a small percentage of the training samples, and the fact that the positive samples of the two tasks are unbalanced. Finally, experiments were conducted on two public road data sets and two large images from Google Earth, and the proposed framework was compared with other state-of-the-art deep learning-based road extraction methods, both quantitatively and qualitatively. The proposed approach outperformed all the compared methods, confirming its advantages in automatic road extraction. Yanfei Zhong, Zhuo Zheng, Ji Zhao 0006, Ailong Ma, Jie Yang 0040 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Multiobjective Sparse Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) of remote sensing imagery is aimed at generating a classification map with a finer spatial resolution based on the abundance maps. The sparse subpixel mapping (SSM) method reformulates the SPM problem into a spatial pattern linear regression problem based on the preconstructed subpixel patch dictionary. However, in the SSM model, the optimization of the L0-norm is a nonconvex NP-hard problem, so the L1-norm is used to replace the L0-norm to obtain an approximate solution, and the selection of the optimal weight parameter between multiple terms is difficult. Thus, in this paper, a novel multiobjective SSM (MOSSM) framework for remote sensing imagery is proposed, which transforms the SSM problem into a multiobjective optimization problem. In MOSSM, first, the sparsity term is accurately modeled using the L0-norm instead of the L1-norm to avoid the potential errors caused by the L1-norm, and an evolutionary algorithm is used to directly optimize the L0-norm. Second, a subfitness-based multiobjective evolutionary algorithm is employed to simultaneously optimize the fidelity term, the sparsity term, and the spatial prior term, and to generate a set of optimal sparse coefficients to balance these three terms. Thus, there is no need to determine sensitive weight parameters. Finally, two spatial prior terms, which can be applied to the overcomplete dictionary, are presented in the proposed MOSSM-TV and MOSSM-L algorithms to incorporate the spatial correlation of subpixels. Experiments were conducted with two synthetic images and two real data sets, and the results were compared with those of ten other SPM algorithms to demonstrate the effectiveness of the proposed method. Mi Song, Yanfei Zhong, Ailong Ma, Ruyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Fully Automatic Spectral-Spatial Fuzzy Clustering Using an Adaptive Multiobjective Memetic Algorithm for Multispectral ImageryabstractClustering of remote sensing imagery is a tough task due to the particular and complex structure of remote sensing images and the shortage of known information. In this paper, we propose a fully automatic spectral-spatial fuzzy clustering method using an adaptive multiobjective memetic algorithm (AMOMA) for multispectral remote sensing imagery. This approach is made up of two automatic layers: an automatic determination layer and an automatic clustering layer. The first layer seeks the optimal number of clusters through a self-adaptive differential evolution algorithm. The second layer then takes advantage of the AMOMA for spectral-spatial clustering using the optimal number of clusters. The knee point from the Pareto front is then selected through the angle-based method in every generation, and we then compare the knee points between generations to output the final optimal solution. The effectiveness of the proposed method is verified by the experimental results obtained with three remote sensing data sets. Yuting Wan, Yanfei Zhong, Ailong Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Blind Hyperspectral Unmixing Considering the Adjacency EffectabstractThis paper focuses on the blind unmixing technique for analyzing hyperspectral images (HSIs). A joint deconvolution and blind hyperspectral unmixing (DBHU) algorithm is proposed, which is aimed at eliminating the impact of the adjacency effect (AE) on unmixing. In remote sensing imagery, the AE occurs in the presence of atmospheric scattering over a heterogeneous surface. The AE leads to blurring and additional mixing of HSIs and makes it difficult to estimate endmembers and abundances accurately. In this paper, we first model the blurred HSIs by the use of a bilinear mixing model (BMM), where a blurring kernel is used to model the mixing caused by the AE. Based on the BMM, the DBHU problem is formulated as a constrained and biconvex optimization problem. Specifically, the minimum-volume simplex (MVS) is incorporated to deal with the additional mixing caused by the AE, and 3-D total variation (TV) priors are adopted to model the spectral-spatial correlation of the data. In DBHU, the biconvex problem is efficiently solved by a nonstandard application of the alternating direction method of multipliers (ADMM) algorithm, where a block coordinate descent scheme is applied by splitting the original problem into two saddle point subproblems, and then minimizing the subproblems alternately via the ADMM until convergence. The experimental results obtained with both simulated and real data confirm the viability of the proposed algorithm, and DBHU works well, even where both blurring and noise are present in the scene. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001, Yanyan Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Blind Spectral Unmixing Considering the Adjacent EffectabstractBlind hyperspectral unmixing (HU) technique aims at identifying pure materials in a hyperspectral image (HSI), called endmembers, and quantifying the corresponding proportions, called abundances, with little prior knowledge. In this paper, the degradation mechanism during data collection - adjacent effect (AE), is considered in the process of blind HU. Since the AE leads to blurring (the loss of sharpness, contrast and apparent resolution) in scene, it blocks the quantitative analysis of HSI in sub-pixel level and makes the estimated endmembers and abundances inaccurate. To solve this problem, a bilinear mixing model is developed to simulate the AE, and a novel algorithm, termed joint deconvolution and blind HU (DBHU) is proposed. In DBHU, the bi-convex optimization problem is efficiently solved by a nonstandard application of the alternating direction method of multipliers (ADMM) algorithm, where a block coordinate descent scheme is applied by splitting the original problem into two saddle-point subproblems and then minimizing the subproblems alternatively via ADMM until convergence. The experimental results on both simulated and real HSI illustrate the viability of the proposed algorithm. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001, Yanyan Xu 0003 |
IGARSS | 2 |
| 2018 | Urban Land Use/Land Cover Classification Based on Feature Fusion Fusing Hyperspectral Image and Lidar DataabstractHyperspectral images have been widely used in classification because of the abundant spectral information. But it can't distinguish the objective with similar spectral character but different elevation. However, LiDAR data can obtain elevation information. Therefore, it will obtain better classification maps if fusing the two data. In recent years, CNN has attracted much attention due to its powerful ability to excavate the potential representation and features of the raw data. However, it's difficult to distinguish the objects with different spectral information but similar surface character. Unlike CNN features, the traditional manual features, such as the normalized vegetation index (NDVI), have a certain characteristic expression significance. In order to consider both the semantic information of traditional manual features and the advanced features of CNN features, this paper proposes a fusion algorithm of hyperspectral and LiDAR fusion based on feature fusion. The proposed algorithm has achieved a good fusion classification effect on the MUUFL Gulfport Hyperspectral and LiDAR Data set. Qiong Cao, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2018 | Land Cover Change Detection Based on Spatial-Temporal Sub-Pixel Evolution Mapping: A Case Study for Urban ExpansionabstractIn the past decades, land cover change detection (LCCD) has been dramatically developed, since it provides corroborative support for policy decision, regulatory actions, and subsequent urban-rural activities. Satellite remote sensing image is the major source of LCCD since it is able to revisit the Earth's surface regularly and provide time series images for monitoring and space-time analysis. However, there is always a trade-off between spatial scale and temporal scale, i.e., finer spatial resolution image generally has a lower revisit frequency, leading to an observation omission; while higher revisit frequency image usually has a lower spatial resolution, resulting in a deficiency in detecting finer scale change information. In this paper, a spatial-temporal sub-pixel mapping (SSM) algorithm is proposed on the premise that one pair of fine spatial resolution image with low frequency revisit period and coarse spatial resolution with high frequently revisit period are available, and SSM is taken to restore the coarse image to a finer scale thematic map which can be then compared to the fine image, realizing a frequency and detailed LCCD. SSM is an extension of traditional mono-temporal sub-pixel mapping (SPM) algorithm, and is improved by incorporating temporally fine distribution patterns for a more appropriate restoration of coarse image. A study case for urban expansion LCCD were carried out to verify the ability of the proposed algorithm to handle change detection based on one pair of china-made Gaofen-2 image (GF-2) and Landsat-8 image, the result demonstrate that the proposed SSM algorithm outperform the other traditional SPM, achieving both fine temporal resolution and spatial resolution LCCD for further applications. Da He, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2018 | Color: Cycling Offline Learning and Online Representing for Remote Sensing DataflowabstractIn recent years, many model driven frameworks have achieved great success in remote sensing. For large scale remote sensing dataflow, however, its ability of processing is far from enough. In this paper, we propose a novel cycling data driven framework based on deep learning, which consists of three components: offline learning, online representing and sample refining. That online representing makes predictions on unlabeled remote sensing images and uses them to retrain the model by offline learning while sample refining as a router refines coarse predictions and forwards refined predictions to offline learning. Because the process is like circling between offline learning and online representing, we refer to our framework as circling offline learning and online representing (COLOR). We show the excellent performance and potentiality on airplane object detection by applying a fairly straightforward implement of COLOR using our GF2 airplane object detection dataset. Code and dataset will be made available. Zhuo Zheng, Yanfei Zhong |
IGARSS | 2 |
| 2018 | Scene Classification Based on Multiscale Convolutional Neural NetworkabstractWith the large amount of high-spatial resolution images now available, scene classification aimed at obtaining high-level semantic concepts has drawn great attention. The convolutional neural networks (CNNs), which are typical deep learning methods, have widely been studied to automatically learn features for the images for scene classification. However, scene classification based on CNNs is still difficult due to the scale variation of the objects in remote sensing imagery. In this paper, a multiscale CNN (MCNN) framework is proposed to solve the problem. In MCNN, a network structure containing dual branches of a fixed-scale net (F-net) and a varied-scale net (V-net) is constructed and the parameters are shared by the F-net and V-net. The images and their rescaled images are fed into the F-net and V-net, respectively, allowing us to simultaneously train the shared network weights on multiscale images. Furthermore, to ensure that the features extracted from MCNN are scale invariant, a similarity measure layer is added to MCNN, which forces the two feature vectors extracted from the image and its corresponding rescaled image to be as close as possible in the training phase. To demonstrate the effectiveness of the proposed method, we compared the results obtained using three widely used remote sensing data sets: the UC Merced data set, the aerial image data set, and the google data set of SIRI-WHU. The results confirm that the proposed method performs significantly better than the other state-of-the-art scene classification methods. Yanfei Zhong, Qianqing Qin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Unsupervised Change Detection Based on Hybrid Conditional Random Field Model for High Spatial Resolution Remote Sensing ImageryabstractHigh spatial resolution (HSR) remote sensing images provide detailed geometric information about land cover. As a result, it is possible to detect more subtle changes with the help of HSR images. However, due to the increased spatial resolution and the limited spectral information, it is difficult to identify the real changes only through the spectral feature of the image. To fully explore the spectral–spatial information and improve the change detection performance for HSR images, this paper proposes the hybrid conditional random field (HCRF) model, which combines the traditional random field method with an object-based technique. In the proposed method, the spectral discriminative information of a single pixel is extracted by the unary potential, which is modeled using a soft clustering method to make an initial separation of changed and unchanged pixels. The pairwise potential then considers the contextual information of adjacent pixels to favor spatial smoothing. An object term is also introduced in the HCRF model to keep the homogeneity of changed objects. By the use of these approaches, the oversmoothing problem of the random field-based methods and the detection error caused by the segmentation strategy in the object-based methods can be relieved. The proposed method was tested on three HSR image data sets and outperformed the compared state-of-the-art techniques. Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Multiobjective Subpixel Land-Cover MappingabstractThe hyperspectral subpixel mapping (SPM) technique can generate a land-cover map at the subpixel scale by modeling the relationship between the abundance map and the spatial distribution image of the subpixels. However, this is an inverse ill-posed problem. The most widely used way to resolve the problem is to introduce additional information as a regularization term and acquire the unique optimal solution. However, the regularization parameter either needs to be determined manually or it cannot be determined in a fully adaptive manner. Thus, in this paper, the multiobjective subpixel land-cover mapping (MOSM) framework for hyperspectral remote sensing imagery is proposed, in which the two function terms [the fidelity term and the prior term (i.e., the regularization term)] can be optimized simultaneously, and there is no need to determine the regularization parameter explicitly. In order to achieve this goal, two strategies are designed in MOSM: 1) a high-resolution distribution image-based individual encoding strategy is designed in order to calculate the prior term accurately and 2) a subfitness-based individual comparison strategy is designed in order to generate subpixel land-cover mapping solutions with a high quality to update the population. Four data sets (one simulated, two synthetic, and one real hyperspectral image) were used to test the proposed method. The experimental results show that MOSM can perform better than the other subpixel land-cover mapping methods, demonstrating the effectiveness of MOSM in balancing the fidelity term and prior term in the SPM model. Ailong Ma, Yanfei Zhong, Da He, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Saliency-Based Endmember Detection for Hyperspectral ImageryabstractThis paper focuses on the endmember extraction (EE) technique for analyzing hyperspectral images. We first prove that the reconstruction errors (REs) and abundance anomalies (AAs) (abundances that fail to satisfy the abundance constraints) are effective in extracting undetected endmembers. Then, according to the spatial continuity of the endmember objects and differing from noise or outliers with a sparse distribution, the endmembers are assumed to be located at some salient areas in the RE and AA maps. A novel EE algorithm termed saliency-based endmember detection (SED) is proposed, where the visual saliency model is introduced to explore and analyze the spatial information that is contained in the AA and RE maps. Specifically, the AA and RE maps are regarded as the visual inputs, whereas the endmembers are treated as the visual stimuli. In SED, we assume that the pure pixel assumption holds. Based on the characteristics of the human visual system, the proposed method can not only extract endmembers in homogenous areas, but it can also highlight the small targets whose abundances may be spatially varied. In addition, since the spatial information is exploited in the reconstruction, the capability of the endmembers to represent the hyperspectral scene is automatically considered in the process of EE, and the detected endmembers are both accurate and reliable. The experimental results obtained on both simulated and real hyperspectral data confirm the merits and viability of the proposed algorithm. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001, Yanyan Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | A New Spectral-Spatial Sub-Pixel Mapping Model for Remotely Sensed Hyperspectral ImageryabstractIn this paper, a new joint spectral-spatial subpixel mapping model is proposed for hyperspectral remotely sensed imagery. Conventional approaches generally use an intermediate step based on the derivation of fractional abundance maps obtained after a spectral unmixing process, and thus the rich spectral information contained in the original hyperspectral data set may not be utilized fully. In this paper, a concept of subpixel abundance map, which calculates the abundance fraction of each subpixel to belong to a given class, was introduced. This allows us to directly connect the original (coarser) hyperspectral image with the final subpixel result. Furthermore, the proposed approach incorporates the spectral information contained in the original hyperspectral imagery and the concept of spatial dependence to generate a final subpixel mapping result. The proposed approach has been experimentally evaluated using both synthetic and real hyperspectral images, and the obtained results demonstrate that the method achieves better results when compared to other seven subpixel mapping methods. The numerical comparisons are based on different indexes such as the overall accuracy and the CPU time. Moreover, the obtained results are statistically significant at 95% confidence. Xiong Xu 0001, Xiaohua Tong, Antonio Plaza, Jun Li 0009, Yanfei Zhong, Huan Xie 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Scene Classification Based on the Sparse Homogeneous-Heterogeneous Topic Feature ModelabstractHigh spatial resolution (HSR) imagery scene classification has been the subject of increased interest in recent years, and has great potential for many applications, such as urban functional analysis. Rooted in natural information processing, the use of the probabilistic topic model (PTM) to capture latent topics to represent HSR images has been an effective way to bridge the semantic gap. However, how to effectively discover discriminative information to recognize the HSR scenes is a challenging task. In this paper, the sparse homogeneous-heterogeneous topic feature model (SHHTFM) is proposed for HSR image scene classification. Differing from the conventional PTM-based scene classification methods, which utilize only heterogeneous features, SHHTFM explores the effect of the homogeneous information. Based on the union of uniform grid sampling and simple linear iterative clustering superpixel sampling, SHHTFM exploits both the heterogeneous and homogeneous information. After separately mining different types of low-level features and latent topics, the sparse topic inference procedure of SHHTFM further improves the fusion of the sparse heterogeneous and homogeneous topics. In addition, multisource geographical data are effectively integrated, where the water and vegetation boundaries define a more accurate way to restrict the boundaries of different scenes, and are then combined with the road network data to further improve the scene annotation performance. This provides more reliable and applicable results for us to better understand the complex scenes. The experimental results obtained with two HSR image classification data sets and an HSR image annotation data set demonstrate that the proposed SHHTFM framework can solve the scene classification problem, with a high classification accuracy as well as a high time efficiency. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Adaptive Deep Sparse Semantic Modeling Framework for High Spatial Resolution Image Scene ClassificationabstractHigh spatial resolution (HSR) imagery scene classification, which involves labeling an HSR image with a specific semantic class according to the geographical properties, has received increased attention, and many algorithms have been proposed for this task. The employment of the probabilistic topic model to acquire latent topics and the convolutional neural networks (CNNs) to capture deep features for representing HSR images has been an effective ways to bridge the semantic gap. However, the midlevel topic features are usually local and significant, whereas the high-level deep features convey more global and detailed information. In this paper, to discover more discriminative semantics for HSR images, the adaptive deep sparse semantic modeling (ADSSM) framework combining sparse topics and deep features is proposed for HSR image scene classification. In ADSSM, the fully sparse topic model and a CNN are integrated. To exploit the multilevel semantics for HSR scenes, the sparse topic features and deep features are effectively fused at the semantic level. Based on the difference between the sparse topic features and the deep features, an adaptive feature normalization strategy is proposed to improve the fusion of the different features. The experimental results obtained with four HSR image classification data sets confirm that the proposed method significantly improves the performance when compared with the other state-of-the-art methods. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Differentiable sparse unmixing based on Bregman divergence for hyperspectral remote sensing imageryabstractSparse unmixing has been successfully applied to hyperspectral remote sensing imagery based on the assumption that the observed image signatures can be expressed in a linear sparse regression with a large standard spectral library. Prior work for sparse unmixing usually utilizes L1norm or Laplacian distribution to promote sparsity. Unfortunately, the L1norm is not differentiable, which may lead to unstable results. In this paper, we adopt Bregman divergence for sparse unmixing, which is a differentiable, smoother prior. Based on the Maximum A Posterior (MAP) estimation, the proposed method has achieved sparse, stable and precise fractional abundances. The experimental results both simulated dataset and the real hyperspectral image demonstrate the effectiveness of the proposed differentiable sparse unmixing algorithm. Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2017 | Robust geospatial object detection based on pre-trained faster R-CNN framework for high spatial resolution imageryabstractGeospatial object detection from high spatial resolution (HSR) imagery is significant and challenging for further analyzing the object-related information in various civil and military applications. Traditional object detection methods based on the handcrafted features are limited by their efficiency in describing the multi-class objects from large-swath and complex-context HSR imagery. Although convolutional neural network (CNN) can extract the features automatically, the feature extraction and detection stages are still separate and time-consuming. In addition, manual labelling information is limited and an efficient real-time one-stage detection framework for HSR imagery is scare. In this paper, a robust pre-trained efficient multi-class geospatial object detection framework - pre-trained Faster R-CNN sharing the convolutional features between region proposal stage and detection stage is proposed for HSR imagery. Extensive experiments and evaluations on a ten-class object detection dataset are conducted for the proposed method. Xiaobing Han, Yanfei Zhong, Ruyi Feng, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2017 | Sub-pixel intelligence mapping considering spatial-temoporal attraction for remote sensing imageryabstractMixed pixel is a ubiquitous phenomenon in remotely sensed imagery, especially in moderate and low spatial resolution imagery, which compromise the hard land cover classification since the dominant class will shadow the information of other vulnerable classes, bringing trouble to imagery interpretation. Since the past decades, sub-pixel mapping (SPM) approaches were developed to deal with the mixture problem, on the basis of soft classification, to retrieval the pure components and its geospatial distribution within mixed pixels. Recently, SPM integrated with auxiliary information is gradually been a state-of-the-art method for mixed pixel problem, and has been proved effectively. However, few works has been dedicated to explore the geostatistic inter-correlation between spatial and temporal among the time sequences imageries. In this paper, a novel SPM algorithm based on swarm intelligence theory, considering spatiotemporal geographical attraction among multi-temporal imageries, called spatiotemporal attraction based sub-pixel evolution mapping (SASEM), is proposed for remote sensing imagery, Experiments were carried out to verify the proposed algorithm, and the result illustrate that the proposed algorithm outperform the traditional SPM, achieving a fine spatial resolution thematic map for further applications. Da He, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2017 | Scene semantic classification based on scale invariance convolutional neural networksabstractConvolutional neural networks (CNNs) has been introduced into remote sensing scene classification, achieving outstanding performance. However, the scale change of objects contained in remote sensing scene image make it difficult to extract feature robust to scale, limiting the further improvement of classification accuracy. In this paper, a scene classification method named Scale Invariance Convolutional Neural Networks (SICNNs) is proposed for remote sensing scene classification. In the proposed method, two images with different scales generated by randomly stretching one image are fed into CNNs simultaneously for training at intervals of several iterations. Then a similarity measure layer was added in SICNN to make the distance of the two feature vectors extracted from the two images as close as possible, leading extracted feature to be robust to scale. Experimental results using two datasets, i.e. the UC Merced dataset, Google dataset of SIRI-WHU, demonstrated the effectiveness of the proposed method. Yanfei Zhong, Ji Zhao 0006, Ailong Ma, Qianqing Qin |
IGARSS | 2 |
| 2017 | Change detection based on structural conditional random field framework for high spatial resolution remote sensing imageryabstractIn this paper, a structural conditional random field framework (SCRF) is proposed to detect the detailed change information from high spatial resolution (HSR) remote sensing imagery. Traditional random field based methods encounter the over-smoothing problem when deal with HSR images and the boundary of changed objects cannot be preserved well. To solve this problem, in SCRF, fuzzy c means (FCM) is used to model the unary potential while avoiding the independent assumption. Pairwise potentials with different shapes are selected as the structural set to model the spatial features of land cover such as buildings and roads. Based on SCRF, a set of change belief maps are generated to describe the observed image from different aspects. An object based fusion strategy is then followed to combine the belief maps to get the refined result. The results of the proposed method on two HSR data sets outperform some state-of-art algorithms. Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2017 | Saliency-based endmember detection for hyperspectral imageryabstractThis paper focuses on the spectral unmixing technique for analyzing hyperspectral image (HSI). In this paper, we first prove that the reconstruction errors and the abundance anomalies (AAs, abundances that are negative or greater than one) are effective in measuring the purity of pixels. Then, due to the continuity of the objects in the space, the endmembers are assumed to be located at some noticeable areas in residual and AA maps. A saliency-based endmember detection (SED) algorithm which aims at iteratively extracting endmembers from the residual and AA maps is proposed, where the visual attention mechanism is developed to understand and analyze the spatial pattern of endmembers. In addition, when searching for new endmembers, the spectral properties are also utilized to promote the robustness of the proposed method. The experimental results on both simulated data and real hyperspectral data illustrate the merits and viability of the proposed algorithm. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001, Yanyan Xu 0003 |
IGARSS | 2 |
| 2017 | MINI-UAV borne hyperspectral remote sensing: A reviewabstractIn recent years, the science of hyperspectral remote sensing has huge development in virtue of the integration of low-cost lightweight hyperspectral sensors and unmanned aerial vehicles (UAVs). As an alternative of manned aircraft, UAV has some unique advantages enabling the researchers acquire the hyperspectral images of their interest area flexibly and promptly. This review focuses on the recent developments of UAV borne hyperspectral remote sensing system, and gives an overview of the corresponding platforms, sensors, data acquisition, processing and current applications. Future challenges and research directions for UAV borne hyperspectral data are also addressed. Yanfei Zhong, Xinyu Wang 0003, Tianyi Jia, Lifei Wei, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2017 | A Novel Approach to Subpixel Land-Cover Change Detection Based on a Supervised Back-Propagation Neural Network for Remotely Sensed Images With Different ResolutionsabstractExtracting subpixel land-cover change detection (SLCCD) information is important when multitemporal remotely sensed images with different resolutions are available. The general steps are as follows. First, soft classification is applied to a low-resolution (LR) image to generate the proportion of each class. Second, the proportion differences are produced by the use of another high-resolution (HR) image and used as the input of subpixel mapping. Finally, a subpixel sharpened difference map can be generated. However, the prior HR land-cover map is only used to compare with the enhanced map of LR image for change detection, which leads to a nonideal SLCCD result. In this letter, we present a new approach based on a back-propagation neural network (BPNN) with a HR map (BPNN_HRM), in which a supervised model is introduced into SLCCD for the first time. The known information of the HR land-cover map is adequately employed to train the BPNN, whether it predates or postdates the LR image, so that a subpixel change detection map can be effectively generated. In order to evaluate the performance of the proposed algorithm, it was compared with four state-of-the-art methods. The experimental results confirm that the BPNN_HRM method outperforms the other traditional methods in providing a more detailed map for change detection. Ke Wu 0004, Yanfei Zhong, Xianmin Wang, Weiwei Sun 0005 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Transfer Learning With Fully Pretrained Deep Convolution Networks for Land-Use ClassificationabstractIn recent years, transfer learning with pretrained convolutional networks (CNets) has been successfully applied to land-use classification with high spatial resolution (HSR) imagery. The commonly used transfer CNets partially use the feature descriptor part of the pretained CNets, and replace the classifier part of the pretrained CNets in the old task with a new one. This causes the separation and asynchrony between the feature descriptor part and the classifier part of the transferred CNets during the learning process, which reduces the effectiveness of the training process. To overcome this weakness, a transfer learning method with fully pretrained CNets is proposed in this letter for the land-use classification of HSR images. In the proposed method, a multilayer perceptron (MLP) classifier is quickly pretrained using the high-level features extracted by the feature descriptor of the pretrained CNets. Fully pretrained CNets can be generated by concatenating the feature descriptor of the pretrained CNets and the pretained MLP. Because both the feature descriptor and the classifier are pretrained, the separation and asynchrony between the two parts can be avoided during the training process. The final transferred CNets are then obtained by fine-tuning the fully pretrained CNets with the random cropping and mirroring strategy. The experiments show that the proposed method can accelerate the convergence of the training process with no loss of accuracy in land-use classification, and its performance is comparable to other latest methods. Bo Huang 0001, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Spatial Group Sparsity Regularized Nonnegative Matrix Factorization for Hyperspectral UnmixingabstractIn recent years, blind source separation (BSS) has received much attention in the hyperspectral unmixing field due to the fact that it allows the simultaneous estimation of both endmembers and fractional abundances. Although great performances can be obtained by the BSS-based unmixing methods, the decomposition results are still unstable and sensitive to noise. Motivated by the first law of geography, some recent studies have revealed that spatial information can lead to an improvement in the decomposition stability. In this paper, the group-structured prior information of hyperspectral images is incorporated into the nonnegative matrix factorization optimization, where the data are organized into spatial groups. Pixels within a local spatial group are expected to share the same sparse structure in the low-rank matrix (abundance). To fully exploit the group structure, image segmentation is introduced to generate the spatial groups. Instead of a predefined group with a regular shape (e.g., a cross or a square window), the spatial groups are adaptively represented by superpixels. Moreover, the spatial group structure and sparsity of the abundance are integrated as a modified mixed-norm regularization to exploit the shared sparse pattern, and to avoid the loss of spatial details within a spatial group. The experimental results obtained with both simulated and real hyperspectral data confirm the high efficiency and precision of the proposed algorithm. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001, Yanyan Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene ClassificationabstractAerial scene classification, which aims to automatically label an aerial image with a specific semantic category, is a fundamental problem for understanding high-resolution remote sensing imagery. In recent years, it has become an active task in the remote sensing area, and numerous algorithms have been proposed for this task, including many machine learning and data-driven approaches. However, the existing data sets for aerial scene classification, such as UC-Merced data set and WHU-RS19, contain relatively small sizes, and the results on them are already saturated. This largely limits the development of scene classification algorithms. This paper describes the Aerial Image data set (AID): a large-scale data set for aerial scene classification. The goal of AID is to advance the state of the arts in scene classification of remote sensing images. For creating AID, we collect and annotate more than 10000 aerial scene images. In addition, a comprehensive review of the existing aerial scene classification techniques as well as recent widely used deep learning methods is given. Finally, we provide a performance analysis of typical aerial scene classification and deep learning approaches on AID, which can be served as the baseline results on this benchmark. Gui-Song Xia, Jingwen Hu 0001, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | Scene Classification Based on the Fully Sparse Semantic Topic ModelabstractIn high spatial resolution (HSR) imagery scene classification, it is a challenging task to recognize the high-level semantics from a large volume of complex HSR images. The probabilistic topic model (PTM), which focuses on modeling topics, has been proposed to bridge the so-called semantic gap. Conventional PTMs usually model the images with a dense semantic representation and, in general, one topic space is generated for all the different features. However, this approach fails to consider the sparsity of the semantic representation, the classification quality, as well as the time consumption. In this paper, to solve the above problems, a fully sparse semantic topic model (FSSTM) framework is proposed for HSR imagery scene classification. FSSTM, with an elaborately designed modeling procedure, is able to represent the image with sparse but representative semantics. Based on this framework, the topic weights of multiple features are exploited by solving a concave maximization problem, which improves the fusion of the discriminative semantic information at the topic level. Meanwhile, the sparsity and representativeness of the topics generated by FSSTM guarantee that the image is adaptive to the change of a topic number. FSSTM can consistently achieve a good performance with a limited number of training samples, and is robust for HSR image scene classification. The experimental results obtained with three different types of HSR image data sets confirm that the proposed algorithm is effective in improving the performance of scene classification, and is highly efficient in discovering the semantics of HSR images when compared with the state-of-the-art PTM methods. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Sparse representation based subpixel information extraction framework for hyperspectral remote sensing imageryabstractSparse representation theory has become a powerful tool since it can obtain the sparsest or the unique solution for the underdetermined problem with the development of linear algebra, optimization, scientific computing and more. As subpixel information extraction encountered in hyperspectral remote sensing, which contains many mixed pixels, are famous under-determined ill-posed problem. In addition, there is no unified model to conquer the problems with the subpixel analysis techniques, i.e., spectral unmixing and subpixel mapping. To cope with this under-determined problem, a unified sparse subpixel information extraction framework was proposed in this paper, which connects sparse unmixing and sparse subpixel mapping methods in a unified theoretical system as a serious of sparse regression problem. The experimental results with hyperspectral images indicate that the proposed sparse representation framework outperforms the previous subpixel analysis approaches, hence, provides an effective option for subpixel information extraction idea for hyperspectral remote sensing imagery. Ruyi Feng, Da He, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2016 | Complete dictionary online learning for sparse unmixingabstractSparse unmixing has been successfully applied to hyperspectral remote sensing imagery, based on an available standard spectral library. However, as the number of hyperspectral remote sensors increases, more and more hyperspectral remote sensing images are requiring analysis without the use of a corresponding standard spectral library. To address this problem, sparse unmixing with a complete dictionary online self-learning technique is proposed in this paper. This paper focuses on complete dictionary, which can tackle the unmixing problem with exactly atoms needed in the dataset and online learning means to process the specific data, or the current single hyperspectral remote sensing imagery, at real time. The proposed method addresses the sparse unmixing problem by considering the physical meaning of atoms in the complete dictionary, as well as a non-negative constraint for the abundance. Compared with the classical dictionary learning approaches in sparse representation theory, the experiments with two simulated hyperspectral datasets and a real dataset confirmed the effectiveness of the proposed method. Ruyi Feng, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2016 | Scene semantic classification based on random-scale stretched convolutional neural network for high-spatial resolution remote sensing imageryabstractConvolutional neural network (CNN) has outstanding performance on nature image classification, such as facial recognition, ImageNet Large Scale Visual Recognition Challenge. However, due to scale variation of the same object in scene, it's difficult to directly utilize CNN for remote sensing image classification. In order to solve this problem, scene classification based on a random-scale stretched convolutional neural network (SRSCNN) for HSR remote sensing imagery is proposed in this paper. In the proposed method, the patches with random scale is cropped from image and stretched to the specified scale as input to train CNN, and in order to further improve the performance of CNN, the proposed method classifies an image multiple times to decide its label by voting. Experimental results using two datasets, i.e. the UC Merced dataset, Google Dataset of SIRI-WHU, show better performance than the traditional scene classification methods. Yanfei Zhong, Feng Fei, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2016 | Unsupervised change detection model based on hybrid conditional random field for high spatial resolution remote sensing imageryabstractIn this paper, an unsupervised change detection model based on hybrid conditional random field model (HCRF) is proposed for high spatial resolution (HSR) remote sensing imagery. Traditional random field based algorithms are mainly based on the analysis of the difference image which ignores the spatial-temporal change information of ground objects which is important in dealing with HSR imagery. Thus in HCRF, a new graph structure is designed to explore the correlation of corresponding ground objects from different times to get a better result. The unary potential is selected as the probabilistic result of change vector analysis (CVA), the pairwise potential is modeled to consider the contextual information of difference image and the similarity between objects from bi-temporal original images is considered using an object term. The proposed method is tested on two HSR data sets (IKONOS and QuickBird) and out performs some state-of-art algorithms. Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2016 | Local spectral-spatial clustering for remote sensing imageryabstractRemote sensing image clustering is a challenging task. Recently, by combining the spectral and spatial information of remote sensing data, the clustering accuracy can be dramatically enhanced. However, it has always been difficult to determine the weight parameter for balancing the spectral and spatial terms of the clustering objective function. In this paper, spectral-spatial clustering with a local weight parameter determination method for remote sensing imagery is proposed, i.e. LSSC. In LSSC, considering the large scale of remote sensing images, the weight parameter is determined locally in a patch image instead of the whole image. The local weight parameter is then used in constructing the objective function of LSSC. Thus, the remote sensing image clustering problem is transformed into an optimization problem. Finally, in order to achieve a better optimization performance, a variant of differential evolution (i.e. jDE) is used as the optimizer due to its powerful optimization capability. Experimental results confirm that the proposed LSSC can acquire a higher clustering accuracy than other spectral-spatial clustering methods. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2016 | Multi-object spatial relationship model for high spatial resolution scene classificationabstractMost of the existed scene classification methods classify the scenes ignoring the semantic ground objects for high spatial resolution imagery. Though lots of methods are proposed to recognize the objects, there are lack of methods modeling the spatial relationship between objects for scene classification. Besides, the frequency vector of the semantic objects in the scenes is inadequate to model the spatial relationship. Therefore, to acquire the semantic relationship among multiple objects or between the objects and the scene, this paper developed a multi-object force histogram (MOFH) to model the topology of multiple objects, and proposed a multi-object spatial relationship model (MOSRM) by combining the frequency vector of the ground objects and MOFH for the high spatial resolution scene classification. The experiments show that the proposed method can outperform the scene classification based on the frequency vector of the semantic objects. Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2016 | Hyperspectral image super resolution reconstruction with a joint spectral-spatial sub-pixel mapping modelabstractHyperspectral image super resolution (SR) reconstruction has been studied widely and many algorithms have been proposed. In this paper, a novel super resolution reconstruction method was designed by employing a joint spectral-spatial sub-pixel mapping model which aims to obtain the probabilities of sub-pixels to belong to different land cover classes by dividing mixed pixels into several sub-pixels. Given these sub-pixel probabilities, the resolution enhanced image can be further generated. The proposed approach has been evaluated using both synthetic and real hyperspectral images and compared with other well-known methods. The visual and quantitative comparisons confirm the effectiveness of the proposed method. Xiong Xu 0001, Xiaohua Tong, Jie Li 0022, Huan Xie 0001, Yanfei Zhong, Liangpei Zhang 0001, Dongmei Song |
IGARSS | 5 |
| 2016 | Thermal anomaly detection based on saliency computation for district heating systemabstractThe leaked heat pipeline can be detected as temperature anomalies from the air-borne thermal image. Existing methods of thermal anomaly detection are prone to generate a large quantity of false alarms. Although supervised classification can reduce the false positive rate, it requires years of accumulated training data. In this paper, we use human visual system to improve the detection capabilities of thermal anomaly in district heating system. Leakage candidates are selected from the saliency map created by the thermal image, then buffer analysis with pipeline GIS layer is used to reject false detections. Experimental results show that the proposed method has better performance in detection rate when prior knowledge is scarce, and it is more adaptable to actual circumstances. Xinyu Wang 0003, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2016 | Feature extraction framework in class space for hyperspectral image classificationabstractIn this paper, a novel feature extraction framework is proposed for hyperspectral image classification. Inspired by the role of discriminant function in classifier, which intends to learn a mapping from the input features to label information in class space, we develop a feature extraction framework to learn the new feature representation of original input features in class space, by establishing the relevance between feature extraction and discriminative classifier. The new learning features integrate the input features and the discrimination information of used classifier with available training samples, which reveal the cues of class in class space. Therefore, the new features are called as the features of class-in-class. Several experiments were conducted to illustrate the availability of the proposed features. Ji Zhao 0006, Yanfei Zhong, Rongrong Gao, Liangpei Zhang 0001, Hong Shu |
IGARSS | 2 |
| 2016 | Change Detection Based on a Multifeature Probabilistic Ensemble Conditional Random Field Model for High Spatial Resolution Remote Sensing ImageryabstractIn this letter, a multifeature probabilistic ensemble conditional random field (MFPECRF) model is proposed to perform the task of change detection for high spatial resolution (HSR) remote sensing imagery. MFPECRF not only considers the spectral feature of single pixels but also the interaction between neighborhood pixels and the structural property of the ground objects in HSR imagery to give a higher detection accuracy than the traditional random field methods, which only utilize spectral and label information. In the unary potential, the spectral and morphological features of the difference image are combined using a probabilistic ensemble strategy, and the pairwise potential considers the contextual information of the observed field. The parameters of MFPECRF are estimated using a piecewise strategy, and the final result is obtained by the use of the loopy belief propagation algorithm. The experimental results of two groups of HSR multispectral images confirm the potential of the proposed method in improving the detection accuracy for HSR imagery. Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Hongzan Jiao, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Bag-of-Visual-Words Scene Classifier With Local and Global Features for High Spatial Resolution Remote Sensing ImageryabstractScene classification has been studied to allow us to semantically interpret high spatial resolution (HSR) remote sensing imagery. The bag-of-visual-words (BOVW) model is an effective method for HSR image scene classification. However, the traditional BOVW model only captures the local patterns of images by utilizing local features. In this letter, a local-global feature bag-of-visual-words scene classifier (LGFBOVW) is proposed for HSR imagery. In LGFBOVW, the shape-based invariant texture index is designed as the global texture feature, the mean and standard deviation values are employed as the local spectral feature, and the dense scale-invariant feature transform (SIFT) feature is employed as the structural feature. The LGFBOVW can effectively combine the local and global features by an appropriate feature fusion strategy at histogram level. Experimental results on UC Merced and Google data sets of SIRI-WHU demonstrate that the proposed method outperforms the state-of-the-art scene classification methods for HSR imagery. Qiqi Zhu, Yanfei Zhong, Gui-Song Xia, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Soft computing in remote sensing image processing
Yanfei Zhong, Zexuan Zhu 0001, Yew-Soon Ong |
Soft Comput. | 1 |
| 2016 | Adaptive Sparse Subpixel Mapping With a Total Variation Model for Remote Sensing ImageryabstractSubpixel mapping, which is a promising technique based on the assumption of spatial dependence, enhances the spatial resolution of images by dividing a mixed pixel into several subpixels and assigning each subpixel to a single land-cover class. The traditional subpixel mapping methods usually utilize the fractional abundance images obtained by a spectral unmixing technique as input and consider the spatial correlation information among pixels and subpixels. However, most of these algorithms treat subpixels separately and locally while ignoring the rationality of global patterns. In this paper, a novel subpixel mapping model based on sparse representation theory, namely, adaptive sparse subpixel mapping with a total variation model (ASSM-TV), is proposed to explore the possible spatial distribution patterns of subpixels by considering these subpixels as an integral patch. In this way, the proposed method can obtain the optimal subpixel mapping result by determining the most appropriate subpixel spatial pattern. However, the number of possible spatial configurations of subpixels can increase sharply with large-scale factors, and therefore, in ASSM-TV, the subpixel mapping is considered as a sparse representation problem. A preconstructed discrete cosine transform dictionary, which consists of piecewise smooth subpixel patches and textured patches, is utilized to express the original subpixel mapping observation in a sparse representation pattern. The total variation prior model is designed as a spatial regularization constraint to characterize the relationship between a subpixel and its neighboring subpixels. In addition, a joint maximum a posteriori model is proposed to adaptively select the regularization parameters. Compared with the other traditional and state-of-the-art subpixel mapping approaches, the experimental results using a simulated image, three synthetic hyperspectral remote sensing images, and two real remote sensing images demonstrate that the proposed algorithm can obtain better results, in both visual and quantitative evaluations. Ruyi Feng, Yanfei Zhong, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Semisupervised Subspace-Based DNA Encoding and Matching Classifier for Hyperspectral Remote Sensing ImageryabstractHyperspectral remote sensing images, which are characterized by their high dimensionality, provide us with the capability to accurately identify objects on the ground. They can also be used to identify subclasses of objects. However, these subclasses are usually embedded in different subspaces due to the complex distribution of pixels in the feature space. In the literature, few hyperspectral image classification methods can take both the subclass and subspace into consideration at the same time. Motivated by the fact that natural DNA can distinguish biological subspecies (subclasses in hyperspectral images) using critical DNA fragments (subspaces in hyperspectral images), a semisupervised subspace-based DNA encoding and matching classifier for hyperspectral remote sensing imagery (SSDNA) is proposed in this paper. First, in the process of DNA encoding, the hyperspectral remote sensing image is transformed into a DNA cube, in which the first-order spectral curve of the hyperspectral remote sensing image is utilized in order to take the gradient information of the spectral curve into consideration. Second, in the process of DNA optimization, evolutionary algorithms are used to obtain the best DNA library of the typical objects, which includes the following: 1) A multicenter individual representation is designed in order to consider the existence of subclasses in the hyperspectral remote sensing image; 2) the unlabeled samples are utilized in the process of population initialization and fitness calculation to enhance the diversity of the population and the generalization of the classification performance; and 3) the different classes are embedded in different subspaces. A semisupervised technique is used to extract the subspaces, including the global subspace for all the classes and the local subspace for each class. Three hyperspectral data sets were tested and confirm that SSDNA performs better than the other supervised or semisupervised classifiers. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Dirichlet-Derived Multiple Topic Scene Classification Model for High Spatial Resolution Remote Sensing ImageryabstractDue to the complex arrangements of the ground objects in high spatial resolution (HSR) imagery scenes, HSR imagery scene classification is a challenging task, which is aimed at bridging the semantic gap between the low-level features and the high-level semantic concepts. A combination of multiple complementary features for HSR imagery scene classification is considered a potential way to improve the performance. However, the different types of features have different characteristics, and how to fuse the different types of features is a classic problem. In this paper, a Dirichlet-derived multiple topic model (DMTM) is proposed to fuse heterogeneous features at a topic level for HSR imagery scene classification. An efficient algorithm based on a variational expectation-maximization framework is developed to infer the DMTM and estimate the parameters of the DMTM. The proposed DMTM scene classification method is able to incorporate different types of features with different characteristics, no matter whether these features are local or global, discrete or continuous. Meanwhile, the proposed DMTM can also reduce the dimension of the features representing the HSR images. In our experiments, three types of heterogeneous features, i.e., the local spectral feature, the local structural feature, and the global textural feature, were employed. The experimental results with three different HSR imagery data sets show that the three types of features are complementary. In addition, the proposed DMTM is able to reduce the dimension of the features representing the HSR images, to fuse the different types of features efficiently, and to improve the performance of the scene classification over that of other scene classification algorithms based on spatial pyramid matching, probabilistic latent semantic analysis, and latent Dirichlet allocation. Yanfei Zhong, Gui-Song Xia, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Multiscale and Multifeature Normalized Cut Segmentation for High Spatial Resolution Remote Sensing ImageryabstractIn this paper, a framework for multiscale and multifeature normalized cut (MMNCut) segmentation is proposed for high spatial resolution (HSR) remote sensing images. Normalized cuts (NCuts), as a widely used segmentation method for natural images, can obtain a globally optimized segmentation result corresponding to the optimized partitions of a graph. However, it is difficult to apply the traditional NCuts directly to HSR images because of the huge computational complexity and the diversity of the characteristics of the land covers. In order to solve these problems, the proposed MMNCuts builds a multiscale graph based on superpixels, which can provide powerful grouping cues to guide the segmentation. Generated by different algorithms with varying parameters, superpixels can capture diverse and multiscale visual patterns of HSR images. In addition, the newly constructed graph integrates the multiscale information by considering various connection relationships. Meanwhile, the successful integration of the multifeature cues, including the spectral information, texture information, and structure information, from a large number of superpixels, helps to enhance the expression ability of the graph. Computationally, this leads to a much more efficient algorithm than the traditional NCuts, and in effect, the proposed method achieves a significantly better performance than the traditional approaches. The experimental results with three HSR image data sets demonstrate that the proposed MMNCut algorithm shows a competitive performance in both qualitative and quantitative evaluations when compared with the other state-of-the-art segmentation algorithms for HSR images. Yanfei Zhong, Rongrong Gao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | High-Resolution Image Classification Integrating Spectral-Spatial-Location Cues by Conditional Random FieldsabstractWith the increase in the availability of high-resolution remote sensing imagery, classification is becoming an increasingly useful technique for providing a large area of detailed land-cover information by the use of these high-resolution images. High-resolution images have the characteristics of abundant geometric and detail information, which are beneficial to detailed classification. In order to make full use of these characteristics, a classification algorithm based on conditional random fields (CRFs) is presented in this paper. The proposed algorithm integrates spectral, spatial contextual, and spatial location cues by modeling the probabilistic potentials. The spectral cues modeled by the unary potentials can provide basic information for discriminating the various land-cover classes. The pairwise potentials consider the spatial contextual information by establishing the neighboring interactions between pixels to favor spatial smoothing. The spatial location cues are explicitly encoded in the higher order potentials. The higher order potentials consider the nonlocal range of the spatial location interactions between the target pixel and its nearest training samples. This can provide useful information for the classes that are easily confused with other land-cover types in the spectral appearance. The proposed algorithm integrates spectral, spatial contextual, and spatial location cues within a CRF framework to provide complementary information from varying perspectives, so that it can address the common problem of spectral variability in remote sensing images, which is directly reflected in the accuracy of each class and the average accuracy. The experimental results with three high-resolution images show the validity of the algorithm, compared with the other state-of-the-art classification algorithms. Ji Zhao 0006, Yanfei Zhong, Hong Shu, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Spectral-spatial DNA encoding discriminative classifier for hyperspectral remote sensing imageryabstractHyperspectral remote sensing image classification is one of the most challenging tasks. In our previous work, motivated by the similarity between the structures of DNA and hyperspectral remote sensing images, a DNA matching mechanism was used to transform the hyperspectral remote sensing image into a DNA cube for classification. However, the above DNA encoding strategy lacks the process of encoding accurate spectral and spatial feature into the DNA cube, resulting in unsatisfying classification performance. In this paper, a spectral-spatial DNA encoding strategy for encoding accurate spectral and spatial feature of hyperspectral remote sensing image is proposed. In the spectral dimension, the first-order spectral curve is encoded into the DNA cube, while in the spatial dimension, the principal components or their corresponding texture feature (GLCM) are encoded into the DNA cube. Finally, different with the previous DNA encoding classifier using genetic algorithm (GA), the paper combines the discriminative classifier (i.e. SVM) with spectral-spatial DNA encoding to improve classification performance for hyperspectral remote sensing imagery. The experimental results confirmed the effectiveness of the newly devised DNA encoding strategy and the discriminative classifier in classifying the DNA cube. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2015 | Sub-pixel mapping based on memetic algorithm for hyperspectral imageryabstractIn this paper, sub-pixel mapping based on memetic algorithm (SMMA) is proposed to perform the sub-pixel mapping task, including a global search and a local search. In the global search, a clonal selection algorithm can be used,. To increase the convergence speed, the local search is designed to search for a feasible solution in the neighborhood of the candidate sub-pixel mapping solutions. In addition, a new post-processing method is designed to decrease the effect of the spectral unmixing errors. The experimental results demonstrate that SMMA is superior to the traditional methods. Yanfei Zhong |
IGARSS | 2 |
| 2015 | Spectral-spatial conditional random field classifier with location cues for high spatial resolution imageryabstractIn this paper, we propose a novel spectral-spatial conditional random field classification algorithm with location cues (CRFSS) for high spatial resolution remote sensing imagery. In the CRFSS algorithm, the spectral and spatial location cues are integrated to provide the complementary information from spectral and spatial location perspectives. The spectral cues of different land-cover types are mainly provided by support vector machine (SVM), because of its excellent spectral classification performance. However, it is difficult to deal with the common spectral variability problem in remote sensing images. To alleviate this dilemma, considering the spectral similarity of the same land-cover in a local region, a point-to-point (P2P) classifier is designed to emphasize the spatial location cues. The P2P classifier considers the nonlocal range of the spatial location interactions between the target pixel and its nearest training samples for all the classes. In addition, the pairwise potential of CRFSS also considers the spatial contextual information to favor spatial smoothing. The experimental results showed that the algorithm has a competitive classification performance, in both the quantitative and qualitative evaluation. Ji Zhao 0006, Yanfei Zhong, Hong Shu, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2015 | An Improved Nonlocal Sparse Unmixing Algorithm for Hyperspectral ImageryabstractAs a result of the spatial consideration of the imagery, spatial sparse unmixing (SU) can improve the unmixing accuracy for hyperspectral imagery, based on the application of a spectral library and sparse representation. To better utilize the spatial information, spatial SU methods such as SU via variable splitting augmented Lagrangian and total variation (SUnSAL-TV) and nonlocal SU (NLSU) have been proposed. However, the spatial information considered in these algorithms comes from the estimated abundance maps, which will change along with the iterations. As the spatial correlations of the imagery are fixed and certain, the spatial relationships obtained from the variable abundances are not reliable during the process of optimization. To obtain more precise and fixed spatial relationships, an improved weight calculation NLSU (I-NLSU) algorithm is proposed in this letter by changing the spatial information acquisition source from the variable estimated abundances to the original hyperspectral imagery. A noise-adjusted principal component analysis strategy is also applied for the feature extraction in the proposed algorithm, and the obtained principal components are the foundation of the spatial relationships. The experimental results of both simulated and real hyperspectral data sets indicate that the proposed I-NLSU algorithm outperforms the previous spatial SU methods. Ruyi Feng, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Change Detection Based on Pulse-Coupled Neural Networks and the NMI Feature for High Spatial Resolution Remote Sensing ImageryabstractIn this letter, a change detection algorithm based on pulse-coupled neural networks (PCNN) and the normalized moment of inertia (NMI) feature is proposed for high spatial resolution (HSR) remote sensing imagery. To better analyze a large remote sensing image, the whole image is divided into blocks by the use of a deblocking mechanism. The PCNN model is utilized to obtain the initial binary image, and the NMI feature is calculated based on the binary image to detect the hot spot changed areas. Finally, the changed areas are processed by expectation–maximization to obtain the final change map. The experimental results using QuickBird and IKONOS images demonstrate that the proposed algorithm has the ability to provide better change detection results for HSR images than the traditional PCNN change detection algorithms. Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Adaptive Multiobjective Memetic Fuzzy Clustering Algorithm for Remote Sensing ImageryabstractDue to the intrinsic complexity of remote sensing images and the lack of prior knowledge, clustering for remote sensing images has always been one of the most challenging tasks in remote sensing image processing. Recently, clustering methods for remote sensing images have often been transformed into multiobjective optimization problems, making them more suitable for complex remote sensing image clustering. However, the performance of the multiobjective clustering methods is often influenced by their optimization capability. To resolve this problem, this paper proposes an adaptive multiobjective memetic fuzzy clustering algorithm (AFCMOMA) for remote sensing imagery. In AFCMOMA, a multiobjective memetic clustering framework is devised to optimize the two objective functions, i.e., Jm and the Xie-Beni (XB) index. One challenging task for memetic algorithms is how to balance the local and global search capabilities. In AFCMOMA, an adaptive strategy is used, which can adaptively achieve a balance between them, based on the statistical characteristic of the objective function values. In addition, in the multiobjective memetic framework, in order to acquire more individuals with high quality, a new population update strategy is devised, in which the updated population is composed of individuals generated in both the local and global searches. Finally, to evaluate the proposed AFCMOMA algorithm, experiments using three remote sensing images were conducted, which confirmed the effectiveness of the proposed algorithm. Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Collaborative Active and Semisupervised Learning for Hyperspectral Remote Sensing Image ClassificationabstractHyperspectral image classification is a challenging problem. Among existing approaches to addressing this problem, the active learning (AL) and semisupervised learning (SSL) techniques have attracted much attention in recent years. AL usually involves a labor-intensive human-labeling process while SSL, although avoiding human labeling by assigning pseudolabels to unlabeled data, may introduce incorrect pseudolabels and thus deteriorate classification performance. To overcome these drawbacks, a novel approach named collaborative active and semisupervised learning (CASSL) is proposed in this paper. CASSL combines AL and SSL to invoke a collaborative labeling process by both human experts and classifiers. Specifically, an AL-based pseudolabel verification procedure is performed for gradually improving the pseudolabeling accuracy to facilitate SSL. Meanwhile, only those unlabeled data with low pseudolabeling confidence in SSL will become the query candidates in AL. We evaluate the performance of CASSL on three hyperspectral data sets and compare it with that of two state-of-the-art hyperspectral image classification methods. Experimental results reveal the superiority of CASSL. Lunjun Wan, Ke Tang 0001, Mingzhi Li, Yanfei Zhong, A. K. Qin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Detail-Preserving Smoothing Classifier Based on Conditional Random Fields for High Spatial Resolution Remote Sensing ImageryabstractIn the field of high spatial resolution (HSR) remote sensing imagery classification, object-oriented classification and conditional random field (CRF) approaches are widely used due to their ability to incorporate the spatial contextual information. However, the selection of the optimal segmentation scale in object-oriented classification is not an easy task, and some pairwise CRF models always show an oversmooth performance. In this paper, a detail-preserving smoothing classifier based on conditional random fields (DPSCRF) for HSR imagery is proposed to apply the object-oriented strategy in the CRF classification framework, thus integrating the merits of both approaches to consider the spatial contextual information and preserve the detail information in the classification. The DPSCRF model defines suitable potential functions based on the CRF model for HSR image classification, which comprise the spatial smoothing and local class label cost terms. Both terms favor spatial smoothing in a local neighborhood to consider the spatial information. In addition, the local class label cost also considers the different label information of neighboring pixels at each iterative step in the classification to preserve the detail information. In order to deal with the spectral variability of HSR imagery, a segmentation prior is used by the object-oriented processing strategy. This models the probability of each pixel based on the segmentation regions obtained by the connected-component labeling algorithm. The experimental results with three HSR images demonstrate that the proposed classification algorithm shows a competitive performance in both the quantitative and the qualitative evaluation when compared to the other state-of-the-art classification algorithms. Ji Zhao 0006, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | An Adaptive Subpixel Mapping Method Based on MAP Model and Class Determination Strategy for Hyperspectral Remote Sensing ImageryabstractThe subpixel mapping technique can specify the spatial distribution of different categories at the subpixel scale by converting the abundance map into a higher resolution image, based on the assumption of spatial dependence. Traditional subpixel mapping algorithms only utilize the low-resolution image obtained by the classification image downsampling and do not consider the spectral unmixing error, which is difficult to account for in real applications. In this paper, to improve the accuracy of the subpixel mapping, an adaptive subpixel mapping method based on a maximum a posteriori (MAP) model and a winner-take-all class determination strategy, namely, AMCDSM, is proposed for hyperspectral remote sensing imagery. In AMCDSM, to better simulate a real remote sensing scene, the low-resolution abundance images are obtained by the spectral unmixing method from the downsampled original image or real low-resolution images. The MAP model is extended by considering the spatial prior models (Laplacian, total variation (TV), and bilateral TV) to obtain the high-resolution subpixel distribution map. To avoid the setting of the regularization parameter, an adaptive parameter selection method is designed to acquire the optimal subpixel mapping results. In addition, in AMCDSM, to take into account the spectral unmixing error in real applications, a winner-take-all strategy is proposed to achieve a better subpixel mapping result. The proposed method was tested on simulated, synthetic, and real hyperspectral images, and the experimental results demonstrate that the AMCDSM algorithm outperforms the traditional subpixel mapping methods and provides a simple and efficient algorithm to regularize the ill-posed subpixel mapping problem. Yanfei Zhong, Yunyun Wu, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Scene Classification Based on the Multifeature Fusion Probabilistic Topic Model for High Spatial Resolution Remote Sensing ImageryabstractScene classification has been proved to be an effective method for high spatial resolution (HSR) remote sensing image semantic interpretation. The probabilistic topic model (PTM) has been successfully applied to natural scenes by utilizing a single feature (e.g., the spectral feature); however, it is inadequate for HSR images due to the complex structure of the land-cover classes. Although several studies have investigated techniques that combine multiple features, the different features are usually quantized after simple concatenation (CAT-PTM). Unfortunately, due to the inadequate fusion capacity of k-means clustering, the words of the visual dictionary obtained by CAT-PTM are highly correlated. In this paper, a semantic allocation level (SAL) multifeature fusion strategy based on PTM, namely, SAL-PTM (SAL-pLSA and SAL-LDA) for HSR imagery is proposed. In SAL-PTM: 1) the complementary spectral, texture, and scale-invariant-featuretransform features are effectively combined; 2) the three features are extracted and quantized separately by k-means clustering, which can provide appropriate low-level feature descriptions for the semantic representations; and 3)the latent semantic allocations of the three features are captured separately by PTM, which follows the core idea of PTM-based scene classification. The probabilistic latent semantic analysis (pLSA) and latent Dirichlet allocation (LDA) models were compared to test the effect of different PTMs for HSR imagery. A U.S. Geological Survey data set and the UC Merced data set were utilized to evaluate SAL-PTM in comparison with the conventional methods. The experimental results confirmed that SAL-PTM is superior to the single-feature methods and CAT-PTM in the scene classification of HSR imagery. Yanfei Zhong, Qiqi Zhu, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Remote sensing imagery clustering using an adaptive bi-objective memetic methodabstractDue to the intrinsic complexity of the remote sensing image and the lack of the prior knowledge, clustering for remote sensing image has always been one of the most challenging works in remote sensing image processing. The proposed algorithm constructs a bi-objective memetic-based framework, exploiting the feature space more efficiently. In the framework, two objective functions, Jm and XB, are used as the objective functions for bi-objective optimization. Furthermore, an adaptive local search method which can dynamically adjust its parameter value according to the selection probability has been developed and incorporated into the proposed algorithm. In order to speed the convergence and obtain more non-dominated solutions in the pareto front, a new strategy is newly devised in the local search process, which considers more solutions as the candidate for the next generation. To evaluate the proposed algorithm, some experiments on two multi-spectral images are conducted. The results show that the proposed algorithm can achieve better performance, compared with related methods. Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Non-local euclidean medians sparse unmixing for hyperspectral remote sensing imageryabstractSparse unmixing based on sparse representation theory has been successfully applied to hyperspectral remote sensing imagery. To better utilize the abundant spatial information and improve the unmixing accuracy, spatial sparse unmixing methods such as non-local sparse unmixing (NLSU) have been proposed. Although the NLSU method utilizes the nonlocal spatial information as its spatial regularization term, and obtains a satisfactory unmixing accuracy, the final abundances are affected by the non-local neighborhoods and drift away from the true abundance values when the hyperspectral images are contaminated by strong noise. To solve this problem, a non-local Euclidean medians sparse unmixing (NLEMSU) method is proposed to improve NLSU by replacing the non-local means total variation spatial consideration with non-local Euclidean medians filtering approach. The experimental results using simulated and real hyperspectral images indicate that NLEMSU outperforms the previous sparse unmixing algorithms and, hence, provides an effective option for the unmixing of hyperspectral remote sensing imagery. Ruyi Feng, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2014 | Multi-feature probability topic scene classifier for high spatial resolution remote sensing imageryabstractScene classification can obtain the high-level semantic information in high spatial resolution (HSR) imagery. Probability topic model as a typical scene semantic representation has been successfully applied to nature scene by utilizing a single feature. However, it is not completely fit for HSR images due to the complexity of land cover classes. To solve the problem, multi-feature probability topic scene classifier based on Latent Dirichlet allocation (LDA), namely MFPTSC, is proposed for HSR imagery. In MFPTSC, the spectral, texture, and SIFT features as three representative features are firstly integrated. If the traditional multi-features fusion method (VIS-LDA) is used, which each feature vector is usually stacked at the visual word level, abundant information is lost, which leads to an undesirable classification performance. In this paper, a novel feature fusion strategy at the semantic allocation level, named SAL-LDA, is proposed to avoid information loss to a large extent by mining the latent semantics in accordance with the distinctive characteristics of each feature. Experiment results using the image dataset of 21 land-use classes demonstrate that the multi-feature fusion strategies of VIS-LDA and SAL-LDA both improve the classification accuracy, but the proposed SAL-LDA strategy is better than VIS-LDA. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2014 | A sub-pixel mapping method based on an attraction model for multiple shifted remotely sensed images
Xiong Xu 0001, Yanfei Zhong, Liangpei Zhang 0001 |
Neurocomputing | 2 |
| 2014 | An Adaptive Differential Evolution Endmember Extraction Algorithm for Hyperspectral Remote Sensing ImageryabstractIn this letter, a new endmember extraction algorithm based on adaptive differential evolution (DE) (ADEE) is proposed for hyperspectral remote sensing imagery. In the proposed algorithm, the endmember extraction is transformed into a combinatorial optimization problem through constructing the objective function by minimizing the root mean square error between the original image and its remixed image. DE is utilized to search for the optimal endmember combination in the feasible solution space by the DE operators, such as crossover and mutation, which have the advantage of high efficiency, rapid convergence, and strong capability for global search. In addition, to avoid the problem of parameter selection, an adaptive strategy without user-defined parameters is utilized to improve the classical DE algorithm. The proposed method was tested and evaluated using both simulated and real hyperspectral remote sensing images, and the experimental results show that ADEE can obtain a higher extraction precision than the traditional endmember extraction algorithms. Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | An Unsupervised Spectral Matching Classifier Based on Artificial DNA Computing for Hyperspectral Remote Sensing ImageryabstractHyperspectral remote sensing image clustering, with the large volume, high dimensions, and temporal-spatial spectral diversity, is a challenging task due to finding interesting clusters in the sparse feature space. In this paper, a novel hyperspectral clustering algorithm, namely, an unsupervised spectral matching classifier based on artificial DNA computing (UADSM), is proposed to perform the task of clustering different ground objects in specific spectral DNA feature encoding subspaces. UADSM builds up the clustering framework with the spectral encoding, optimizing, and matching mechanism by introducing the basic notions and operators of artificial DNA computing. By discretized spectral DNA feature encoding processing, the spectral shape, amplitude, and slope features of the hyperspectral data are extracted. Furthermore, the optimal clustering centers in the form of DNA strands can be found by recombining the DNA strands in the spectral DNA encoding subspace. Finally, a reasonable category for each spectral signature is automatically identified by the normalized spectral DNA similarity norm. The traditional clustering methods of k-means, ISODATA, fuzzy c-means classifier, and FCM and MoDEFC after principal component analysis transformation are provided to compare with the UADSM classifier, using Hyperspectral Digital Imagery Collection Experiment and Reflective Optics System Imaging Spectrometer hyperspectral images. The experimental results show that the UADSM classifier can achieve the best classification accuracy; hence, it is considered that the UADSM classifier is an effective clustering method for hyperspectral remote sensing imagery. Hongzan Jiao, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Adaptive Subpixel Mapping Based on a Multiagent System for Remote-Sensing ImageryabstractThe existence of mixed pixels is a major problem in remote-sensing image classification. Although the soft classification and spectral unmixing techniques can obtain an abundance of different classes in a pixel to solve the mixed pixel problem, the subpixel spatial attribution of the pixel will still be unknown. The subpixel mapping technique can effectively solve this problem by providing a fine-resolution map of class labels from coarser spectrally unmixed fraction images. However, most traditional subpixel mapping algorithms treat all mixed pixels as an identical type, either boundary-mixed pixel or linear subpixel, leading to incomplete and inaccurate results. To improve the subpixel mapping accuracy, this paper proposes an adaptive subpixel mapping framework based on a multiagent system for remote-sensing imagery. In the proposed multiagent subpixel mapping framework, three kinds of agents, namely, feature detection agents, subpixel mapping agents and decision agents, are designed to solve the subpixel mapping problem. Experiments with artificial images and synthetic remote-sensing images were performed to evaluate the performance of the proposed subpixel mapping algorithm in comparison with the hard classification method and other subpixel mapping algorithms: subpixel mapping based on a back-propagation neural network and the spatial attraction model. The experimental results indicate that the proposed algorithm outperforms the other two subpixel mapping algorithms in reconstructing the different structures in mixed pixels. Xiong Xu 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Multiagent Object-Based Classifier for High Spatial Resolution ImageryabstractObject-based classification, including object-based segmentation and classification, has been applied for the classification of high spatial resolution imagery due to the increase in the spatial resolution and the limited spectral resolution. Because of the independent design of the object-based segmentation and classification in many of the traditional object-based classification methods, additional work is required to select the appropriate segmentation algorithms to match the classification algorithms. The object-based segmentation algorithms, e.g., the fractal net evolution approach (FNEA), have been successfully utilized to provide the homogeneous regions, and are the basis of object-based classification. However, the traditional FNEA algorithm is greatly influenced by the global control strategy of the region-growing procedure. In addition, the existing object classification methods take little account of the object context information, which is important for high spatial-resolution image interpretation. To improve the accuracy of the object-based classification, in this paper, a multiagent object-based classification framework (MAOCF) for high-resolution remote sensing imagery is proposed. The proposed approach avoids the issue of segmentation algorithm selection by unifying the processing of object-based segmentation and classification through the use of a 4-tuple agent model. In the uniform framework, a multiagent object-based segmentation (MAOS) algorithm is proposed to optimally control the procedure of object merging. In addition, a MAOC is proposed to utilize the contextual information from the surrounding objects by taking advantage of the benefits of a multiagent system, e.g., strong interaction, high flexibility, and parallel global control capability. Due to the characteristics of a multiagent system, MAOCF has the potential for a parallel computing ability. Three experiments with different types of images were performed to evaluate the performance of MAOS and MAOC in comparison to other segmentation and classification algorithms: 1) mean-shift segmentation; 2) FNEA; 3) recursive hierarchical segmentation; and 4) the majority voting object-based classification method. The experimental results demonstrate that MAOS and MAOC give a stable performance with high spatial resolution remote-sensing imagery, and are competitive with the other methods. Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | A Hybrid Object-Oriented Conditional Random Field Classification Framework for High Spatial Resolution Remote Sensing ImageryabstractHigh spatial resolution (HSR) remote sensing imagery provides abundant geometric and detailed information, which is important for classification. In order to make full use of the spatial contextual information, object-oriented classification and pairwise conditional random fields (CRFs) are widely used. However, the segmentation scale choice is a challenging problem in object-oriented classification, and the classification result of pairwise CRF always has an oversmooth appearance. In this paper, a hybrid object-oriented CRF classification framework for HSR imagery, namely, CRF$+$OO, is proposed to address these problems by integrating object-oriented classification and CRF classification. In CRF$+$OO, a probabilistic pixel classification is first performed, and then, the classification results of two CRF models with different potential functions are used to obtain the segmentation map by a connected-component labeling algorithm. As a result, an object-level classification fusion scheme can be used, which integrates the object-oriented classifications using a majority voting strategy at the object level to obtain the final classification result. The experimental results using two multispectral HSR images (QuickBird and IKONOS) and a hyperspectral HSR image (HYDICE) demonstrate that the proposed classification framework has a competitive quantitative and qualitative performance for HSR image classification when compared with other state-of-the-art classification algorithms. Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Adaptive Differential Evolution Fuzzy Clustering Algorithm with Spatial Information and Kernel Metric for Remote Sensing Imagery
Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IDEAL | 2 |
| 2013 | Subspace clustering based on decision fusion strategy for hyperspectral imageryabstractIn this paper, a novel hyperspectral subspace clustering algorithm based on decision fusion strategy (SCDFS) is proposed. Because the different clusters are contained in different subspace of the same hyper-dimensional data, the clustering processing in different subspace is conducted by genetic K-means algorithm (KGA). The clustering results from different subspace can be combined into decision string. The proposed subspace clustering based on decision fusion strategy is conducted on decision string. Considering the selection of subspace, the decision results may be inaccurate. So by the majority voting processing for different subspace, the steady subspace combination can be determined. Finally, the weighted strategy is introduced into SCDFS algorithm to evaluate the distance of different decision string, and determine the fusion clustering result. Hongzan Jiao, Yanfei Zhong, Liangpei Zhang 0001, Pingxiang Li |
IGARSS | 2 |
| 2013 | Hybrid generative/discriminative scene classification strategy based on latent dirichlet allocation for high spatial resolution remote sensing imageryabstractIn order to capture the high-level concepts in high spatial resolution remote sensing (HSR) imagery, scene classification based on a latent Dirichlet allocation (LDA) model, a generative topic model, is a practical method to bridge the semantic gaps between the low-level features and the high-level concepts of HSR imagery. In the previous work, LDA has been considered as a scene classifier, namely C-LDA, and multiple LDA models for each scene class are built separately, where the scene class is determined by a maximum likelihood rule. The C-LDA strategy disregards the correlations between the generative topic spaces of the different scene classes. In this paper, two novel strategies of scene classification based on LDA are proposed to consider the correlations between the generative topic spaces of the different scene classes by sharing the topic spaces for all the scene classes. One of the proposed strategies utilizes LDA as part of the classifier, namely P-LDA, which generates the topic space from all the training images. A discriminative classifier (e.g., support vector machine, SVM) is also employed as the other classification part of P-LDA. The other proposed strategy employs LDA as the topic feature extractor, namely F-LDA, which generates the topic space from all the training and test images, and utilizes a discriminative classifier to classify the topic features. The experimental results using aerial orthophotographs show that the performances of the two proposed strategies for scene classification based on LDA are better than the traditional C-LDA method. Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2013 | Sub-pixel mapping based on artificial immune systems for remote sensing imagery
Yanfei Zhong, Liangpei Zhang 0001 |
Pattern Recognit. | 1 |
| 2012 | Research on image reconstruction based and pixel unmixing based sub-pixel mapping methodsabstractThe sub-pixel mapping technique, which can provide a fine-resolution map of class labels, has attracted more and more attention in recent years. Generally speaking, there are two kinds of methods used to realize the sub-pixel labeling. The first kind are image reconstruction based methods, which first improve the spatial resolution of an image by the super-resolution technique, and then perform a hard classification on the super-resolved image. The second kind are pixel unmixing based methods, where the sub-pixel mapping is implemented based on the results of image unmixing. In this paper, we present a sparse representation method and a back-propagation (BP) neural network method for image reconstruction based and pixel unmixing based mapping, respectively. The advantages and disadvantages of both kinds of methods are analyzed and discussed. Liangpei Zhang 0001, Xiong Xu 0001, Jie Li 0022, Huanfeng Shen, Yanfei Zhong, Xin Huang 0002 |
IGARSS | 5 |
| 2012 | Artificial DNA Computing-Based Spectral Encoding and Matching Algorithm for Hyperspectral Remote Sensing DataabstractIn this paper, a spectral encoding and matching algorithm inspired by biological deoxyribonucleic acid (DNA) computing is proposed to perform the task of spectral signature classification for hyperspectral remote sensing data. As a novel branch of computational intelligence, DNA computing has the strong computing and matching capability to discriminate the tiny differences in DNA strands by DNA encoding and matching in the molecule layer. Similar to DNA discrimination, a hyperspectral remote sensing data matching approach is used to recognize the land cover material from a spectral library or image, according to the rich spectral information. However, it is difficult to apply DNA computing to hyperspectral remote sensing data processing because traditional DNA computing often relies on biochemical reactions of DNA molecules and may result in incorrect or undesirable computations. To utilize the advantages and avoid the problems of biological DNA computing, an artificial DNA computing approach is proposed for spectral encoding and matching for hyperspectral remote sensing data. A DNA computing-based spectral matching approach is used to first transform spectral signatures into DNA codewords by capturing the key spectral features with a spectral feature encoding operation. After DNA encoding, the typical DNA database for interesting classes is constructed and saved by DNA evolutionary operating mechanisms such as crossover, mutation, and structured mutation. During the course of spectral matching, each pixel of the hyperspectral image, or each signature measured in the field, is input to the constructed DNA database. By computing the distance between an unclassified spectrum and the typical DNA codewords from the database, the class property of each pixel is set as the minimum distance class. Experiments using different hyperspectral data sets were performed to evaluate the performance of the proposed artificial DNA computing-based spectral matching algorithm by comparing it with other traditional hyperspectral classifiers, including spectral matching classifiers (binary coding, spectral angle mapper and spectral derivative feature coding (SDFC) matching methods) and a novel statistical method of machine learning termed support vector machine (SVM). Experimental results demonstrate that the proposed algorithm is distinctly superior to the three traditional hyperspectral data classification algorithms. It presents excellent processing efficiency, compared to SVM, with high-dimensional data captured by the Hyperspectral Digital Imagery Collection Experiment sensor, and hence provides an effective option for spectral matching classification of hyperspectral remote sensing data. Hongzan Jiao, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | An Adaptive Artificial Immune Network for Supervised Classification of Multi-/Hyperspectral Remote Sensing ImageryabstractThe artificial immune network (AIN), a computational intelligence model based on artificial immune systems inspired by the vertebrate immune system, has been widely utilized for pattern recognition and data analysis. However, due to the inherent complexity of current AIN models, their application to multi-/hyperspectral remote sensing image classification has been severely restricted. This paper presents a novel supervised AIN-namely, the artificial antibody network (ABNet), based on immune network theory-aimed at performing multi-/hyperspectral image classification. To construct the ABNet, the artificial antibody population (AB) model was utilized. AB is the set of antibodies where each antibody has two attributes-its center vector and recognizing radius-thus each can recognize all antigens within its recognizing radius. In contrast to the traditional AIN model, ABNet can adaptively obtain these two parameters by evolving the antigens without relying on user-defined parameters in the training step. During the process of training, to enlarge the recognizing range, the immune operators (such as clone, mutation, and selection) were used to enhance the AB model to find better antibody in the feature space, which may recognize as much antigen as possible. After the training process, the trained ABNet was utilized to classify the remote sensing image, exhibiting superior learning abilities. Three experiments with different types of images were performed to evaluate the performance of the proposed algorithm in comparison to other supervised classification algorithms: minimum distance, Gaussian maximum likelihood, back-propagation neural network, and our previously developed artificial immune classifiers-resource-limited classification of remote sensing image and multiple-valued immune network classifier. The experimental results demonstrate that ABNet has remarkable recognizing accuracy and ability to provide effective classification for multi-/hyperspectral remote sensing imagery, superior to other methods. Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | Remote Sensing Image Subpixel Mapping Based on Adaptive Differential EvolutionabstractIn this paper, a novel subpixel mapping algorithm based on an adaptive differential evolution (DE) algorithm, namely, adaptive-DE subpixel mapping (ADESM), is developed to perform the subpixel mapping task for remote sensing images. Subpixel mapping may provide a fine-resolution map of class labels from coarser spectral unmixing fraction images, with the assumption of spatial dependence. In ADESM, to utilize DE, the subpixel mapping problem is transformed into an optimization problem by maximizing the spatial dependence index. The traditional DE algorithm is an efficient and powerful population-based stochastic global optimizer in continuous optimization problems, but it cannot be applied to the subpixel mapping problem in a discrete search space. In addition, it is not an easy task to properly set control parameters in DE. To avoid these problems, this paper utilizes an adaptive strategy without user-defined parameters, and a reversible-conversion strategy between continuous space and discrete space, to improve the classical DE algorithm. During the process of evolution, they are further improved by enhanced evolution operators, e.g., mutation, crossover, repair, exchange, insertion, and an effective local search to generate new candidate solutions. Experimental results using different types of remote images show that the ADESM algorithm consistently outperforms the previous subpixel mapping algorithms in all the experiments. Based on sensitivity analysis, ADESM, with its self-adaptive control parameter setting, is better than, or at least comparable to, the standard DE algorithm, when considering the accuracy of subpixel mapping, and hence provides an effective new approach to subpixel mapping for remote sensing imagery. Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2011 | Sub-pixel mapping algorithm based on adaptive differential evolution for remote sensing imageryabstractIn this paper, a sub-pixel mapping algorithm based on differential evolution is proposed, namely adaptive differential evolution sub-pixel mapping algorithm (ADESM). In ADESM, the sub-pixel mapping problem becomes one of assigning land cover classes to the sub-pixels while maximizing the spatial dependence index (SDI). In the proposed ADESM algorithm, individuals are represented as a discrete sub-pixel mapping solution by discrete encoding strategy. These individuals are improved by enhanced evolution operators, e.g., mutation, crossover, repair, exchange, insertion, and an effective local search to generate new candidate solutions. Furthermore, ADESM may adaptively obtain the optimal sub-pixel mapping result without user-defined parameters by adaptive parameter-setting strategy. The proposed method was tested using the synthetic real imagery. Experimental results demonstrate that the proposed approach outperform traditional sub-pixel mapping algorithms, and hence provide an effective option for sub-pixel mapping of remote sensing imagery. Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2010 | Hybrid Detectors Based on Selective EndmembersabstractSubpixel target detection is a challenge in hyperspectral image analysis. As the spatial resolution of hyperspectral imagery is usually limited, subpixel targets only occupy part of the pixel area. In such cases, the spatial characteristics of the targets are hard to acquire, and the only information we can use comes from spectral characteristics. Several kinds of method based on spectral characteristics have been proposed in the past. One is the linear unmixing method, which can provide the abundances of different endmembers in the hyperspectral imagery, including the target abundance. Another focuses on providing statistically reliable rules to separate subpixel targets from their backgrounds. Recently, hybrid detectors combining the aforementioned two methods were put forward, which cannot only figure out the quantitative information of the endmembers but also put this quantitative information into an adaptive matched subspace detector or adaptive cosine/coherent estimate detector to separate the target pixels from the background with statistically reliable rules. However, in these methods, all the endmembers are used to construct the statistical rule, while in most cases only some of the endmembers are actually contained in the pixels. This paper proposes hybrid endmembers selective detectors in which different kinds of endmembers are used according to different pixels to ensure that the true composition of endmembers in each pixel is applied in the detection procedure. Three different types of hyperspectral data were used in our experiments, and our proposed hybrid endmember selective detectors showed better performances than the current hybrid detectors in all the experiments. Liangpei Zhang 0001, Bo Du 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2009 | A Sub-pixel Mapping Algorithm based on Artificial Immune Systems for Remote Sensing ImageryabstractIn this paper, a new sub-pixel mapping method inspired by the clonal selection algorithm (CSA) in artificial immune systems (AIS) is proposed, namely clonal selection subpixel mapping (CSSM). In CSSM, the sub-pixel mapping problem becomes one of assigning land cover classes to the sub-pixels while maximizing the spatial dependence by clonal selection algorithm. CSSM inherits the biologic properties of human immune systems, i.e. clone, mutation, memory, to build a memory-cell population with a diverse set of local optimal solutions. Based on the memory-cell population, CSSM outputs the value of the memory cell and find the optimal sub-pixel mapping result. The proposed method was tested using the synthetic and degraded real imagery. Experimental results demonstrate that the proposed approach outperform traditional sub-pixel mapping algorithms, and hence provide an effective option for sub-pixel mapping of remote sensing imagery. Yanfei Zhong, Liangpei Zhang 0001, Pingxiang Li, Huanfeng Shen |
IGARSS (3) | 1 |
| 2008 | A new sub-pixel mapping algorithm based on a BP neural network with an observation model
Liangpei Zhang 0001, Ke Wu 0004, Yanfei Zhong, Pingxiang Li |
Neurocomputing | 3 |
| 2007 | Dimensionality Reduction Based on Clonal Selection for Hyperspectral ImageryabstractA new stochastic search strategy inspired by the clonal selection theory in an artificial immune system is proposed for dimensionality reduction of hyperspectral remote-sensing imagery. The clonal selection theory is employed to describe the basic features of an immune response to an antigenic stimulus in order to meet the requirement of diversity in the antibody population. In our proposed strategy, dimensionality reduction is formulated as an optimization problem that searches an optimum with less number of features in a feature space. In line with this novel strategy, a feature subset search algorithm, clonal selection Feature-Selection (CSFS) algorithm, and a feature-weighting algorithm, Clonal-Selection Feature-Weighting (CSFW) algorithm, have been developed. In the CSFS, each solution is evolved in binary space, and the value of each bit is either 0 or 1, which indicates that the corresponding feature is either removed or selected, respectively. In CSFW, each antibody is directly represented by a string consisting of integer numbers and their corresponding weights. These algorithms are compared with the following four well-known algorithms: sequential forward selection, sequential forward floating selection, genetic-algorithm-based feature selection, and decision-boundary feature extraction using the hyperspectral remote-sensing imagery acquired by the Pushbroom Hyperspectral Imager and the Airborne Visible/Infrared Imaging Spectrometer, respectively. Experimental results demonstrate that CSFS and CSFW outperform other algorithms and hence provide effective new options for dimensionality reduction of hyperspectral remote-sensing imagery. Liangpei Zhang 0001, Yanfei Zhong, Bo Huang 0001, Jianya Gong, Pingxiang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2007 | A Supervised Artificial Immune Classifier for Remote-Sensing ImageryabstractThe artificial immune network (AIN), which is a new computational intelligence model based on artificial immune systems inspired by the vertebrate immune system, has been widely utilized for pattern recognition and data analysis. However, due to the inherent complexity of current AIN models, their application to remote-sensing image classification has been rather limited. This paper presents a novel supervised classification algorithm based on a multiple-valued immune network, which is a novel AIN model, to perform remote-sensing image classification. The proposed method trains the immune network using the samples of regions of interest and obtains an immune network with memory to classify the remote-sensing imagery. Two experiments with different types of images are performed to evaluate the performance of the proposed algorithm in comparison with other traditional image classification algorithms: Parallelepiped, Minimum Distance, Maximum Likelihood, and Back-Propagation Neural Network. The results evince that the proposed algorithm consistently outperforms the traditional algorithms in all the experiments and, hence, provides an effective option for processing remote-sensing imagery. Yanfei Zhong, Liangpei Zhang 0001, Jianya Gong, Pingxiang Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2006 | An unsupervised artificial immune classifier for multi/hyperspectral remote sensing imageryabstractA new method in computational intelligence namely artificial immune systems (AIS), which draw inspiration from the vertebrate immune system, have strong capabilities of pattern recognition. Even though AIS have been successfully utilized in several fields, few applications have been reported in remote sensing. Modern commercial imaging satellites, owing to their large volume of high-resolution imagery, offer greater opportunities for automated image analysis. Hence, we propose a novel unsupervised machine-learning algorithm namely unsupervised artificial immune classifier (UAIC) to perform remote sensing image classification. In addition to their nonlinear classification properties, UAIC possesses biological properties such as clonal selection, immune network, and immune memory. The implementation of UAIC comprises two steps: initially, the first clustering centers are acquired by randomly choosing from the input remote sensing image. Then, the classification task is carried out. This assigns each pixel to the class that maximizes stimulation between the antigen and the antibody. Subsequently, based on the class, the antibody population is evolved and the memory cell pool is updated by immune algorithms until the stopping criterion is met. The classification results are evaluated by comparing with four known algorithms: K-means, ISODATA, fuzzy K-means, and self-organizing map. It is shown that UAIC is an adaptive clustering algorithm, which outperforms other algorithms in all the three experiments we carried out. Yanfei Zhong, Liangpei Zhang 0001, Bo Huang 0001, Pingxiang Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2005 | Multispectral remote sensing image classification based on simulated annealing clonal selection algorithm
Yanfei Zhong, Liangpei Zhang 0001, Pingxiang Li |
IGARSS | 1 |