VLDB 2026 Research / reviewers in the wild / expert
Ailong Ma
dblp:136/1675
· DBLP profile ↗
78ranked-venue papers
11as first author
50since 2021 · last 2026
0000-0003-3692-6473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 63 · 10 first-author · 37 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DisasterKD: Frequency-guided cross-decoder knowledge distillation for UAV real-time disaster damage assessment
Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
Pattern Recognit. | 4 |
| 2025 | A Global-Local Collaborative and Decomposition-Based Multiobjective Evolutionary Optimization Method for UAV 3-D Path PlanningabstractIn the context of the widespread application of unmanned aerial vehicles (UAVs) across various industries, effective path planning in three-dimensional (3-D) environments has emerged as a crucial challenge in their deployment. In real-world applications, UAV path planning missions are usually converted into multi-objective tasks and solved using evolutionary computation, where the optimal flight path should consider both the overall flight route length and potential terrain threat. However, the existing methods usually treat complete paths as individuals, and this modeling approach lacks the evaluation of track points and is unable to fully reflect the quality of the path. In addition, as the quantity of track points increases, it is difficult for the traditional genetic crossover operator to quickly converge to the global optimum in complex high dimensional objective space. Thus, in this paper, we propose a UAV 3-D path planning method utilizing the global-local collaborative modeling approach with a decomposition-based method (P2GLCM). In the P2GLCM method, the global objective functions and the local objective functions are used to evaluate the path and track points, respectively, to achieve accurate modeling. In addition, to efficiently utilize the high-quality track points in the candidate paths, a dominance relationship approach is introduced to guide the generation of offsprings in a point-by-point manner, improving the search capability in complex objective space. The experimental results on 3-D environments with unified representation of voxels demonstrate that P2GLCM outperforms current methods in convergence and effectiveness. Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Internet Things J. | 3 |
| 2025 | Learning Global Context and Fine Structures for Enhanced Hyperspectral Subpixel MappingabstractSubpixel mapping (SPM) is a crucial technique in remote sensing imagery analysis, aimed at characterizing subpixel distribution within the mixed pixels. Traditional SPM methods and convolutional neural network (CNN)-based SPM methods primarily rely on local spatial autocorrelation, which limits their ability to capture long-range dependencies between distant locations or objects. To address this limitation, we propose a global-local spatial dependence integrator for the SPM method (GLSDSPM) that employs both CNN and the vision transformer as dual-path structures to model global context and local spatial dependencies efficiently. Besides, the previous SPM methods often struggle to accurately reconstruct high-quality spatial patterns for linear features, such as slender rivers and roads, due to insensitivity to textures and sharp, high-frequency details. To overcome this challenge, we integrate a linear pattern refinement module (LPRM) into GLSDSPM, which adaptively focuses on thin and long local structures to accurately capture high-frequency features and detailed information. Two experiments conducted on the Pavia and Houston hyperspectral images prove that the proposed method achieves superior performance, outperforming the state-of-the-art by 4.13% and 3.95% in overall accuracy (OA), respectively. Wen Zhou 0018, Ailong Ma, Da He, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Layout-Anchored Prioritizing Continual Learning for Continuous Building Footprint Extraction From High-Resolution Remote Sensing ImageryabstractContinuous building footprint extraction requires learning new building patterns from remote sensing imagery without forgetting old knowledge. It is inherently challenging due to the spatial layout heterogeneity, which leads to the problem of knowledge forgetting in two aspects: complex background could have distinct patterns (background diversity) and buildings could have similar patterns to the background (foreground-background similarity). To solve the issues, we propose a domain-incremental continual learning algorithm named layout-anchored prioritizing learning network (LAPNet), including a latent layout anchoring module and layout-aware prioritizing learning module. The latent layout anchoring aggregates background information into latent layout features and employs a herding strategy to select representative layout anchors iteratively. This module maintains a memory buffer to narrow the background differences by dynamically discarding unrepresentative experiences and storing layout-anchored experiences. Furthermore, layout-aware prioritizing learning uses these experiences to identify and emphasize the most valuable knowledge for maximizing interclass distance. This module leverages the layout variance metric to measure interclass discrepancies and employs prioritizing learning to reweight the optimization function based on this layout prior. We established a Global-CL dataset to validate the proposed LAPNet framework, containing six study areas across four continents with different remote sensing sensors. Experiments showed that LAPNet achieves state-of-the-art performance in continuous building footprint extraction by effectively correlating knowledge across various domains. The code is available at:https://github.com/Dingyuan-Chen/LAPNet. Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | BASHVS: A Multispectral and SAR Image Fusion Method Based on Bidirectional Aggregation of Saliency in Human Visual SystemabstractThe limitations of remote sensing sensor technology make it difficult to simultaneously capture earth observation information presented in different forms within a single remote sensing image. Acquiring more plentiful target information through fusion technology has therefore remained a research hotspot. The fusion of multispectral (MS) and synthetic aperture radar (SAR) imagery integrates spectral and backscatter information, thereby improving land cover (LC) classification effects. However, current pixel-level fusion methods often fail to adequately account for the model differences between SAR and MS images, leading to spectral-spatial inconsistencies and severe degradation from speckle noise. To address this problem, A fusion method is proposed based on Bidirectional Aggregation of Saliency in the Human Visual System (BASHVS). First, the SAR and MS images are decomposed into base and detail layers using a synchronized anisotropic diffusion algorithm. Subsequently, the detail layer is fused using a New Sum of Modified Anisotropic Laplacian (NSMAL) algorithm. Finally, for base layer fusion, pixel saliency and structural saliency are extracted bidirectionally. The BASHVS is compared with 16 existing fusion methods using 10 evaluation metrics. The results demonstrate that BASHVS achieves the best comprehensive performance and significantly improves the visual quality of the fused images. LC classification using BASHVS fused images shows an average increase of 1.050% in overall accuracy and 0.014 in Kappa coefficient compared to the original MS images, confirming its advantage for LC classification. The source code of BASHVS is shared at https://github.com/CHUANGL8346/BASHVS. Xunqiang Gong, Yichuang Luo, Yonglei Chang, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Disaster-Aware Path Planning Based on Reinforcement Learning for Postearthquake Emergency ResponseabstractAfter earthquake disasters, ensuring that emergency rescue operations reach affected areas as quickly as possible, with the aim of maximizing the rescue of lives and property and minimizing secondary disaster losses, is the primary task of post-earthquake emergency response. Remote sensing technology, characterized by its large-scale coverage, non-contact nature, and rapid response capabilities, provides valuable information for disaster area assessment and post-earthquake emergency response path planning. However, existing post-disaster emergency path planning studies often fail to utilize this rich information. Traditional path planning algorithms insufficiently consider the disaster situation, hindering the efficient utilization of rescue forces and resources. To address these challenges, this study proposes a disaster-aware path planning method based on reinforcement learning for post-earthquake emergency response (P2DARL). The P2DARL method utilizes real disaster information provided by high-resolution remote sensing images and models it in a reinforcement learning environment. This allows the intelligent agent to learn optimal post-earthquake emergency response path planning strategies through continuous trial-and-error interactions with the environment. This approach facilitates the optimal allocation of rescue resources and enhances rescue efficiency. Experiments using real earthquake disaster images demonstrate that the P2DARL method effectively plans paths that cover more affected centers in mid-short distance scenarios, maximizing rescue efficiency and significantly reducing casualties and economic losses resulting from earthquake disasters. Jianchong Guo, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Lightweight Multiscale and Multiattention Hyperspectral Image Classification Network Based on Multistage SearchabstractHyperspectral image (HSI) classification has become a core task in hyperspectral remote sensing interpretation, with deep learning dominating due to its ability to learn hierarchical features without manual engineering. As the model complexity has grown, manual design limitations have prompted a shift to automated approaches such as differentiable architecture search (DARTS), where the architectures are optimized for greater accuracy and efficiency. However, applying gradient-based neural architecture search (NAS) methods directly to hyperspectral classification presents several challenges. Regarding search space design, there is a lack of lightweight operators that can mitigate the spectral variability, spatial heterogeneity, and scale differences inherent in hyperspectral imagery. In terms of search strategy, the traditional DARTS approach directly derives the topology from operation weights, which can lead to suboptimal topological structures, and thus affects the performance of the network in HSI classification. In this article, to address these issues, we propose L3M, which is a lightweight multiscale and multiattention HSI classification network based on multistage search. The proposed approach introduces a novel lightweight operator to address the spectral variability, spatial heterogeneity, and scale differences in HSIs. The operation search and topology search are also decomposed into a multistage process to prevent a suboptimal network by searching for and determining the topological order of the candidate operations in a predefined operation space. L3M was validated on four public datasets, where the proposed model demonstrated a superior classification performance, compared to other lightweight models, while maintaining a low parameter count, low model complexity, and high inference speed. Kefan Li, Yuting Wan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Learning Temporal Consistency for High Spatial Resolution Remote Sensing Imagery Semantic Change DetectionabstractSemantic change detection (SCD) is a crucial task in remote sensing imagery interpretation, which identifies where the changes are and what categories the objects before and after the changes belong to. Post-classification comparison (PCC), as the simplest method, suffers from severe false alarms. While multi-task methods face another issue, in that the changed areas in the binary change map have same object categories in the semantic maps, i.e., semantic change inconsistency. In this study, we leveraged the commutative property of temporal semantic labels in unchanged areas to generate temporal consistency embedding and propose temporal-semantic feature blending (TSFB), which is a feature interaction operation that can dynamically adjust the distance between bi-temporal features by controlling a blending weight. In order to realize adaptive temporal consistency learning, we propose the TEmporal-Semantic feature CalibratiOn (TESCO) module, which can estimate the optimal value of the blending weight for TSFB automatically and make the segmentation network learn temporal consistency from coarse to fine through recurrence and the end-to-end training process. The TESCO module is a plug-and-play module that can be combined with any SCD method. In this study, we added the TESCO module to both PCC and multi-task based SCD methods and conducted comprehensive experiments on three large-scale and different application scenario SCD datasets with different semantic change numbers and time series. The experimental results show that the TESCO module can effectively learn temporal consistency, resulting in a significant improvement in the performance of PCC methods and the semantic consistency of the multi-task based methods. The implementation of TESCO module will be available at https://github.com/Daisy-7/TESCO. Shiqi Tian, Ailong Ma, Zhuo Zheng, Xicheng Tan, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | MSAHiFiC: A Super-Prior Driven High-Fidelity Spectral Attention Network for Hyperspectral Image CompressionabstractTo address the limitations of existing deep learning-based hyperspectral image compression methods in accurately modeling the rate-distortion problem, we propose a coupled multi-scale attention spatial-spectral high-fidelity compression network (MSAHiFiC). MSAHiFiC employs a super-prior network to estimate bitrate and guide rate-distortion optimization, enhancing performance under constrained bitrate conditions. A multi-scale spectral attention module is introduced to capture spectral dependencies across varying inter-band distances and preserve key spectral features during downscaling. A spectral fidelity term is further incorporated into the loss function to improve reconstruction accuracy. Experiments on three benchmark hyperspectral datasets—HySpecNet-11k, XiongAn, and WHU-Hi—demonstrate that MSAHiFiC outperforms state-of-the-art methods by achieving 5% higher spectral fidelity and 6% improvement in reconstruction accuracy under a low bitrate of 0.5 bpp. Yuting Wan, Chao Chen 0029, Ailong Ma, Xunqiang Gong, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Low-Light and Infrared Multimodal Remote Sensing in Nighttime Rescue Mission: A Review of Anomaly Detection MethodsabstractWhen facing natural disasters like sudden floods at night, due to sudden nature disasters, high timeliness of rescue information and complexity and diversity of rescue environment, it is difficult for ground personnel to perform rescue operations in disaster area in time. Remote sensing UAV technology plays an escalating role in disaster relief due to fast response and high flexibility advantages. However, with low nighttime visibility level, complex post-disaster environment, and numerous obstructions, traditional UAV’s visible light remote sensing struggles to achieve accurate rescue detection at night. Therefore, the multimodal detection methods are investigated using low-light and infrared modalities, exploring the integration of data fusion and detection in night rescue applications, and examine the advantages and disadvantages of different anomaly detection methods here. This paper provides the following contributions: 1) a fully-annotated low-light infrared co-observation multimodal remote sensing image dataset for nighttime emergency rescue, termed MRSI-NERD; 2) a benchmark test for most state-of-the-art unsupervised anomaly detection methods to thoroughly explore their capability in extracting useful information from normal samples; and 3) a low-light infrared bimodal fusion method based on frequency domain feature decomposition, which enhances the performance of detectors. The performance of eight types of traditional or deep learning-based detection methods on the MRSI-NERD dataset is reported, including metrics such as ROC-AUC, FPR, TPR, etc. Additionally, a comprehensive analysis of the principles and performance of each category of detection methods is provided. Yuting Wan, Haoyu Yao, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Progressive Symmetric Registration for Multimodal Remote Sensing ImageryabstractImage registration forms the foundation of collaborative processing in multimodal remote sensing imagery (MRSI). However, high-resolution MRSIs frequently display complex distortions due to imaging characteristics and terrain variations, with both global and local distortions present. Effectively addressing these complex distortions necessitates the identification of uniformly and densely distributed corresponding points across the entire image. Existing methods primarily focus on global affine distortions and often extract only sparse and unevenly distributed corresponding points, which makes the effective handling of these coexisting distortions a significant challenge. To address this problem, we propose a progressive symmetric registration learning network (PSRNet) for MRSIs. In PSRNet, multimodal remote sensing image registration (MRSIR) is redefined as a symmetric dense regression task, differing from the traditional pipeline that concentrates on unidirectional sparse transformation parameter prediction. Specifically, PSRNet consists of three primary components: 1) a multiscale feature projector (MFP), which employs a dual-branch structure with nonshared weights to achieve modality-specific representation of different modal images across multiple scales, 2) a progressive cross-modal transformer (PCMT) to further mine modality-invariant features and progressively predict symmetric deformation fields, and 3) a symmetric consistency loss (SCL) function capable of elegantly achieving high-precision reversible alignment of image pairs, encompassing endpoint error loss, bidirectional alignment loss, and smoothness loss. Experimental results demonstrate that PSRNet achieves more comprehensive and advanced registration performance on our self-constructed large-scale high-resolution MRSIR dataset, which includes complex global-local geometric distortions and significant nonlinear radiometric differences (NRD). Heng Yan, Ailong Ma, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question AnsweringabstractEarth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we develop a multi-modal multi-task VQA dataset (EarthVQA) to advance relational reasoning-based judging, counting, and comprehensive analysis. The EarthVQA dataset contains 6000 images, corresponding semantic masks, and 208,593 QA pairs with urban and rural governance requirements embedded. As objects are the basis for complex relational reasoning, we propose a Semantic OBject Awareness framework (SOBA) to advance VQA in an object-centric way. To preserve refined spatial locations and semantics, SOBA leverages a segmentation network for object semantics generation. The object-guided attention aggregates object interior features via pseudo masks, and bidirectional cross-attention further models object external relations hierarchically. To optimize object counting, we propose a numerical difference loss that dynamically adds difference penalties, unifying the classification and regression tasks. Experimental results show that SOBA outperforms both advanced general and remote sensing methods. We believe this dataset and framework provide a strong benchmark for Earth vision's complex analysis. The project page is at https://Junjue-Wang.github.io/homepage/EarthVQA. Zhuo Zheng, Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
AAAI | 4 |
| 2024 | Self-Supporting Adaptive Prototype Learning for Remote Sensing Few-Shot Semantic SegmentationabstractThe segmentation of remote sensing images with few shots is valuable both in theory and application. Many existing few-shot segmentation methods rely on prototype learning, where a single support prototype guides predictions for the query set. However, due to visual differences between the support and query sets, a single support prototype struggles to capture all semantic information effectively. This paper introduces an adaptive self-supporting prototype learning network to address these challenges. We propose adaptive hyper prototype representation (HPR), comprising hyper prototype clustering (HPC) and guided prototype matching (GPM). HPC, a parameter-free and adaptive method, extracts more representative prototypes by clustering similar feature vectors using superpixel features. GPM selects matched prototypes to offer more accurate guidance, ensuring a uniform representation of multiple prototypes and complex semantic information. We also present self-supporting matching (SSM) prototype learning, guiding query set segmentation by obtaining query set prototypes. SSM generates initial pseudo-labels for the query set based on support set prototypes and further guides the query set using its own features, avoiding visual differences between support and query sets. Our adaptive self-supporting prototype learning network significantly enhances prototype quality and outperforms on object-level remote sensing datasets. Weihao Shen, Yanfei Zhong, Ailong Ma |
IGARSS | 3 |
| 2024 | Amalgamating Convolutional and Graph Neural Networks for Fast Multimodal Remote Sensing Image RegistrationabstractMultimodal remote sensing image registration is crucial for the comprehensive utilization of multimodal data. Existing methods demonstrate limited robustness in addressing significant nonlinear radiometric differences and geometric distortions in multimodal remote sensing images. This paper presents ACGNet, a fast registration method that amalgamates convolutional and graph neural networks to address these challenges. ACGNet first uses a deep convolutional network with a dual-branch encoder-decoder architecture to simultaneously extract feature points and corresponding advanced descriptors. Subsequently, a graph neural network based on attention mechanisms is employed for feature matching and outlier rejection, aiming to obtain a large number of high-confidence correspondences, which are crucial for estimating transformation parameters. Experimental results demonstrate that the proposed method exhibits robustness to variations in scale and rotation, and can efficiently and stably perform multimodal remote sensing image registration, outperforming other methods in terms of speed and accuracy. Heng Yan, Ailong Ma, Yanfei Zhong |
IGARSS | 2 |
| 2024 | Historical Product Driven Large-Scale High-Resolution Land Cover and Wetland ClassificationabstractUnderstanding land cover dynamics, especially in wetlands, is crucial for environmental monitoring and ecosystem management, given the threats posed by climate change and human activities. Despite the existence of global-scale land cover maps and detailed wetland thematic maps, a gap remains in their integration, limiting comprehensive ecosystem analysis. Typically, these methods utilize medium-resolution data, which fail to capture the finer details of land cover and wetland interfaces, highlighting the need for high-resolution mapping. As high-resolution mapping becomes more prevalent, the classification of land cover and wetland classes grows increasingly complex. Deep learning methods, particularly semantic segmentation, have emerged as solutions, yet they require extensive training labels, which are challenging and resource-intensive to acquire. This paper proposes a novel framework utilizing historical land cover products as training labels for high-resolution mapping, addressing the challenges of resolution mismatches and semantic errors. The developed cross-resolution mapping dataset, evaluated against a manually labeled test set, demonstrates the framework's effectiveness in providing detailed classification at a provincial scale. Siqi Zeng 0005, Xueying Huang, Yinhe Liu, Ailong Ma, Yanfei Zhong |
IGARSS | 6 |
| 2024 | Decoupling Features for Remote Sensing Missing Modality LearningabstractMissing modality learning aims to enhance the robustness of the model when certain modalities are missing during the testing phase, by transferring knowledge between different modalities during the training stage. Existing works directly reduce the distance between different modalities feature through auxiliary tasks. While simple, this approach shows limited capability in some complex scenarios. It disrupts the distribution of modality-specific features, leading to a decline in modality-discriminative capability. Therefore, we prosed the DFNet. DFNet is designed to decouple the feature into modality-invariant and modality-variant features, by whitening transformation loss. Only the content features of different modalities are brought closer through auxiliary tasks. Experiments are conducted on MSAW datasets, with results indicating that our model outperforms competing methods. Yiheng Zhou, Ailong Ma, Zihang Chen 0001, Yanfei Zhong |
IGARSS | 2 |
| 2024 | Single-Temporal Supervised Learning for Universal Remote Sensing Change Detection
Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Explicable Fine-Grained Aircraft Recognition Via Deep Part Parsing Prior Framework for High-Resolution Remote Sensing ImageryabstractAircraft recognition is crucial in both civil and military fields, and high-spatial resolution remote sensing has emerged as a practical approach. However, existing data-driven methods fail to locate discriminative regions for effective feature extraction due to limited training data, leading to poor recognition performance. To address this issue, we propose a knowledge-driven deep learning method called the explicable aircraft recognition framework based on a part parsing prior (APPEAR). APPEAR explicitly models the aircraft's rigid structure as a pixel-level part parsing prior, dividing it into five parts: 1) the nose; 2) left wing; 3) right wing; 4) fuselage; and 5) tail. This fine-grained prior provides reliable part locations to delineate aircraft architecture and imposes spatial constraints among the parts, effectively reducing the search space for model optimization and identifying subtle interclass differences. A knowledge-driven aircraft part attention (KAPA) module uses this prior to achieving a geometric-invariant representation for identifying discriminative features. Part features are generated by part indexing in a specific order and sequentially embedded into a compact space to obtain a fixed-length representation for each part, invariant to aircraft orientation and scale. The part attention module then takes the embedded part features, adaptively reweights their importance to identify discriminative parts, and aggregates them for recognition. The proposed APPEAR framework is evaluated on two aircraft recognition datasets and achieves superior performance. Moreover, experiments with few-shot learning methods demonstrate the robustness of our framework in different tasks. Ablation analysis illustrates that the fuselage and wings of the aircraft are the most effective parts for recognition. Yanfei Zhong, Ailong Ma, Zhuo Zheng, Liangpei Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Adaptive Self-Supporting Prototype Learning for Remote Sensing Few-Shot Semantic SegmentationabstractThe semantic segmentation of remote sensing images with few shots has important theoretical and application value. Most of the existing few-shot semantic segmentation frameworks are based on prototype learning methods, in which a single support prototype is designed to guide the query set for prediction. However, the visual differences between the support set and the query set make it difficult for a single support prototype, generated from the support set, to comprehensively encapsulate the semantic information of all the query images. This article introduces an adaptive self-supporting prototype learning network designed for few-shot segmentation (FSS), in order to tackle the challenges mentioned earlier. We propose adaptive hyperprototype representation (HPR), which consists of hyperprototype clustering (HPC) and guided prototype matching (GPM), to generate and assign multiple representative prototypes to compensate for the limitations of a single prototype in representing the semantic information of the query images. Specifically, HPC is a parameter-free and adaptive approach, which can extract more representative prototypes by aggregating similar feature vectors utilizing superpixel feature clustering. Meanwhile, GPM can select matched prototypes to provide more accurate guidance, allowing for uniformly aligned representation of multiple prototypes and complex image semantic information. We also introduce self-supporting matching (SSM) prototype learning, which can accurately guide the query set segmentation by acquiring query set prototypes. SSM generates initial pseudo labels for the query set based on the support set prototypes, and further guides the query set using the pseudo labels, along with the query prototypes generated by its own features, thus effectively avoiding visual differences between the support set and query set. The proposed adaptive self-supporting prototype learning network substantially improves the prototype quality and achieves a superior performance on object-level remote sensing datasets. Weihao Shen, Ailong Ma, Zhuo Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Multiobjective Spatiotemporal Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) aims to reconstruct a subpixel-level class distribution map from the pixel-level abundance maps, which is an under-determined problem that has nonunique solutions. To address this, the spatiotemporal SPM uses the abundance, spatial, and temporal constraints to reduce the uncertainty of the mapping solutions, so the spatiotemporal SPM is essentially a constrained optimization problem. However, it is hard to find the optimal weighting parameters to combine the three joint constraints. In addition, the existing spatiotemporal SPM methods mainly use the temporal information either for the unchanged subpixels detection or for the subpixel classification, which is insufficient in the utilization of the temporal information. In this article, a novel spatiotemporal SPM algorithm based on multiobjective optimization (STSPM_MO) is proposed. STSPM_MO is composed of an unchanged subpixels detection stage and a multiobjective spatiotemporal mapping stage. In the former stage, the historical thematic map is used for identifying the unchanged subpixels. In the latter stage, the historical thematic map is further used for providing the temporal dependence, so that the temporal information can be more fully utilized. Moreover, to solve the constrained optimization problem of the spatiotemporal SPM, the abundance, spatial, and temporal constraints are modeled as three objectives and are dynamically fused through the subfitness-based multiobjective evolution, to generate the optimal subpixel classification map. Both synthetic and real-data experiments have been conducted, and the results show the proposed method, and its two variants are superior, stable, and effective. Mi Song, Yanfei Zhong, Ailong Ma, Da He, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | E2SCNet: Efficient Multiobjective Evolutionary Automatic Search for Remote Sensing Image Scene Classification Network ArchitectureabstractRemote sensing image scene classification methods based on deep learning have been widely studied and discussed. However, most of the network architectures are directly reliant on natural image processing methods and are fixed. A few studies have focused on automatic search mechanisms, but they cannot weigh the interpretation accuracy and the parameter quantity for practical application. As a result, automatic global search methods based on multiobjective evolutionary computation have more advantages. However, in the ranking process, the network individuals with large parameter quantities are easy to eliminate, but a higher accuracy may be obtained after full training. In addition, evolutionary neural architecture search methods often take several days. In this article, in order to solve the above concerns, we propose an efficient multiobjective evolutionary automatic search framework for remote sensing image scene classification deep learning network architectures (E2SCNet). In E2SCNet, eight kinds of lightweight operators are used to build a diversified search space, and the coding connection mode is flexible. In the search process, a large model retention mechanism is implemented through two-step multiobjective modeling and evolutionary search, where one step involves the "parameter quantity and accuracy," and the other step involves the "parameter quantity and accuracy growth quantity." Moreover, a super network is constructed to share the weight in the process of individual network evaluation and promote the search speed. The effectiveness of E2SCNet is proven by comparison with several networks designed by human experts and networks obtained by gradient and evolutionary computing-based search methods. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Scalable Multi-Temporal Remote Sensing Change Data Generation via Simulating Stochastic Change ProcessabstractUnderstanding the temporal dynamics of Earth’s surface is a mission of multi-temporal remote sensing image analysis, significantly promoted by deep vision models with its fuel—labeled multi-temporal images. However, collecting, preprocessing, and annotating multi-temporal remote sensing images at scale is non-trivial since it is expensive and knowledge-intensive. In this paper, we present a scalable multi-temporal remote sensing change data generator via generative modeling, which is cheap and automatic, alleviating these problems. Our main idea is to simulate a stochastic change process over time. We consider the stochastic change process as a probabilistic semantic state transition, namely generative probabilistic change model (GPCM), which decouples the complex simulation problem into two more trackable sub-problems, i.e., change event simulation and semantic change synthesis. To solve these two problems, we present the change generator (Changen), a GAN-based GPCM, enabling controllable object change data generation, including customizable object property, and change event. The extensive experiments suggest that our Changen has superior generation capability, and the change detectors with Changen pre-training exhibit excellent transferability to real-world change datasets. Zhuo Zheng, Shiqi Tian, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 3 |
| 2023 | FarSeg++: Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, a fundamental Earth vision task, always suffers from scale variation, the larger intra-class variance of background, and foreground-background imbalance in high spatial resolution (HSR) remote sensing imagery. Generic semantic segmentation methods mainly focus on the scale variation in natural scenarios. However, the other two problems are insufficiently considered in large area Earth observation scenarios. In this paper, we propose a foreground-aware relation network (FarSeg++) from the perspectives of relation-based, optimization-based, and objectness-based foreground modeling, alleviating the above two problems. From the perspective of the relations, the foreground-scene relation module improves the discrimination of the foreground features via the foreground-correlated contexts associated with the object-scene relation. From the perspective of optimization, foreground-aware optimization is proposed to focus on foreground examples and hard examples of the background during training to achieve a balanced optimization. Besides, from the perspective of objectness, a foreground-aware decoder is proposed to improve the objectness representation, alleviating the objectness prediction problem that is the main bottleneck revealed by an empirical upper bound analysis. We also introduce a new large-scale high-resolution urban vehicle segmentation dataset to verify the effectiveness of the proposed method and push the development of objectness prediction further forward. The experimental results suggest that FarSeg++ is superior to the state-of-the-art generic semantic segmentation methods and can achieve a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | An Accurate UAV 3-D Path Planning Method for Disaster Emergency Response Based on an Improved Multiobjective Swarm Intelligence AlgorithmabstractPlanning a practical three-dimensional (3-D) flight path for unmanned aerial vehicles (UAVs) is a key challenge for the follow-up management and decision making in disaster emergency response. The ideal flight path is expected to balance the total flight path length and the terrain threat, to shorten the flight time and reduce the possibility of collision. However, in the traditional methods, the tradeoff between these concerns is difficult to achieve, and practical constraints are lacking in the optimized objective functions, which leads to inaccurate modeling. In addition, the traditional methods based on gradient optimization lack an accurate optimization capability in the complex multimodal objective space, resulting in a nonoptimal path. Thus, in this article, an accurate UAV 3-D path planning approach in accordance with an enhanced multiobjective swarm intelligence algorithm is proposed (APPMS). In the APPMS method, the path planning mission is converted into a multiobjective optimization task with multiple constraints, and the objectives based on the total flight path length and degree of terrain threat are simultaneously optimized. In addition, to obtain the optimal UAV 3-D flight path, an accurate swarm intelligence search approach based on improved ant colony optimization is introduced, which can improve the global and local search capabilities by using the preferred search direction and random neighborhood search mechanism. The effectiveness of the proposed APPMS method was demonstrated in three groups of simulated experiments with different degrees of terrain threat, and a real-data experiment with 3-D terrain data from an actual emergency situation. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Accurate Multiobjective Low-Rank and Sparse Model for Hyperspectral Image Denoising MethodabstractDue to the unavoidable influence of sparse and Gaussian noise during the process of data acquisition, the quality of hyperspectral images (HSIs) is degraded and their applications are greatly limited. It is therefore necessary to restore clean HSIs. In the traditional methods, low-rank and sparse matrix decomposition methods are usually applied to restore the pure data matrix from the observed data matrix. However, due to the fact that the optimization of the${l}_{0}$-norm for the sparse modeling is a nonconvex and NP-hard problem, convex relaxation and regularization parameters are usually introduced. However, convex relaxation often leads to inaccurate sparse modeling results, and the sensitive regularization parameters can lead to unstable results. Thus, in this article, to address these issues, an accurate multiobjective low-rank and sparse denoising framework is proposed for HSIs to achieve accurate modeling. The${l}_{0}$-norm is directly modeled as the sparse noise and is optimized by an evolutionary algorithm, and the denoising problem is converted into a multiobjective optimization problem through simultaneously optimizing the low-rank term, the sparse term, and the data fidelity term, without sensitive regularization parameters. However, since the low-rank clean image and sparse noise of the HSI are encoded into a solution, the length of the solution is too long to be optimized. In this article, a subfitness strategy is constructed to achieve effective optimization by comparing the objective function values corresponding to each band for each solution. The experiments undertaken with simulated images in 11 noise cases and four real noisy images confirm the effectiveness of the proposed method. Yuting Wan, Ailong Ma, Wei He 0003, Yanfei Zhong |
IEEE Trans. Evol. Comput. | 2 |
| 2023 | Multiobjective Memetic Spatiotemporal Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) technology is an effective way to account for the distribution of the component objects within the mixed pixels and can alleviate the mixed pixel problem to some degree. However, the traditional SPM methods rely on only a single coarse-resolution image, and the limited information source can lead to great uncertainty. The rapid development of Earth observation systems has resulted in many fine spatial resolution remote sensing images now being accessible. Spatiotemporal SPM approaches utilize the finer spatial distribution of a historical thematic map to provide temporal prior information for the mapping process. However, spatiotemporal SPM is essentially a constrained optimization problem that aims to predict the optimal class distribution map subject to the abundance, spatial, and temporal constraints. It is a challenging task to properly model and optimize the multiple constraints. In this article, a multiobjective memetic spatiotemporal SPM (MOMSPM) framework is proposed. This model transforms the data fidelity term, spatial prior term, and temporal prior term into a multiobjective optimization problem (MOP) to discard the sensitive regularization parameters. The multiobjective model realizes the fusion of abundance, spatial, and temporal information. To optimize the three objective functions simultaneously, MOMSPM provides a multiobjective memetic algorithm framework in which the global multiobjective search method [multiobjective evolutionary algorithm based on decomposition (MOEA/D)] combines two commonly employed single-objective local search operators (geospatial distribution preference (GSDP) local search and maximum a posteriori (MAP)-based local search). The GSDP and MAP operators are employed to refine the solution and achieve an improved outcome. The hybridized method provides powerful search ability and achieves a good balance between the three objective functions. Experiments on synthetic and real datasets prove that the proposed method is superior to the state-of-the-art spatiotemporal SPM methods. Ailong Ma, Wen Zhou 0018, Mi Song, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Domain Adaptive Land-Cover Classification via Local Consistency and Global DiversityabstractUnsupervised domain adaptive (UDA) land-cover classification has recently gained more and more attention. UDA aimed to learn a model from the annotated source data and the unlabeled target data that can perform well on the target domain. The existing UDA frameworks based on adversarial training and self-training methods have boosted this field a lot. However, these methods almost all originate from the computer vision field, and they ignore the very nature of high-resolution remote sensing (HRS) images. The core insight of this paper is that a good land-cover classification result always has strong local consistency and good global diversity, which makes it possible to construct a metric representing the properties of good land-cover mapping, to improve the existing UDA algorithms. Firstly, based on this finding, we prove that local consistency and global diversity can be measured by the Frobenius norm and nuclear norm, respectively. Secondly, we propose a novel local consistency and global diversity metric (LCGDM), which can be easily integrated into the existing UDA frameworks. Finally, the experiments conducted on the LoveDA data set prove the validity of the proposed metric, which can not only improve the overall land-cover mapping but also the category-wise prediction. Ailong Ma, Chenyu Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Adaptive Multistrategy Particle Swarm Optimization for Hyperspectral Remote Sensing Image Band SelectionabstractHyperspectral remote sensing band selection picks out characteristic feature combination to weaken the strong correlation caused by spectral continuity. However, it is difficult for traditional methods with fixed strategies to search the entire space and make adjustments for the optimization process. Thus, the solutions obtained can be mostly local optima. In this paper, a novel adaptive multi-strategy particle swarm optimization for hyperspectral image remote sensing band selection (AMSPSO_BS) is introduced to obtain a subset solution suitable for classification. The problem is modeled as an effective fitness function, and the quotient of the linear discriminant value and the mean mutual information (LD/MMI) is used to remove the redundancy between bands. The randomly generated solutions are then encoded to form a population, which rely on various particle update strategies (PUS) with different reference positions for updating. During the particle motion, the effect of each strategy on population evolution is considered comprehensively and reflected in the change of selection probability. And the motion parameters are dynamically adjusted to balance the global and local capabilities. Four hyperspectral remote sensing image datasets were utilized to conduct band selection experiments, to confirm the effectiveness of AMSPSO_BS. Yuting Wan, Chao Chen 0029, Ailong Ma, Liangpei Zhang 0001, Xunqiang Gong, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | AutoLC: Search Lightweight and Top-Performing Architecture for Remote Sensing Image Land-Cover ClassificationabstractLand-cover classification has long been a hot and difficult challenge in remote sensing community. With massive High-resolution Remote Sensing (HRS) images available, manually and automatically designed Convolutional Neural Networks (CNNs) have already shown their great latent capacity on HRS land-cover classification in recent years. Especially, the former can achieve better performance while the latter is able to generate lightweight architecture. Unfortunately, they both have shortcomings. On the one hand, because manual CNNs are almost proposed for natural image processing, it becomes very redundant and inefficient to process HRS images. On the other hand, nascent Neural Architecture Search (NAS) techniques for dense prediction tasks are mainly based on encoder-decoder architecture, and just focus on the automatic design of the encoder, which makes it still difficult to recover the refined mapping when confronting complicated HRS scenes.To overcome their defects and tackle the HRS land-cover classification problems better, we propose AutoLC which combines the advantages of two methods. First, we devise a hierarchical search space and gain the lightweight encoder underlying gradient-based search strategy. Second, we meticulously design a lightweight but top-performing decoder that is adaptive to the searched encoder of itself. Finally, experimental results on the LoveDA land-cover dataset demonstrate that our AutoLC method outperforms the state-of-art manual and automatic methods with much less computational consumption. Chenyu Zheng, Ailong Ma, Yanfei Zhong |
ICPR | 3 |
| 2022 | Mae-Net: A Micro Network Architecture Evolutionary Search Method for Remote Sensing Image Scene ClassificationabstractDeep learning based remote sensing scene classification methods have become a research hotspot, but they can not fully mine the image information due to the architecture comes directly from natural image. The automatic search method-based network architecture has then attracted a lot of attention benefits by its ability to independently learn the network structure suitable for remote sensing data. However, in the process of search and sorting, slightly larger models with better performance after full training are often eliminated due to insufficient training. Moreover, the methods often spend a lot of time searching. In this paper, a micro network architecture evolutionary search method is proposed (MAE-Net), the contributions are reflected in the slightly larger model retention mechanism by two-layer multi-objective functions and the super network mechanism used to reduce search time through weight sharing. The effectiveness is proved by comparison with human expert and search based networks on NWPU45 dataset. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2022 | Capformer: Pure Transformer for Remote Sensing Image CaptionabstractAccurately describing high-spatial resolution remote sensing images requires the understanding the inner attributes of the objects and the outer relations between different objects. The existing image caption algorithms lack the ability of global representation, which are not fit for the summarization of complex scenes. To this end, we propose a pure transformer (CapFormer) architecture for remote sensing image caption. Specifically, a scalable vision transformer is adopted for image representation, where the global content can be captured with multi-head self-attention layers. A transformer decoder is designed to successively translate the image features into comprehensive sentences. The transformer decoder explicitly model the historical words and interact with the image features using cross-attention layers. The comprehensive and ablation experiments on RSICD dataset demonstrate that the CapFormer outperforms the state-of-the-art image caption methods. Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
IGARSS | 3 |
| 2022 | HFGAN: A Heterogeneous Fusion Generative Adversarial Network for Sar-to-Optical Image TranslationabstractDue to the influence of the imaging mechanism of SAR images, it is difficult to interpret ground information through SAR images without expert knowledge. On the contrary, optical images have rich spatial and color information, so it is necessary to conduct research on the translation of SAR to optical remote sensing images. In this end, we propose a heterogeneous fusion generative adversarial network (HFGAN) for SAR-to-optical image translation. There are two main improvements: (1) Complementary generation of global structure and texture information. A heterogeneous fusion generator and a multi-scale discriminator are proposed to ensure that the global and detailed features of the generated image are more accurate and rich. (2) Color fidelity. Chromatic aberration loss are introduced to reduce the color difference between the generated image and the real optical image. Through qualitative and quantitative experiments, it is proved that the proposed method not only obtains better visual effects, but also has certain progress in the evaluation metrics, which proves that the proposed method is superior to the previous advanced methods in SAR-to-optical image translation. Ailong Ma, Yanfei Zhong, Xiaodong Gong |
IGARSS | 2 |
| 2022 | TypeFormer: Multiscale Transformer With Type Controller for Remote Sensing Image CaptionabstractImage captioning in remote sensing can help us understandthe inner attributes of the objects and the outer relations between different objects. However, the existing image captioning algorithms lack the ability of global representation, and cannot obtain object relations over long distances. In addition, these algorithmics generate captions randomly without consideration of the specific demands. To this end, we propose a pure transformer architecture with caption type controller for remote sensing image captioning. Specifically, a multi-scale vision transformer is adopted for the image representation, where the global and detailed content can be captured with multi-head self-attention layers. A transformer decoder is then introduced to successively translate the image features into comprehensive sentences. The optional block called the caption type controller is designed to consider the types of captions through caption type matrix sets according to the demands, embedding the learnable sentence feature with the required type. The comparison and ablation experiments conducted on the Remote Sensing Image Captioning Dataset (RSICD) dataset demonstrate that the proposed framework outperforms the current state-of-the-art image captioning methods. The experiments conducted on the FloodNet caption dataset further illustrate that the proposed methods can effectively generate specific types of captions. Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Multiobjective Sine Cosine Algorithm for Remote Sensing Image Spatial-Spectral ClusteringabstractRemote sensing image data clustering is a tough task, which involves classifying the image without any prior information. Remote sensing image clustering, in essence, belongs to a complex optimization problem, due to the high dimensionality and complexity of remote sensing imagery. Therefore, it can be easily affected by the initial values and trapped in locally optimal solutions. Meanwhile, remote sensing images contain complex and diverse spatial-spectral information, which makes them difficult to model with only a single objective function. Although evolutionary multiobjective optimization methods have been presented for the clustering task, the tradeoff between the global and local search abilities is not well adjusted in the evolutionary process. In this article, in order to address these problems, a multiobjective sine cosine algorithm for remote sensing image data spatial-spectral clustering (MOSCA_SSC) is proposed. In the proposed method, the clustering task is converted into a multiobjective optimization problem, and the Xie-Beni (XB) index and Jeffries-Matusita (Jm) distance combined with the spatial information term (SI_Jm measure) are utilized as the objective functions. In addition, for the first time, the sine cosine algorithm (SCA), which can effectively adjust the local and global search capabilities, is introduced into the framework of multiobjective clustering for continuous optimization. Furthermore, the destination solution in the SCA is automatically selected and updated from the current Pareto front through employing the knee-point-based selection approach. The benefits of the proposed method were demonstrated by clustering experiments with ten UCI datasets and four real remote sensing image datasets. Yuting Wan, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Cybern. | 2 |
| 2022 | A Decomposition-Based Multiobjective Clonal Selection Algorithm for Hyperspectral Image Feature SelectionabstractFeature selection is an effective way to handle the strong correlation of hyperspectral image data by screening the significant features, and is generally accepted to be a multiobjective optimization problem. Nevertheless, due to the randomness of the strategies and the ambiguity of the optimization directions, the existing multiobjective evolutionary optimization based feature selection methods can suffer from inefficient search and loss of search space with promising solutions when faced with the high-dimensional and multi-peak search space. The multiobjective evolutionary algorithm based on decomposition (MOEA/D) employs a decomposition framework to provide exact guidance for the optimization directions. Unfortunately, random operators are still used, leading to inadequate local optimization. Thus, evolutionary strategies with search preference such as clonal selection may be necessary for local search. In this paper, a novel decomposition-based multiobjective clonal selection algorithm for feature selection (MOCSA/D_FS) is proposed to obtain a feature subset with a superior classification performance. In MOCSA/D_FS, the information entropy and the ratio of the relative scatter value and mutual information are utilized as two objective functions to evaluate the information amount and redundancy. A series of subproblems are then obtained by decomposing the multiobjective problem through weight vectors, with anl2-norm constraint used to balance the search space. Subsequently, a clonal selection method with search space preference performs a detailed local search on each subproblem, which can fully exploit the potential optimal space. The effectiveness and generalizability of the proposed method was confirmed by experiments on four hyperspectral remote sensing image datasets. Chao Chen 0029, Yuting Wan, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Cross-Modality Image Matching Network With Modality-Invariant Feature Representation for Airborne-Ground Thermal Infrared and Visible DatasetsabstractThermal infrared (TIR) remote-sensing imagery can allow objects to be imaged clearly at night through the long-wave infrared, so that the fusion of thermal infrared and visible (VIS) imagery is a way to improve the remote-sensing interpretation ability. However, due to the large radiation difference between the two kinds of images, it is very difficult to match them. One of the most important issues is the lack of comprehensive consideration of the modality-specific information and modality-shared information, which makes it difficult for the existing methods to obtain a modality-invariant feature representation. In this article, a cross-modality image matching network, which we refer to as CMM-Net, is proposed to realize thermal infrared and visible image matching by learning a modality-invariant feature representation. First, in order to extract the modality-specific features of the imagery, the framework constructs a shallow two-branch network to make full use of the modality-specific information, without sharing parameters. Second, in order to extract the high-level semantic information between the different modalities, modality-shared layers are embedded into the deep layers of the network. In addition, three novel loss functions are designed and combined to learn the modality-invariant feature representation, that is, the discriminative loss of the non-corresponding features in the same modality, the cross-modality loss of the corresponding features between different modalities, and the cross-modality triplet (CMT) loss. The multimodal matching experiments conducted with ground- and airborne-based thermal infrared images and visible images showed that the proposed method outperforms the existing image matching methods by about 2% and 6% for the ground and airborne images, respectively. Ailong Ma, Yuting Wan, Yanfei Zhong, Bin Luo 0005, Miaozhong Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | MAP-Net: SAR and Optical Image Matching via Image-Based Convolutional Network With Attention Mechanism and Spatial Pyramid Aggregated PoolingabstractThe complementarity of synthetic aperture radar (SAR) and optical images allows remote sensing observations to “see” unprecedented discoveries. Image matching plays a fundamental role in the fusion and application of SAR and optical images. However, both the geometric imaging pattern and the physical radiation mechanism of these two sensors are significantly different, so that the images show complex geometric distortion and nonlinear radiation differences. This phenomenon brings great challenges to image matching, which neither the handcrafted descriptors nor the deep learning-based methods have adequately addressed. In this article, a novel image-based matching method for SAR to optical images via an image-based convolutional network with spatial pyramid aggregated pooling (SPAP) and an attention mechanism is proposed, namely MAP-Net. The original image is embedded through the convolutional neural network to generate the feature map. Through the information extraction and abstraction of the original imagery, the embedded features containing the high-level semantic information are more robust to the geometric distortion and radiation variation among the different modal images, which is beneficial to the matching of cross-modal images. The adoption of the SPAP module makes the network more capable of integrating global and local contextual information. The attention block weights the dense features generated from the network to extract the key features that are invariant, distinguishable, repeatable, and suitable for the image matching task. In the experiments, five sets of multisource and multiresolution SAR and optical images with wide and varied ground coverage were used to evaluate the accuracy of MAP-Net, compared to both handcrafted and deep learning-based methods. The experimental results show that the MAP-Net method is superior to the current state-of-the-art image matching methods for SAR to optical images. Ailong Ma, Liangpei Zhang 0001, Miaozhong Xu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Cascaded Multi-Task Road Extraction Network for Road Surface, Centerline, and Edge ExtractionabstractRoad extraction from very high-resolution (VHR) remote sensing imagery remains a huge challenge, due to the shadows and occlusions of trees and buildings. Such complex backgrounds result in deep networks often producing fragmented roads with poor connectivity. Road extraction has three typical tasks: road surface segmentation (SS), centerline extraction (CE), and edge detection (ED), which are conducted in a wide range of real applications. Also, the three tasks have a symbiotic relationship, i.e., the road SS determines the location of the centerline and edges, and the CE and ED can allow the generation of more continuous road surfaces. However, most of the previous works have completed these three tasks separately, without exploiting the symbiotic relationship between them to boost the road connectivity. In this article, in order to improve road connectivity, a cascaded multitask (CasMT) road extraction framework for simultaneously extracting the road surface, centerline, and edges is proposed. In the proposed framework, topology-aware learning is applied to capture the long-distance topological relationships, and hard example mining (HEM) loss is employed to focus more on hard samples, to further enhance the road completeness. Extensive experiments were conducted on the DeepGlobe road dataset and a large-scale road dataset (called the LSCC dataset) from the three Chinese cities of Beijing, Shanghai, and Wuhan. The experimental results obtained on the public DeepGlobe dataset demonstrate that the proposed CasMT framework can significantly outperform the current state-of-the-art method. Moreover, the generalization capability of the model was verified on the LSCC dataset, where the proposed CasMT framework achieved the best performance in the average path length similarity (APLS) road topology metric, which further confirms the superiority of the proposed framework. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | FactSeg: Foreground Activation-Driven Small Object Semantic Segmentation in Large-Scale Remote Sensing ImageryabstractThe small object semantic segmentation task is aimed at automatically extracting key objects from high-resolution remote sensing (HRS) imagery. Compared with the large-scale coverage areas for remote sensing imagery, the key objects, such as cars and ships, in HRS imagery often contain only a few pixels. In this article, to tackle this problem, the foreground activation (FA)-driven small object semantic segmentation (FactSeg) framework is proposed from perspectives of structure and optimization. In the structure design, FA object representation is proposed to enhance the awareness of the weak features in small objects. The FA object representation framework is made up of a dual-branch decoder and collaborative probability (CP) loss. In the dual-branch decoder, the FA branch is designed to activate the small object features (activation) and suppress the large-scale background, and the semantic refinement (SR) branch is designed to further distinguish small objects (refinement). The CP loss is proposed to effectively combine the activation and refinement outputs of the decoder under the CP hypothesis. During the collaboration, the weak features of the small objects are enhanced with the activation output, and the refined output can be viewed as the refinement of the binary outputs. In the optimization stage, small object mining (SOM)-based network optimization is applied to automatically select effective samples and refine the direction of the optimization while addressing the imbalanced sample problem between the small objects and the large-scale background. The experimental results obtained with two benchmark HRS imagery segmentation datasets demonstrate that the proposed framework outperforms the state-of-the-art semantic segmentation methods and achieves a good tradeoff between accuracy and efficiency. Code will be available at:http://rsidea.whu.edu.cn/FactSeg.htm Ailong Ma, Yanfei Zhong, Zhuo Zheng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Supervised Progressive Growing Generative Adversarial Network for Remote Sensing Image Scene ClassificationabstractRemote sensing image scene classification is a challenging task. With the development of deep learning, methods based on convolutional neural networks (CNNs) have made great achievements in remote sensing image scene classification. Since the training of a CNN requires a large number of labeled samples, a generative adversarial network (GAN) for sample generation represents a new opportunity to solve the problem of the limited samples. However, most of the existing GAN-based sample generation methods can only generate unlabeled samples, instead of samples labeled with the corresponding scene category. In this article, to solve the problem, a supervised progressive growing generative adversarial network (SPG-GAN) is proposed for remote sensing image scene classification. The proposed method can generate labeled samples for the remote sensing image scene classification, significantly improving the classification accuracy in the case of limited samples. The SPG-GAN method has two main improvements. First, a conditional generative framework for labeled samples is proposed, in which the label information is added in the channel dimension as the input. By considering the constraints of the label information in the loss function, the network can be trained in the direction of a specific category. As a result, the network can generate remote sensing image scene classification samples with label categories. Second, a progressive growing sample generation method is introduced. In order to ensure that the generated samples have more spatial details, they are generated by progressively adding modules to the generator and discriminator, thereby ensuring that the generated sample is of better quality. After testing on two benchmark datasets and carrying out a large-scale experiment in the central area of the city of Wuhan in China, it was found that the proposed method can obtain a superior scene classification accuracy in the case of limited samples. Ailong Ma, Zhuo Zheng, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Joint Spectral Unmixing and Subpixel Mapping Framework Based on Multiobjective OptimizationabstractConventional subpixel mapping (SPM) is performed based on the abundance maps obtained by spectral unmixing (SU), to interpret the mixed pixels and improve the mapping resolution for hyperspectral remote-sensing imagery. However, the SU and SPM tasks are separately conducted, so that the unmixing error is propagated to SPM, and the mapping result is strongly reliant on the quality of the abundance maps. In this article, a novel joint SPM and SU framework (MO_SUSM) based on multiobjective optimization is proposed to simultaneously perform unmixing and mapping. Specifically, the multiobjective joint optimization model with a data fidelity term and a Laplacian prior term is constructed for SU and SPM. For the data fidelity term, since the unmixing result can be recovered by downsampling the mapping result, the unmixing model is joined with the mapping model by the downsampling matrix, so that the reconstruction errors of the unmixing and mapping results can be minimized together. Meanwhile, the Laplacian prior term is used to maximize the spatial dependence of the mapping result and provide the spatial constraint for SU. In addition, the multiobjective optimization algorithm with local search is designed to search for the optimal unmixing and mapping results that can balance the objective terms. Since the two objective terms are dynamically integrated during optimization, there is no need to set sensitive weights for the objectives combination. Four experiments were conducted on hyperspectral images of various data sources, including ground, airborne, and satellite images. The unmixing results show that MO_SUSM can reduce the unmixing error and can improve the quality of the abundance maps. The mapping results show that MO_SUSM can alleviate the dependence of SPM on the abundance maps and can improve the mapping accuracy. Mi Song, Yanfei Zhong, Ailong Ma, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Three-Dimensional Change Detection in Urban Areas Based on Complementary Evidence FusionabstractWith the acceleration of urbanization, it is essential to carry out change detection (CD) and obtain surface change information in urban areas. In the early stages, the spectral information of remote sensing images was used as a change index to capture the spectral and texture changes of ground objects in a two-dimensional plane. However, due to the dense buildings in urban areas, shadows and occlusions can be easily formed, and spectral information is sensitive to the imaging environment, such as the illumination, atmospheric conditions, and imaging angles, so that the detection results based on remote sensing images can often be incomplete. Most changes include not only 2-D plane changes but also 3-D elevation changes. Compared with spectral information, the elevation is more stable and more resistant to interference. Therefore, the fusion of remote sensing image and digital surface model (DSM) data has the potential to be used to detect the changes in urban areas. In this article, we propose a complementary evidence fusion 3-D CD framework based on the Dempster–Shafer theory (CDST). In this framework, DSM and normalized difference vegetation index (NDVI) data are combined using a complementary evidence combination rule. The DSM data can effectively overcome the impact of shadows, and the NDVI data can capture the relevant changes of height-insensitive ground objects, such as vegetation and water. When mapping the basic probability assignment (BPA) of the difference image (DI), prior knowledge is used to ensure that the BPA is not affected by the data distribution. Since DSM and remote sensing image data are heterogeneous data, there is a high degree of conflict when representing the change information characteristics of specific areas. For example, the change between grassland and road is small in elevation but significant in the spectral details, and the traditional Dempster’s combination rule no longer applies. The proposed CDST framework uses a complementary evidence combination rule, which can effectively alleviate the conflicts between the evidence sources and improve the integrity of the detected changes. The experimental results obtained on real datasets confirm that the proposed method does indeed perform well. Shiqi Tian, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Domain Adaptation via a Task-Specific Classifier Framework for Remote Sensing Cross-Scene ClassificationabstractThe scene classification of high spatial resolution (HSR) imagery involves labeling an HSR image with a specific high-level semantic class according to the composition of the semantic objects and their spatial relationships. As such, scene classification has attracted increased attention in recent years, and many different algorithms have now been proposed for the cross-scene classification task. However, the recently proposed scene classification methods based on deep convolutional neural networks (CNNs) still suffer from domain shift problems, because of the training data and validation data not following the assumption of independent and identical distributions. The employment of generative adversarial networks has been found to be an effective way to bridge the domain shift/gap. However, the existing cross-scene classification methods do not use the classification information in the target domain, and the domain classifier is task-independent for different scene classification tasks. In this article, to solve this problem, domain adaptation via a task-specific classifier (DATSNET) framework is proposed for HSR image scene classification. Task-specific classifiers and minimizing and maximizing, “ i.e., minimaxing,” of the classifier discrepancy are integrated in the DATSNET framework. The task-specific classifiers are proposed to align the distributions of the source domain features and target domain features by utilizing task-specific decision boundaries in the target domain. In order to align the two task-specific classifiers’ feature distributions, minimaxing the defined discrepancy between the different classifiers in an adversarial manner is proposed to obtain better task-specific classifier boundaries in the target domain and a better-aligned feature distribution in both domains. The experimental results obtained with different remote sensing cross-scene classification tasks demonstrate that the proposed method achieves a significantly improved performance compared with the other state-of-the-art remote sensing cross-scene classification algorithms. Zhendong Zheng, Yanfei Zhong, Ailong Ma |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing ImageryabstractFor high spatial resolution (HSR) remote sensing images, bitemporal supervised learning always dominates change detection using many pairwise labeled bitemporal images. However, it is very expensive and time-consuming to pairwise label large-scale bitemporal HSR remote sensing images. In this paper, we propose single-temporal supervised learning (STAR) for change detection from a new perspective of exploiting object changes in unpaired images as supervisory signals. STAR enables us to train a high-accuracy change detector only using unpaired labeled images and generalize to real-world bitemporal images. To evaluate the effectiveness of STAR, we design a simple yet effective change detector called ChangeStar, which can reuse any deep semantic segmentation architecture by the ChangeMixin module. The comprehensive experimental results show that ChangeStar outperforms the baseline with a large margin under single-temporal super-vision and achieves superior performance under bitemporal supervision. Code is available at https://github.com/Z-Zheng/ChangeStar. Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 2 |
| 2021 | Large-Scale Urban Road Vectorization Mapping Via A Road Node Proposal Network for High-Resolution Remote Sensing ImageryabstractUrban road vectorization mapping can reflect the urban development of cities, which consists of two separate tasks: road extraction and road vectorization. Most of the current road vectorization mapping methods focus on road extraction yet ignoring the importance of road vectorization, facing the problem of road connectivity. In this work, to implement urban road vectorization mapping in a unified way, a novel road vectorization mapping framework is proposed. The proposed framework consists of a node proposal network (NPN) module and a node connectivity based road refinement module. In the NPN module, a node proposal head is adopted, which improves the connectivity of the road mask by providing supervision of the road nodes, which are actually part of the road mask. The road mask is then converted into a road vector map by vectorization. In the node connectivity based road refinement module, road nodes are inserted into the road vector map to improve the connectivity. The experimental results of two public datasets (SpaceNet 3 and DeepGlobe) confirm the advantages of the proposed framework. The experiments for two areas in Shanghai and Wuhan demonstrate the generalization ability of the proposed framework. Yanfei Zhong, Ailong Ma |
IGARSS | 3 |
| 2021 | Pointnet: Learning Point Representation for High-Resolution Remote Sensing Imagery Land-Cover ClassificationabstractWith the development of remote sensing, a large amount of high-spatial-resolution (HSR) images is available, which makes refined land-cover mapping possible. However, the details of ground objects in HSR images are complex, especially in edges, therefore brings new challenges in land-cover classification. Existing deep learning method views it as a semantic segmentation task based on the fully convolutional networks (FCN), taking no account of complex details recognition. In this paper, we tackle this problem by proposing a point representation network (PointNet) for HSR land-cover classification Specifically, the uncertain point selection is designed for finding the most uncertain details at the end of ResNet encoder. According to these points, the coarse and fine features in the encoder are fused, followed by a multilayer perceptron (MLP). Different from convolutional sampling, the MLP focuses on the recognition of uncertain points, which modifies the network optimization during training. During the prediction, the coarse features are successively upsampled with the points refining, improving the performance on land-cover details recognition. Experimental results on a Nanjing land-cover dataset demonstrate that the PointNet outperforms the state-of-the-art methods. Longyuan Ding, Chenyu Zheng, Ailong Ma |
IGARSS | 8 |
| 2021 | Sensor-Specific Adversarial Network for Transferable Land-Cover ClassificationabstractAs the multi-source high-spatial-resolution (HSR) images are being daily acquired from different sensors, it brings the challenge of transferring the recognition model from labeled images to new unlabelled images obtained from other sensors. Existing deep transfer learning methods encode the land-cover features in the same architecture, which ignores the sensor divergence. In this paper, we tackle this problem by proposing a sensor-specific adversarial network for HSR land-cover classification. Specifically, the sensor-specific normalization (SN) is designed for decoupling the sensor divergence in different normalization weights. Moreover, the transferable adversarial optimization is proposed for effectively optimizing the source-related, target-related, and discriminator weights. Considering the sensor-specific characteristics, our proposed method improves the transferability of deep learning models between airborne and spaceborne sensors. The mutual transferability experiments on a self-constructed cross-sensor land-cover dataset demonstrate that the proposed method outperforms the state-of-the-art deep transfer learning methods. Yanfei Zhong, Zhuo Zheng, Ailong Ma |
IGARSS | 4 |
| 2021 | Weakly Supervised Semantic Change Detection via Label Refinement FrameworkabstractSemantic change detection is a meaningful but challenging task in the remote sensing community. The currently dominant approaches are mainly based on deep learning. However, the lack of high-resolution annotations is the main bottleneck for semantic change detection at scale when using these state-of-the-art deep learning models. In this paper, the label refinement framework is proposed for weakly-supervised semantic change detection, which allows the deep network to learn from low-resolution labels and produce high-resolution semantic change maps, thus alleviating the data-hungry problem. This framework contains four parts: coarse label training’ pseudo-label refinement, multitask change detection and post-process. The experimental results on 2021 IEEE GRSS Data Fusion Contest Track MSD dataset confirmed the effectiveness of the proposed method. Additionally, our method wins 4th place in the 2021 IEEE GRSS Data Fusion Contest Track MSD (DFC21-MSD). Zhuo Zheng, Yinhe Liu, Shiqi Tian, Ailong Ma, Yanfei Zhong |
IGARSS | 5 |
| 2021 | Multiscale U-Shaped CNN Building Instance Extraction Framework With Edge Constraint for High-Spatial-Resolution Remote Sensing ImageryabstractBuilding extraction based on high-resolution remote sensing imagery has been widely used in automatic surveying and mapping. However, few methods have been developed for building instance extraction, i.e., extracting each building's footprint separately, which is required in a number of applications, such as the smallest unit of a cadastral database. In building instance extraction, there are two challenges: 1) buildings with various scales exist in the imagery and 2) precise building footprints are difficult to extract due to the blurry boundaries. In this article, to solve these problems, a multiscale U-shaped convolutional neural network building instance extraction framework with edge constraint (EMU-CNN) for high-spatial-resolution remote sensing imagery is proposed. The proposed framework consists of three components: 1) a multiscale fusion U-shaped network (MFUN); 2) a region proposal network (RPN); and 3) an edge-constrained multitask network (ECMN). First, in the proposed method, the MFUN includes three parallel branches to learn multiple building features with different scales. The RPN then detects the positions of the building instances, even for buildings that are connected with each other. Moreover, according to the instance positions, the ECMN is proposed to extract a precise mask and suppress overfitting. The experiments conducted on a self-annotated data set and two public data sets (the ISPRS Vaihingen semantic labeling contest data set and the WHU aerial image data set) show that the EMU-CNN method can achieve excellent performance and shows great robustness at different scales. Yuanyuan Liu 0004, Ailong Ma, Yanfei Zhong, Fang Fang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | RSNet: The Search for Remote Sensing Deep Neural Networks in Recognition TasksabstractDeep learning algorithms, especially convolutional neural networks (CNNs), have recently emerged as a dominant paradigm for high spatial resolution remote sensing (HRS) image recognition. A large amount of CNNs have already been successfully applied to various HRS recognition tasks, such as land-cover classification and scene classification. However, they are often modifications of the existing CNNs derived from natural image processing, in which the network architecture is inherited without consideration of the complexity and specificity of HRS images. In this article, the remote sensing deep neural network (RSNet) framework is proposed using an automatically search strategy to find the appropriate network architecture for HRS image recognition tasks. In RSNet, the hierarchical search space is first designed to include module- and transition-level spaces. The module-level space defines the basic structure block, where a series of lightweight operations as candidates, including depthwise separable convolutions, is proposed to ensure the efficiency. The transition-level space controls the spatial resolution transformations of the features. In the hierarchical search space, a gradient-based search strategy is used to find the appropriate architecture. In RSNet, the task-driven architecture training process can acquire the optimal model parameters of the switchable recognition module for HRS image recognition tasks. The experimental results obtained using four benchmark data sets for land-cover classification and scene classification tasks demonstrate that the searched RSNet can achieve a satisfactory accuracy with a high computational efficiency and, hence, provides an effective option for the processing of HRS imagery. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, as a particular semantic segmentation task, always faces with larger-scale variation, larger intra-class variance of background, and foreground-background imbalance in the high spatial resolution (HSR) remote sensing imagery. However, general semantic segmentation methods mainly focus on scale variation in the natural scene, with inadequate consideration of the other two problems that usually happen in the large area earth observation scene. In this paper, we argue that the problems lie on the lack of foreground modeling and propose a foreground-aware relation network (FarSeg) from the perspectives of relation-based and optimization-based foreground modeling, to alleviate the above two problems. From perspective of relation, FarSeg enhances the discrimination of foreground features via foreground-correlated contexts associated by learning foreground-scene relation. Meanwhile, from perspective of optimization, a foreground-aware optimization is proposed to focus on foreground examples and hard examples of background during training for a balanced optimization. The experimental results obtained using a large scale dataset suggest that the proposed method is superior to the state-of-the-art general semantic segmentation methods and achieves a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma |
CVPR | 4 |
| 2020 | Dense Greenhouse Extraction in High Spatial Resolution Remote Sensing ImageryabstractGreenhouses are densely distributed across the cultivated land in high-resolution remote sensing imagery, resulting in the problem of dense object extraction. On the one hand, objects tend to be wrongly merged into one object since objects are connected; on the other hand, the existed random sampling scheme does not make good use of the object distribution density to improve the training effect. To meet the demand for dense greenhouse extraction, this paper proposes a novel deep learning-based greenhouse extraction algorithm. To solve the problem that dense objects tend to be wrongly merged, this paper proposes a dual-task learning module, which uses the discriminant object boundary of the pixel-based branch to distinguish the objects of the object-based branch; to take advantage of the object distribution density for effective training, a high-density biased sampler is proposed. Moreover, this paper provides a dataset of manually labeled imagery to train and develop the proposed method. Results on greenhouse extraction in six regions in China achieve peak mIoU and mAP value, surpassing the state-of-the-art methods. Finally, a product of a greenhouse map in China is provided for analysis. Yanfei Zhong, Ailong Ma, Liqin Cao |
IGARSS | 3 |
| 2020 | RSSM-Net: Remote Sensing Image Scene Classification Based on Multi-Objective Neural Architecture SearchabstractThe deep learning (DL)-based scene classification methods have been obtained the remarkable attention for the high spatial resolution remote sensing (HRS) imagery. However, from one aspect, the existing DL methods in HRS image scene classification are usually the variations of the natural image processing methods and often the inherent network structures; from another aspect, the strenuous and significant efforts have been devoted to the design of relevant network structures by human experts. In this paper, learning from the natural evolution, the deep neural network is expected to be globally evolved by the machine for automatically adapting the structure of the HRS imagery, a multi-objective neural architecture search based HRS image scene classification method is proposed (RSSM-Net). The two objectives of minimizing a classification error and the computational complexity have been simultaneously optimized through the evolutionary multi-objective method, the competitive neural architectures in a Pareto solution set are then obtained. The effectiveness is proved by the experiment of the UC Merced dataset with several networks designed by human experts. Yuting Wan, Yanfei Zhong, Ailong Ma, Ruyi Feng |
IGARSS | 3 |
| 2020 | Mapping Local Climate Zones with Circled Similarity Propagation Based Domain AdaptationabstractLocal climate zone (LCZ) has the potential to describe the urban form and function among global cities, which is vital for urban-related researches. However, lacking high quality labeled samples makes it difficult to map global LCZ, and how to transfer labeled samples from one city to another becomes an urgent issue. In this paper, a circled similarity propagation-based domain adaptation method named CPDA for LCZ classification is proposed. In CPDA, the existed labeled samples and unseen samples are called the source and target domain, respectively. The circled similarity matrix of features from source to target and back to the source domain is measured using cosine similarity. Furthermore, an adaptation loss is proposed to make the feature distribution of these two domains as similar as possible, which is beneficial for predicting the target areas with no need for labels. Experiments on the LCZ42 dataset demonstrated the effectiveness of the proposed method. Yanfei Zhong, Ailong Ma |
IGARSS | 3 |
| 2020 | Multiobjective Subpixel Mapping With Multiple Shifted Hyperspectral ImagesabstractSubpixel mapping (SPM) is a useful technique that can interpret the spatial distribution inside mixed pixels and produce a finer-resolution classification map for hyperspectral remote-sensing imagery. However, SPM is essentially an ill-posed problem that requires additional information to produce the unique solution. The limited information of a single image is insufficient to make the mapping problem well posed, whereas the complementary spatial information of multiple shifted images is able to reduce the uncertainty and generate an accurate map. The maximum a posteriori model is a feasible way to incorporate auxiliary information for SPM with multiple shifted images, but it introduces a sensitive regularization parameter, which is difficult to preset. Furthermore, the fixed parameter in the iterations influences the incorporation of the multiple images and the spatial prior. In this article, to address these issues, a multiobjective SPM framework for use with multiple shifted hyperspectral images (MOMSM) is proposed. In the proposed algorithm, a multiobjective model consisting of two objective functions, i.e., data fidelity and spatial prior terms, is constructed to transform the SPM into a multiobjective optimization problem, to get rid of the sensitive regularization parameter. To simultaneously optimize the two objective functions, a multiobjective memetic algorithm with a local search operator and an adaptive global replacement strategy is proposed. The multiple images and spatial information can be dynamically fused and the optimal mapping solution with a good balance between the two objectives can be finally obtained. Experiments conducted on both synthetic and real data sets confirm that the proposed method outperforms the other tested SPM algorithms. Mi Song, Yanfei Zhong, Ailong Ma, Xiong Xu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Multiobjective Hyperspectral Feature Selection Based on Discrete Sine Cosine AlgorithmabstractFeature selection is an effective way to reduce the data dimensionality of hyperspectral imagery and obtain a better performance in the subsequent applications, such as classification. The ideal approach is to obtain the optimal tradeoff between two criteria for hyperspectral image feature selection: 1) information preservation and 2) redundancy reduction. However, constructing a hyperspectral feature selection model for the above two criteria is difficult due to the complexity of hyperspectral imagery. Although evolutionary multiobjective optimization methods have been recently presented to simultaneously optimize the above criteria, they cannot control the global exploration versus local exploitation capabilities in the search space for the hyperspectral feature selection problem. Thus, in this article, a novel discrete sine cosine algorithm (SCA)-based multiobjective feature selection (MOSCA_FS) approach is proposed for hyperspectral imagery. In the proposed method, a novel and effective framework of multiobjective hyperspectral feature selection is designed. In the framework, the ratio between the Jeffries-Matusita (JM) distance and mutual information (MI) is modeled to minimize the redundancy and maximize the relevance of the selected feature subset. In addition, another measurement - the variance (Var) - is applied for maximizing the information amount. Furthermore, to resolve the discrete hyperspectral feature selection problem, a novel discrete SCA is first proposed, which enhances the selection of the ideal feature subset. The effectiveness and universality of the proposed method was verified by experiments with ten University of California at Irvine (UCI) data sets, five hyperspectral image data sets, and one spectral data set of typical surface features. Yuting Wan, Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Multi-Objective Sparse Subspace Clustering for Hyperspectral ImageryabstractHyperspectral images (HSIs) are typical high-dimensional and complex data. As such, the clustering of HSIs is a challenging task. Out of the motivation to find the low-dimensional structure representation of the high-dimensional data, sparse subspace clustering (SSC) methods have been proposed in recent studies. Sparse representation is an important technique in SSC, which is aimed at obtaining the sparse coefficient matrix of the HSI data. Generally speaking, the acquisition of the sparse coefficient matrix is an ill-posed problem, and the existing methods introduce an extra condition as a regularization term to resolve it. However, the regularization parameter is determined manually, which is difficult and lacks self-adaptability. Hence, in this article, a multi-objective SSC method for hyperspectral imagery is proposed, which simultaneously optimizes the sparse term and the data fidelity term. In addition, the spatial structure information of the HSIs is often neglected in the processing model, and thus, a spatial prior term, as the third optimization objective function, is also tested in this article. As a result, there is no need to manually set a regularization parameter. Furthermore, by using the l0norm as the sparse term, this reduces the error caused by the convex relaxation of the other norms. In the proposed method, a multi-objective optimization model is first used to acquire the sparse coefficient matrix, in which a strategy for constructing the dictionary is proposed for more precise and efficient multi-objective optimization. In addition, a knee point-based selection method is utilized to automatically select the optimal sparse representation solution from the Pareto front. The adjacency matrix is then constructed according to the sparse coefficient matrix. Finally, a spectral clustering method is used to obtain clustering results. Experiments undertaken with four HSI data sets confirm the effectiveness of the proposed method. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image ClassificationabstractDeep learning techniques have provided significant improvements in hyperspectral image (HSI) classification. The current deep learning-based HSI classifiers follow a patch-based learning framework by dividing the image into overlapping patches. As such, these methods are local learning methods, which have a high computational cost. In this article, a fast patch-free global learning (FPGA) framework is proposed for HSI classification. The proposed framework consists of three main parts: 1) a designed sampling strategy; 2) an encoder-decoder-based fully convolutional network (FCN); and 3) lateral connections between the encoder and decoder. In FPGA, an encoder-decoder-based FCN is utilized to consider the global spatial information by processing the whole image, which results in fast inference. However, it is difficult to directly utilize the encoder-decoder-based FCN for HSI classification as it always fails to converge due to the insufficiently diverse gradients caused by the limited training samples. To solve the divergence problem and maintain the FCNs abilities of fast inference and global spatial information mining, a global stochastic stratified (GS2) sampling strategy is first proposed by transforming all the training samples into a stochastic sequence of stratified samples. This strategy can obtain diverse gradients to guarantee the convergence of the FCN in the FPGA framework. For a better design of FCN architecture, FreeNet, which is a fully end-to-end network for HSI classification, is proposed to maximize the exploitation of the global spatial information and boost the performance via a spectral attention-based encoder and a lightweight decoder. A lateral connection module is also designed to connect the encoder and decoder, fusing the spatial details in the encoder and the semantic features in the decoder. The experimental results obtained using three public benchmark data sets suggest that the FPGA framework is superior to the patch-based framework in both speed and accuracy for HSI classification. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | COLOR: Cycling, Offline Learning, and Online Representation Framework for Airport and Airplane Detection Using GF-2 Satellite ImagesabstractMonitoring airports using remote sensing imagery require us to first detect the airports and then perform airplane detection. Detecting airports and airplanes with large-scale remote sensing imagery are significant and challenging tasks in the field of remote sensing. Although many detection algorithms have been developed for detecting airports and airplanes in remote sensing imagery, the efficiency of the processing does not meet the needs of real applications in large-scale remote sensing imagery. In recent years, deep learning techniques, such as deep convolutional neural networks (DCNNs), have achieved great progress in image recognition. However, training a DCNN needs a large number of training examples to accurately fit the data distribution. Annotating training examples in large-scale remote sensing imagery is time-consuming, which makes the pipeline inefficient. In this article, to overcome the above two weaknesses, we propose a novel cycling data-driven framework for efficient and robust airport localization and airplane detection. The proposed method consists of three modules: cycling by example refinement (C), offline learning (OL), and online representation (OR), namely cycling, offline learning, and online representation (COLOR). The OR module is a coarse-to-fine cascaded convolutional neural network, which is used to detect airports and airplanes. The example refinement (ER) module implements the cycling and makes use of the unlabeled remote sensing images and the corresponding predictions obtained by the OR module, to generate training examples. The OL module aims to use the training examples from the ER module to update the OR module, to further improve the performance. The whole workflow involves COLOR. The COLOR framework was used to detect airplanes and airports in 512 large-scale Gaofen-2 (GF-2) remote sensing images with 29$200\times27$ 620 pixels. The results showed that the proposed method obtained a mean average precision (mAP) of 88.32% for the airplane detection. In addition due to the proposed coarse-to-fine cascaded OR module the proposed method is much faster than the traditional approaches in real-world applications. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | A Novel Robust Feature Descriptor for Multi-Source Remote Sensing Image RegistrationabstractNon-linear radiation difference (NRD) will lead to the corresponding features cannot be mapped one by one, so the traditional image feature matching methods based on intensity or gradient fail to be directly applied to the multisource remote sensing image registration. In this paper, a new robust feature descriptor is proposed, which has the invariance of radiation, scale and rotation. The nonlinear diffusion function which is insensitive to the radiation difference is used to construct the scale space so that the descriptors can be used in images with different resolutions. A pixel-by-pixel local phase congruency algorithm is used to extract the corresponding points, and then the features are described by means of rotation invariance description. Feature matching is completed based on feature vector constructed by the descriptor, thus to realize image registration. In the experimental part, three kinds of multisource remote sensing images with large radiation differences were used to test the descriptors. The results showed that the proposed method effectively extracted the corresponding features, and achieved the best effect in the quantitative evaluation of image registration. Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2019 | Sub-Pixel Mapping with Multiple Shifted Hyperspectral Images Based on Multiobjective Evolutionary AlgorithmabstractSub-pixel mapping (SPM) can interpret the sub-pixel spatial distribution of land-cover classes in hyperspectral image, which is an ill-posed problem due to the inadequate information of a single image. Auxiliary information provided by multiple shifted (MS) images can make SPM problem well-posed and improve mapping accuracy. The maximum a posteriori (MAP) technique can incorporate the auxiliary information of MS images, but it introduces a fixed weight parameter to fuse the auxiliary information and spatial prior information, heavily influencing the mapping result. This paper proposed a novel SPM method to model the auxiliary information and spatial prior information into two objective functions, which can be simultaneously optimized by the devised multiobjective evolutionary algorithm. Therefore, there is no need of weight parameter, and the two objective functions can be intelligently integrated during the evolution. Experimental results and parameter analysis have indicated the superiority of the proposed method. Mi Song, Yanfei Zhong, Ailong Ma, Qiqi Zhu, Liqin Cao, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2019 | Tailings Reservoir Disaster and Environmental Monitoring Using the UAV-ground Hyperspectral Joint Observation and Processing: A Case of Study in Xinjiang, the Belt and RoadabstractThe tailings reservoir is an inevitable part of the production of metal mines, and due to it is usually the accumulation of waste residue and waste water, the risk source of artificial debris flow with high potential energy has been formed and the environmental risk cannot be underestimated. Thus, it is an important disaster and environmental protection project for the mining enterprises. However, the existing methods cannot conduct a comprehensive disaster and environmental monitoring, considering the remote sensing technology is an effective method for the large-scale monitoring, thus, this global monitoring will be carried out through a novel UAV-ground hyper-spectral joint observation and processing, where the UAV hyper-spectral image, the ground hyper-spectral data of the water and waste residue, and water quality testing report will be used. In addition, the study area is in Xinjiang, the Belt and Road. Yuting Wan, Yanfei Zhong, Ailong Ma, Lifei Wei, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2019 | Hyperspectral Remote Sensing Image Band Selection Via Multi-Objective Sine Cosine AlgorithmabstractFor hyper-spectral image band selection, there are two main key concerns, which are the curial information preservation and the redundancy information reduction. Since the two objectives are contradictory, the single-objective based band selection methods are usually unsatisfactory, thus, a superior approach which can obtain a trade-off between them is needed. In order to address this problem, an evolutionary computation method called sine cosine algorithm which has the capabilities of the global exploration and the local exploitation is applied, and its multi-objective discrete version is designed for multi-objective band selection. In addition, in this paper, two novel measures are utilized for meeting the requirements. The effectiveness of the proposed method is confirmed by the experimental results obtained with two real hyper-spectral images. Yuting Wan, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2019 | Multi-Scale and Multi-Task Deep Learning Framework for Automatic Road ExtractionabstractRoad detection and centerline extraction from very high-resolution (VHR) remote sensing imagery are of great significance in various practical applications. Road detection and centerline extraction operations depend on each other, to a certain extent. The road detection constrains the appearance of the centerline, and the centerline enhances the linear features of the road detection. However, most of the previous works have addressed these two tasks separately and have not considered the symbiotic relationship between them, making it difficult to obtain smooth and complete roads. In this paper, a novel multi-scale and multi-task deep learning framework for automatic road extraction (MSMT-RE) is proposed to build the relationship between them and simultaneously complete the road detection and centerline extraction tasks. U-Net is selected as the basic network for multi-task learning due to its strong ability to preserve spatial details. Multi-scale feature integration is also applied in the framework to increase the robustness of the feature extraction. Meanwhile, an adaptive loss function is introduced to solve the problems of roads taking up a small percentage of the training samples, and the fact that the positive samples of the two tasks are unbalanced. Finally, experiments were conducted on two public road data sets and two large images from Google Earth, and the proposed framework was compared with other state-of-the-art deep learning-based road extraction methods, both quantitatively and qualitatively. The proposed approach outperformed all the compared methods, confirming its advantages in automatic road extraction. Yanfei Zhong, Zhuo Zheng, Ji Zhao 0006, Ailong Ma, Jie Yang 0040 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Multiobjective Sparse Subpixel Mapping for Remote Sensing ImageryabstractSubpixel mapping (SPM) of remote sensing imagery is aimed at generating a classification map with a finer spatial resolution based on the abundance maps. The sparse subpixel mapping (SSM) method reformulates the SPM problem into a spatial pattern linear regression problem based on the preconstructed subpixel patch dictionary. However, in the SSM model, the optimization of the L0-norm is a nonconvex NP-hard problem, so the L1-norm is used to replace the L0-norm to obtain an approximate solution, and the selection of the optimal weight parameter between multiple terms is difficult. Thus, in this paper, a novel multiobjective SSM (MOSSM) framework for remote sensing imagery is proposed, which transforms the SSM problem into a multiobjective optimization problem. In MOSSM, first, the sparsity term is accurately modeled using the L0-norm instead of the L1-norm to avoid the potential errors caused by the L1-norm, and an evolutionary algorithm is used to directly optimize the L0-norm. Second, a subfitness-based multiobjective evolutionary algorithm is employed to simultaneously optimize the fidelity term, the sparsity term, and the spatial prior term, and to generate a set of optimal sparse coefficients to balance these three terms. Thus, there is no need to determine sensitive weight parameters. Finally, two spatial prior terms, which can be applied to the overcomplete dictionary, are presented in the proposed MOSSM-TV and MOSSM-L algorithms to incorporate the spatial correlation of subpixels. Experiments were conducted with two synthetic images and two real data sets, and the results were compared with those of ten other SPM algorithms to demonstrate the effectiveness of the proposed method. Mi Song, Yanfei Zhong, Ailong Ma, Ruyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Fully Automatic Spectral-Spatial Fuzzy Clustering Using an Adaptive Multiobjective Memetic Algorithm for Multispectral ImageryabstractClustering of remote sensing imagery is a tough task due to the particular and complex structure of remote sensing images and the shortage of known information. In this paper, we propose a fully automatic spectral-spatial fuzzy clustering method using an adaptive multiobjective memetic algorithm (AMOMA) for multispectral remote sensing imagery. This approach is made up of two automatic layers: an automatic determination layer and an automatic clustering layer. The first layer seeks the optimal number of clusters through a self-adaptive differential evolution algorithm. The second layer then takes advantage of the AMOMA for spectral-spatial clustering using the optimal number of clusters. The knee point from the Pareto front is then selected through the angle-based method in every generation, and we then compare the knee points between generations to output the final optimal solution. The effectiveness of the proposed method is verified by the experimental results obtained with three remote sensing data sets. Yuting Wan, Yanfei Zhong, Ailong Ma |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Urban Land Use/Land Cover Classification Based on Feature Fusion Fusing Hyperspectral Image and Lidar DataabstractHyperspectral images have been widely used in classification because of the abundant spectral information. But it can't distinguish the objective with similar spectral character but different elevation. However, LiDAR data can obtain elevation information. Therefore, it will obtain better classification maps if fusing the two data. In recent years, CNN has attracted much attention due to its powerful ability to excavate the potential representation and features of the raw data. However, it's difficult to distinguish the objects with different spectral information but similar surface character. Unlike CNN features, the traditional manual features, such as the normalized vegetation index (NDVI), have a certain characteristic expression significance. In order to consider both the semantic information of traditional manual features and the advanced features of CNN features, this paper proposes a fusion algorithm of hyperspectral and LiDAR fusion based on feature fusion. The proposed algorithm has achieved a good fusion classification effect on the MUUFL Gulfport Hyperspectral and LiDAR Data set. Qiong Cao, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2018 | Multiobjective Subpixel Land-Cover MappingabstractThe hyperspectral subpixel mapping (SPM) technique can generate a land-cover map at the subpixel scale by modeling the relationship between the abundance map and the spatial distribution image of the subpixels. However, this is an inverse ill-posed problem. The most widely used way to resolve the problem is to introduce additional information as a regularization term and acquire the unique optimal solution. However, the regularization parameter either needs to be determined manually or it cannot be determined in a fully adaptive manner. Thus, in this paper, the multiobjective subpixel land-cover mapping (MOSM) framework for hyperspectral remote sensing imagery is proposed, in which the two function terms [the fidelity term and the prior term (i.e., the regularization term)] can be optimized simultaneously, and there is no need to determine the regularization parameter explicitly. In order to achieve this goal, two strategies are designed in MOSM: 1) a high-resolution distribution image-based individual encoding strategy is designed in order to calculate the prior term accurately and 2) a subfitness-based individual comparison strategy is designed in order to generate subpixel land-cover mapping solutions with a high quality to update the population. Four data sets (one simulated, two synthetic, and one real hyperspectral image) were used to test the proposed method. The experimental results show that MOSM can perform better than the other subpixel land-cover mapping methods, demonstrating the effectiveness of MOSM in balancing the fidelity term and prior term in the SPM model. Ailong Ma, Yanfei Zhong, Da He, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Sub-pixel intelligence mapping considering spatial-temoporal attraction for remote sensing imageryabstractMixed pixel is a ubiquitous phenomenon in remotely sensed imagery, especially in moderate and low spatial resolution imagery, which compromise the hard land cover classification since the dominant class will shadow the information of other vulnerable classes, bringing trouble to imagery interpretation. Since the past decades, sub-pixel mapping (SPM) approaches were developed to deal with the mixture problem, on the basis of soft classification, to retrieval the pure components and its geospatial distribution within mixed pixels. Recently, SPM integrated with auxiliary information is gradually been a state-of-the-art method for mixed pixel problem, and has been proved effectively. However, few works has been dedicated to explore the geostatistic inter-correlation between spatial and temporal among the time sequences imageries. In this paper, a novel SPM algorithm based on swarm intelligence theory, considering spatiotemporal geographical attraction among multi-temporal imageries, called spatiotemporal attraction based sub-pixel evolution mapping (SASEM), is proposed for remote sensing imagery, Experiments were carried out to verify the proposed algorithm, and the result illustrate that the proposed algorithm outperform the traditional SPM, achieving a fine spatial resolution thematic map for further applications. Da He, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2017 | Scene semantic classification based on scale invariance convolutional neural networksabstractConvolutional neural networks (CNNs) has been introduced into remote sensing scene classification, achieving outstanding performance. However, the scale change of objects contained in remote sensing scene image make it difficult to extract feature robust to scale, limiting the further improvement of classification accuracy. In this paper, a scene classification method named Scale Invariance Convolutional Neural Networks (SICNNs) is proposed for remote sensing scene classification. In the proposed method, two images with different scales generated by randomly stretching one image are fed into CNNs simultaneously for training at intervals of several iterations. Then a similarity measure layer was added in SICNN to make the distance of the two feature vectors extracted from the two images as close as possible, leading extracted feature to be robust to scale. Experimental results using two datasets, i.e. the UC Merced dataset, Google dataset of SIRI-WHU, demonstrated the effectiveness of the proposed method. Yanfei Zhong, Ji Zhao 0006, Ailong Ma, Qianqing Qin |
IGARSS | 4 |
| 2017 | Change detection based on structural conditional random field framework for high spatial resolution remote sensing imageryabstractIn this paper, a structural conditional random field framework (SCRF) is proposed to detect the detailed change information from high spatial resolution (HSR) remote sensing imagery. Traditional random field based methods encounter the over-smoothing problem when deal with HSR images and the boundary of changed objects cannot be preserved well. To solve this problem, in SCRF, fuzzy c means (FCM) is used to model the unary potential while avoiding the independent assumption. Pairwise potentials with different shapes are selected as the structural set to model the spatial features of land cover such as buildings and roads. Based on SCRF, a set of change belief maps are generated to describe the observed image from different aspects. An object based fusion strategy is then followed to combine the belief maps to get the refined result. The results of the proposed method on two HSR data sets outperform some state-of-art algorithms. Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2017 | MINI-UAV borne hyperspectral remote sensing: A reviewabstractIn recent years, the science of hyperspectral remote sensing has huge development in virtue of the integration of low-cost lightweight hyperspectral sensors and unmanned aerial vehicles (UAVs). As an alternative of manned aircraft, UAV has some unique advantages enabling the researchers acquire the hyperspectral images of their interest area flexibly and promptly. This review focuses on the recent developments of UAV borne hyperspectral remote sensing system, and gives an overview of the corresponding platforms, sensors, data acquisition, processing and current applications. Future challenges and research directions for UAV borne hyperspectral data are also addressed. Yanfei Zhong, Xinyu Wang 0003, Tianyi Jia, Lifei Wei, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 7 |
| 2016 | Local spectral-spatial clustering for remote sensing imageryabstractRemote sensing image clustering is a challenging task. Recently, by combining the spectral and spatial information of remote sensing data, the clustering accuracy can be dramatically enhanced. However, it has always been difficult to determine the weight parameter for balancing the spectral and spatial terms of the clustering objective function. In this paper, spectral-spatial clustering with a local weight parameter determination method for remote sensing imagery is proposed, i.e. LSSC. In LSSC, considering the large scale of remote sensing images, the weight parameter is determined locally in a patch image instead of the whole image. The local weight parameter is then used in constructing the objective function of LSSC. Thus, the remote sensing image clustering problem is transformed into an optimization problem. Finally, in order to achieve a better optimization performance, a variant of differential evolution (i.e. jDE) is used as the optimizer due to its powerful optimization capability. Experimental results confirm that the proposed LSSC can acquire a higher clustering accuracy than other spectral-spatial clustering methods. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2016 | Semisupervised Subspace-Based DNA Encoding and Matching Classifier for Hyperspectral Remote Sensing ImageryabstractHyperspectral remote sensing images, which are characterized by their high dimensionality, provide us with the capability to accurately identify objects on the ground. They can also be used to identify subclasses of objects. However, these subclasses are usually embedded in different subspaces due to the complex distribution of pixels in the feature space. In the literature, few hyperspectral image classification methods can take both the subclass and subspace into consideration at the same time. Motivated by the fact that natural DNA can distinguish biological subspecies (subclasses in hyperspectral images) using critical DNA fragments (subspaces in hyperspectral images), a semisupervised subspace-based DNA encoding and matching classifier for hyperspectral remote sensing imagery (SSDNA) is proposed in this paper. First, in the process of DNA encoding, the hyperspectral remote sensing image is transformed into a DNA cube, in which the first-order spectral curve of the hyperspectral remote sensing image is utilized in order to take the gradient information of the spectral curve into consideration. Second, in the process of DNA optimization, evolutionary algorithms are used to obtain the best DNA library of the typical objects, which includes the following: 1) A multicenter individual representation is designed in order to consider the existence of subclasses in the hyperspectral remote sensing image; 2) the unlabeled samples are utilized in the process of population initialization and fitness calculation to enhance the diversity of the population and the generalization of the classification performance; and 3) the different classes are embedded in different subspaces. A semisupervised technique is used to extract the subspaces, including the global subspace for all the classes and the local subspace for each class. Three hyperspectral data sets were tested and confirm that SSDNA performs better than the other supervised or semisupervised classifiers. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Spectral-spatial DNA encoding discriminative classifier for hyperspectral remote sensing imageryabstractHyperspectral remote sensing image classification is one of the most challenging tasks. In our previous work, motivated by the similarity between the structures of DNA and hyperspectral remote sensing images, a DNA matching mechanism was used to transform the hyperspectral remote sensing image into a DNA cube for classification. However, the above DNA encoding strategy lacks the process of encoding accurate spectral and spatial feature into the DNA cube, resulting in unsatisfying classification performance. In this paper, a spectral-spatial DNA encoding strategy for encoding accurate spectral and spatial feature of hyperspectral remote sensing image is proposed. In the spectral dimension, the first-order spectral curve is encoded into the DNA cube, while in the spatial dimension, the principal components or their corresponding texture feature (GLCM) are encoded into the DNA cube. Finally, different with the previous DNA encoding classifier using genetic algorithm (GA), the paper combines the discriminative classifier (i.e. SVM) with spectral-spatial DNA encoding to improve classification performance for hyperspectral remote sensing imagery. The experimental results confirmed the effectiveness of the newly devised DNA encoding strategy and the discriminative classifier in classifying the DNA cube. Ailong Ma, Yanfei Zhong, Hongzan Jiao, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2015 | Adaptive Multiobjective Memetic Fuzzy Clustering Algorithm for Remote Sensing ImageryabstractDue to the intrinsic complexity of remote sensing images and the lack of prior knowledge, clustering for remote sensing images has always been one of the most challenging tasks in remote sensing image processing. Recently, clustering methods for remote sensing images have often been transformed into multiobjective optimization problems, making them more suitable for complex remote sensing image clustering. However, the performance of the multiobjective clustering methods is often influenced by their optimization capability. To resolve this problem, this paper proposes an adaptive multiobjective memetic fuzzy clustering algorithm (AFCMOMA) for remote sensing imagery. In AFCMOMA, a multiobjective memetic clustering framework is devised to optimize the two objective functions, i.e., Jm and the Xie-Beni (XB) index. One challenging task for memetic algorithms is how to balance the local and global search capabilities. In AFCMOMA, an adaptive strategy is used, which can adaptively achieve a balance between them, based on the statistical characteristic of the objective function values. In addition, in the multiobjective memetic framework, in order to acquire more individuals with high quality, a new population update strategy is devised, in which the updated population is composed of individuals generated in both the local and global searches. Finally, to evaluate the proposed AFCMOMA algorithm, experiments using three remote sensing images were conducted, which confirmed the effectiveness of the proposed algorithm. Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Remote sensing imagery clustering using an adaptive bi-objective memetic methodabstractDue to the intrinsic complexity of the remote sensing image and the lack of the prior knowledge, clustering for remote sensing image has always been one of the most challenging works in remote sensing image processing. The proposed algorithm constructs a bi-objective memetic-based framework, exploiting the feature space more efficiently. In the framework, two objective functions, Jm and XB, are used as the objective functions for bi-objective optimization. Furthermore, an adaptive local search method which can dynamically adjust its parameter value according to the selection probability has been developed and incorporated into the proposed algorithm. In order to speed the convergence and obtain more non-dominated solutions in the pareto front, a new strategy is newly devised in the local search process, which considers more solutions as the candidate for the next generation. To evaluate the proposed algorithm, some experiments on two multi-spectral images are conducted. The results show that the proposed algorithm can achieve better performance, compared with related methods. Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2013 | Adaptive Differential Evolution Fuzzy Clustering Algorithm with Spatial Information and Kernel Metric for Remote Sensing Imagery
Ailong Ma, Yanfei Zhong, Liangpei Zhang 0001 |
IDEAL | 1 |