Lei Ding 0008

dblp:59/2353-8 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0003-0653-8373ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 S2C: A Noise-Resistant Difference Learning Framework for Unsupervised Change Detection in VHR Remote Sensing Images
abstract
Unsupervised Change Detection (UCD) in Very High Resolution (VHR) Remote Sensing (RS) images remains to be a difficult challenge due to the inherent spatio-temporal complexity within data. Inspired by recent advancements in Visual Foundation Models (VFMs) and Contrastive Learning (CL), this research aims to develop CL methodologies to translate implicit knowledge in VFM into change representations, thus eliminating the need for explicit supervision. To this end, we introduce a Semantic-to-Change (S2C) learning framework for UCD in VHR RS images. Differently from existing CL methodologies that typically focus on learning multi-temporal similarities, we introduce a novel triplet learning strategy that explicitly models temporal differences, which are crucial to the CD task. Furthermore, random spatial and spectral perturbations are introduced during training to enhance robustness to temporal noise. In addition, a grid sparsity regularization is defined to suppress insignificant changes, and an IoU-matching algorithm is developed to refine the CD results. Experiments on three benchmark CD datasets demonstrate that the proposed S2C learning framework achieves significant improvements in accuracy, surpassing current state-of-the-art by over 31%, 9% and 23%, respectively. It also demonstrates robustness and sample efficiency, suitable for training and adaptation of various VFMs or backbone neural networks.
Lei Ding 0008, Xibing Zuo, Haitao Guo, Jun Lu 0005, Zhihui Gong, Xuanguang Liu, Jicang Lu
AAAI1
2026 Propagating spatio-temporal state and progressively associating trajectory for satellite video multi-object tracking
abstract
High-resolution video satellites enable large-view dynamic monitoring for earth observation. Among satellite video interpretation techniques, multi-object tracking (MOT) receives growing attention for its foundational role. Rigid targets in satellite videos exhibit strong inter-frame appearance and posture consistency, showing quasi-linear trajectories with constrained displacements. Inspired by these kinematic characteristics, this paper introduces an online end-to-end P ropagating S patio-temporal S tate and P rogressively A ssociating T rajectory MOT (PS2PAT-MOT) framework. It consists of a detection branch for locating multi-category and multi-target objects in current frame, and a correlation branch using an inter-frame spatio-temporal state propagation (STSP) module to propagate location and appearance information and encode same-target correlations between adjacent frames. Detection and correlation outputs from both branches undergo affinity calculation, while the Progressively Associating Trajectory (PAT) strategy generates continuous tracklets using differentiated association thresholds for distinct trajectory segments. Experimental results on two publicly available AIR-MOT, SAT-MTB, and a self-built LV-SatMOT ( L arge- V iew Sat ellite video MOT ) datasets demonstrate the effectiveness of the proposed PS2PAT-MOT framework. Codes are available at: https://github.com/HELOBILLY/PS2PAT-MOT .
Kun Zhu 0003, Haitao Guo, Guanzhou Chen 0001, Xiaodong Zhang 0027, Lei Ding 0008, Xiangyun Liu, Jun Lu 0005
Expert Syst. Appl.5
2025 Recurrent Semantic Change Detection in VHR Remote Sensing Images Using Visual Foundation Models
abstract
Semantic change detection (SCD) involves the simultaneous extraction of changed regions and their corresponding semantic classifications (pre- and post-change) in remote sensing images (RSIs). Despite recent advancements in vision foundation models (VFMs), the fast-segment anything model has demonstrated insufficient performance in SCD. In this article, we propose a novel VFMs architecture for SCD, designated as VFM-ReSCD. This architecture integrates a side adapter (SA) into the VFM-ReSCD to fine-tune the fast segment anything model (FastSAM) network, enabling zero-shot transfer to novel image distributions and tasks. This enhancement facilitates the extraction of spatial features from very high-resolution (VHR) RSIs. Moreover, we introduce a recurrent neural network (RNN) to model semantic correlation and capture feature changes. We evaluated the proposed methodology on two benchmark datasets. Extensive experiments show that our method achieves state-of-the-art (SOTA) performances over existing approaches and outperforms other CNN-based methods on two RSI datasets.
Jing Zhang 0023, Lei Ding 0008, Tingyuan Zhou, Jian Wang 0138, Peter M. Atkinson, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.2
2025 HALO: Hierarchical Adaptive Feature Learning and Cross-View Interaction for Object-Level Geo-Localization
abstract
Cross-view object localization (CVOGL) offers fine-grained geographic information regarding the specified object of interest, in comparison to typical cross-view geo-location at the image level. However, it faces more challenges, primarily due to significant visual appearance changes induced by viewpoint discrepancies and the difficulty of accurately identifying corresponding objects within reference images containing multiple targets. To address these challenges, we propose a novel CVOGL model termed HALO. First, we introduce a cascade cross-view feature interaction module that enables effective local-to-global feature fusion across different viewpoints, thereby enhancing the feature representation of objects in the reference image. Second, to mitigate the scale variations and feature distribution discrepancies between the query and reference images, we propose an adaptive feature hierarchization and aggregate module. It extracts more representative global descriptors by adaptively hierarchizing and aggregating features. Lastly, utilizing the learned global descriptors, we perform cross-view CL and introduce a hard sample mining strategy to further enhance the discriminative ability of the network. Through these advancements, we significantly improve the discriminative representations of the query objects at both views. Extensive experiments on the CVOGL dataset demonstrate the effectiveness and robustness of the proposed method. In the CVOGL task of drone→satellite, HALO improves 2.36% in [email protected] on the test sets. Similarly, HALO shows an increase of 2.49% in [email protected] on the validation set for the CVOGL task of ground→satellite. Our codes are publicly available at https://github.com/ZehaoZhang-Uestc/HALO.
Zehao Zhang, Lei Ding 0008, Yufeng Wang 0004, Wenrui Ding
IEEE Trans. Geosci. Remote. Sens.3
2025 Integrating Segment Anything Model With Instance-Level Change Generation for Single-Temporal Unsupervised Change Detection
abstract
Implementing change detection (CD) with bi-temporal remote sensing images (RSIs) is essential for gaining insights into the dynamic evolution of the Earth’s surface. Recent advancements in deep learning methodologies have demonstrated considerable success in CD applications. However, the efficacy of supervised CD networks is significantly constrained by their dependence on high-quality change labels and precisely registered bi-temporal RSIs, which imposes significant limitations in real-world scenarios. To address this issue, we propose a novel single-temporal unsupervised CD method, STU-SAMI, which integrates the Segment Anything Model (SAM) with instance-level change generation. This method aims to facilitate the training of high-performance CD networks using unpaired and unlabeled single-temporal RSIs. The proposed approach comprises three main steps. First, the SAM is introduced to extract morphologically complete object instances from single-temporal RSIs. Second, an online instance-level change generation method is proposed, which transforms single-temporal RSIs into bi-temporal pseudo change samples based on the object instances obtained by SAM. Third, a streamlined and efficient deep Siamese CD network is constructed to support the training and inference processes. Extensive experiments conducted on three benchmark datasets demonstrate that the proposed method outperforms the existing state-of-the-art unsupervised CD methods. The code is available at https://github.com/IceStreams/STU-SAMI.
Xibing Zuo, Jie Rui, Lei Ding 0008, Fei Jin, Yuzhun Lin, Shuxiang Wang, Xiao Liu 0050, Juan Lei
IEEE Trans. Geosci. Remote. Sens.3
2024 Joint Spatio-Temporal Modeling for Semantic Change Detection in Remote Sensing Images
abstract
Semantic Change Detection (SCD) refers to the task of simultaneously extracting the changed areas and the semantic categories (before and after the changes) in Remote Sensing Images (RSIs). This is more meaningful than Binary Change Detection (BCD) since it enables detailed change analysis in the observed areas. Previous works established triple-branch Convolutional Neural Network (CNN) architectures as the paradigm for SCD. However, it remains challenging to exploit semantic information with a limited amount of change samples. In this work, we investigate to jointly consider the spatio-temporal dependencies to improve the accuracy of SCD. First, we propose a Semantic Change Transformer (SCanFormer) to explicitly model the ’from-to’ semantic transitions between the bi-temporal RSIs. Then, we introduce a semantic learning scheme to leverage the spatio-temporal constraints, which are coherent to the SCD task, to guide the learning of semantic changes. The resulting network (SCanNet) significantly outperforms the baseline method in terms of both detection of critical semantic changes and semantic consistency in the obtained bi-temporal results. It achieves the SOTA accuracy on two benchmark datasets for the SCD.
Lei Ding 0008, Jing Zhang 0023, Haitao Guo, Kai Zhang 0010, Bing Liu 0018, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2024 Adapting Segment Anything Model for Change Detection in VHR Remote Sensing Images
abstract
Vision Foundation Models (VFMs) such as the Segment Anything Model (SAM) allow zero-shot or interactive segmentation of visual contents, thus they are quickly applied in a variety of visual scenes. However, their direct use in many Remote Sensing (RS) applications is often unsatisfactory due to the special imaging properties of RS images. In this work, we aim to utilize the strong visual recognition capabilities of VFMs to improve change detection (CD) in very high-resolution (VHR) remote sensing images (RSIs). We employ the visual encoder of FastSAM, a variant of the SAM, to extract visual representations in RS scenes. To adapt FastSAM to focus on some specific ground objects in RS scenes, we propose a convolutional adaptor to aggregate the task-oriented change information. Moreover, to utilize the semantic representations that are inherent to SAM features, we introduce a task-agnostic semantic learning branch to model the semantic latent in bi-temporal RSIs. The resulting method, SAM-CD, obtains superior accuracy compared to the SOTA fully-supervised CD methods and exhibits a sample-efficient learning ability that is comparable to semi-supervised CD methods. To the best of our knowledge, this is the first work that adapts VFMs to CD in VHR RS images.
Lei Ding 0008, Kun Zhu 0003, Daifeng Peng, Hao Tang 0005, Kuiwu Yang, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2023 Bi-Directional Temporal Modelling for Semantic Change Detection in Remote Sensing Images
abstract
Semantic change detection (SCD) is a branch of change detection (CD) that provides detailed land-cover/land-use (LCLU) change information. It presents not only the changed information but also the bi-temporal LCLU semantic maps (in the changed areas). Studies have recently highlighted [1] that SCD can be addressed through a triple-branch Convolutional Neural Network (CNN), which contains two multi-temporal branches and a change detection branch. However, in this architecture, the two temporal branches learn insufficient LCLU transition information. In this paper, we present a novel architecture that combines CNN and RNN for the SCD of remote sensing images. It employs a Siamese CNN to learn semantic information from two temporal images, followed by a Bidirectional Recurrent Neural Network (Bi-directional RNN) to learn temporal dependencies of the LCLU classes. The resulting CNN-RNN architecture can model better the LCLU transitions, thus enhancing the semantic representation of the bi-temporal features. Experimental results on a benchmark dataset show that the proposed method obtains significant accuracy improvements over the existing approaches. It also shows advantages in recognizing the minority LCLU changes.
Jing Zhang 0023, Lei Ding 0008, Lorenzo Bruzzone
IGARSS2
2023 An Anchor-Free and Angle-Free Detector for Oriented Object Detection Using Bounding Box Projection
abstract
The detection and recognition of oriented objects in remote sensing images is a challenging task due to their complex backgrounds, various sizes, diverse aspect ratios, and especially arbitrary orientations. Many oriented object detection algorithms need to obtain accurate angles or adopt anchors to predict the oriented bounding boxes. When directly predicting the angles of objects’ oriented bounding boxes, the loss of angle is discontinuous during training, which makes it difficult to obtain accurate boundary of oriented objects. And the anchors also aggravate the problems of class imbalance and computational cost. To address the above problems, this paper proposes an anchor-free and angle-free detector called AF2Det. AF2Det adopts the information of the bounding box projection instead of the angle to represent and reconstruct the object’s oriented bounding boxes, which could avoid the problem of boundary discontinuity. To predict the information of the bounding box projection, an anchor-free architecture is built to predict objects as points based on a simple but strong U-shaped architecture. And the deformable convolution and the bottom-up feature fusion method are integrated effectively to enhance AF2Det’ s capacity for objects’ shapes, orientations, and scales. The extensive experiments are conducted on multiple datasets, i.e., HRSC2016, FGSD2021, DOTA, and RSDD-SAR to validate the effectiveness of our method. The experimental results demonstrate that the proposed AF2Det outperforms other anchor-free algorithms and obtains competitive results on oriented object detection.
Donghang Yu, Haitao Guo, Xiangyun Liu, Qing Xu 0005, Yuzhun Lin, Lei Ding 0008
IEEE Trans. Geosci. Remote. Sens.7
2023 Relation Changes Matter: Cross-Temporal Difference Transformer for Change Detection in Remote Sensing Images
abstract
Thanks to their capability of modeling global information, transformers have been recently applied to change detection in remote sensing images. Generally, the changes in terms of shape and appearance of objects lead to relation changes among these objects in multi-temporal images. However, in this context, the attention mechanism in transformers has not been fully explored yet to learn relation changes in the observed scenes. In this paper, we analyze the relation changes in multi-temporal images and propose a cross-temporal difference (CTD) attention to capture these changes efficiently. Through the CTD attention, the changed areas are distinguished better from the unchanged areas. Based on the CTD attention, two CTD-transformer encoders are constructed to extract the features of changed areas from the embedded tokens of multi-temporal images in a cross manner. Then, the extracted features at the coarse scale are further improved to the fine-scale by the corresponding CTD-transformer decoders. In addition, consistency-perception blocks (CPBs) are designed to preserve the structures and contours of changed areas. Finally, all extracted features from multi-temporal images are concatenated to produce the desired change map. Compared to state-of-the-art methods, experimental results on LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that the proposed method produces better performance. The source code is available at https://github.com/RSMagneto/CTD-Former.
Kai Zhang 0010, Feng Zhang 0028, Lei Ding 0008, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2023 Deep Unsupervised Key Frame Extraction for Efficient Video Classification
abstract
Video processing and analysis have become an urgent task, as a huge amount of videos (e.g., YouTube, Hulu) are uploaded online every day. The extraction of representative key frames from videos is important in video processing and analysis since it greatly reduces computing resources and time. Although great progress has been made recently, large-scale video classification remains an open problem, as the existing methods have not well balanced the performance and efficiency simultaneously. To tackle this problem, this work presents an unsupervised method to retrieve the key frames, which combines the convolutional neural network and temporal segment density peaks clustering. The proposed temporal segment density peaks clustering is a generic and powerful framework, and it has two advantages compared with previous works. One is that it can calculate the number of key frames automatically. The other is that it can preserve the temporal information of the video. Thus, it improves the efficiency of video classification. Furthermore, a long short-term memory network is added on the top of the convolutional neural network to further elevate the performance of classification. Moreover, a weight fusion strategy of different input networks is presented to boost performance. By optimizing both video classification and key frame extraction simultaneously, we achieve better classification performance and higher efficiency. We evaluate our method on two popular datasets (i.e., HMDB51 and UCF101), and the experimental results consistently demonstrate that our strategy achieves competitive performance and efficiency compared with the state-of-the-art approaches.
Hao Tang 0005, Lei Ding 0008, Songsong Wu, Bin Ren 0005, Nicu Sebe, Paolo Rota
ACM Trans. Multim. Comput. Commun. Appl.2
2022 MP-ResNet: Multipath Residual Network for the Semantic Segmentation of High-Resolution PolSAR Images
abstract
There are limited studies on the semantic segmentation of high-resolution polarimetric synthetic aperture radar (PolSAR) images due to the scarcity of training data and the complexity of managing speckle noise. The Gaofen contest has provided open access a high-quality PolSAR semantic segmentation dataset. Taking this opportunity, we propose a multipath residual network (MP-ResNet) architecture for the semantic segmentation of high-resolution PolSAR images. Compared to conventional U-shape encoder–decoder convolutional neural network (CNN) architectures, the MP-ResNet learns semantic context with its parallel multiscale branches, which greatly enlarges its valid receptive fields and improves the embedding of local discriminative features. In addition, MP-ResNet adopts a multilevel feature fusion design in its decoder to effectively exploit the features learned from its different branches. Comparisons with the baseline method of fully connected network (FCN with ResNet34) show that the MP-ResNet has achieved significant accuracy improvements. It also surpasses several state-of-the-art methods in terms of overall accuracy (OA),$\text{m}F_{1}$and frequency weighted intersection over union (fwIoU), with only a limited increase of computational costs. This CNN architecture can be used as a baseline method for future studies on the semantic segmentation of PolSAR images. The code is available at:https://github.com/ggsDing/SARSeg.
Lei Ding 0008, Dong Lin, Yuxing Chen 0002, Bing Liu 0018, Jiansheng Li, Lorenzo Bruzzone
IEEE Geosci. Remote. Sens. Lett.1
2022 Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images
abstract
Semantic change detection (SCD) extends the multiclass change detection (MCD) task to provide not only the change locations but also the detailed land-cover/land-use (LCLU) categories before and after the observation intervals. This fine-grained semantic change information is very useful in many applications. Recent studies indicate that the SCD can be modeled through a triple-branch convolutional neural network (CNN), which contains two temporal branches and a change branch. However, in this architecture, the communications between the temporal branches and the change branch are insufficient. To overcome the limitations in existing methods, we propose a novel CNN architecture for the SCD, where the semantic temporal features are merged in a deep CD unit. Furthermore, we elaborate on this architecture to reason the bi-temporal semantic correlations. The resulting bi-temporal semantic reasoning network (Bi-SRNet) contains two types of semantic reasoning blocks to reason both single-temporal and cross-temporal semantic correlations, as well as a novel loss function to improve the semantic consistency of change detection results. Experimental results on a benchmark dataset show that the proposed architecture obtains significant accuracy improvements over the existing approaches, while the added designs in the Bi-SRNet further improve the segmentation of both semantic categories and the changed areas. The codes in this article are accessible athttps://github.com/ggsDing/Bi-SRNet.
Lei Ding 0008, Haitao Guo, Sicong Liu 0001, Lichao Mou, Jing Zhang 0023, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2022 Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Long-range contextual information is crucial for the semantic segmentation of high-resolution (HR) remote sensing images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the perception of long-range contexts in large RSIs. To overcome this limitation, we propose a wide-context network (WiCoNet) for the semantic segmentation of HR RSIs. Apart from extracting local features with a conventional convolutional neural network (CNN), the WiCoNet has an extra context branch to aggregate information from a larger image area. Moreover, we introduce a context transformer to embed contextual information from the context branch and selectively project it onto the local features. The context transformer extends the vision transformer, an emerging kind of neural networks, to model the dual-branch semantic correlations. It overcomes the locality limitation of CNNs and enables the WiCoNet to see the bigger picture before segmenting the land-cover/land-use (LCLU) classes. Ablation studies and comparative experiments conducted on several benchmark datasets demonstrate the effectiveness of the proposed method. In addition, we present a new Beijing Land-Use (BLU) dataset. This is a large-scale HR satellite dataset with high-quality and fine-grained reference labels, which can facilitate future studies in this field.
Lei Ding 0008, Dong Lin, Shaofu Lin, Jing Zhang 0023, Xiaojie Cui, Yuebin Wang, Hao Tang 0005, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2022 REMNet: Recurrent Evolution Memory-Aware Network for Accurate Long-Term Weather Radar Echo Extrapolation
abstract
Weather radar echo extrapolation, which predicts future echoes based on historical observations, is one of the complicated spatial–temporal sequence prediction tasks and plays a prominent role in severe convection and precipitation nowcasting. However, existing extrapolation methods mainly focus on a defective echo-motion extrapolation paradigm based on finite observational dynamics, neglecting that the actual echo sequence has a more complicated evolution process that contains both nonlinear motions and the lifecycle from initiation to decay, resulting in poor prediction precision and limited application ability. To complement this paradigm, we propose to incorporate a novel long-term evolution regularity memory (LERM) module into the network, which can memorize long-term echo-evolution regularities during training and be recalled for guiding extrapolation. Moreover, to resolve the blurry prediction problem and improve forecast accuracy, we also adopt a coarse–fine hierarchical extrapolation strategy and compositive loss function. We separate the extrapolation task into coarse and fine two levels which can reduce the downsampling loss and retain echo fine details. Except for the average reconstruction loss, we additionally employ adversarial loss and perceptual similarity loss to further improve the visual quality. Experimental results from two real radar echo datasets demonstrate the effectiveness of our methodology and show that it can accurately extrapolate the echo evolution while ensuring the echo details are realistic enough, even for the long term. Our method can further be improved in the future by integrating multimodal radar variables or introducing certain domain prior knowledge of physical mechanisms. It can also be applied to other spatial–temporal sequence prediction tasks, such as the prediction of satellite cloud images and wind field figures.
Jinrui Jing, Qian Li 0014, Leiming Ma, Lei Ding 0008
IEEE Trans. Geosci. Remote. Sens.5
2022 Perceiving Spectral Variation: Unsupervised Spectrum Motion Feature Learning for Hyperspectral Image Classification
abstract
In recent years, deep-learning-based hyperspectral image (HSI) classification methods have achieved significant development. The superior capability of feature extraction from these data-driven methods dramatically improves the classification performance. However, the previous methods usually require to retrain the network from scratch to obtain the capability of feature extraction adaptive for the target image when facing a new HSI to be classified, which is a time-consuming and redundant process. In this paper, we consider putting this process ahead and making the network have a robust capability of feature extraction with generalization through pre-training. Therefore, the network enables to directly extract features of the target HSI without re-training. For this purpose, we rethink the three-dimension (3D) HSI data from a perspective of spectral sequence, and we attempt to extract the spectral variation information as the spectrum motion feature. Then, we construct an unsupervised spectrum motion feature learning framework (SMF-UL), which can be pre-trained on mass unlabeled HSI data to learn the knowledge about perceiving spectral variation. Furthermore, to achieve the expansion of source data for pre-training, we develop an extendable training dataset construction method, which can integrate HSIs of different sizes, number of bands and sensors into a unified training set to utilize the rapidly growing mass unlabeled HSI data effectively. Finally, we use the trained network to directly extract the spectrum motion feature of the target HSI for classification, so the laborious re-training of the network can be avoided. Extensive experiments show that the proposed SMF-UL acquires the robust capability of feature extraction with generalization through unsupervised learning on mass unlabeled HSI data, and the classification performance of extracted spectrum motion feature is competitive to advanced in-domain and cross-domain methods, which shows its flexibility and superiority. The code of SMF-UL will be open at: https://github.com/sssssyf/SMF-UL.
Yifan Sun 0008, Bing Liu 0018, Xuchu Yu, Anzhu Yu, Kuiliang Gao, Lei Ding 0008
IEEE Trans. Geosci. Remote. Sens.6
2022 Adversarial Shape Learning for Building Extraction in VHR Remote Sensing Images
abstract
Building extraction in VHR RSIs remains a challenging task due to occlusion and boundary ambiguity problems. Although conventional convolutional neural networks (CNNs) based methods are capable of exploiting local texture and context information, they fail to capture the shape patterns of buildings, which is a necessary constraint in the human recognition. To address this issue, we propose an adversarial shape learning network (ASLNet) to model the building shape patterns that improve the accuracy of building segmentation. In the proposed ASLNet, we introduce the adversarial learning strategy to explicitly model the shape constraints, as well as a CNN shape regularizer to strengthen the embedding of shape features. To assess the geometric accuracy of building segmentation results, we introduced several object-based quality assessment metrics. Experiments on two open benchmark datasets show that the proposed ASLNet improves both the pixel-based accuracy and the object-based quality measurements by a large margin. The code is available at: https://github.com/ggsDing/ASLNet.
Lei Ding 0008, Hao Tang 0005, Yilei Shi, Xiao Xiang Zhu 0001, Lorenzo Bruzzone
IEEE Trans. Image Process.1
2021 SDFL-FC: Semisupervised Deep Feature Learning With Feature Consistency for Hyperspectral Image Classification
abstract
Semisupervised deep learning methods (DLMs) can mitigate the dependence on large amounts of labeled samples using a small number of labeled samples. However, for semisupervised deep feature learning (SDFL), the quality of extracted features cannot be well ensured without a certain amount of labeled samples. To address this issue, we develop the SDFL method with feature consistency (SDFL-FC) for the hyperspectral image (HSI) classification. The SDFL-FC first adopts the convolutional neural network (CNN) to extract spectral–spatial features of HSI and then uses the fully connected layers (FCLs) to model the feature consistency. Moreover, two constraints that enforce both the feature consistency of single pixel (FCS) and feature consistency of group pixels (FCG) are introduced to obtain the representative and discriminative features. The FCS is achieved by the generative adversarial network (GAN) regularization, which can reconstruct the original data from extracted features. The FCG is based on the assumption that the features of group pixels should have similar characteristics within a superpixel, which is embedded in each FCL. The final FCL outputs the class labels, and the cross-entropy (CE) loss is calculated with the labeled samples, while the two losses of FCS and FCG are calculated with all the training samples (both labeled and unlabeled). SDFL-FC integrates the FCS, FCG, and CE loss into a unified objective function and uses a customized iterative optimization algorithm to optimize it. Experiments demonstrate that the SDFL-FC can outperform the related state-of-the-art HSI classification methods.
Yuebin Wang, Junhuan Peng, Chunping Qiu, Lei Ding 0008, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 DiResNet: Direction-Aware Residual Network for Road Extraction in VHR Remote Sensing Images
abstract
The binary segmentation of roads in very high resolution (VHR) remote sensing images (RSIs) has always been a challenging task due to factors such as occlusions (caused by shadows, trees, buildings, etc.) and the intraclass variances of road surfaces. The wide use of convolutional neural networks (CNNs) has greatly improved the segmentation accuracy and made the task end-to-end trainable. However, there are still margins to improve in terms of the completeness and connectivity of the results. In this article, we consider the specific context of road extraction and present a direction-aware residual network (DiResNet) that includes three main contributions: 1) an asymmetric residual segmentation network with deconvolutional layers and a structural supervision to enhance the learning of road topology (DiResSeg); 2) a pixel-level supervision of local directions to enhance the embedding of linear features; and 3) a refinement network to optimize the segmentation results (DiResRef). Ablation studies on two benchmark data sets (the Massachusetts data set and the DeepGlobe data set) have confirmed the effectiveness of the presented designs. Comparative experiments with other approaches show that the proposed method has advantages in both overall accuracy and F1-score. The code is available at:https://github.com/ggsDing/DiResNet.
Lei Ding 0008, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2021 LANet: Local Attention Embedding to Improve the Semantic Segmentation of Remote Sensing Images
abstract
The trade-off between feature representation power and spatial localization accuracy is crucial for the dense classification/semantic segmentation of remote sensing images (RSIs). High-level features extracted from the late layers of a neural network are rich in semantic information, yet have blurred spatial details; low-level features extracted from the early layers of a network contain more pixel-level information but are isolated and noisy. It is therefore difficult to bridge the gap between high- and low-level features due to their difference in terms of physical information content and spatial distribution. In this article, we contribute to solve this problem by enhancing the feature representation in two ways. On the one hand, a patch attention module (PAM) is proposed to enhance the embedding of context information based on a patchwise calculation of local attention. On the other hand, an attention embedding module (AEM) is proposed to enrich the semantic information of low-level features by embedding local focus from high-level features. Both proposed modules are lightweight and can be applied to process the extracted features of convolutional neural networks (CNNs). Experiments show that, by integrating the proposed modules into a baseline fully convolutional network (FCN), the resulting local attention network (LANet) greatly improves the performance over the baseline and outperforms other attention-based methods on two RSI data sets.
Lei Ding 0008, Hao Tang 0005, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2020 Semantic Segmentation of Large-Size VHR Remote Sensing Images Using a Two-Stage Multiscale Training Architecture
abstract
Very-high resolution (VHR) remote sensing images (RSIs) have significantly larger spatial size compared to typical natural images used in computer vision applications. Therefore, it is computationally unaffordable to train and test classifiers on these images at a full-size scale. Commonly used methodologies for semantic segmentation of RSIs perform training and prediction on cropped image patches. Thus, they have the limitation of failing to incorporate enough context information. In order to better exploit the correlations between ground objects, we propose a deep architecture with a two-stage multiscale training strategy that is tailored to the semantic segmentation of large-size VHR RSIs. In the first stage of the training strategy, a semantic embedding network is designed to learn high-level features from downscaled images covering a large area. In the second training stage, a local feature extraction network is designed to introduce low-level information from cropped image patches. The resulting training strategy is able to fuse complementary information learned from multiple levels to make predictions. Experimental results on two data sets show that it outperforms local-patch-based training models in terms of both accuracy and stability.
Lei Ding 0008, Jing Zhang 0023, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2019 A Deep Architecture Based on a Two-Stage Learning for Semantic Segmentation of Large-Size Remote Sensing Images
abstract
Remote sensing images (RSIs) usually have much larger size compared to typical natural images used in computer vision applications. This makes the computational cost of training convolutional neural networks with full-size images unaffordable. Commonly used methodologies for semantic segmentation of RSIs perform training and prediction on cropped local image patches. Thus they fail to model the potential dependencies between ground objects at a higher level of abstraction. In order to better exploit global context information in RSIs, a deep architecture based on a two-stage training approach that is specially tailored to training large-size RSIs is proposed. In the first training stage, down-scaled images are used as input to learn high-level features from a large image area. In the second training stage, a local feature extraction network is designed to extract low-level information from cropped image patches. The complementary information learned from different levels is fused to make the prediction. As a result, the proposed two-stage training approach is able to exploit the context information of RSIs from a larger perspective without losing spatial details. Experimental results on a benchmark remote sensing dataset demonstrate the effectiveness of the proposed approach.
Lei Ding 0008, Lorenzo Bruzzone
IGARSS1