EDBT 2026 Demo / reviewers in the wild / expert
Li Shen 0004
dblp:91/3680-4
· DBLP profile ↗
17ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0003-4638-3329ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Lightweight Dual-Branch Network for Building Change Detection in Remote Sensing Images Integrating Cross-Scale Coupling and Boundary ConstraintabstractCapturing spatial details and contextual semantics is crucial for building change detection (BCD) in remote sensing images. However, achieving both aspects within traditional single-branch feature extraction networks faces challenges in computation costs and model sizes. To tackle these challenges, this article introduces a lightweight neural model tailored for BCD in remote sensing images. Its primary contribution lies in the design of a lightweight dual branch for efficient feature extraction, a cross-scale coupling module (CSCM) for effective multiscale feature enhancement, and a boundary constraint module (BCM) for edge details compensation. Specifically, the lightweight dual branch can efficiently extract spatial details and contextual semantics in two independent branches, thereby generating changed detail-semantics feature maps. The CSCM further enriches the semantic and scale representation ability of changed feature maps, by adapting to the multiscale characteristics of changed buildings. In addition, the BCM can improve the model’s sensitivity to the boundaries of changed buildings, mitigating the loss of edge information caused by convolution and pooling operations. As a result, the fine-grained and high-level semantic feature maps for BCD are obtained. We comprehensively evaluate the effectiveness of our proposed method against numerous state-of-the-art lightweight and non lightweight change detection models on the LEVIR and WHU datasets. The results demonstrate that our method not only achieves remarkable accuracy but also stands out in efficiency. Yanshuai Dai, Li Shen 0004, Shichuan Liu, Zhilin Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Weakly Supervised Bitemporal Scene Change Detection Approach for Pixel-Level Building Damage Assessment Using Pre- and Post-Disaster High-Resolution Remote Sensing ImagesabstractSudden-onset natural disasters, such as destructive earthquakes, pose significant threats to human life and property. The use of high-resolution remote sensing (HRRS) images for automated assessment of building damage can rapidly and accurately provide spatial distribution information and statistical data on building damage, assisting in disaster response and relief efforts. However, the task is exceedingly challenging due to the diverse and intricate appearance of damaged buildings in HRRS images, coupled with interference from surrounding areas that exhibit certain damage characteristics as a result of the disaster. To overcome these issues, this article proposes a weakly supervised building damage assessment method based on scene change detection in pre- and post-disaster bitemporal images. This method fully leverages visual information of building boundaries and deeper semantic information of building scenes from pre-disaster images to guide the identification of building damage in post-disaster images. Specifically, the method first generates fine-grained subbuilding objects with detailed boundaries from pre-disaster images by combining semantic segmentation of buildings with superpixel segmentation. Then, bitemporal image scene blocks obtained using sub-building objects as clues are input into our proposed Siamese local-global visual transformer (SLgViT) network, enabling scene change detection guided by deep semantic information from pre-disaster images. Finally, the change detection results serve as the basis to depict pixel-level building damage in post-disaster images. The proposed SLgViT network is primarily composed of a specially designed local-global visual transformer (LgViT) module and a cross-Siamese interaction fusion (CSIF) module, both of which play a crucial role in the deep mining and integrated interaction of local and global semantic features from pre- and post-disaster images. It is noteworthy that our method operates in a weakly supervised manner. The training of the SLgViT network requires only scene patches centered around building objects from pre- and post-disaster bitemporal images, along with image-level annotations. Experiments conducted with satellite images from the 2010 Port-au-Prince, Haiti earthquake and unmanned aerial vehicle (UAV) images from the 2019 Changning, China earthquake have demonstrated the effectiveness and superior performance of the proposed method. Wenfan Qiao, Li Shen 0004, Zhilin Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Weakly Supervised Semantic Segmentation Approach for Damaged Building Extraction From Postearthquake High-Resolution Remote-Sensing ImagesabstractQuick and accurate building damage assessment following a disaster is critical to making a preliminary estimate of losses. Remote-sensing image analysis based on convolutional neural networks (CNNs) and their relatives has shown a growing potential in this task, but faces the challenge of collecting dense pixel-level annotations. In this letter, we propose a novel weakly supervised semantic segmentation (WSSS) method based on image-level labels for pixel-wise damaged building extraction from postearthquake high-resolution remote-sensing (HRRS) images. The proposed method aims to improve the quality of the class activation map (CAM) to boost model performance. To be specific, a multiscale dependence (MSD) module and a spatial correlation refinement (SCR) module are designed by considering the special characteristics of the damaged building and are integrated into an encoder–decoder network. The former is used for complete and dense localization of damaged buildings in CAM, and the latter contributes to noise suppression. Extensive experimental evaluations over three datasets are conducted to confirm the effectiveness of the proposed approach. Both generated CAMs and extracted damaged building results of our methods are better than that of current state-of-the-art methods. Wenfan Qiao, Li Shen 0004, Zhilin Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | ALNet: Auxiliary Learning-Based Network for Weakly Supervised Building Extraction From High-Resolution Remote Sensing ImagesabstractWeakly supervised semantic segmentation (WSSS) based on image-level labels can reduce the expensive costs of annotating pixel-level labels, and it has achieved great progress by generating class activation maps (CAMs) as pseudo labels for training a segmentation model. However, it is challenging to generate high-quality CAMs due to the large supervision gap between classification and segmentation. To aid in mitigating the supervision gap, we firstly design a simple but effective Feature-To-Image Restoration (FTIR) branch to acquire supplementary pixel-wise supervisory information. Secondly, to effectively utilize FTIR supervision in addressing the supervision gap, we have developed an auxiliary learning-based network (ALNet) that incorporates the FTIR as an auxiliary task to support the primary WSSS task. By integrating them into a unified framework, the proposed network can leverage the FTIR supervision to enhance the performance of WSSS. Thirdly, we propose a semantic-aware multi-level feature progressive aggregation module (SMPA) to enhance the fusion of multi-level features, accommodating the varying feature requirements of the two distinct tasks at different levels of granularity. Finally, we also introduce a global reasoning augmentation module (GloRe) into our model to further strengthen global correlations between similar building objects. Extensive experiments on public HRSI building datasets, including the WHU dataset and the InriaAID dataset, demonstrate that with the assistance of the FTIR, the proposed ALNet with the SMPA and the GloRe can produce more integral and accurate CAMs for weakly supervised building extraction in HRSIs, and achieve state-of-the-art performance that significantly surpasses other WSSS methods. Li Shen 0004, Zhilin Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | PANet: Pixelwise Affinity Network for Weakly Supervised Building Extraction From High-Resolution Remote Sensing ImagesabstractTo save large human efforts to annotate pixel-level labels, weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted increasing attention. For WSSS, generating high-quality class activation maps (CAMs) is crucial to obtain pseudo labels for training an accurate building extraction model. To generate high-quality CAMs, many existing methods make use of multiscale context fusion of individual entities. Although these methods have shown an improvement on weakly supervised building extraction, they do not take account of the global interrelations beyond individual entities, resulting in inconsistent activated values in CAMs for different building objects. In this study, we develop a pixelwise affinity network (PANet) for weakly supervised building extraction based on image-level labels. We model and enhance the interrelations between building objects by leveraging reliable interpixel affinities, thus optimizing the generation of the CAMs. Moreover, we propose a consistency regularization loss to further refine the generated CAMs on the accuracy of boundary regions. Experiments on two public datasets (InriaAID dataset and WHU dataset) verify the effectiveness of the proposed PANet. Experimental results also show that our method achieves excellent results with over 0.57 points in intersection-over-union (IOU) score and over 0.73 points in F1 score on both datasets and outperforms the comparing methods. Li Shen 0004, Zhilin Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Building Damage Detection in Vhr Satellite Images Via Multi-Scale Scene Change DetectionabstractDetecting damaged buildings quickly and accurately is of great significance for the disaster area to carry out emergency rescue and decision-making. However, the majority of traditional methods based on scene classification uses a fixed-size sliding window to traverse the very-high-resolution (VHR) satellite images, ignoring the scale difference among the geographical objects, which may lead to poor objects delineation and large time consumption. In this paper, we propose a combined multi-scale segmentation and scene change detection (MSSCD) strategy for detecting damaged buildings from pre- and post-disaster VHR satellite images. We compared the performance of four different scene classification structures in detecting damaged buildings in the 2010 Haiti earthquake, and the experimental results show that the proposed strategy is better than traditional methods. Li Shen 0004, Wenfan Qiao |
IGARSS | 2 |
| 2019 | Detection of Vegetation Areas Attacked By Pests and Diseases Based on Adaptively Weighted Enhanced Global and Local Deep FeaturesabstractVery high resolution (VHR) remote sensing images offer the potential for efficient identification of vegetation areas attacked by pests and diseases without using the near-infrared band. In order to achieve this purpose, various object detection methods are required to perform the extraction task from the VHR images. Traditional pixel-based and object-based classification methods capture little high-level semantic concepts, making it intractable to extract unhealthy vegetation areas. To overcome these problems, this paper proposes a novel method based on the scene level in the framework of the CNN by using adaptively weighted enhanced global and local deep features to extract vegetation areas attacked by pests and diseases from unmanned aerial vehicle (UAV) VHR remote sensing images. To be specific, the last convolutional feature maps of VGG-16 network are firstly extracted as the original features. Then, the learned original features are used for clustering to obtain the scene features, which are regarded as the global features to estimate whether the target object appears in the scene image. Next, each pixel features at different locations of feature maps are used for clustering again to form the object features, which are regarded as local features to determine what the target object is and supplement the global feature with the local information. Afterwards, the shape information of central object in scene image, quantified by the roundness value, is utilized as the foundation of weighting, which can distinguish the different objects with similar spectral and texture feature to some extent. Finally, make use of order of roundness values to adaptively weight both of features, and connect them to form adaptively weighted enhanced global and local deep features. Experimental result indicates that the proposed method outperforms the existing benchmark methods. Yanshuai Dai, Li Shen 0004, Yungang Cao, Tianjie Lei, Wenfan Qiao |
IGARSS | 2 |
| 2019 | Landslide Image Classification Using Semi-Supervised LearningabstractMany researchers focus on the problem of accurately and rapidly classifying regional landslide hazard, whose spectrum and shape are complicated and varied in remote sensing images. A large number of training examples with labels are necessary to construct predictive models in supervised classification, which are difficult to get strong supervision information due to the high cost of data labeling process for rapid regional landslides identification. Our methods use pre- and post-event MODIS NDVI products, and post-event SPOT-5 images to classify the landslide image during the 2008 Wenchuan Earthquake based on semi-supervised learning model, which means that only a subset of training data are given with labels to train a good learner. To examine the effectiveness of the proposed method, the results are compared with state-of-the-art support vector machine (SVM). Experimental results demonstrate that the proposed method is an accurate and rapid way to classify landslide images. Shi He, Haitao Jing, Hong Tang 0002, Li Shen 0004, Liangliang Tao, Jiehai Cheng |
IGARSS | 4 |
| 2019 | A Fine-Grained Fully Convolutional Network For Extraction of Building Along High-Speed Rail Lines from VHR Remote Sensing ImageabstractVery high resolution (VHR) remote sensing offers the potential for efficient identification of building areas along the high-speed rail (HSR) lines. However, high intra-class and low inter-class variances in this type of images and pose serious challenges. Fully convolutional network (FCN) has shown impressive performance and great potential for image classification and object detection. Nevertheless, only classification results with coarse resolutions can be obtained from traditional FCN-based methods. To address this issue, this paper presents a fine-grained fully convolutional network to conduct the building extraction task along the HSR lines. The proposed method introduce a fine-grained upsampling module in the encoder-decoder paradigm to upsample the low-resolution feature maps to higher resolution ones instead of the traditional interpolation processing, thus can reserve and recover more fine-detailed information. Furthermore, an enhanced loss function is designed to impose smoothness prior for building extraction. Experiments on the remote sensing image with 0.5 m spatial resolution from Google Earth covering the part area of Zhengzhou-Xi'an high-speed rail line indicates the proposed method achieve competitive results compared with other state-of-the-art FCN based methods. The study demonstrates the potential application of using high-resolution remote sensing technique to investigate the building areas on both sides of HSR lines regularly, in order to detect the hazardous building areas. Wenfan Qiao, Li Shen 0004, Yungang Cao, Shi He, Yanshuai Dai |
IGARSS | 2 |
| 2017 | Fast and robust structure-based multimodal geospatial image matchingabstractThis paper presents a fast and robust framework integrating local features for the matching of multimodal geospatial data (e.g., optical, LiDAR, SAR and map). In the proposed framework, local feature descriptors, such as Histogram of Oriented Gradient (HOG) and Local Self Similarity (LSS), are first extracted for every pixel to form a pixel-wise structural feature representation of an image. Then we define a similarity metric based on the feature representation in frequency domain using the 3 Dimensional Fast Fourier Transform (3DFFT) technique, followed by a template matching scheme to detect control points between multimodal data. The proposed framework is based on the hypothesis that structural similarity between images is preserved across different modalities. The major advantages of this framework include (1) structural similarity representation using pixel-wise feature description and (2) high computational efficiency due to the use of 3DFFT. Experimental results on different types of multimodal geospatial data show more accurate matching performance of the proposed framework than the state-of-the-art methods. Yuanxin Ye, Lorenzo Bruzzone, Jie Shan, Li Shen 0004 |
IGARSS | 4 |
| 2017 | Robust Optical-to-SAR Image Matching Based on Shape PropertiesabstractAlthough image matching techniques have been developed in the last decades, automatic optical-to-synthetic aperture radar (SAR) image matching is still a challenging task due to significant nonlinear intensity differences between such images. This letter addresses this problem by proposing a novel similarity metric for image matching using shape properties. A shape descriptor named dense local self-similarity (DLSS) is first developed based on self-similarities within images. Then a similarity metric (named DLSC) is defined using the normalized cross correlation (NCC) of the DLSS descriptors, followed by a template matching strategy to detect correspondences between images. DLSC is robust against significant nonlinear intensity differences because it captures the shape similarity between images, which is independent of intensity patterns. DLSC has been evaluated with four pairs of optical and SAR images. Experimental results demonstrate its advantage over the state-of-the-art similarity metrics (such as NCC and mutual information), and show the superior matching performance. Yuanxin Ye, Li Shen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Robust Registration of Multimodal Remote Sensing Images Based on Structural SimilarityabstractAutomatic registration of multimodal remote sensing data [e.g., optical, light detection and ranging (LiDAR), and synthetic aperture radar (SAR)] is a challenging task due to the significant nonlinear radiometric differences between these data. To address this problem, this paper proposes a novel feature descriptor named the histogram of orientated phase congruency (HOPC), which is based on the structural properties of images. Furthermore, a similarity metric named HOPCncc is defined, which uses the normalized correlation coefficient (NCC) of the HOPC descriptors for multimodal registration. In the definition of the proposed similarity metric, we first extend the phase congruency model to generate its orientation representation and use the extended model to build HOPCncc. Then, a fast template matching scheme for this metric is designed to detect the control points between images. The proposed HOPCncc aims to capture the structural similarity between images and has been tested with a variety of optical, LiDAR, SAR, and map data. The results show that HOPCncc is robust against complex nonlinear radiometric differences and outperforms the state-of-the-art similarities metrics (i.e., NCC and mutual information) in matching performance. Moreover, a robust registration method is also proposed in this paper based on HOPCncc, which is evaluated using six pairs of multimodal remote sensing images. The experimental results demonstrate the effectiveness of the proposed method for multimodal image registration. Yuanxin Ye, Jie Shan, Lorenzo Bruzzone, Li Shen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Research on relationship between remote sensing image quality and performance of interest point detectionabstractThe paper researches on the relationship between image quality and the performance of interest point detection. In this paper, we use the image quality metrics and interest point repeatability as the measures of image quality and the performance of interest point detection respectively. Considering the differences of image's scene and quality degradation factor, nine images covering three kinds of classic scenes are selected as the experiment data, and they are respectively contaminated by Gaussian blur and Gaussian noise. The results show that image quality metrics and the repeatability have a certain quantitative relationship, and the repeatability is reduced with the decrease of image quality metrics. Moreover, the relationship can be simulated by some simple functions such as linear and exponential model when the image's scene and degradation factor are fixed. Yuanxin Ye, Li Shen 0004, Songbo Wu |
IGARSS | 3 |
| 2015 | Automatic matching of optical and SAR imagery through shape propertyabstractAutomatic matching of optical and SAR images could be challenging due to the significant non-linear intensity differences caused by radiometric variations among such images. To address this problem, this paper utilize the Shape Property to detect the correspondences between images. A new shape descriptor is first built on the basis of the local self-similarity descriptor. Then, the normalized correlation coefficient of the descriptor is used as the similarity metric (named DLSS), which represents the shape similarity among images. Finally, a template-matching strategy is applied to achieve the correspondences between images. The proposed method has been evaluated with four pairs of optical and SAR images. The experimental results demonstrate that the proposed method is robust to non-linear intensity differences, and improves the matching performance compared with the state-of-the-art methods. Yuanxin Ye, Li Shen 0004 |
IGARSS | 2 |
| 2014 | A Semisupervised Latent Dirichlet Allocation Model for Object-Based Classification of VHR Panchromatic Satellite ImagesabstractTypically, object-based classification methods are learned using training samples with labels attached to image objects. In this letter, a semisupervised object-based method in the framework of topic modeling is proposed to classify very high resolution panchromatic satellite images using partially labeled pixels. In the stage of training, both topics and their co-occurred distributions are learned in an unsupervised manner from segmented satellite images. Meanwhile, unlabeled pixels are allocated user-provided geo-object class labels based on the learned model. In the stage of classification, each segment is classified as a user-provided geo-object class label with the maximum posterior probability. Experimental results show that the proposed method outperforms several SVM-based supervised classification methods in terms of both spatial consistency and semantic consistency. Li Shen 0004, Hong Tang 0002, Adu Gong, Jing Li 0018, Wenbin Yi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | A Multiscale Latent Dirichlet Allocation Model for Object-Oriented Clustering of VHR Panchromatic Satellite ImagesabstractA novel model is presented to address the problem of semantic clustering of geo-objects in very high resolution panchromatic satellite images. The proposed model combines a probabilistic topic model with a multiscale image representation into an automatic framework by embedding both document and scale selections. The probabilistic topic model is used to characterize the statistical distributions of both intraclass appearance and inter-class coherence of geo-objects within documents, i.e., squared sub-images. Because the bag-of-words assumption involved in the probabilistic topic models does not consider the spatial coherence between topic labels, the multiscale image representation is designed to provide a self-adaptive spatial regularization for various geo-object categories. By introducing scale and document selections, the automatic framework integrates the probabilistic topic model and the multiscale image representation to ensure that words on a site should be allocated the same topic label no matter what documents they reside in. Consequently, unlike the traditional method of applying topic models for analyzing satellite images, the process of explicitly generating a set of documents before modeling and then combining multiple labels for a word on a given site is unnecessary. Gibbs sampling is adopted for parameter estimation and image clustering. Extensive experimental evaluations are designed to first analyze the effect of parameters in the proposed model and then compare the results of our model with those of some state-of-the-art methods for three different types of images. The results indicate that the proposed algorithm consistently outperforms these exiting state-of-the-art methods in all of the experiments. Hong Tang 0002, Li Shen 0004, Yinfeng Qi, Yang Shu 0002, Jing Li 0018, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | An object-oriented clustering algorithm for VHR panchromatic images using nonparametric latent Dirichlet allocationabstractIn this paper, we present a novel object-oriented semantic clustering algorithm for VHR panchromatic satellite images using a variant of latent Dirichlet allocation model. Firstly, an image collection is implicitly generated by partitioning a large satellite image into densely overlapped sub-images. Then, the Latent Dirichlet Allocation with a hierarchy Dirichlet process is employed to model the image collection. Gibbs sampling is adopted for parameter estimation and image clustering. Specifically, the introduction of Dirichlet process is purposed to extend the LDA to an infinite mixtures model which can estimate the number of components (e.g. clusters in image analysis) automatically. Finally, the effect of the proposed algorithm is analyzed through experiments, and the results of it with the traditional K-means method over a QUICKBIRD image are compared. Yinfeng Qi, Hong Tang 0002, Yang Shu 0002, Li Shen 0004, Jianwei Yue, Weiguo Jiang |
IGARSS | 4 |