EDBT 2026 Demo / reviewers in the wild / expert
Xiao Li 0017
dblp:66/2069-17
· DBLP profile ↗
19ranked-venue papers
7as first author
17since 2021 · last 2023
0000-0002-2406-3781ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TCD: Task-Collaborated Detector for Oriented Objects in Remote Sensing ImagesabstractOriented object detection (OOD) in remote sensing image interpretation is challenging due to the difficulty of locating objects with arbitrary orientations. Existing methods have made considerable progress based on oriented heads or anchors. However, most of them follow the classical detection paradigm, such as assigning samples based on Intersection-over-Unions (IoU) and predicting through two independent tasks. These fixed strategies impair the consistency between classification and localization predictions, resulting in the prediction with optimal localization accuracy being suppressed by the nonoptimal ones during nonmaximum suppression (NMS). To address this problem, a task-collaborated detector (TCD) is proposed. Compared with current single-stage methods, its improvements include two aspects: task-collaborated assignment (TCA) and task-collaborated head (TCH). Specifically, to better pull closer the best anchors for two tasks, TCA introduces classification and localization confidence into sample assignment and tends to select the anchors with accurate and consistent predictions as positive during training. TCH provides a better balance for learning interactive and discriminative features. It can flexibly adjust the spatial feature distribution of classification and localization tasks by learning the joint features from the aggregation layer. Extensive experiments are conducted on HRSC2016, DOTA, and DIOR-R, and the proposed TCD achieves the state-of-the-art performance [90.60, 80.89, and 65.04 mean average precision (mAP), respectively]. Consistency analysis also demonstrates that TCD can significantly improve prediction consistency. Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Locality-Constrained Bilinear Network for Land Cover Classification Using Heterogeneous ImagesabstractOptical and SAR modalities provide complementary information of land properties, which can lead to outstanding classification performance. Recently, factorized bilinear coding (FBC) as an extension of bilinear pooling in respect of coding-pooling perspective, which extracted compact bilinear fusion features with second-order interaction information in the form of sparse representation, brought the performance improvements on multimodal learning tasks. However, it lost locality attributes among similar samples to be encoded. In this letter, we propose a novel locality-constrained bilinear network (LC-BNet) for land cover classification with heterogeneous remote sensing (RS) images. Specifically, the locality-constrained bilinear coding (LC-BC) introduces locality information to generate compact and discriminative fusion features for land cover classification. Extensive experimental results show superior performances of our work on two broad coregistered optical and SAR datasets. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Multilevel Adaptive-Scale Context Aggregating Network for Semantic Segmentation in High-Resolution Remote Sensing ImagesabstractHigh-resolution remote sensing (HR2S) images contain complex land objects of difference sizes, and it is important for semantic segmentation of the HR2S images to extract multiscale information. In this letter, we introduce a novel multilevel adaptive-scale context aggregating network (MACANet) for semantic segmentation of the HR2S images, which mainly consists of two parts—adaptive-scale context extraction block (AS-CEB) and sequential aggregation block (SAB). In particular, the AS-CEB introduces an inflexible strategy to obtain the features with appropriate scale information based on different asymmetric convolutions and the gated mechanism. Meanwhile, the SAB progressively aggregates multilevel adaptive-scale features, which are used to relieve the semantic gap between different-level features and generate precise score maps. Experimental results on representative HR2S datasets show the advantages of our method. The code is available athttps://github.com/RSIP-NUDT/MACANet. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Cross-Layer Nonlocal Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) is a fundamental yet challenging task in the domain of remote sensing (RS). Currently, the methods based on deep features from convolutional neural networks (CNNs) have significantly improved the scene classification accuracy (ACC). However, the standard convolution operations have limited capacity to model the long-range correlations and cannot effectively obtain global contextual understanding ability. In this letter, we propose a novel scene classification framework, termed cross-layer nonlocal network (CL-NL-Net), consisting of a backbone network, a cross-layer nonlocal (CL-NL) module, and a classifier. Among them, the backbone network is used to obtain multilayer convolutional features. The CL-NL module is the core of the proposed method, which captures the long-range correlations between different layers, so as to achieve a better global scene understanding ability. To verify the effectiveness of the proposed CL-NL-Net, we conduct experiments on four benchmark datasets, and the results demonstrate that the proposed method achieves competitive classification ACC and outperforms some state-of-the-art methods. Ming Li 0066, Lin Lei, Yuli Sun, Xiao Li 0017, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Aspect-Ratio-Guided Detection for Oriented Objects in Remote Sensing ImagesabstractAlthough existing oriented object detection methods have made considerable progress based on oriented heads or anchors, the training process itself is not perfect. In this letter, we point out the inconsistency problem between the fixed network setting and varying aspect ratios, which greatly limits the performance. For example, the fixed parameters in label assignment and regression loss cannot fit the changes of aspect ratios and, thus, are harmful to the training process. Considering the prior information about objects’ aspect ratios, the aspect-ratio-guided (ARG) methods are proposed. Specifically, the ARG label assignment is used to adjust the label assignment criteria (intersection over union (IoU) threshold) automatically, and the ARG IoU loss can change the weights of angle regression dynamically. This ARG design makes better use of training samples and pushes the detector more robust to the change of aspect ratios. With no additional cost, our method improves upon the ResNet-50-feature pyramid network (FPN) baseline with 3.99% AP50 and 6.09% AP75 on HRSC2016. Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Dynamic-Hierarchical Attention Distillation With Synergetic Instance Selection for Land Cover Classification Using Missing Heterogeneity ImagesabstractOptical and SAR modalities can provide the complementary information on the land properties, which usually lead to more robust and better classification performance. However, due to the restriction of imaging condition, not all modalities included into the training data sets could be available in real testing samples. Therefore, it is important to explore how to learn discriminative representations using multimodal data during the training stage, while achieving fine land cover classification using missing modalities at test time. In this article, we propose a novel dynamic-hierarchical attention distillation network (DH-ADNet) with multimodal synergetic instance selection (MSIS) for land cover classification using missing data modalities. First, the MSIS realizes the selection of the most representative multimodal instances to enhance the DH-ADNet’s ability of discriminative feature extraction. Then, the DH-ADNet is training on the basis of the curriculum learning strategy and promotes the hallucination stream to learn the privileged information. In particular, a novel dynamic-hierarchical attention distillation module (DH-ADM) is introduced, which adaptively highlights different contributions of multilayer attention distillation by carefully exploring the classification losses of multilayer features over the training iterations. Comprehensive evaluations on two coregistered optical and SAR data sets and report state-of-the-art results in the privileged information scenario. Xiao Li 0017, Lin Lei, Yuli Sun, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Dense Adaptive Grouping Distillation Network for Multimodal Land Cover Classification With Privileged ModalityabstractMultimodal land cover classification (MLCC) is a fundamental problem in remote sensing interpretation, which can obtain excellent performance on account of the complementary information between the optical and SAR modalities. However, it is usually impossible to obtain multimodal data at the same time, due to the restriction of imaging conditions. When one of the modalities data is completely missing during test phase, classical multimodal learning methods might not be able to handle the MLCC task with privileged modality. In this paper, we propose an efficient Dense Adaptive Grouping Distillation Network (DAGDNet), which learns privileged information from available modalities in the train sets, and improves the classification performance in the test sets when one modality data is scarce. More specifically, to relieve the heterogeneous gaps between different modalities and then transfer the privileged information, we propose an Interactive Gated-based Feature Grouping Module (IG-FGM), which decomposes multimodal features into modalities-shared and modality-specific components to realize the decoupling of multimodal features and grouping distillation. Furthermore, the IG-FGM is inserted into different layers of the “teacher" network to implement progressive blending of multi-modalities. Then, to adaptively highlight the importance of hierarchical features distillation and grouping distillation, we propose a Multi-stage Adaptive Distillation Learning (MS-ADL) strategy so that the weights of different distillation losses are required to change continuously along with the training process. Finally, we evaluate the superior performances of our model on representative co-registered optical and SAR datasets. Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Multimodal Semantic Consistency-Based Fusion Architecture Search for Land Cover ClassificationabstractMultimodal Land Cover Classification (MLCC) using the optical and Synthetic Aperture Radar (SAR) modalities has resulted in outstanding performances over using only unimodal data due to their complementary information on land properties. Previous multimodal deep learning (MDL) methods have relied on handcrafted multi-branch convolutional neural networks (CNN) to extract the features of different modalities and merged them for land cover classification. However, natural images-oriented handcrafted CNN models may not the optimal strategies to handle Remote Sensing (RS) image interpretation problems, due to the huge difference in terms of imaging angles and imaging ways. Furthermore, few MDL methods have analyzed optimal combinations of hierarchical features from different modalities. In this article, we propose an efficient multimodal architecture search framework, namely Multimodal Semantic Consistency-Based Fusion Architecture Search (M2SC-FAS) in continuous search space with the gradient-based optimization method, which can not only discover optimal optical- and SAR-specific architectures according to the different characteristics of the optical and SAR images, respectively, but also realizes the search of optimal multimodal dense fusion architecture. Specifically, the semantic-consistency constraint is introduced to guarantee dense fusion between hierarchical optical and SAR features with high semantic consistency and then capture the complementary performance on land properties. Finally, the basis of curriculum learning strategy is adopted on the M2SC-FAS. Extensive experiments show superior performances of our work on three broad co-registered optical and SAR datasets. Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Structure Consistency-Based Graph for Unsupervised Change Detection With Homogeneous and Heterogeneous Remote Sensing ImagesabstractChange detection (CD) of remote sensing (RS) images is one of the important problems in earth observation, which has been extensively studied in recent years. However, with the development of RS technology, the specific characteristics of remotely sensed images, including sensor characteristics, resolutions, noises, and distortions in imagery, make the CD more complex. In this article, we propose a structure consistency-based method for CD, which detects changes by comparing the structures of two images, rather than comparing the pixel values of images. Because the image structure is imaging modality-invariant and not sensitive to noise, illumination, and other interference factors, the proposed method can be applied to a variety of CD scenarios and has strong robustness. Structural comparison is realized by constructing and mapping an improved nonlocal patch-based graph (NLPG) to avoid the data leakage of two images. First, we demonstrate the effectiveness of the method in homogeneous and heterogeneous CD, which shows that the proposed method can be used as a unified CD framework. Second, we extend the method to the heterogeneous CD with multichannel synthetic aperture radar (SAR) image, which can provide a reference for future research as the heterogeneous CD with multichannel SAR is rarely studied. Third, through the decomposition and in-depth analysis of NLPG, we modify the graph construction process, structure difference calculation, and the difference image fusion to make it more robust and accurate. Experiments on six scenarios 12 data sets demonstrate the effectiveness of the proposed method. Yuli Sun, Lin Lei, Xiao Li 0017, Xiang Tan, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Fine-grained visual classification via multilayer bilinear pooling with object localization
Ming Li 0066, Lin Lei, Hao Sun 0042, Xiao Li 0017, Gangyao Kuang |
Vis. Comput. | 4 |
| 2021 | A Multi-Scale Feature Aggregation Network Based on Channel-Spatial Attention for Remote Sensing Scene ClassificationabstractConvolutional Neural Networks (CNNs) have been shown remarkable performance in the task of remote sensing image scene classification. Recent works demonstrate that aggregating multi-scale convolutional features can significantly improve the classification accuracy. However, existing methods either use some unsupervised feature encoding methods or based on feature aggregation methods to aggregate multi -scale convolutional features, ignoring the information redundancy and semantic ambiguity between of them. To address the above-mentioned limitations, an end-to-end multi-scale feature aggregation network (MSF A) based on channel-spatial attention module is proposed to learn discriminative scene representation for remote sensing scene classification. The experimental results on the aerial image data set (AID) demonstrate that the proposed method achieves competitive classification performance compared with other state-of-the-art methods. Ming Li 0066, Lin Lei, Xiao Li 0017, Yuli Sun |
IGARSS | 3 |
| 2021 | Multi-Modal Fusion Architecture Search for Land Cover Classification Using Heterogeneous Remote Sensing ImagesabstractOptical and SAR modalities can provide the complementary information on land properties for better land cover classification. Most of existing multi-modal land cover classification methods based on two-streams convolutional neural networks (CNNs), which obtained fusion features by merging optical and SAR features that come from manually selective layer of different streams. However, they ignored different semantic between manually selective optical and SAR features, which might result in suboptimal fusion features. We tackle the problem of finding good fusion architectures for multimodal land cover classification inspired by the network architecture search (NAS), and introduces the multi-modal fusion architecture search network (M2PASNet). Extensive experimental results show superior performances of our work on a broad co-registered optical and SAR dataset. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IGARSS | 1 |
| 2021 | Nonlocal patch similarity based heterogeneous remote sensing change detection
Yuli Sun, Lin Lei, Xiao Li 0017, Hao Sun 0042, Gangyao Kuang |
Pattern Recognit. | 3 |
| 2021 | Sparse signal recovery via infimal convolution based penalty
Lin Lei, Yuli Sun, Xiao Li 0017 |
Signal Process. Image Commun. | 3 |
| 2021 | Collaborative Attention-Based Heterogeneous Gated Fusion Network for Land Cover ClassificationabstractExisting land cover classification methods mostly rely on either the optical or synthetic aperture radar (SAR) features alone, which ignore the mutual complementary effects between optical and SAR sources. In this article, we compare the distribution histograms of deep semantic features extracted from optical and SAR modalities within land cover categories, which intuitively demonstrates that there are the large complementary potentials between the optical and SAR features. Therefore, we propose a novel collaborative attention-based heterogeneous gated fusion network (CHGFNet), which hierarchically fuses both optical and SAR features for land cover classification. More specifically, the CHGFNet consists of three main components: two-stream feature extractor, multimodal collaborative attention module (MCAM), and the gated heterogeneous fusion module (GHFM). Given optical and SAR patch pairs, two-stream feature extractor introduces multistage feature learning methodology to acquire discriminative optical and SAR features. Then, to explore the inherent complementarity between optical and SAR features, MCAM is embedded into CHGFNet, which provides an efficient stage to capture the correlation between optical and SAR features by jointly calculating the collaborative attention in joint feature space. Finally, to automatically learn the varying contributions of both optical and SAR features for classifying different land categories, GHFM is used to fuse both optical and SAR features. Extensive comparative evaluations demonstrate the advantages of CHGFNet within land cover classification over the state-of-the-art methods on three co-registered optical and SAR data sets. Xiao Li 0017, Lin Lei, Yuli Sun, Ming Li 0066, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | SAR Image Speckle Reduction Based on Nonconvex Hybrid Total Variation ModelabstractSpeckle noise inherent in synthetic aperture radar (SAR) images seriously affects the visual effect and brings great difficulties to the postprocessing of the SAR image. Due to the edge-preserving feature, total variation (TV) regularization-based techniques have been extensively utilized to reduce the speckle. However, the strong scatters in SAR image with radiometry several orders of magnitude larger than their surrounding regions limit the effectiveness of TV regularization. Meanwhile, the ℓ1-norm first-order TV regularization sometimes causes staircase artifacts as it favors solutions that are piecewise constant, and it usually underestimates high-amplitude components of image gradient as the ℓ1-norm uniformly penalizes the amplitude. To overcome these shortcomings, a new hybrid variation model, called Fisher-Tippett (FT) distribution-ℓp-norm first-and second-order hybrid TVs (HTpVs), is proposed to reduce the speckle after removing the strong scatters. Especially, the FT-HTpV inherits the advantages of the distribution based data fidelity term, the nonconvex regularization, and the higher order TV regularization. Therefore, it can effectively remove the speckle while preserving point scatters and edges and reducing staircase artifacts well. To efficiently solve the nonconvex minimization problem, an iterative framework with a nonmonotone-accelerated proximal gradient (nmAPG) method and a matrix-vector acceleration strategy are used. Extensive experiments on both the simulated and real SAR images demonstrate the effectiveness of the proposed method. Yuli Sun, Lin Lei, Dongdong Guan, Xiao Li 0017, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Patch Similarity Graph Matrix-Based Unsupervised Remote Sensing Change Detection With Homogeneous and Heterogeneous SensorsabstractChange detection (CD) of remote sensing images is an important and challenging topic, which has found a wide range of applications in many fields. In particular, one of the main challenges is to detect changes between heterogeneous images, where the difference in imaging mechanism makes it difficult to carry out a direct comparison. In this article, we propose an unsupervised CD framework based on the patch similarity graph matrix (PSGM), which assumes that the patch similarity graph structure of each homogeneous or heterogeneous image is consistent if no change occurs. First, it learns the PSGM of one image based on the self-expressive property, which can be interpreted as containing the edges of the fully connected graphs with each image patch as a vertex. Then, the change level depends on how much one image still conforms to the similarity graph structure learned from the other image. Meanwhile, the change map can be further optimized by using the prior sparse knowledge that only a small part of the image changed and most areas remain unchanged. Experiments with both homogeneous and heterogeneous data sets demonstrate the effective performance of the proposed PSGM-based CD method. Yuli Sun, Lin Lei, Xiao Li 0017, Xiang Tan, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | A robust recovery algorithm with smoothing strategies
Yuli Sun, Lin Lei, Xiao Li 0017, Ming Li 0066, Gangyao Kuang |
Neurocomputing | 3 |
| 2020 | Sparse optimization problem with s-difference regularization
Yuli Sun, Xiang Tan, Xiao Li 0017, Lin Lei, Gangyao Kuang |
Signal Process. | 3 |