EDBT 2026 Demo / reviewers in the wild / expert
Rui Li 0036
dblp:96/4282-36
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0001-7858-3160ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LSwinSR: UAV Imagery Super-Resolution Based on Linear Swin TransformerabstractSuper-resolution, which aims to reconstruct high-resolution (HR) images from low-resolution (LR) images, has drawn considerable attention and has been intensively studied in computer vision and remote sensing communities. Super-resolution technology is especially beneficial for unmanned aerial vehicles (UAVs), as the number and resolution of images captured by UAVs are highly limited by physical constraints such as flight altitude and load capacity. In the wake of the successful application of deep learning methods in the super-resolution task, in recent years, a series of super-resolution algorithms have been developed. In this article, for the super-resolution of UAV images, a novel network based on the state-of-the-art Swin Transformer is proposed with better efficiency and competitive accuracy. Meanwhile, as one of the essential applications of the UAV is land cover and land use monitoring, simple image quality assessments such as the peak-signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) are not enough to comprehensively measure the performance of an algorithm. Therefore, we further investigate the effectiveness of super-resolution methods using the accuracy of semantic segmentation. The code is available athttps://github.com/lironui/GeoSR. Rui Li 0036, Xiaowei Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Efficient large-scale oblique image matching based on cascade hashing and match data scheduling
Shunyi Zheng, Ce Zhang 0005, Xiqi Wang, Rui Li 0036 |
Pattern Recognit. | 5 |
| 2022 | MACU-Net for Semantic Segmentation of Fine-Resolution Remotely Sensed ImagesabstractSemantic segmentation of remotely sensed images plays an important role in land resource management, yield estimation, and economic assessment. U-Net, a deep encoder–decoder architecture, has been used frequently for image segmentation with high accuracy. In this letter, we incorporate multiscale features generated by different layers of U-Net and design a multiscale skip connected and asymmetric-convolution-based U-Net (MACU-Net), for segmentation using fine-resolution remotely sensed images. Our design has the following advantages: (1) the multiscale skip connections combine and realign semantic features contained in both low-level and high-level feature maps; (2) the asymmetric convolution block strengthens the feature representation and feature extraction capability of a standard convolution layer. Experiments conducted on two remotely sensed data sets captured by different satellite sensors demonstrate that the proposed MACU-Net transcends the U-Net, U-Netpyramid pooling layers (PPL), U-Net 3+, among other benchmark approaches. Code is available athttps://github.com/lironui/MACU-Net. Rui Li 0036, Chenxi Duan, Shunyi Zheng, Ce Zhang 0005, Peter M. Atkinson |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Multistage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing ImagesabstractThe attention mechanism can refine the extracted feature maps and boost the classification performance of the deep network, which has become an essential technique in computer vision and natural language processing. However, the memory and computational costs of the dot-product attention mechanism increase quadratically with the spatiotemporal size of the input. Such growth hinders the usage of attention mechanisms considerably in application scenarios with large-scale inputs. In this letter, we propose a linear attention mechanism (LAM) to address this issue, which is approximately equivalent to dot-product attention with computational efficiency. Such a design makes the incorporation between attention mechanisms and deep networks much more flexible and versatile. Based on the proposed LAM, we refactor the skip connections in the raw U-Net and design a multistage attention ResU-Net (MAResU-Net) for semantic segmentation from fine-resolution remote sensing images. Experiments conducted on the Vaihingen data set demonstrated the effectiveness and efficiency of our MAResU-Net. Our code is available athttps://github.com/lironui/MAResU-Net. Rui Li 0036, Shunyi Zheng, Chenxi Duan, Jianlin Su, Ce Zhang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Class-Guided Swin Transformer for Semantic Segmentation of Remote Sensing ImageryabstractSemantic segmentation of remote sensing images plays a crucial role in a wide variety of practical applications, including land cover mapping, environmental protection, and economic assessment. In the last decade, convolutional neural network (CNN) is the mainstream deep learning-based method of semantic segmentation. Compared with conventional methods, CNN-based methods learn semantic features automatically, thereby achieving strong representation capability. However, the local receptive field of the convolution operation limits CNN-based methods from capturing long-range dependencies. In contrast, Vision Transformer (ViT) demonstrates its great potential in modeling long-range dependencies and obtains superior results in semantic segmentation. Inspired by this, in this letter, we propose a class-guided Swin Transformer (CG-Swin) for semantic segmentation of remote sensing images. Specifically, we adopt a Transformer-based encoder–decoder structure, which introduces the Swin Transformer backbone as the encoder and designs a class-guided Transformer block to construct the decoder. The experimental results on ISPRS Vaihingen and Potsdam datasets demonstrate the significant breakthrough of the proposed method over ten benchmarks, outperforming both advanced CNN-based and recent Transformer-based approaches. Xiaoliang Meng, Yuechi Yang, Rui Li 0036, Ce Zhang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing ImagesabstractThe fully convolutional network (FCN) with an encoder-decoder architecture has been the standard paradigm for semantic segmentation. The encoder-decoder architecture utilizes an encoder to capture multilevel feature maps, which are incorporated into the final prediction by a decoder. As the context is crucial for precise segmentation, tremendous effort has been made to extract such information in an intelligent fashion, including employing dilated/atrous convolutions or inserting attention modules. However, these endeavors are all based on the FCN architecture with ResNet or other backbones, which cannot fully exploit the context from the theoretical concept. By contrast, we introduce the Swin Transformer as the backbone to extract the context information and design a novel decoder of densely connected feature aggregation module (DCFAM) to restore the resolution and produce the segmentation map. The experimental results on two remotely sensed semantic segmentation datasets demonstrate the effectiveness of the proposed scheme. Rui Li 0036, Chenxi Duan, Ce Zhang 0005, Xiaoliang Meng, Shenghui Fang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multiattention Network for Semantic Segmentation of Fine-Resolution Remote Sensing ImagesabstractSemantic segmentation of remote sensing images plays an important role in a wide range of applications, including land resource management, biosphere monitoring, and urban planning. Although the accuracy of semantic segmentation in remote sensing images has been increased significantly by deep convolutional neural networks, several limitations exist in standard models. First, for encoder–decoder architectures such as U-Net, the utilization of multiscale features causes the underuse of information, where low-level features and high-level features are concatenated directly without any refinement. Second, long-range dependencies of feature maps are insufficiently explored, resulting in suboptimal feature representations associated with each semantic class. Third, even though the dot-product attention mechanism has been introduced and utilized in semantic segmentation to model long-range dependencies, the large time and space demands of attention impede the actual usage of attention in application scenarios with large-scale input. This article proposed a multiattention network (MANet) to address these issues by extracting contextual dependencies through multiple efficient attention modules. A novel attention mechanism of kernel attention with linear complexity is proposed to alleviate the large computational demand in attention. Based on kernel attention and channel attention, we integrate local feature maps extracted by ResNet-50 with their corresponding global dependencies and reweight interdependent channel maps adaptively. Numerical experiments on two large-scale fine-resolution remote sensing datasets demonstrate the superior performance of the proposed MANet. Code is available athttps://github.com/lironui/Multi-Attention-Network. Rui Li 0036, Shunyi Zheng, Ce Zhang 0005, Chenxi Duan, Jianlin Su, Peter M. Atkinson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Building Extraction With Vision TransformerabstractAs an important carrier of human productive activities, the extraction of buildings is not only essential for urban dynamic monitoring but also necessary for suburban construction inspection. Nowadays, accurate building extraction from remote sensing images remains a challenge due to the complex background and diverse appearances of buildings. The convolutional neural network (CNN) based building extraction methods, although increased the accuracy significantly, are criticized for their inability for modelling global dependencies. Thus, this paper applies the Vision Transformer for building extraction. However, the actual utilization of the Vision Transformer often comes with two limitations. First, the Vision Transformer requires more GPU memory and computational costs compared to CNNs. This limitation is further magnified when encountering large-sized inputs like fine-resolution remote sensing images. Second, spatial details are not sufficiently preserved during the feature extraction of the Vision Transformer, resulting in the inability for fine-grained building segmentation. To handle these issues, we propose a novel Vision Transformer (BuildFormer), with a dual-path structure. Specifically, we design a spatial-detailed context path to encode rich spatial details and a global context path to capture global dependencies. Besides, we develop a window-based linear multi-head self-attention to make the complexity of the multi-head self-attention linear with the window size, which strengthens the global context extraction by using large windows and greatly improves the potential of the Vision Transformer in processing large-sized remote sensing images. The proposed method yields state-of-the-art performance (75.74% IoU) on the Massachusetts building dataset. Code will be available at https://github.com/WangLibo1995/BuildFormer. Shenghui Fang, Xiaoliang Meng, Rui Li 0036 |
IEEE Trans. Geosci. Remote. Sens. | 4 |