Baikai Sui

dblp:264/0766 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0003-1244-5771ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Detail-Optimized Super-Resolution Reconstruction-Based Multistage Training Strategy for Remote Sensing Semantic Segmentation
abstract
Low resolution is a major factor that negatively impacts the accuracy of remote sensing (RS) interpretation. High-quality super-resolution reconstruction (SRR) can help alleviate this problem. In this study, we present a multistage semantic segmentation training strategy called SRSSTS, which is based on detail-optimized SRR. SRSSTS addresses the challenges of poor reconstruction quality and low semantic segmentation (SS) accuracy of RS images. Our approach involves constructing a generative adversarial network (GAN)-based perceptual-loss-dominated SRR network, which generates texture-realistic, high-quality high-resolution images from low-resolution RS images. We then input the generated images into an SS network in steps and propose a three-stage joint loss to generate segmentation results at different resolutions. SRSSTS is straightforward and efficient, and it can be applied to any SS backbone networks. Additionally, this strategy does not increase computational resources compared to other SR-based segmentation methods, since the SRR network is trained separately. We evaluated the effectiveness of our designed SRR network and SRSSTS on five RS datasets (ISPRS Potsdam and Vaihingen, WHDLD, LoveDA Rural and Urban) with different resolutions, achieving excellent performance compared to an SS model for a single task and other SR-based segmentation methods. Moreover, we observed that the learned perceptual image patch similarity (LPIPS) of the super-resolution (SR) reconstructed images, i.e., the stronger the perception ability, the more helpful it is for the SS task.
Baikai Sui, Yungang Cao, Jun Zhu 0007, Yakun Xie
IEEE Trans. Geosci. Remote. Sens.1
2023 Cloud Detection From High-Resolution Remote Sensing Images Based on Convolutional Neural Networks With Geographic Features and Contextual Information
abstract
The diversity and complexity of the subsurface, as well as the similarity of cloud features to those of highlighted surface objects (especially snowy areas), are the main reasons for the current limitations in cloud detection accuracy. This letter proposes a semantic segmentation network fusing geographic features with contextual information for cloud detection named GCI_CD. We design three modules: geographical-cloud information fusion module (GCIFM) (to distinguish between clouds and snow), separable ResNet with CSAM encoder (to enhance channel and spatial association information extraction), and contextual information integration module (CIIM) (to extract multiscale features). To the best of our knowledge, this is the first time that geographic information has been fused in a cloud detection method. The validity of our proposed method is demonstrated using a high-resolution GF-1 dataset containing the near-infrared band and dividing the snow-covered and non-snow-covered areas. The experimental results show that GCI_CD has a better performance not only in snow-covered regions but also in the non-snow-covered regions. Especially for snow-covered images, the IOU of GCI_CD’s cloud detection results are improved by 5% on average compared to other cloud detection methods.
Yungang Cao, Baikai Sui, Hui Qin
IEEE Geosci. Remote. Sens. Lett.2
2023 EGDSR: Encoder-Generator-Decoder Network for Remote Sensing Super-Resolution Reconstruction
abstract
Remote sensing single-image super-resolution reconstruction process detail information is easy to be lost and prone to defocus phenomenon, in order to solve these problems, we are based on the idea of potential spatial vector mapping, focusing on the attribute control of the feature recovery process, to improve the quality of remote sensing image reconstruction. This letter proposes an encoder-generator-decoder super-resolution reconstruction network for remote sensing named EGDSR. We design three modules: multiscale feature extraction and latent code generation module, multi-attribute control of resolution progression module (recovery of latent encoding, fusion of multi-scale features, and generation of high-resolution features), high-resolution image reconstruction module. The experimental results show that the high-resolution images generated by our proposed EGDSR network have a stronger sense of truth, richer texture, and more realistic details, and have a better perceptual effect compared with other state-of-the-art super-resolution networks. In addition, we also combine the semantic segmentation task to assist in verifying the quality of the high-resolution remote sensing images generated by EGDSR, and successfully verify the higher application value of our proposed method.
Baikai Sui, Yungang Cao
IEEE Geosci. Remote. Sens. Lett.1
2023 DTHNet: Dual-Stream Network Based on Transformer and High-Resolution Representation for Shadow Extraction from Remote Sensing Imagery
abstract
Shadow extraction from remote sensing images is critical work. However, the inter-class similarity of shadows with dark water, trees, and roads and the dependence on other objects make accurate shadow extraction still challenging. This letter proposes a dual-stream network based on the transformer and high-resolution representation (DTHNet) for multi-scale shadow extraction. The DTHNet utilizes two streams: the high-resolution representation stream extracts deep features while maintaining detail information, and the clustering representation stream provides clustering feature constraints to enhance the ability to distinguish between foreground and background. Additionally, the designed multi-scale auxiliary predictor and hybrid loss aid in extracting multi-scale shadows. We evaluated our proposed method on the AISD dataset and compared it against six state-of-the-art generic semantic segmentation models and shadow extraction methods. The experimental results demonstrate that the DTHNet outperforms the existing methods.
Yungang Cao, Baikai Sui
IEEE Geosci. Remote. Sens. Lett.3
2023 DF-Mask R-CNN: Direction Field-Based Optimized Instance Segmentation Network for Building Instance Extraction
abstract
Extracting building instances from remote sensing images has various applications. However, existing instance segmentation methods have difficulty maintaining regular and angular building boundaries, corrupting the mask quality. This letter proposes a boundary-optimized instance segmentation method based on the direction field to cope with the low boundary quality. The proposed method develops a DF-Mask head to improve the mask quality of the primary instance segmentation network (Mask R-CNN) through boundary optimization. Within the DF-Mask head, first, a gated directional context-aware module (GDCAM) improves the mask features through the directional context and a designed gating mechanism. Then a direction field rectification module (DFRM) predicts the direction field and iteratively rectifies the mask. To the best of our knowledge, this is the first time that direction field rectification has been developed to the instance segmentation method and applied to building instance extraction. Experimental results on the Vaihingen and Massachusetts datasets suggest that our method outperforms the eight state-of-the-art methods. Compared with Mask R-CNN, our method improves APs by 4.7% and 2.2% on the two datasets, respectively.
Yungang Cao, Baikai Sui
IEEE Geosci. Remote. Sens. Lett.3
2023 IBCO-Net: Integrity-Boundary-Corner Optimization in a General Multistage Network for Building Fine Segmentation From Remote Sensing Images
abstract
Building extraction is a significant topic in high-resolution remote sensing. Insufficient integrity, irregular boundaries, and inaccurate corners remain a problem for existing methods. However, individually optimizing one of these aspects may leave problems in others. Unfortunately, few methods consider integrity, boundary, and corner simultaneously. In this study, we propose a three-stage network (IBCO-Net) incorporating integrity-boundary-corner optimization for fine segmentation of buildings. First, long-range dependent and spatial-continuous blocks (LDSCs) are plugged into the decoder to enhance building integrity. Second, the direction field correction module (DFCM) controls the overall shape of the building by learning the direction field and executing an iterative correction algorithm. Finally, the multi-strategy point refinement module (MSPRM) selects boundary and corner points for re-classification to further refine the boundary and relocate corners. And a hybrid loss function supervises IBCO-Net to optimize each stage. Comparative experiments were conducted on three datasets: the Massachusetts building dataset, the ISPRS Potsdam dataset, and the dataset of building instances of typical cities in China. We evaluated common pixel-level metrics and object-level boundary and corner metrics, with experimental results showing that IBCO-Net outperforms 8 state-of-the-art CNN and Transformer-based methods. In addition, the generality of the proposed method is demonstrated via its performance by applying 9 existing backbone networks.
Yungang Cao, Baikai Sui, Yakun Xie, Jun Zhu 0007
IEEE Trans. Geosci. Remote. Sens.3