Yungang Cao

dblp:121/0852 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0003-0495-594XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 A flood knowledge-constrained large language model interactable with GIS: enhancing public risk perception of floods
abstract
Public’s rational flood mitigation behaviors depend on accurate perception of flood risks. The use of natural language for flood risk perception is an effective approach, and it is critical to ensure the accuracy and comprehensibility of the flood information provided by the system in natural language dialogues. This study presents a framework for large language model (LLM) that is constrained by flood knowledge and can interact with geographic information system (GIS), aimed at enhancing the public’s perception of flood risks. We tested the performance of LLM within this framework and the results demonstrate that LLM can generate accurate information about floods under the constraints of entities and relationships in the knowledge graph, and interact with GIS to produce personalized knowledge through real-time coding. Furthermore, we conducted flood risk perception experiments on users with different cognitive levels. The results indicate that using natural language dialogue can narrow the differences brought about by cognitive levels, allowing the public to equally access knowledge related to flood events.
Jun Zhu 0007, Pei Dang, Yungang Cao, Jianbo Lai, Yukun Guo, Ping Wang 0085, Weilian Li
Int. J. Geogr. Inf. Sci.3
2024 The impact of spatial scale on layout learning and individual evacuation behavior in indoor fires: single-scale learning perspectives
abstract
The detail and representation of a spatial layout varies with scale. This affects an individual’s learning effectiveness and understanding, in turn directly influencing their behavior in a fire evacuation. However, the impact of layout learning methods with different spatial scales on fire evacuation behavior, and the relationship between spatial cognition and evacuation effects, remains unclear. We conducted spatial layout learning across three scales with 81 participants and simulated a fire evacuation scenario in a mobile virtual reality for groups. We collected evacuation decision-making and user experience questionnaires as supplementary data. The results demonstrate that small-scale learning objects are the easiest for participants to understand in terms of spatial layout and relationships, but their performance in fire evacuation is poor. Large-scale learning objects significantly improve participants’ evacuation efficiency. Spatial layout learning plays a crucial role in fire evacuation outcomes, but traditional spatial knowledge acquisition measurement methods cannot predict fire evacuation performance. This study sheds light on how spatial cognition influences fire evacuation behavior and provides a more reliable fire evacuation simulation method based on mobile virtual reality (MVR).
Jun Zhu 0007, Pei Dang, Jinbin Zhang, Yungang Cao, Jianlin Wu, Weilian Li, Ya Hu, Jigang You
Int. J. Geogr. Inf. Sci.4
2024 Detail-Optimized Super-Resolution Reconstruction-Based Multistage Training Strategy for Remote Sensing Semantic Segmentation
abstract
Low resolution is a major factor that negatively impacts the accuracy of remote sensing (RS) interpretation. High-quality super-resolution reconstruction (SRR) can help alleviate this problem. In this study, we present a multistage semantic segmentation training strategy called SRSSTS, which is based on detail-optimized SRR. SRSSTS addresses the challenges of poor reconstruction quality and low semantic segmentation (SS) accuracy of RS images. Our approach involves constructing a generative adversarial network (GAN)-based perceptual-loss-dominated SRR network, which generates texture-realistic, high-quality high-resolution images from low-resolution RS images. We then input the generated images into an SS network in steps and propose a three-stage joint loss to generate segmentation results at different resolutions. SRSSTS is straightforward and efficient, and it can be applied to any SS backbone networks. Additionally, this strategy does not increase computational resources compared to other SR-based segmentation methods, since the SRR network is trained separately. We evaluated the effectiveness of our designed SRR network and SRSSTS on five RS datasets (ISPRS Potsdam and Vaihingen, WHDLD, LoveDA Rural and Urban) with different resolutions, achieving excellent performance compared to an SS model for a single task and other SR-based segmentation methods. Moreover, we observed that the learned perceptual image patch similarity (LPIPS) of the super-resolution (SR) reconstructed images, i.e., the stronger the perception ability, the more helpful it is for the SS task.
Baikai Sui, Yungang Cao, Jun Zhu 0007, Yakun Xie
IEEE Trans. Geosci. Remote. Sens.2
2023 Cloud Detection From High-Resolution Remote Sensing Images Based on Convolutional Neural Networks With Geographic Features and Contextual Information
abstract
The diversity and complexity of the subsurface, as well as the similarity of cloud features to those of highlighted surface objects (especially snowy areas), are the main reasons for the current limitations in cloud detection accuracy. This letter proposes a semantic segmentation network fusing geographic features with contextual information for cloud detection named GCI_CD. We design three modules: geographical-cloud information fusion module (GCIFM) (to distinguish between clouds and snow), separable ResNet with CSAM encoder (to enhance channel and spatial association information extraction), and contextual information integration module (CIIM) (to extract multiscale features). To the best of our knowledge, this is the first time that geographic information has been fused in a cloud detection method. The validity of our proposed method is demonstrated using a high-resolution GF-1 dataset containing the near-infrared band and dividing the snow-covered and non-snow-covered areas. The experimental results show that GCI_CD has a better performance not only in snow-covered regions but also in the non-snow-covered regions. Especially for snow-covered images, the IOU of GCI_CD’s cloud detection results are improved by 5% on average compared to other cloud detection methods.
Yungang Cao, Baikai Sui, Hui Qin
IEEE Geosci. Remote. Sens. Lett.1
2023 EGDSR: Encoder-Generator-Decoder Network for Remote Sensing Super-Resolution Reconstruction
abstract
Remote sensing single-image super-resolution reconstruction process detail information is easy to be lost and prone to defocus phenomenon, in order to solve these problems, we are based on the idea of potential spatial vector mapping, focusing on the attribute control of the feature recovery process, to improve the quality of remote sensing image reconstruction. This letter proposes an encoder-generator-decoder super-resolution reconstruction network for remote sensing named EGDSR. We design three modules: multiscale feature extraction and latent code generation module, multi-attribute control of resolution progression module (recovery of latent encoding, fusion of multi-scale features, and generation of high-resolution features), high-resolution image reconstruction module. The experimental results show that the high-resolution images generated by our proposed EGDSR network have a stronger sense of truth, richer texture, and more realistic details, and have a better perceptual effect compared with other state-of-the-art super-resolution networks. In addition, we also combine the semantic segmentation task to assist in verifying the quality of the high-resolution remote sensing images generated by EGDSR, and successfully verify the higher application value of our proposed method.
Baikai Sui, Yungang Cao
IEEE Geosci. Remote. Sens. Lett.2
2023 DTHNet: Dual-Stream Network Based on Transformer and High-Resolution Representation for Shadow Extraction from Remote Sensing Imagery
abstract
Shadow extraction from remote sensing images is critical work. However, the inter-class similarity of shadows with dark water, trees, and roads and the dependence on other objects make accurate shadow extraction still challenging. This letter proposes a dual-stream network based on the transformer and high-resolution representation (DTHNet) for multi-scale shadow extraction. The DTHNet utilizes two streams: the high-resolution representation stream extracts deep features while maintaining detail information, and the clustering representation stream provides clustering feature constraints to enhance the ability to distinguish between foreground and background. Additionally, the designed multi-scale auxiliary predictor and hybrid loss aid in extracting multi-scale shadows. We evaluated our proposed method on the AISD dataset and compared it against six state-of-the-art generic semantic segmentation models and shadow extraction methods. The experimental results demonstrate that the DTHNet outperforms the existing methods.
Yungang Cao, Baikai Sui
IEEE Geosci. Remote. Sens. Lett.2
2023 DF-Mask R-CNN: Direction Field-Based Optimized Instance Segmentation Network for Building Instance Extraction
abstract
Extracting building instances from remote sensing images has various applications. However, existing instance segmentation methods have difficulty maintaining regular and angular building boundaries, corrupting the mask quality. This letter proposes a boundary-optimized instance segmentation method based on the direction field to cope with the low boundary quality. The proposed method develops a DF-Mask head to improve the mask quality of the primary instance segmentation network (Mask R-CNN) through boundary optimization. Within the DF-Mask head, first, a gated directional context-aware module (GDCAM) improves the mask features through the directional context and a designed gating mechanism. Then a direction field rectification module (DFRM) predicts the direction field and iteratively rectifies the mask. To the best of our knowledge, this is the first time that direction field rectification has been developed to the instance segmentation method and applied to building instance extraction. Experimental results on the Vaihingen and Massachusetts datasets suggest that our method outperforms the eight state-of-the-art methods. Compared with Mask R-CNN, our method improves APs by 4.7% and 2.2% on the two datasets, respectively.
Yungang Cao, Baikai Sui
IEEE Geosci. Remote. Sens. Lett.2
2023 IBCO-Net: Integrity-Boundary-Corner Optimization in a General Multistage Network for Building Fine Segmentation From Remote Sensing Images
abstract
Building extraction is a significant topic in high-resolution remote sensing. Insufficient integrity, irregular boundaries, and inaccurate corners remain a problem for existing methods. However, individually optimizing one of these aspects may leave problems in others. Unfortunately, few methods consider integrity, boundary, and corner simultaneously. In this study, we propose a three-stage network (IBCO-Net) incorporating integrity-boundary-corner optimization for fine segmentation of buildings. First, long-range dependent and spatial-continuous blocks (LDSCs) are plugged into the decoder to enhance building integrity. Second, the direction field correction module (DFCM) controls the overall shape of the building by learning the direction field and executing an iterative correction algorithm. Finally, the multi-strategy point refinement module (MSPRM) selects boundary and corner points for re-classification to further refine the boundary and relocate corners. And a hybrid loss function supervises IBCO-Net to optimize each stage. Comparative experiments were conducted on three datasets: the Massachusetts building dataset, the ISPRS Potsdam dataset, and the dataset of building instances of typical cities in China. We evaluated common pixel-level metrics and object-level boundary and corner metrics, with experimental results showing that IBCO-Net outperforms 8 state-of-the-art CNN and Transformer-based methods. In addition, the generality of the proposed method is demonstrated via its performance by applying 9 existing backbone networks.
Yungang Cao, Baikai Sui, Yakun Xie, Jun Zhu 0007
IEEE Trans. Geosci. Remote. Sens.1
2019 Detection of Vegetation Areas Attacked By Pests and Diseases Based on Adaptively Weighted Enhanced Global and Local Deep Features
abstract
Very high resolution (VHR) remote sensing images offer the potential for efficient identification of vegetation areas attacked by pests and diseases without using the near-infrared band. In order to achieve this purpose, various object detection methods are required to perform the extraction task from the VHR images. Traditional pixel-based and object-based classification methods capture little high-level semantic concepts, making it intractable to extract unhealthy vegetation areas. To overcome these problems, this paper proposes a novel method based on the scene level in the framework of the CNN by using adaptively weighted enhanced global and local deep features to extract vegetation areas attacked by pests and diseases from unmanned aerial vehicle (UAV) VHR remote sensing images. To be specific, the last convolutional feature maps of VGG-16 network are firstly extracted as the original features. Then, the learned original features are used for clustering to obtain the scene features, which are regarded as the global features to estimate whether the target object appears in the scene image. Next, each pixel features at different locations of feature maps are used for clustering again to form the object features, which are regarded as local features to determine what the target object is and supplement the global feature with the local information. Afterwards, the shape information of central object in scene image, quantified by the roundness value, is utilized as the foundation of weighting, which can distinguish the different objects with similar spectral and texture feature to some extent. Finally, make use of order of roundness values to adaptively weight both of features, and connect them to form adaptively weighted enhanced global and local deep features. Experimental result indicates that the proposed method outperforms the existing benchmark methods.
Yanshuai Dai, Li Shen 0004, Yungang Cao, Tianjie Lei, Wenfan Qiao
IGARSS3
2019 A Fine-Grained Fully Convolutional Network For Extraction of Building Along High-Speed Rail Lines from VHR Remote Sensing Image
abstract
Very high resolution (VHR) remote sensing offers the potential for efficient identification of building areas along the high-speed rail (HSR) lines. However, high intra-class and low inter-class variances in this type of images and pose serious challenges. Fully convolutional network (FCN) has shown impressive performance and great potential for image classification and object detection. Nevertheless, only classification results with coarse resolutions can be obtained from traditional FCN-based methods. To address this issue, this paper presents a fine-grained fully convolutional network to conduct the building extraction task along the HSR lines. The proposed method introduce a fine-grained upsampling module in the encoder-decoder paradigm to upsample the low-resolution feature maps to higher resolution ones instead of the traditional interpolation processing, thus can reserve and recover more fine-detailed information. Furthermore, an enhanced loss function is designed to impose smoothness prior for building extraction. Experiments on the remote sensing image with 0.5 m spatial resolution from Google Earth covering the part area of Zhengzhou-Xi'an high-speed rail line indicates the proposed method achieve competitive results compared with other state-of-the-art FCN based methods. The study demonstrates the potential application of using high-resolution remote sensing technique to investigate the building areas on both sides of HSR lines regularly, in order to detect the hazardous building areas.
Wenfan Qiao, Li Shen 0004, Yungang Cao, Shi He, Yanshuai Dai
IGARSS4