VLDB 2026 Research / reviewers in the wild / expert
Daoyuan Zheng
dblp:260/5384
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-0344-1760ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Self-Prompt Calibration Network Based on Segment Anything Model 2 for High-Resolution Remote Sensing Image SegmentationabstractRemote sensing image segmentation is particularly difficult due to the coexistence of large-scale variations and fine-grained structures in very high-resolution imagery. Conventional CNN-based or Transformer-based networks often struggle to capture global context while preserving boundary details, leading to degraded performance on small or thin objects. To address these challenges, we propose a Self-Prompt Calibration Network based on Segment Anything Model 2 (SC-SAM). The SC-SAM achieves self-prompt by feeding mask prompts from a lightweight decoder into frozen prompt encoder. Output calibration is achieved through the proposed Cross-Probability Guided Calibration module, which employs cross-probability uncertainty as complementary guidance to refine final predictions via self-prompted outputs. Furthermore, to better preserve contextual and structural information across multiple scales, a scale-decoupled kernel mixture module is designed. Experimental results on the ISPRS Vaihingen and Potsdam dataset demonstrate that the proposed approach surpasses state-of-the-art methods by 1.02% and 1.34% in mIoU, highlighting its effectiveness. This study provides new insights into adapting SAM for domain-specific remote sensing segmentation tasks. Yizhou Lan, Daoyuan Zheng, Ke Shang 0002, Feizhou Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | A non-overlapping image stitching method for reconstruction of page in ancient Chinese books
Yizhou Lan, Daoyuan Zheng, Qingwu Hu, Shaohua Wang 0003, Shunli Wang 0003, Tong Yue, Jiayuan Li 0001 |
Comput. Vis. Image Underst. | 2 |
| 2025 | FTransDeepLab: Multimodal Fusion Transformer-Based DeepLabv3+ for Remote Sensing Semantic SegmentationabstractHigh-resolution remote sensing images contain rich color and texture information, but due to the inherent limitations of 2-D data, achieving high-quality semantic segmentation remains a challenge. Multimodal data fusion technology has emerged as an effective approach to overcome this issue. To accurately capture the semantic information in remote sensing images, this study designs a multimodal fusion Transformer-based DeepLabv3+ model for remote sensing semantic segmentation, named FTransDeepLab. Specifically, the network learns features from two modalities and is inspired by the DeepLab architecture. We extended the encoder by stacking the multiscale Segformer, encoding the input images into highly representative spatial features. Additionally, we introduced the multimodal feature rectification (MFR) module and the multimodal feature fusion (MFF) module. The MFR, composed of a channel attention module and a spatial attention module, enhances the model’s ability to capture essential features and improves performance by focusing on both global and local contexts. The MFF module utilizes a cross-attention mechanism to optimize the feature fusion process, which enhances representation learning by facilitating the interaction between diverse information and integrates features from different modalities. Finally, in the decoding path, the extracted high-level features are concatenated with low-level features to optimize the feature representation and upsampled to restore the size of input image. Extensive results on two datasets, the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam, have confirmed that the proposed FTransDeepLab can achieve superior performance compared to the state-of-the-art segmentation methods. Haixia Feng, Qingwu Hu, Shunli Wang 0003, Mingyao Ai, Daoyuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | CLC²SIE: Cross-Level Consistency Constraint for Semisupervised Building Instance Extraction From High-Resolution Remote Sensing ImageryabstractSemi-supervised building instance extraction aims to learn from limited labeled data alongside an extensive collection of unlabeled data, offering a promising approach for extracting building instances from high-resolution (HR) remote sensing images (RSIs). However, complex RSIs frequently encounter challenges such as background interference and intricate noise, struggling in generating reliable pseudo labels. To alleviate this, we propose a novel cross-level consistency constraint semi-supervised building instance extraction method (CLC2SIE) to enhance pseudo label generation. Specifically, CLC2SIE contains two core modules: object-level dynamic consistency (OLDC) and pixel-level saliency consistency (PLSC). The OLDC module dynamically converts building features from background into valuable supplementary information, enhancing the model’s perception of building instances in complex scenes. Additionally, the PLSC module is designed to mitigate the boundary noise in pseudo labels by saliency guidance, which improves model’s awareness of building contours. By co-learning these modules in an end-to-end manner, CLC2SIE facilitates pseudo label generation and improves extraction performance. Experiments were conducted on the three public building datasets, i.e., WHU, CrowdAI and TCC, demonstrate that CLC2SIE achieves superior performance compared to state-of-the-art semi-supervised instance extraction methods at different labeling ratios. This study explores a novel semi-supervised learning (SSL) framework that exploits cross-level consistency to improve pseudo label generation, offering a methodological reference for various SSL applications in RSIs. Jun Pan 0001, Fang Fang 0008, Daoyuan Zheng, Shengwen Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | CSNet: Change Selection of Activations and Pseudomasks for Image-Level Weakly Supervised Change DetectionabstractWeakly supervised change detection (WSCD) of bi-temporal remote sensing (RS) images has gained attention for its ability to reduce reliance on labor-intensive pixel-level change masks. Recent methods leverage image-level weak supervision to generate change pseudomasks for detecting changed objects, typically using class activation map (CAM) technique combined with DenseCRF or the Segment Anything Model (SAM). However, these methods still face two main challenges: first, CAMs tend to produce weak or false activations for changed objects, and second, DenseCRF and SAM lead to unreliable pseudomask generation, particularly when complex variations occur within objects in bi-temporal images. To address these challenges, a change selection network (CSNet) is proposed to enhance the quality of change activation maps and pseudomasks, improving their ability to accurately extract changed regions in bi-temporal RS images. First, a change activation selection (CAS) module is designed to generate a weight mask that selects and aggregates change-representing features, effectively highlighting missed change activations and strengthening weak activations. Second, a bi-temporal image selection (BIS) strategy is developed, incorporating two selection rules to filter out image pairs with poor-quality mask derived from SAM, while retaining those with high-quality results. Finally, a change pseudomask generation (CPG) module integrated with an atrous-spatial pyramid pooling (ASPP) classifier is developed to predict accurate change pixels for final pseudomask generation. Experimental results demonstrate that the proposed CSNet outperforms existing WSCD methods, achieving 79.32% IoU in change pseudomasks for the WHU-CD dataset, 68.12% for the GZ-CD dataset and 75.44% for the GVLM dataset. This study proposes a novel method that enhances the performance of the weakly supervised paradigm in RS CD. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | AMIANet: Asymmetric Multimodal Interactive Augmentation Network for Semantic Segmentation of Remote Sensing ImageryabstractIn recent years, the inherent 2-D characteristics of optical images have led to a plateau in semantic segmentation performance. The complementary nature of light detection and ranging (LiDAR) point clouds and camera images can effectively enhance semantic segmentation capabilities, and thus, research into multimodal joint semantic segmentation is garnering increasing attention. However, the domain gaps between different dimensions present challenges for the fusion of multimodal data. In this article, we introduce a novel asymmetric multimodal interaction augmented network (AMIANet), which directly processes heterogeneous data from images and point clouds. The treatment of the disparities in modal data ensures consistency in the features of both modes. Through the newly developed synergistic multimodal interaction module (SMI Module), AMIANet is capable of combining the complementary characteristics of cross-modal data. This is achieved by interactively fusing and extracting precise and rich structural information from point cloud features to enhance image characteristics. The experimental results on the N3C-California, WHU-RRDSD, and ISPRS Vaihingen datasets demonstrate that AMIANet surpasses benchmark methods and current state-of-the-art (SOTA) approaches. The code will be available athttps://github.com/2012153946/AMIANet. Qingwu Hu, Wenlei Fan, Haixia Feng, Daoyuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Weakly Supervised Building Extraction From High-Resolution Remote Sensing Images Based on Building-Aware Clustering and Activation Refinement NetworkabstractWeakly supervised building extraction methods, utilizing image-level labels, offer a cost-effective solution by significantly reducing the need for pixel-level annotation in high-resolution (HR) remote sensing (RS) images. These methods often focus on class activation map (CAM) optimization based on features extracted from individual images, missing out on the benefits of associating building features from multiple RS images (i.e., n images) to improve CAMs. This limitation leaves room for improvement in both CAM optimization and pseudo-mask generation. To address this, we propose the building-aware clustering and activation refinement network (BAC-AR-Net), a novel weakly supervised network to enhance weakly supervised building extraction performance. The building-aware clustering (BAC) module aggregates and clusters feature maps from multiple building samples to obtain common features of buildings. The common features are subsequently used to extract regions with similar building semantics, thereby enhancing the accuracy and completeness of building coverage in CAMs. Additionally, the activation refinement module is designed to generate pseudo-masks with clear boundaries and an effective separation of buildings and background. Experiments were conducted on the ISPRS Potsdam and Vaihingen datasets as well as a self-built building dataset to verify the effectiveness of our proposed method. The results show the proposed method outperforms both the weakly supervised semantic segmentation and weakly supervised building extraction methods that use image-level labels, achieving IoU accuracies of 0.8556, 0.8163, and 0.7797 on the respective datasets. This study introduces a novel weakly supervised learning framework to the RS application, with a particular focus on building extraction and semantic segmentation tasks. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Jiayuan Li 0001, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Spatial Extent-Aware Multimodal Fusion Method for Measuring Urban Socioeconomic StatusabstractThe automatic measurement of socioeconomic status (SES), such as household income, provides fundamental data for policymakers and business applications. Although several urban data sources were employed in previous studies, the spatial extent of the measured objects was ignored, which left room for further improvement in the accuracy of measuring SES. This study develops a multimodal semantic segmentation framework to fuse the spatial extent of ground objects, remote sensing, and near-sensing images for predicting SES levels. The framework first constructs ground feature layer (GFL) tiles by projecting ground-level features from sparse ground images. Then, the ground feature tiles aggregate regional ground-level features with the help of the spatial extents of ranges. Lastly, an improved deep semantic segmentation network is employed to fuse GFL and remote sensing images to predict SES levels. Experimental results on the London dataset show that the proposed method outperforms SOTA models and is robust. The framework provides a rewarding exploration in fusing image data and spatial vector data for image-based intelligent applications and can be applied to a series of socioeconomic applications. Fang Fang 0008, Shengwen Li, Daoyuan Zheng, Linyun Zeng, Bo Wan 0006 |
IGARSS | 4 |
| 2023 | Utilizing Bounding Box Annotations for Weakly Supervised Building Extraction From Remote-Sensing ImagesabstractImage-level weakly supervised semantic segmentation (WSSS) methods have greatly facilitated the extraction of buildings from remote sensing (RS) images. However, the lack of the locations and extents of individual buildings in image-level labels results in some limitations of the methods, especially in the cases of cluttered backgrounds, diverse building shapes and sizes. By utilizing bounding box annotations, a novel WSSS model is developed to improve building extraction from RS images in this paper. Specifically, during the training phase, a multiscale feature retrieval (MFR) module is designed to learn multiscale building features and suppress the background noise inside the bounding box. In the inference phase, multiscale class activation maps (CAM) are generated from multiscale features to achieve accurate building localization. Finally, a pseudo mask generation and correction (PGC) module refines the CAMs to generate and correct the building pseudo masks. Experiments are conducted to examine the proposed model in three datasets, namely, the WHU aerial building dataset, the CrowdAI building dataset, and a self-annotated building dataset. Experimental results demonstrate that the proposed method outperforms baselines, achieving 76.99%, 75.51% and 67.35% in terms of IoU scores on the three challenging datasets, respectively. This paper provides a methodological reference for the application of weakly supervised learning on RS images. Daoyuan Zheng, Shengwen Li, Fang Fang 0008, Bo Wan 0006, Yuanyuan Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | DSM-Assisted Unsupervised Domain Adaptive Network for Semantic Segmentation of Remote Sensing Imagery
Shunping Zhou, Shengwen Li, Daoyuan Zheng, Fang Fang 0008, Yuanyuan Liu 0004, Bo Wan 0006 |
IEEE Trans. Geosci. Remote. Sens. | 4 |