EDBT 2026 Demo / reviewers in the wild / expert
Yupeng Deng 0001
dblp:283/2592-1
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-8396-0338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynamicEarth: How Far Are We from Open-Vocabulary Change Detection?abstractMonitoring Earth's evolving land covers requires methods capable of detecting changes across a wide range of categories and contexts. Existing change detection methods are hindered by their dependency on predefined classes, reducing their effectiveness in open-world applications. To address this issue, we introduce open-vocabulary change detection (OVCD), a novel task that bridges vision and language to detect changes across any category. Considering the lack of high-quality data and annotation, we propose two training-free frameworks, M-C-I and I-M-C, which leverage and integrate off-the-shelf foundation models for the OVCD task. The insight behind the M-C-I~framework is to discover all potential changes and then classify these changes, while the insight of I-M-C~framework is to identify all targets of interest and then determine whether their states have changed. Based on these two frameworks, we instantiate to obtain several methods, e.g., SAM-DINOv2-SegEarth-OV, Grounding-DINO-SAM2-DINO, etc. Extensive evaluations on 4 benchmark datasets demonstrate the superior generalization and robustness of our OVCD methods over existing supervised and unsupervised methods. To support continued exploration, we release DynamicEarth, a dedicated codebase designed to advance research and application of OVCD. Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Chao Pang 0001, Zepeng Xin, Tieliang Gong, Deyu Meng, Zhi Wang 0002 |
AAAI | 3 |
| 2025 | Open-CD: A Comprehensive Toolbox for Change DetectionabstractWe present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd. Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao |
ACM Multimedia | 4 |
| 2025 | SemiCD-VL: Visual-Language Model Guidance Makes Better Semi-Supervised Change DetectorabstractChange detection (CD) aims to identify pixels with semantic changes between images. However, annotating massive numbers of pixel-level images is labor-intensive and costly, especially for multitemporal images, which require pixel-wise comparisons by human experts. Considering the excellent performance of visual-language models (VLMs) for zero-shot, OV, etc., with prompt-based reasoning, it is promising to utilize VLMs to make better CD under limited labeled data. In this article, we propose a VLM guidance-based semi-supervised CD method, namely SemiCD-VL. The insight of SemiCD-VL is to synthesize free change labels using VLMs to provide additional supervision signals for unlabeled data. However, almost all current VLMs are designed for single-temporal images and cannot be directly applied to bi- or multitemporal images. Motivated by this, we first propose a VLM-based mixed change event generation (CEG) strategy to yield pseudo-labels for unlabeled CD data. Since the additional supervised signals provided by these VLM-driven pseudo-labels may conflict with the original pseudo-labels from the consistency regularization paradigm (e.g., FixMatch), we propose the dual projection head for de-entangling different signal sources. Further, we explicitly decouple the bitemporal images semantic representation through two auxiliary segmentation decoders, which are also guided by VLM. Finally, to make the model more adequately capture change representations, we introduce contrastive consistency regularization (CCR) by constructing feature-level contrastive loss in auxiliary branches. Extensive experiments show the advantage of SemiCD-VL. For instance, SemiCD-VL improves the FixMatch baseline by$+ 5.3~\text {IoU}^{c}$on WHU-CD and by$+ 2.4~\text {IoU}^{c}$on LEVIR-CD with 5% labels, and SemiCD-VL requires only 5%–10% of the labels to achieve performance similar to the supervised methods. In addition, our CEG strategy, in an unsupervised manner, can achieve performance far superior to state-of-the-art (SOTA) unsupervised CD methods (e.g., IoU improved from 18.8% to 46.3% on LEVIR-CD dataset). The code is available athttps://github.com/likyoo/SemiCD-VL. Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Junmin Liu, Deyu Meng, Zhi Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | PolyFootNet: Extracting Polygonal Building Footprints in Off-Nadir Remote Sensing ImagesabstractExtracting polygonal building footprints from off-nadir imagery is crucial for diverse applications. Current deep-learning-based extraction approaches predominantly rely on semantic segmentation paradigms and postprocessing algorithms, limiting their boundary precision and applicability. However, existing polygonal extraction methodologies are inherently designed for near-nadir imagery and fail under the geometric complexities introduced by off-nadir viewing angles. To address these challenges, this article introduces the polygonal footprint network (PolyFootNet), a novel deep-learning framework that directly outputs polygonal building footprints without requiring external postprocessing steps. The PolyFootNet employs a high-quality mask prompter to generate precise roof masks, which guide polygonal vertex extraction in a unified model pipeline. A key contribution of the PolyFootNet is introducing the self-offset attention (SOFA) mechanism, grounded in Nadaraya–Watson regression, to effectively mitigate the accuracy discrepancy observed between low-rise and high-rise buildings. This approach allows low-rise building predictions to leverage angular corrections learned from high-rise building offsets, significantly enhancing overall extraction accuracy. Additionally, motivated by the inherent ambiguity of building footprint extraction (BFE) tasks, we systematically investigate alternative extraction paradigms and demonstrate that a combined approach of building masks and offsets achieves superior polygonal footprint results. Extensive experiments validate PolyFootNet’s effectiveness, illustrating its promising potential as a robust, generalizable, and precise polygonal BFE method from challenging off-nadir imagery. To facilitate further research, we will release pretrained weights of our offset prediction module athttps://github.com/likaiucas/PolyFootNet Kai Li 0025, Yupeng Deng 0001, Jingbo Chen, Yu Meng 0002, Zhihao Xi, Junxian Ma, Maolin Wang 0001, Xiangyu Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | IRSAMap: Toward Large-Scale, High-Resolution Land Cover Map VectorizationabstractWith the continuous enhancement of remote sensing image resolution and the rapid advancement of deep learning techniques, land cover mapping is undergoing a significant transformation from pixel-level segmentation to object-based vector modeling. This shift imposes higher demands on deep learning models, requiring not only precise delineation of object boundaries but also the preservation of topological consistency among geographic elements. However, existing public datasets face three major limitations: limited class annotations, restricted data scale, and the lack of spatial structural information, which severely hinder the development of breakthrough methods in high-resolution remote sensing vectorization. To address these challenges, we present IRSAMap, the first global remote sensing dataset designed for large-scale, high-resolution, multi-feature land cover vector mapping. This dataset offers four key advantages: First, a comprehensive element vector annotation system that includes over 1.8 million instances of 10 typical natural and man-made objects, such as buildings, roads, rivers, and trees, employing a unified vector annotation standard framework that ensures both semantic integrity and spatial structural accuracy. Second, an intelligent annotation workflow incorporating “manual pre-annotation + AI-based training and inference + manual review and correction,” which enhances annotation efficiency while ensuring consistency. Third, a global coverage that spans 67 regions across six continents, representing diverse terrain types, including urban and rural areas, with a total coverage area exceeding 1,000 square kilometers. Fourth, multi-task adaptability, supporting various tasks such as pixel-level land cover classification, building outline regularization extraction, road centerline extraction, and panoramic segmentation. As a fundamental resource for remote sensing intelligent interpretation, IRSAMap provides a standardized benchmark for the paradigm shift from pixels to objects, which will significantly advance the development of high-precision geographic feature automation, collaborative modeling, and other cutting-edge research directions. The dataset is of great value for applications such as global geographic information updating and digital twin construction. IRSAMap is publicly available at https://github.com/ucas-dlg/IRSAMap. Yu Meng 0002, Ligao Deng, Zhihao Xi, Jingbo Chen, Anzhi Yue, Diyou Liu, Kai Li 0025, Kaiyu Li 0001, Yupeng Deng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 11 |
| 2025 | PCP: A Prompt-Based Cartographic-Level Polygonal Vector Extraction Framework for Remote Sensing ImagesabstractAccurate cartographic-level polygonal vector extraction from remote sensing images is crucial for land survey and land cover mapping. However, current deep learning-based segmentation models often generate raster outputs with irregular boundaries, making them unsuitable for applications requiring precise vector data. This study presents the Prompt-based Cartographic-level Polygonal (PCP) Vector Extraction framework, which leverages segmentation results as prompts to guide the extraction process and enhance regularity. Built upon the Segment Anything Model (SAM), the PCP extracts the segmentation mask and an additional vertex map. To improve vertex extraction accuracy and enable more regular polygon generation, two key modules are introduced: the Iterative Upscaling Refinement (IUR) module, which addresses challenges related to low-resolution feature maps, and the Shape Rule-Based Vertex Filtering and Connecting (SRVFC) module, which enhances the vertex filtering process by learning from shape features. Experimental results on the Land Survey Vector (LSV), LoveDA and WHU-Mix (Vector) datasets demonstrate that the PCP outperforms existing methods in terms of vertex precision and recall, reflecting geometric accuracy, as well as in complexity-aware Intersection over Union (C-IoU), which balances overall accuracy and simplicity. Thus, the PCP framework provides a promising solution for cartographic-level vector extraction in remote sensing applications. The source code is available at https://github.com/wchh-2000/PCP. Zhihao Xi, Diyou Liu, Yuman Feng, Yupeng Deng 0001, Kai Li 0025, Jingbo Chen, Yu Meng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Prompt-Driven Building Footprint Extraction in Aerial Images With Offset-Building ModelabstractMore accurate extraction of invisible building footprints from very-high-resolution (VHR) aerial images relies on roof segmentation and roof-to-footprint offset extraction. Existing methods based on instance segmentation suffer from poor generalization when extended to large-scale data production and fail to achieve low-cost human interaction. This prompt paradigm inspires us to design a promptable framework for roof and offset extraction, and transforms end-to-end algorithms into promptable methods. Within this framework, we propose a novel offset-building model (OBM). Based on prompt prediction, we first discover a common pattern of predicting offsets and tailored Distance-NMS (DNMS) algorithms for offset optimization. To rigorously evaluate the algorithm’s capabilities, we introduce a prompt-based evaluation method, where our model reduces offset errors by 16.6% and improves roof Intersection over Union (IoU) by 10.8% compared to other models. Leveraging the common patterns in predicting offsets, DNMS algorithms enable models to further reduce offset vector loss (VL) by 6.5%. To further validate the generalization of models, we tested them using a newly proposed test set, Huizhou test set, with over 7,000 manually annotated instance samples. Our algorithms and dataset will be available athttps://github.com/likaiucas/OBM. Kai Li 0025, Yupeng Deng 0001, Yunlong Kong, Diyou Liu, Jingbo Chen, Yu Meng 0002, Junxian Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Multilevel-Guided Curriculum Domain Adaptation Approach to Semantic Segmentation for High-Resolution Remote Sensing ImagesabstractThe semantic segmentation of high-resolution (HR) remote sensing images (RSIs) has been extensively researched in various applications. However, segmentation networks are prone to significant performance degradation on unlabeled data due to domain shift, such as data distribution shifts arising from distinct geographic locations. To address this issue, we propose a multilevel-guided curriculum domain adaptation (MuGCDA) approach for joint samplewise, categorywise, and pixelwise tasks, which facilitates the final fine-grained segmentation task by guiding the target domain to acquire samplewise and categorywise domain-robust properties. Concretely, at the sample level, we formulate a sample spatial relationship consistency guidance (SSCG) loss that guides the target domain to acquire similar sample spatial relationship properties to the source domain. At the category level, we propose a category layout structure consistency guidance (CLCG) module that guides the target domain to acquire consistent layout properties. At the pixel level, we design an adaptive hierarchical pseudolabel weight setting (AHPWS) method with a self-training (ST) paradigm to reduce the effect of label noise while improving the quality of the generated pseudolabels. Furthermore, to improve the stability of the training process, we use a momentum network (MN) as the teacher network to obtain the property knowledge and pseudolabels, and then guide the whole domain transfer process of the segmentation network, which acts as the student network. Extensive comparison and ablation experiments are conducted in several cross-space and cross-spectral scenes, and the results show that our method achieves significant performance improvements in cross-domain scenes for HR RSIs. Zhihao Xi, Yu Meng 0002, Anzhi Yue, Jingbo Chen, Yupeng Deng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |