EDBT 2026 Demo / reviewers in the wild / expert
Kaiyu Li 0001
dblp:131/2839-1
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-8015-6245ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynamicEarth: How Far Are We from Open-Vocabulary Change Detection?abstractMonitoring Earth's evolving land covers requires methods capable of detecting changes across a wide range of categories and contexts. Existing change detection methods are hindered by their dependency on predefined classes, reducing their effectiveness in open-world applications. To address this issue, we introduce open-vocabulary change detection (OVCD), a novel task that bridges vision and language to detect changes across any category. Considering the lack of high-quality data and annotation, we propose two training-free frameworks, M-C-I and I-M-C, which leverage and integrate off-the-shelf foundation models for the OVCD task. The insight behind the M-C-I~framework is to discover all potential changes and then classify these changes, while the insight of I-M-C~framework is to identify all targets of interest and then determine whether their states have changed. Based on these two frameworks, we instantiate to obtain several methods, e.g., SAM-DINOv2-SegEarth-OV, Grounding-DINO-SAM2-DINO, etc. Extensive evaluations on 4 benchmark datasets demonstrate the superior generalization and robustness of our OVCD methods over existing supervised and unsupervised methods. To support continued exploration, we release DynamicEarth, a dedicated codebase designed to advance research and application of OVCD. Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Chao Pang 0001, Zepeng Xin, Tieliang Gong, Deyu Meng, Zhi Wang 0002 |
AAAI | 1 |
| 2025 | SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesabstractCurrent remote sensing semantic segmentation methods are mostly built on the close-set assumption, meaning that the model can only recognize pre-defined categories that exist in the training set. However, in practical Earth observation, there are countless new categories, and manual annotation is impractical. To address this challenge, we first attempt to introduce training-free1open-vocabulary semantic segmentation (OVSS) into the remote sensing context. However, due to the sensitivity of remote sensing images to low-resolution features, distorted target shapes and ill-fitting boundaries are exhibited in the prediction mask. To tackle these issues, we propose a simple and universal upsampler, i.e. SimFeatUp, to restore lost spatial information of deep features. Specifically, SimFeatUp only needs to learn from a few unlabeled images, and can upsample arbitrary remote sensing image features. Furthermore, based on the observation of the abnormal response Kaiyu Li 0001, Ruixun Liu, Xiangyong Cao, Xueru Bai, Feng Zhou 0001, Deyu Meng, Zhi Wang 0002 |
CVPR | 1 |
| 2025 | Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel MethodabstractRecently, road graph extraction has garnered increasing attention due to its crucial role in autonomous driving, navigation, etc. However, accurately and efficiently extracting road graphs remains a persistent challenge, primarily due to the severe scarcity of labeled data. To address this limitation, we collect a global-scale satellite road graph extraction dataset, i.e. Global-Scale dataset. Specifically, the Global-Scale dataset is ∼ 20× larger than the largest existing public road extraction dataset and spans over 13,800 km2globally. Additionally, we develop a novel road graph extraction model, i.e. SAM-Road++, which adopts a node-guided resampling method to alleviate the mismatch issue between training and inference in SAM-Road [17], a pioneering state-of-the-art road graph extraction model. Furthermore, we propose a simple yet effective "extended-line" strategy in SAM-Road++ to mitigate the occlusion issue on the road. Extensive experiments demonstrate the validity of the collected Global-Scale dataset and the proposed SAM-Road++ method, particularly highlighting its superior predictive power in unseen regions. The dataset and code are available at https://github.com/earth-insights/samroadplus. Pan Yin, Kaiyu Li 0001, Xiangyong Cao, Jing Yao 0002, Lei Liu 0014, Xueru Bai, Feng Zhou 0001, Deyu Meng |
CVPR | 2 |
| 2025 | Open-CD: A Comprehensive Toolbox for Change DetectionabstractWe present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd. Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao |
ACM Multimedia | 1 |
| 2025 | SemiCD-VL: Visual-Language Model Guidance Makes Better Semi-Supervised Change DetectorabstractChange detection (CD) aims to identify pixels with semantic changes between images. However, annotating massive numbers of pixel-level images is labor-intensive and costly, especially for multitemporal images, which require pixel-wise comparisons by human experts. Considering the excellent performance of visual-language models (VLMs) for zero-shot, OV, etc., with prompt-based reasoning, it is promising to utilize VLMs to make better CD under limited labeled data. In this article, we propose a VLM guidance-based semi-supervised CD method, namely SemiCD-VL. The insight of SemiCD-VL is to synthesize free change labels using VLMs to provide additional supervision signals for unlabeled data. However, almost all current VLMs are designed for single-temporal images and cannot be directly applied to bi- or multitemporal images. Motivated by this, we first propose a VLM-based mixed change event generation (CEG) strategy to yield pseudo-labels for unlabeled CD data. Since the additional supervised signals provided by these VLM-driven pseudo-labels may conflict with the original pseudo-labels from the consistency regularization paradigm (e.g., FixMatch), we propose the dual projection head for de-entangling different signal sources. Further, we explicitly decouple the bitemporal images semantic representation through two auxiliary segmentation decoders, which are also guided by VLM. Finally, to make the model more adequately capture change representations, we introduce contrastive consistency regularization (CCR) by constructing feature-level contrastive loss in auxiliary branches. Extensive experiments show the advantage of SemiCD-VL. For instance, SemiCD-VL improves the FixMatch baseline by$+ 5.3~\text {IoU}^{c}$on WHU-CD and by$+ 2.4~\text {IoU}^{c}$on LEVIR-CD with 5% labels, and SemiCD-VL requires only 5%–10% of the labels to achieve performance similar to the supervised methods. In addition, our CEG strategy, in an unsupervised manner, can achieve performance far superior to state-of-the-art (SOTA) unsupervised CD methods (e.g., IoU improved from 18.8% to 46.3% on LEVIR-CD dataset). The code is available athttps://github.com/likyoo/SemiCD-VL. Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Junmin Liu, Deyu Meng, Zhi Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | IRSAMap: Toward Large-Scale, High-Resolution Land Cover Map VectorizationabstractWith the continuous enhancement of remote sensing image resolution and the rapid advancement of deep learning techniques, land cover mapping is undergoing a significant transformation from pixel-level segmentation to object-based vector modeling. This shift imposes higher demands on deep learning models, requiring not only precise delineation of object boundaries but also the preservation of topological consistency among geographic elements. However, existing public datasets face three major limitations: limited class annotations, restricted data scale, and the lack of spatial structural information, which severely hinder the development of breakthrough methods in high-resolution remote sensing vectorization. To address these challenges, we present IRSAMap, the first global remote sensing dataset designed for large-scale, high-resolution, multi-feature land cover vector mapping. This dataset offers four key advantages: First, a comprehensive element vector annotation system that includes over 1.8 million instances of 10 typical natural and man-made objects, such as buildings, roads, rivers, and trees, employing a unified vector annotation standard framework that ensures both semantic integrity and spatial structural accuracy. Second, an intelligent annotation workflow incorporating “manual pre-annotation + AI-based training and inference + manual review and correction,” which enhances annotation efficiency while ensuring consistency. Third, a global coverage that spans 67 regions across six continents, representing diverse terrain types, including urban and rural areas, with a total coverage area exceeding 1,000 square kilometers. Fourth, multi-task adaptability, supporting various tasks such as pixel-level land cover classification, building outline regularization extraction, road centerline extraction, and panoramic segmentation. As a fundamental resource for remote sensing intelligent interpretation, IRSAMap provides a standardized benchmark for the paradigm shift from pixels to objects, which will significantly advance the development of high-precision geographic feature automation, collaborative modeling, and other cutting-edge research directions. The dataset is of great value for applications such as global geographic information updating and digital twin construction. IRSAMap is publicly available at https://github.com/ucas-dlg/IRSAMap. Yu Meng 0002, Ligao Deng, Zhihao Xi, Jingbo Chen, Anzhi Yue, Diyou Liu, Kai Li 0025, Kaiyu Li 0001, Yupeng Deng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2024 | A New Learning Paradigm for Foundation Model-Based Remote-Sensing Change DetectionabstractChange detection (CD) is a critical task to observe and analyze dynamic processes of land cover. Although numerous deep learning-based CD models have performed excellently, their further performance improvements are constrained by the limited knowledge extracted from the given labelled data. On the other hand, the foundation models that emerged recently contain a huge amount of knowledge by scaling up across data modalities and proxy tasks. In this paper, we propose a Bi-Temporal Adapter Network (BAN), which is a universal foundation model-based CD adaptation framework aiming to extract the knowledge of foundation models for CD. The proposed BAN contains three parts, i.e. frozen foundation model (e.g., CLIP), bi-temporal adapter branch (Bi-TAB), and bridging modules between them. Specifically, BAN extracts general features through a frozen foundation model, which are then selected, aligned, and injected into Bi-TAB via the bridging modules. Bi-TAB is designed as a model-agnostic concept to extract task/domain-specific features, which can be either an existing arbitrary CD model or some hand-crafted stacked blocks. Beyond current customized models, BAN is the first extensive attempt to adapt the foundation model to the CD task. Experimental results show the effectiveness of our BAN in improving the performance of existing CD methods (e.g., up to 4.08% IoU improvement) with only a few additional learnable parameters. More importantly, these successful practices show us the potential of foundation models for remote sensing CD. The code is available at https://github.com/likyoo/BAN and will be supported in our Open-CD. Kaiyu Li 0001, Xiangyong Cao, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Changer: Feature Interaction is What You Need for Change DetectionabstractChange detection is an important tool for long-term earth observation missions. It takes bi-temporal images as input and predicts “where” the change has occurred. Different from other dense prediction tasks, a meaningful consideration for change detection is the interaction between bi-temporal features. With this motivation, in this paper we propose a novel general change detection architecture, MetaChanger, which includes a series of alternative interaction layers in the feature extractor. To verify the effectiveness of MetaChanger, we propose two derived models, ChangerAD and ChangerEx with simple interaction strategies: Aggregation-Distribution (AD) and feature “exchange”. AD is abstracted from some complex interaction methods, and feature “exchange” is a completely parameter&computation-free operation by exchanging bi-temporal features. In addition, for better alignment of bi-temporal features, we propose a Flow-based Dual-Alignment Fusion (FDAF) module which allows interactive alignment and feature fusion. Crucially, we observe Changer series models achieve competitive performance on different scale change detection datasets. Further, our proposed ChangerAD and ChangerEx could serve as a starting baseline for future MetaChanger design. Code and weights are made available at https://github.com/likyoo/open-cd. Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | S²ENet: Spatial-Spectral Cross-Modal Enhancement Network for Classification of Hyperspectral and LiDAR DataabstractThe effective utilization of multimodal data (e.g., hyperspectral and light detection and ranging (LiDAR) data) has profound implications for further development of the remote sensing (RS) field. Many studies have explored how to effectively fuse features from multiple modalities; however, few of them focus on information interactions that can effectively promote the complementary semantic content of multisource data before fusion. In this letter, we propose a spatial–spectral enhancement module (S2EM) for cross-modal information interaction in deep neural networks. Specifically, S2EM consists of SpAtial Enhancement Module (SAEM) for enhancing spatial representation of hyperspectral data by LiDAR features and SpEctral Enhancement Module (SEEM) for enhancing spectral representation of LiDAR data by hyperspectral features. A series of experiments and ablation studies on the Houston2013 dataset show that S2EM can effectively facilitate the interaction and understanding between multimodal data. Our source code is available athttps://github.com/likyoo/Multimodal-Remote-Sensing-Toolkit, contributing to the RS community. Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR ImagesabstractChange detection is an important task in remote sensing (RS) image analysis. It is widely used in natural disaster monitoring and assessment, land resource planning, and other fields. As a pixel-to-pixel prediction task, change detection is sensitive about the utilization of the original position information. Recent change detection methods always focus on the extraction of deep change semantic feature, but ignore the importance of shallow-layer information containing high-resolution and fine-grained features, this often leads to the uncertainty of the pixels at the edge of the changed target and the determination miss of small targets. In this letter, we propose a densely connected siamese network for change detection, namely SNUNet-CD (the combination of Siamese network and NestedUNet). SNUNet-CD alleviates the loss of localization information in the deep layers of neural network through compact information transmission between encoder and decoder, and between decoder and decoder. In addition, Ensemble Channel Attention Module (ECAM) is proposed for deep supervision. Through ECAM, the most representative features of different semantic levels can be refined and used for the final classification. Experimental results show that our method improves greatly on many evaluation criteria and has a better tradeoff between accuracy and calculation amount than other state-of-the-art (SOTA) change detection methods. Sheng Fang 0001, Kaiyu Li 0001, Jinyuan Shao, Zhe Li 0015 |
IEEE Geosci. Remote. Sens. Lett. | 2 |