VLDB 2026 Research / reviewers in the wild / expert
Yi Wang 0072
dblp:17/221-72
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-3096-6610ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards a Unified Copernicus Foundation Model for Earth VisionabstractAdvances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus solely on the Earth's surface, and overlook valuable metadata beyond imagery. In this work, we take a step towards next-generation EO foundation models with three key components: 1) Copernicus-Pretrain, a massive-scale pretraining dataset that integrates 18.7M aligned images from all major Copernicus Sentinel missions, spanning from the Earth's surface to its atmosphere; 2) Copernicus-FM, a unified foundation model capable of processing any spectral or non-spectral sensor modality using extended dynamic hypernetworks and flexible metadata encoding; and 3) Copernicus-Bench, a systematic evaluation benchmark with 15 hierarchical downstream tasks ranging from preprocessing to specialized applications for each Sentinel mission. Our dataset, model, and benchmark greatly improve the scalability, versatility, and multimodal adaptability of EO foundation models, while also creating new opportunities to connect EO, weather, and climate research. Codes, datasets and models are available at https://github.com/zhu-xlab/Copernicus-FM. Yi Wang 0072, Zhitong Xiong, Chenying Liu 0001, Adam J. Stewart, Thomas Dujardin, Nikolaos-Ioannis Bountos, Angelos Zavras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taixé, Xiao Xiang Zhu 0001 |
ICCV | 1 |
| 2025 | CromSS: Cross-Modal Pretraining With Noisy Labels for Remote Sensing Image SegmentationabstractWe explore the potential of large-scale noisily labeled data to enhance feature learning by pretraining semantic segmentation models within a multimodal framework for geospatial applications. We propose a novel cross-modal sample selection (CromSS) method, a weakly supervised pretraining strategy designed to improve feature representations through cross-modal consistency and noise mitigation techniques. Unlike conventional pretraining approaches, CromSS exploits massive amounts of noisy and easy-to-come-by labels for improved feature learning beneficial to semantic segmentation tasks. We investigate middle and late fusion strategies to optimize the multimodal pretraining architecture design. We also introduce a cross-modal sample selection module to mitigate the adverse effects of label noise, which employs a cross-modal entangling strategy to refine the estimated confidence masks within each modality to guide the sampling process. Additionally, we introduce a spatial–temporal label smoothing technique to counteract overconfidence for enhanced robustness against noisy labels. To validate our approach, we assembled the multimodal dataset, NoLDO-S12, which consists of a large-scale noisy label subset from Google’s Dynamic World (DW) dataset for pretraining and two downstream subsets with high-quality labels from Google DW and OpenStreetMap (OSM) for transfer learning. Experimental results on two downstream tasks and the publicly available DFC2020 dataset demonstrate that when effectively utilized, the low-cost noisy labels can significantly enhance feature learning for segmentation tasks. The data, codes, and pretrained weights are freely available athttps://github.com/zhu-xlab/CromSS. Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
Yi Wang 0072, Conrad M. Albrecht, Nassim Ait Ali Braham, Chenying Liu 0001, Zhitong Xiong, Xiao Xiang Zhu 0001 |
ECCV (29) | 1 |
| 2024 | Task Specific Pretraining with Noisy Labels for Remote Sensing Image SegmentationabstractCompared to supervised deep learning, self-supervision provides remote sensing a tool to reduce the amount of exact, human-crafted geospatial annotations. While image-level information for unsupervised pretraining efficiently works for various classification downstream tasks, the performance on pixel-level semantic segmentation lags behind in terms of model accuracy. On the contrary, many easily available label sources (e.g., automatic labeling tools and land cover land use products) exist, which can provide a large amount of noisy labels for segmentation model training. In this work, we propose to exploit noisy semantic segmentation maps for model pretraining. Our experiments provide insights on robustness per network layer. The transfer learning settings test the cases when the pretrained encoders are fine-tuned for different label classes and decoders. The results from two datasets indicate the effectiveness of task-specific supervised pretraining with noisy labels. Our findings pave new avenues to improved model accuracy and novel pretraining strategies for efficient remote sensing image segmentation. Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2024 | Post-Earthquake SAR-Optical Dataset for Quick Damaged-Building DetectionabstractThis work introduces a dataset for automated earthquake-damaged building detection from post-event satellite imagery. Using very high-resolution Synthetic Aperture Radar (SAR) and optical data from the 2023 Turkey-Syria earthquakes, the dataset includes over four thousand co-registered building footprints and patches. The task is framed as a binary image classification problem, serving as a reference for researchers to expedite algorithm development for rapid damaged building detection in future events. The dataset and codes together with detailed explanations will be made publicly available at https://github.com/ya0-sun/PostEQ-SARopt-BuildingDamage. Yao Sun 0005, Yi Wang 0072, Michael Eineder |
IGARSS | 2 |
| 2024 | Multi-Label Guided Supervised Contrastive Learning for Earth Observation PretrainingabstractPretraining foundation models on large-scale satellite imagery has raised great interest in Earth observation. While most pretraining is conducted purely self-supervised, many land cover land use products that provide free and global annotations tend to be overlooked. To bridge this gap, we propose to exploit land-cover-generated multi-label annotations to guide supervised contrastive learning for Earth observation. We match the SSL4EO-S12 dataset with Dynamic World land cover maps and integrate image-level multi-label annotations. During pretraining, the label similarities between different images are calculated, and those with high similarity scores are pulled together in the embedding space. Experimental results on classification and segmentation downstream tasks demonstrate the effectiveness of the proposed method. Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2024 | One for All: Toward Unified Foundation Models for Earth VisionabstractFoundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation models typically specialize in a single modality or a specific spatial resolution range, limiting their versatility for downstream datasets. While there have been attempts to develop multi-modal remote sensing foundation models, they typically employ separate vision encoders for each modality or spatial resolution, necessitating a switch in backbones contingent upon the input data. To address this issue, we introduce a simple yet effective method, termed OFA-Net (One-For-All Network): employing a single, shared Transformer backbone for multiple data modalities with different spatial resolutions. Using the masked image modeling mechanism, we pre-train a single Transformer backbone on a curated multi-modal dataset with this simple design. Then the backbone model can be used in different downstream tasks, thus forging a path towards a unified foundation backbone model in Earth vision. The proposed method is evaluated on 12 distinct downstream tasks and demonstrates promising performance. Zhitong Xiong, Yi Wang 0072, Fahong Zhang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2024 | QuickQuakeBuildings: Post-Earthquake SAR-Optical Dataset for Quick Damaged-Building DetectionabstractQuick and automated earthquake-damaged building detection from post-event satellite imagery is crucial, yet it is challenging due to the scarcity of training data required for developing robust algorithms. This letter presents the first dataset dedicated to detecting earthquake-damaged buildings from post-event very high resolution (VHR) Synthetic Aperture Radar (SAR) and optical imagery. Utilizing open satellite imagery and annotations acquired after the 2023 Turkey–Syria earthquakes, we deliver a dataset of co-registered building footprints and satellite image patches of both SAR and optical data, encompassing more than four thousand buildings. The task of damaged building detection is formulated as a binary image classification problem, that can also be treated as an anomaly detection problem due to extreme class imbalance. We provide baseline methods and results to serve as references for comparison. Researchers can utilize this dataset to expedite algorithm development, facilitating the rapid detection of damaged buildings in response to future events. The dataset and codes together with detailed explanations and visualization will be made publicly available at https://github.com/ya0-sun/PostEQ-SARopt-BuildingDamage. Yao Sun 0005, Yi Wang 0072, Michael Eineder |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | AIO2: Online Correction of Object Labels for Deep Learning With Incomplete Annotation in Remote Sensing Image SegmentationabstractWhile the volume of remote sensing data is increasing daily, deep learning in Earth Observation faces lack of accurate annotations for supervised optimization. Crowdsourcing projects such as OpenStreetMap distribute the annotation load to their community. However, such annotation inevitably generates noise due to insufficient control of the label quality, lack of annotators, frequent changes of the Earth’s surface as a result of natural disasters and urban development, among many other factors. We presentAdaptively trIggered Online Object-wise correction (AIO2)to address annotation noise induced by incomplete label sets. AIO2 features anAdaptive Correction Trigger (ACT)module that avoids label correction when the model training under- or overfits, and anOnline Object-wise Correction (O2C)methodology that employs spatial information for automated label modification. AIO2 utilizes a mean teacher model to enhance training robustness with noisy labels to both stabilize the training accuracy curve for fitting in ACT and provide pseudo labels for correction in O2C. Moreover, O2C is implementedonlinewithout the need to store updated labels every training epoch. We validate our approach on two building footprint segmentation datasets with different spatial resolutions. Experimental results with varying degrees of building label noise demonstrate the robustness of AIO2. Source code will be available at https://github.com/zhu-xlab/AIO2.git. Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Qingyu Li 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multilabel-Guided Soft Contrastive Learning for Efficient Earth Observation PretrainingabstractSelf-supervised pretraining on large-scale satellite data has raised great interest in building Earth observation (EO) foundation models. However, many important resources beyond pure satellite imagery, such as land-cover-land-use products that provide free global semantic information, as well as vision foundation models that hold strong knowledge of the natural world, are not widely studied. In this work, we show these free additional resources not only help resolve common contrastive learning bottlenecks but also significantly boost the efficiency and effectiveness of EO pretraining. Specifically, we first propose soft contrastive learning (SoftCon) that optimizes cross-scene soft similarity based on land-cover-generated multilabel supervision, naturally solving the issue of multiple positive samples and too strict positive matching in complex scenes. Second, we revisit and explore cross-domain continual pretraining for both multispectral and synthetic aperture radar (SAR) imagery, building efficient EO foundation models from strongest vision models such as DINOv2. Adapting simple weight-initialization and Siamese masking strategies into our SoftCon framework, we demonstrate impressive continual pretraining performance even when the input modalities are not aligned. Without prohibitive training, we produce multispectral and SAR foundation models that achieve significantly better results in 10 out of 11 downstream tasks than most existing SOTA models. For example, our ResNet50/ViT-S achieve 84.8/85.0 linear probing mAP scores on BigEarthNet-10%, which are better than most existing ViT-L models; under the same setting, our ViT-B sets a new record of 86.8 in multispectral, and 82.5 in SAR, the latter even better than many multispectral models. Dataset and models are available athttps://github.com/zhu-xlab/softcon. Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Semi-Supervised Learning for Hyperspectral Images by Non Parametrically Predicting View AssignmentCRediTabstractHyperspectral image (HSI) classification is gaining a lot of momentum in present time because of high inherent spectral information within the images. However, these images suffer from the problem of curse of dimensionality and usually require a large number samples for tasks such as classification, especially in supervised setting. Recently, to effectively train the deep learning models with minimal labelled samples, the unlabeled samples are also being leveraged in self-supervised and semi-supervised setting. In this work, we leverage the idea of semi-supervised learning to assist the discriminative self-supervised pretraining of the models. The proposed method takes different augmented views of the unlabeled samples as input and assigns them the same pseudo-label corresponding to the labelled sample from the downstream task. We train our model on two HSI datasets, anemly Houston dataset (from data fusion contest, 2013) and Pavia university dataset, and show that the proposed approach performs better than self-supervised approach and supervised training. Shivam Pande, Nassim Ait Ali Braham, Yi Wang 0072, Conrad M. Albrecht, Biplab Banerjee, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2022 | Peaks Fusion assisted Early-stopping Strategy for Overhead Imagery Segmentation with Noisy LabelsabstractAutomatic label generation systems, which are capable to generate huge amounts of labels with limited human efforts, enjoy lots of potential in the deep learning era. These easy-to-come-by labels inevitably bear label noises due to a lack of human supervision and can bias model training to some inferior solutions. However, models can still learn some plausible features, before they start to overfit on noisy patterns. Inspired by this phenomenon, we propose a new Peaks fusion assisted EArly-Stopping (PEAS) approach for imagery segmentation with noisy labels, which is mainly composed of two parts. First, a fitting based early-stopping criterion is used to detect the turning phase from which models are about to mimic noise details. After that, a peaks fusion strategy is applied to select reliable models in the detection zone to generate final fusion results. Here, validation accuracies are utilized as indicators in model selection. The proposed method was evaluated on New York City dataset whose labels were automatically collected by a rule-based label generation system, thus noisy to some extent due to a lack of human supervision. The experimental results showed that the proposed PEAS method can achieve both promising statistical and visual results when trained with noisy labels. Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001 |
IEEE Big Data | 3 |
| 2022 | Deep Semantic Model Fusion for Ancient Agricultural Terrace DetectionabstractDiscovering ancient agricultural terraces in desert regions is important for the monitoring of long-term climate changes on the Earth’s surface. However, traditional ground surveys are both costly and limited in scale. With the increasing accessibility of aerial and satellite data, machine learning techniques bear large potential for the automatic detection and recognition of archaeological landscapes. In this paper, we propose a deep semantic model fusion method for ancient agricultural terrace detection. The input data includes aerial images and LiDAR generated terrain features in the Negev desert. Two deep semantic segmentation models, namely DeepLabv3+ and UNet, with EfficientNet backbone, are trained and fused to provide segmentation maps of ancient terraces and walls. The proposed method won the first prize in the International AI Archaeology Challenge. Codes are available at https://github.com/wangyi111/international-archaeologyai-challenge. Yi Wang 0072, Chenying Liu 0001, Arti Tiwari, Micha Silver, Arnon Karnieli, Xiao Xiang Zhu 0001, Conrad M. Albrecht |
IEEE Big Data | 1 |
| 2022 | Monitoring Urban Forests from Auto-Generated Segmentation MAPSabstractWe present and evaluate a weakly-supervised methodology to quantify the spatiotemporal distribution of urban forests based on remotely sensed data with close-to-zero human interaction. Successfully training machine learning models for semantic segmentation typically depends on the availability of high-quality labels. We evaluate the benefit of high-resolution, three-dimensional point cloud data (LiDAR) as source of noisy labels in order to train models for the localization of trees in orthophotos. As proof of concept we sense Hurricane Sandy's impact on urban forests in Coney Island, New York City (NYC) and reference it to less impacted urban space in Brooklyn, NYC. Conrad M. Albrecht, Chenying Liu 0001, Yi Wang 0072, Levente J. Klein, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2022 | Self-Supervised Vision Transformers for Joint SAR-Optical Representation LearningabstractSelf-supervised learning (SSL) has attracted much interest in remote sensing and Earth observation due to its ability to learn task-agnostic representations without human annotation. While most of the existing SSL works in remote sensing utilize ConvNet backbones and focus on a single modality, we explore the potential of vision transformers (ViTs) for joint SAR-optical representation learning. Based on DINO, a state-of-the-art SSL algorithm that distills knowledge from two augmented views of an input image, we combine SAR and optical imagery by concatenating all channels to a unified input. Subsequently, we randomly mask out channels of one modality as a data augmentation strategy. While training, the model gets fed optical-only, SAR-only, and SAR-optical image pairs learning both inner- and intra-modality representations. Experimental results employing the BigEarthNet-MM dataset demonstrate the benefits of both, the ViT backbones and the proposed multimodal SSL algorithm DINO-MM. Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001 |
IGARSS | 1 |