Marc Rußwurm

dblp:215/3440 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-6612-5744ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
abstract
Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or distillation from massive global imagery datasets. To address this challenge, we introduce Satellite Contrastive Location-Image Pretraining (SatCLIP). This global, general-purpose geographic location encoder learns an implicit representation of locations by matching CNN and ViT inferred visual patterns of openly available satellite imagery with their geographic coordinates. The resulting SatCLIP location encoder efficiently summarizes the characteristics of any given location for convenient use in downstream tasks. In our experiments, we use SatCLIP embeddings to improve performance on nine diverse geospatial prediction tasks including temperature prediction, animal recognition, and population density estimation. Across tasks, SatCLIP consistently outperforms alternative location encoders and shows promise for improving geographic domain adaptation. These results demonstrate the potential of vision-location models to learn meaningful representations of our planet from the vast, varied, and largely untapped modalities of geospatial data.
Konstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey, Marc Rußwurm
AAAI5
2025 SAMSelect: A Spectral Index Search for Marine Debris Visualization Using Segment Anything
abstract
This work proposes SAMSelect, an algorithm to obtain a salient three-channel visualization for multispectral images. We develop SAMSelect and show its use for marine scientists visually interpreting floating marine debris in Sentinel-2 imagery. These debris are notoriously difficult to visualize due to their compositional heterogeneity in medium-resolution imagery. Out of these difficulties, a visual interpretation of imagery showing marine debris remains a common practice by domain experts, who select bands and spectral indices on a case-by-case basis informed by common practices and heuristics. SAMSelect selects the band or index combination that achieves the best classification accuracy on a small annotated dataset through the Segment Anything Model. Its central assumption is that the three-channel visualization achieves the most accurate segmentation results also provide good visual information for photo-interpretation. We evaluate SAMSelect in three Sentinel-2 scenes containing generic marine debris in Accra, Ghana, and Durban, South Africa, and deployed plastic targets from the Plastic Litter Project. This reveals the potential of new previously unused band combinations (e.g., a normalized difference index of B8, B2), which demonstrate improved performance compared to literature-based indices. We describe the algorithm in this paper and provide an open-source code repository that will be helpful for domain scientists doing visual photo interpretation, especially in the marine field.
Joost van Dalen, Yuki Markus Asano, Marc Rußwurm
IEEE Geosci. Remote. Sens. Lett.3
2024 Geographic Location Encoding with Spherical Harmonics and Sinusoidal Representation Networks
abstract
Learning representations of geographical space is vital for any machine learning model that integrates geolocated data, spanning application domains such as remote sensing, ecology, or epidemiology. Recent work embeds coordinates using sine and cosine projections based on Double Fourier Sphere (DFS) features. These embeddings assume a rectangular data domain even on global data, which can lead to artifacts, especially at the poles. At the same time, little attention has been paid to the exact design of the neural network architectures with which these functional embeddings are combined. This work proposes a novel location encoder for globally distributed geographic data that combines spherical harmonic basis functions, natively defined on spherical surfaces, with sinusoidal representation networks (SirenNets) that can be interpreted as learned Double Fourier Sphere embedding. We systematically evaluate positional embeddings and neural network architectures across various benchmarks and synthetic evaluation datasets. In contrast to previous approaches that require the combination of both positional encoding and neural networks to learn meaningful representations, we show that both spherical harmonics and sinusoidal representation networks are competitive on their own but set state-of-the-art performances across tasks when combined. The model code and experiments are available at https://github.com/marccoru/locationencoder.
Marc Rußwurm, Konstantin Klemmer, Esther Rolf, Robin Zbinden, Devis Tuia
ICLR1
2024 WildCLIP: Scene and Animal Attribute Retrieval from Camera Trap Data with Domain-Adapted Vision-Language Models
abstract
Abstract Wildlife observation with camera traps has great potential for ethology and ecology, as it gathers data non-invasively in an automated way. However, camera traps produce large amounts of uncurated data, which is time-consuming to annotate. Existing methods to label these data automatically commonly use a fixed pre-defined set of distinctive classes and require many labeled examples per class to be trained. Moreover, the attributes of interest are sometimes rare and difficult to find in large data collections. Large pretrained vision-language models, such as contrastive language image pretraining (CLIP), offer great promises to facilitate the annotation process of camera-trap data. Images can be described with greater detail, the set of classes is not fixed and can be extensible on demand and pretrained models can help to retrieve rare samples. In this work, we explore the potential of CLIP to retrieve images according to environmental and ecological attributes. We create WildCLIP by fine-tuning CLIP on wildlife camera-trap images and to further increase its flexibility, we add an adapter module to better expand to novel attributes in a few-shot manner. We quantify WildCLIP’s performance and show that it can retrieve novel attributes in the Snapshot Serengeti dataset. Our findings outline new opportunities to facilitate annotation processes with complex and multi-attribute captions. The code is available at https://github.com/amathislab/wildclip .
Valentin Gabeff, Marc Rußwurm, Devis Tuia, Alexander Mathis
Int. J. Comput. Vis.2
2023 Semi-Supervised Deep Learning Representations in Earth Observation Based Forest Management
abstract
In this study, we examine the potential of several self-supervised deep learning models in predicting forest attributes and detecting forest changes using ESA Sentinel-1 and Sentinel-2 images. The performance of the proposed deep learning models is compared to established conventional machine learning approaches. Studied use-cases include mapping of forest disturbance (windthrown forests, snowload damages) using deep change vector analysis, forest height mapping using UNet+ based models, Momentum contrast and regression modeling. Study areas were represented by several boreal forest sites in Finland. Our results indicate that developed methods allow to achieve superior classification and prediction accuracies compared to traditional methodologies and mimimize the amount of necessary in-situ forestry data.
Oleg Antropov, Matthieu Molinier, Ridvan Salih Kuzu, Lloyd H. Hughes, Marc Rußwurm, Devis Tuia, Corneliu Octavian Dumitru, Shaojia Ge, Sudipan Saha, Xiao Xiang Zhu 0001
IGARSS5
2023 Improving Few-Shot Object Detection with Object Part Proposals
abstract
Few-Shot Object Detection (FSOD) allows fast adaptation of an object detection model to new classes of objects using few examples per class. This has many applications, in particular in satellite and aerial observation, as it allows learning from experts who can only annotate a few examples for new classes and helps migrate models across tasks. In this work, we present a technique to improve the performance of FSOD in remote sensing by defining a contrastive loss that utilizes parts of objects. For this, we generate, what we call, Object Parts Proposals (OPPs) on the fly for each novel class, and use them to learn more robust features with an additional contrastive objective. We observe that training with OPPs brings a consistent improvement over the state-of-the-art when evaluating on the DIOR dataset.The code is available at https://github.com/arthurchevalley/Improving-FSOD-on-RSI-using-Sub-Parts.
Arthur Chevalley, Ciprian Tomoiaga, Marcin Detyniecki, Marc Rußwurm, Devis Tuia
IGARSS4
2023 Classification of Tropical Deforestation Drivers with Machine Learning and Satellite Image Time Series
abstract
Tropical deforestation is a major environmental problem with severe consequences such as carbon emissions or biodiversity loss. While much research focuses on monitoring and mapping deforestation, less attention is paid to understanding the various reasons and motivations behind it, known as deforestation drivers. Drivers can typically be identified from optical satellite imagery, but it is often necessary to view the deforested site at multiple points in time to determine the driver, making manual annotation of drivers laborious. In this work, we propose a deep learning model that classifies drivers from time series of Sentinel-2 images. The model combines convolutional, LSTM, and attention layers. To train the model, we use a large crowd-sourced dataset spanning across the tropics. We compare its results to other architectures and show that using time series can bring significant improvement in accuracy compared to single images, especially if a suitable architecture is used. Additionally, we analyze the attention scores produced by our model and show that it learns different strategies for different classes.
Jan Pisl, Lloyd H. Hughes, Marc Rußwurm, Devis Tuia
IGARSS3
2023 Detection of Settlements in Tanzania and Mozambique by Many Regional Few-Shot Models
abstract
In this work, we propose an approach to aid in mapping small settlements, which are often misclassified by models trained on a large-scale context (global or regional). We leverage pre-trained land cover models and few-shot learning to enhance the detection of these settlements. The backbone models are trained globally, but their application is localized through a spatial sampling strategy to address the challenge of detecting missed or unlabelled settlements. The proposed sampling strategy is based on the distance around a test patch and allows for the sampling of both backgrounds (non-settlements) points and settlements. Following this strategy results in a balanced dataset for model fine-tuning and ensures that the model is well-adapted to the local context. The idea is that nearby settlements share more similar properties, which is leveraged in our approach. We evaluate these transferred models by measuring the number of previously unmapped settlements detected by the fine-tuned classifier. For this, we manually annotated over two thousand buildings across two regions of Tanzania, previously unmapped in the original urban landcover product. Our results indicate the potential of the sampling approach, particularly when combined with a model pretrained with Momentum Contrast (MoCo). However, we also highlight the limitations in terms of spatial resolution of Sentinel-2 data for the detection of small settlements.
Marc Rußwurm, Lloyd H. Hughes, Giorgio Pasquali, Corneliu Octavian Dumitru, Devis Tuia
IGARSS1
2022 Humans are Poor Few-Shot Classifiers for Sentinel-2 Land Cover
abstract
Learning to predict accurately from a few data samples is a central challenge in modern data-hungry machine learning. On natural images, human vision typically outperforms deep learning approaches on few-shot learning. However, we hypothesize that aerial and satellite images are more challenging to the human eye. This applies particularly when the image resolution is comparatively low, as with the 10m ground sampling distance of Sentinel-2. In this study, we benchmark model-agnostic meta-learning (MAML) algorithms against human participants on few-shot land cover classification with Sentinel-2 imagery on the Sen12MS dataset. We find that categorization of land cover from globally distributed regions is a difficult task for the participants, who classified the given images less accurately than the MAML-trained model and with a highly variable success rate. This suggest that hand-labeling land cover directly on Sentinel-2 imagery is not optimal when tackling a new land cover classification problem. Labeling only a few images and employing a trained meta-learning model to this task may lead to more accurate and consistent solutions compared to hand labeling by multiple individuals.
Marc Rußwurm, Sherrie Wang, Devis Tuia
IGARSS1
2021 InSAR Displacement Time Series Mining: A Machine Learning Approach
abstract
Interferometric Synthetic Aperture Radar (InSAR)-derived surface displacement time series enable a wide range of applications from urban structural monitoring to geohazard assessment. With systematic data acquisitions becoming the new norm for SAR missions, millions of time series are continuously generated. Machine Learning provides a framework for the efficient mining of such big data. Here, we focus on unsupervised mining of the data via clustering the similar temporal patterns and data-driven displacement signal reconstruction from the InSAR time series. We propose a deep Long Short Term Memory (LSTM) autoencoder model which can exploit temporal relations in contrast to the commonly used shallow learning methods, such as Uniform Manifold Approximation and Projection (UMAP). We also modify the loss function to allow the quantification of uncertainties in the time series data. The two approaches are applied to the Lazufre Volcanic Complex located at the central volcanic zone of the Andes and thereby compared.
Homa Ansari, Marc Rußwurm, Sina Montazeri, Alessandro Parizzi, Xiao Xiang Zhu 0001
IGARSS2
2020 Model and Data Uncertainty for Satellite Time Series Forecasting with Deep Recurrent Models
abstract
Deep Learning is often criticized as being a black-box method that provides accurate predictions, but a limited explanation of the underlying processes and no indication when to not trust those predictions. Equipping existing deep learning models with an (general) notion of uncertainty can help mitigate both these issues. The Bayesian deep learning community has developed model-agnostic methodology to estimate both data and model uncertainty that can be implemented on top of existing deep learning models. In this work, we test this methodology for deep recurrent satellite time series forecasting and test its assumptions on data and model uncertainty. We tested its effectiveness on an application on climate change where the activity of seasonal vegetation decreased over multiple years.
Marc Rußwurm, Xiao Xiang Zhu 0001, Yarin Gal, Marco Körner 0001
IGARSS1
2020 Tslearn, A Machine Learning Toolkit for Time Series Data
abstract
tslearn is a general-purpose Python machine learning library for time series that offers tools for pre-processing and feature extraction as well as dedicated models for clustering, classification and regression. It follows scikit-learn's Application Programming Interface for transformers and estimators, allowing the use of standard pipelines and model selection tools on top of tslearn objects. It is distributed under the BSD-2-Clause license, and its source code is available at https://github.com/tslearn-team/tslearn.
Romain Tavenard, Johann Faouzi, Gilles Vandewiele, Felix Divo, Guillaume Androz, Chester Holtz, Marie Payne, Roman Yurchak, Marc Rußwurm, Kushal Kolar, Eli Woods
J. Mach. Learn. Res.9
2019 Multi3Net: Segmenting Flooded Buildings via Fusion of Multiresolution, Multisensor, and Multitemporal Satellite Imagery
abstract
We propose a novel approach for rapid segmentation of flooded buildings by fusing multiresolution, multisensor, and multitemporal satellite imagery in a convolutional neural network. Our model significantly expedites the generation of satellite imagery-based flood maps, crucial for first responders and local authorities in the early stages of flood events. By incorporating multitemporal satellite imagery, our model allows for rapid and accurate post-disaster damage assessment and can be used by governments to better coordinate medium- and long-term financial assistance programs for affected areas. The network consists of multiple streams of encoder-decoder architectures that extract spatiotemporal information from medium-resolution images and spatial information from high-resolution images before fusing the resulting representations into a single medium-resolution segmentation map of flooded buildings. We compare our model to state-of-the-art methods for building footprint segmentation as well as to alternative fusion approaches for the segmentation of flooded buildings and find that our model performs best on both tasks. We also demonstrate that our model produces highly accurate segmentation maps of flooded buildings using only publicly available medium-resolution data instead of significantly more detailed but sparsely available very high-resolution data. We release the first open-source dataset of fully preprocessed and labeled multiresolution, multispectral, and multitemporal satellite images of disaster sites along with our source code.
Tim G. J. Rudner, Marc Rußwurm, Jakub Fil, Ramona Pelich, Benjamin Bischke, Veronika Kopacková, Piotr Bilinski
AAAI2
2019 Sequential Recurrent Encoders for Land Cover Mapping in The Brazilian Amazon Using Modis Imagery and Auxiliary Datasets
abstract
We explore a state-of-art architecture based on convolutional recurrent networks, designed to ingest and extract information of a sequence of satellite images, for large area LULC classification in the Brazilian Amazon biome. We fine-tune and evaluate it according to multiple combinations between MODIS archives at 250 m for a single year, providing surface reflectance time series data and derived indices, and external environmental data. Qualitative and quantitative differences between the trained models were evident according the input features assessed. The combination of surface reflectance and auxiliary data yield slightly better performance than only using surface reflectance data. In specific, the auxiliary data contributed to the accuracy of Mixed Forest, Savanna, Grassland and Cropland. In contrast, the arrangement including derived indices showed lower performance and therefore they negatively contribute to the classification task in the area of study.
Alejandro Coca-Castro, Marc Rußwurm, Louis Reymondin, Mark Mulligan
IGARSS2