VLDB 2026 Research / reviewers in the wild / expert
Linus Scheibenreif
dblp:246/1722
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-5580-8910ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Tune Smarter, Not Harder: Parameter-Efficient Fine-Tuning for Geospatial Foundation Models
Francesc Marti Escofet, Benedikt Blumenstiel, Linus Scheibenreif, Paolo Fraccaro, Konrad Schindler |
ECML/PKDD (6) | 3 |
| 2024 | Parameter Efficient Self-Supervised Geospatial Domain AdaptationabstractAs large-scale foundation models become publicly available for different domains, efficiently adapting them to individual downstream applications and additional data modalities has turned into a central challenge. For example, foun-dation models for geospatial and satellite remote sensing applications are commonly trained on large optical RGB or multi-spectral datasets, although data from a wide variety of heterogeneous sensors are available in the remote sensing domain. This leads to significant discrepancies between pre-training and downstream target data distributions for many important applications. Fine-tuning large foundation models to bridge that gap incurs high computational cost and can be infeasible when target datasets are small. In this paper, we address the question of how large, pre-trained foundational transformer models can be efficiently adapted to downstream remote sensing tasks involving different data modalities or limited dataset size. We present a self-supervised adaptation method that boosts downstream linear evaluation accuracy of different foundation models by 4-6% (absolute) across 8 remote sensing datasets while outperforming full fine-tuning when training only 1-2% of the model parameters. Our method significantly improves label efficiency and increases few-shot accuracy by 6-10% on different datasets11Code available at github. com/HSG-AIML/GDA . Linus Scheibenreif, Michael Mommert, Damian Borth |
CVPR | 1 |
| 2024 | Merging Patches and Tokens: A VQA System for Remote SensingabstractIn this paper, we investigate the integration of transformer-based feature extractors in a Remote Sensing Visual Question Answering (RSVQA) framework. Our findings demonstrate an improvement to the baseline achieved through additional attention modules after feature extraction and using MUTAN (Multimodal Tucker Fusion). Further, we delve into the potential of multi-task learning, observing a considerable boost in performance when feature extractors are trained. Our results suggest a promising future research avenue in multitask learning for RSVQA, while also emphasizing the need for careful selection of hyperparameters per question type as well as finding the proper balance for training the shared backbone and individual classifiers simultaneously to further improve performance. Damian Falk, Kaan Aydin, Linus Scheibenreif, Damian Borth |
IGARSS | 3 |
| 2024 | Less is More: Active Self-Supervised Learning in Remote SensingabstractActive learning (AL) has shown effectiveness in supervised learning studies in computer vision (CV), while its integration with self-supervised learning (SSL) remains underexplored. In our study, we establish the "SSL+AL" sampling framework in remote sensing, incorporating active learning strategies with self-supervised pretraining (SSP) to identify pre-training samples that improve downstream task performance. Our findings indicate that in the context of remote sensing image classification, different pre-training sampling methods can affect the downstream performance results: when freezing features, uncertainty sampling outperforms random sampling when the budget size is larger than 30% of the full dataset, whereas diversity sampling does not demonstrate a significant advantage over other sampling methods, particularly when the pre-training budget size is low. Linus Scheibenreif, Damian Borth |
IGARSS | 2 |
| 2023 | Self Supervised Learning in Remote Sensing: Quantifying Approaches Effectiveness Across Downstream TasksabstractIn the remote sensing field, vast amounts of data are available. However, labeling such data is expensive. Self-supervised learning makes it possible to leverage unlabeled data for the training of deep neural network models. This work focuses on the effectiveness of self-supervised pretext tasks for different supervised downstream tasks. Therefore, we compare generative, contrastive, and generative-contrastive pretext tasks across classification and semantic segmentation downstream tasks. Our results show that the contrastive setup is beneficial for remote sensing image classification, whereas the generative-contrastive setup shows the best results for the semantic segmentation downstream task. Therefore, our work indicates that the choice of self-supervised pretext task is an important consideration to optimize downstream task performance. Marc Grau, Alexander Lontke, Linus Scheibenreif |
IGARSS | 4 |
| 2023 | Eurosat Model Zoo: A Dataset and Benchmark on Populations of Neural Networks and Its Sparsified Model TwinsabstractThe availability of large-scale labeled datasets in remote sensing and Earth observation accelerated the use of deep neural networks in this domain. In the standard workflow, data is first downloaded from satellites and then processed or analyzed locally on Earth. However, the advent of commercial satellite mega-constellations in low-earth orbit, from corporations such as SpaceX or Planet, renders the downstream link of the data for analysis as the limiting factor. Downstream capacity thus becomes a crucial bottleneck for Earth observation. Given this situation, potential future paradigms would require to run deep neural networks on the satellite itself and only download the results of data processing or analysis to Earth. In this scenario, data processing and analysis capabilities are limited by the constraints in compute and power supply of the on-board computer. This work investigates the potential of neural network sparsification techniques to reduce model size and improve efficiency in the Earth observation domain. We study neural network sparsification through the use of populations of neural networks, i.e., thousands of models trained on satellite imagery, to derive insights into the performance of sparsified neural networks for Earth observation. Dominik Honegger, Konstantin Schürholt, Linus Scheibenreif, Damian Borth |
IGARSS | 3 |
| 2023 | Dataset Distillation for EurosatabstractIn supervised learning, which is commonly used in Remote Sensing applications, the performance of a model trained on a larger dataset is generally better than or equal to a model trained on a smaller dataset. Dataset distillation is a method that extracts the discriminative features from a larger dataset to a smaller one. By doing so, the important characteristics of the original dataset that are critical for learning can be isolated. This has implications regarding computational efficiency and understanding underlying representation learning dynamics. These implications are of particular interest in the context of remote sensing, where large amounts of data are being generated and processed every day. We use dataset distillation across multiple network architectures on the RGB bands of the EuroSAT dataset to test how it would behave in a Remote Sensing scenario with real-world data. Our distilled dataset leads to a consistent out-performance of 5%-10% compared to random sampling for downstream classification tasks. Julius Lautz, Daniel Leal, Linus Scheibenreif, Damian Borth, Michael Mommert |
IGARSS | 3 |
| 2023 | Ben-Ge: Extending Bigearthnet with Geographical and Environmental DataabstractDeep learning methods have proven to be a powerful tool in the analysis of large amounts of complex Earth observation data. However, while Earth observation data are multi-modal in most cases, only single or few modalities are typically considered. In this work, we present the ben-ge dataset, which supplements the BigEarthNet-MM dataset by compiling freely and globally available geographical and environmental data. Based on this dataset, we showcase the value of combining different data modalities for the downstream tasks of patch-based land-use/land-cover classification and land-use/land-cover segmentation. ben-ge is freely available and expected to serve as a test bed for fully supervised and self-supervised Earth observation applications. Michael Mommert, Nicolas Kesseli, Joëlle Hanna, Linus Scheibenreif, Damian Borth, Begüm Demir |
IGARSS | 4 |
| 2022 | A Multimodal Approach for Event Detection: Study of UK Lockdowns in the Year 2020abstractSatellites allow spatially precise monitoring of the Earth, but provide only limited information on events of societal impact. Subjective societal impact, however, may be quantified at a high frequency by monitoring social media data. In this work, we propose a multi-modal data fusion framework to accurately identify periods of COVID-19-related lockdown in the United Kingdom using satellite observations (NO2measurements from Sentinel-5P) and social media (textual content of tweets from Twitter) data. We show that the data fusion of the two modalities improves the event detection accuracy on a national level and for large cities such as London. Joëlle Hanna, Linus Scheibenreif, Michael Mommert, Damian Borth |
IGARSS | 2 |
| 2022 | Toward Global Estimation of Ground-Level NO2 Pollution With Deep Learning and Remote SensingabstractAir pollution is a central environmental problem in countries around the world. It contributes to climate change through the emission of greenhouse gases, and adversely impacts the health of billions of people. Despite its importance, detailed information about the spatial and temporal distribution of pollutants is complex to obtain. Ground-level monitoring stations are sparse, and approaches for modeling air pollution rely on extensive datasets which are unavailable for many locations. We introduce three techniques for the estimation of air pollution to overcome these limitations: 1) a baseline localized approach that mimics conventional land-use regression through gradient boosting; 2) an OpenStreetMap (OSM) approach with gradient boosting that is applicable beyond regions covered by detailed geographic datasets; and 3) a remote sensing-based deep learning method utilizing multiband imagery and trace-gas column density measurements from satellites. We focus on the estimation of nitrogen dioxide (NO2), a common anthropogenic air pollutant with adverse effects on the environment and human health. Our local baseline model achieves strong results with a mean absolute error (MAE) of 5.18 ±$0.16~\mu \text {g/m}^{3}$NO2. Substituting localized inputs with OSM leads to a degraded performance (MAE 7.22 ± 0.14) but enables NO2estimation at a global scale. The proposed deep learning model on remote sensing data combines high accuracy (MAE 5.5 ± 0.14) with global coverage and heteroscedastic uncertainty quantification. Our results enable the estimation of surface-level NO2pollution with high spatial resolution for any location on Earth. We illustrate this capability with an out-of-distribution test set on the US westcoast. Code1and data2are publicly available. Linus Scheibenreif, Michael Mommert, Damian Borth |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Power Plant Classification from Remote Imaging with Deep LearningabstractSatellite remote imaging enables the detailed study of land use patterns on a global scale. We investigate the possibility to improve the information content of traditional land use classification by identifying the nature of industrial sites from medium-resolution remote sensing images. In this work, we focus on classifying different types of power plants from Sentinel-2 imaging data. Using a ResNet-50 deep learning model, we are able to achieve a mean accuracy of 90.0% in distinguishing 10 different power plant types and a background class. Furthermore, we are able to identify the cooling mechanisms utilized in thermal power plants with a mean accuracy of 87.5%. Our results enable us to qualitatively investigate the energy mix from Sentinel-2 imaging data, and prove the feasibility to classify industrial sites on a global scale from freely available satellite imagery. Michael Mommert, Linus Scheibenreif, Joëlle Hanna, Damian Borth |
IGARSS | 2 |
| 2021 | A Novel Dataset and Benchmark for Surface No2 Prediction from Remote Sensing Data Including Covid Lockdown MeasuresabstractNO2is an atmospheric trace gas that contributes to global warming as a precursor of greenhouse gases and has adverse effects on human health. Surface NO2concentrations are commonly measured through strictly localized networks of air quality stations on the ground. This work presents a novel dataset of surface NO2measurements aligned with atmospheric column densities from Sentinel-5P, as well as geographic and meteorological variables and lockdown information11Available at https://github.com/HSG-AIML/NO2-dataset. The dataset provides access to data from a variety of sources through a common format and will foster data-driven research into the causes and effects of NO2pollution. We showcase the value of the new dataset on the task of surface NO2estimation with gradient boosting. The resulting models enable daily estimates and confident identification of EU NO2exposure limit breaches. Additionally, we investigate the influence of COVID-19 lockdowns on air quality in Europe and find a significant decrease in NO2levels. Linus Scheibenreif, Michael Mommert, Damian Borth |
IGARSS | 1 |
| 2019 | FunFam protein families improve residue level molecular function predictionabstractBACKGROUND: The CATH database provides a hierarchical classification of protein domain structures including a sub-classification of superfamilies into functional families (FunFams). We analyzed the similarity of binding site annotations in these FunFams and incorporated FunFams into the prediction of protein binding residues. RESULTS: FunFam members agreed, on average, in 36.9 ± 0.6% of their binding residue annotations. This constituted a 6.7-fold increase over randomly grouped proteins and a 1.2-fold increase (1.1-fold on the same dataset) over proteins with the same enzymatic function (identical Enzyme Commission, EC, number). Mapping de novo binding residue prediction methods (BindPredict-CCS, BindPredict-CC) onto FunFam resulted in consensus predictions for those residues that were aligned and predicted alike (binding/non-binding) within a FunFam. This simple consensus increased the F1-score (for binding) 1.5-fold over the original prediction method. Variation of the threshold for how many proteins in the consensus prediction had to agree provided a convenient control of accuracy/precision and coverage/recall, e.g. reaching a precision as high as 60.8 ± 0.4% for a stringent threshold. CONCLUSIONS: The FunFams outperformed even the carefully curated EC numbers in terms of agreement of binding site residues. Additionally, we assume that our proof-of-principle through the prediction of protein binding residues will be relevant for many other solutions profiting from FunFams to infer functional information at the residue level. Linus Scheibenreif, Maria Littmann, Christine A. Orengo, Burkhard Rost |
BMC Bioinform. | 1 |