VLDB 2026 Research / reviewers in the wild / expert
Damian Borth
dblp:48/1492
· DBLP profile ↗
48ranked-venue papers
6as first author
25since 2021 · last 2025
0000-0002-4660-2627ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) is a challenging problem that has been extensively studied in recent years. Traditional approaches often rely on external modules like Class Activation Maps to highlight regions of interest and generate pseudo segmentation masks. In this work, we propose an end-to-end method that directly utilizes the attention maps learned by a Vision Transformer (ViT) for WSSS. We propose training a sparse ViT with multiple [CLS] tokens (one for each class), using a random masking strategy to promote [CLS] token - class assignment. At inference time, we aggregate the different self-attention maps of each [CLS] token corresponding to the predicted labels to generate pseudo segmentation masks. Our proposed approach enhances the interpretability of self-attention maps and ensures accurate class assignments. Extensive experiments on two standard benchmarks and three specialized datasets demonstrate that our method generates accurate pseudo-masks, outperforming related works. Those pseudo-masks can be used to train a segmentation model which achieves results comparable to fully-supervised models, significantly reducing the need for fine-grained labeled data. Joëlle Hanna, Damian Borth |
ICCV | 2 |
| 2025 | Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular DataabstractAnomaly detection in tabular data remains challenging due to complex feature interactions and the scarcity of anomalous examples. Denoising autoencoders rely on fixed-magnitude noise, limiting adaptability to diverse data distributions. Diffusion models introduce scheduled noise and iterative denoising, but lack explicit reconstruction mappings. We propose the Diffusion-Scheduled Denoising Autoencoder (DDAE), a framework that integrates diffusion-based noise scheduling and contrastive learning into the encoding process to improve anomaly detection. We evaluated DDAE on 57 datasets from ADBench. Our method outperforms in semi-supervised settings and achieves competitive results in unsupervised settings, improving PR-AUC by up to 65% (9%) and ROC-AUC by 16% (6%) over state-of-the-art autoencoder (diffusion) model baselines. We observed that higher noise levels benefit unsupervised training, while lower noise with linear scheduling is optimal in semi-supervised settings. These findings underscore the importance of principled noise strategies in tabular anomaly detection. Timur Sattarov, Marco Schreyer, Damian Borth |
KDD (2) | 3 |
| 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsabstractModern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of
prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to
\textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs.
We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research. Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036 |
NeurIPS | 10 |
| 2024 | Parameter Efficient Self-Supervised Geospatial Domain AdaptationabstractAs large-scale foundation models become publicly available for different domains, efficiently adapting them to individual downstream applications and additional data modalities has turned into a central challenge. For example, foun-dation models for geospatial and satellite remote sensing applications are commonly trained on large optical RGB or multi-spectral datasets, although data from a wide variety of heterogeneous sensors are available in the remote sensing domain. This leads to significant discrepancies between pre-training and downstream target data distributions for many important applications. Fine-tuning large foundation models to bridge that gap incurs high computational cost and can be infeasible when target datasets are small. In this paper, we address the question of how large, pre-trained foundational transformer models can be efficiently adapted to downstream remote sensing tasks involving different data modalities or limited dataset size. We present a self-supervised adaptation method that boosts downstream linear evaluation accuracy of different foundation models by 4-6% (absolute) across 8 remote sensing datasets while outperforming full fine-tuning when training only 1-2% of the model parameters. Our method significantly improves label efficiency and increases few-shot accuracy by 6-10% on different datasets11Code available at github. com/HSG-AIML/GDA . Linus Scheibenreif, Michael Mommert, Damian Borth |
CVPR | 3 |
| 2024 | Towards Scalable and Versatile Weight Space LearningabstractLearning representations of well-trained neural network models holds the promise to provide an understanding of the inner workings of those models. However, previous work has either faced limitations when processing larger networks or was task-specific to either discriminative or generative tasks. This paper introduces the SANE approach to weight-space learning. SANE overcomes previous limitations by learning task-agnostic representations of neural networks that are scalable to larger models of varying architectures and that show capabilities beyond a single task. Our method extends the idea of *hyper-representations* towards sequential processing of subsets of neural network weights, thus allowing one to embed larger neural networks as a set of tokens into the learned representation space. SANE reveals global model information from layer-wise embeddings, and it can sequentially generate unseen neural network models, which was unattainable with previous *hyper-representation* learning methods. Extensive empirical evaluation demonstrates that SANE matches or exceeds state-of-the-art performance on several weight representation learning benchmarks, particularly in initialization for new tasks and larger ResNet architectures. Konstantin Schürholt, Michael W. Mahoney, Damian Borth |
ICML | 3 |
| 2024 | Merging Patches and Tokens: A VQA System for Remote SensingabstractIn this paper, we investigate the integration of transformer-based feature extractors in a Remote Sensing Visual Question Answering (RSVQA) framework. Our findings demonstrate an improvement to the baseline achieved through additional attention modules after feature extraction and using MUTAN (Multimodal Tucker Fusion). Further, we delve into the potential of multi-task learning, observing a considerable boost in performance when feature extractors are trained. Our results suggest a promising future research avenue in multitask learning for RSVQA, while also emphasizing the need for careful selection of hyperparameters per question type as well as finding the proper balance for training the shared backbone and individual classifiers simultaneously to further improve performance. Damian Falk, Kaan Aydin, Linus Scheibenreif, Damian Borth |
IGARSS | 4 |
| 2024 | Less is More: Active Self-Supervised Learning in Remote SensingabstractActive learning (AL) has shown effectiveness in supervised learning studies in computer vision (CV), while its integration with self-supervised learning (SSL) remains underexplored. In our study, we establish the "SSL+AL" sampling framework in remote sensing, incorporating active learning strategies with self-supervised pretraining (SSP) to identify pre-training samples that improve downstream task performance. Our findings indicate that in the context of remote sensing image classification, different pre-training sampling methods can affect the downstream performance results: when freezing features, uncertainty sampling outperforms random sampling when the budget size is larger than 30% of the full dataset, whereas diversity sampling does not demonstrate a significant advantage over other sampling methods, particularly when the pre-training budget size is low. Linus Scheibenreif, Damian Borth |
IGARSS | 3 |
| 2024 | Multi-Modal Diffusion for Self-Supervised PretrainingabstractSelf-supervised pretraining has been shown to greatly improve the performance and label-efficiency of Deep Learning models for multi-modal remote sensing applications. In this work, we explore the use of diffusion methods in a multi-modal setup as a means to pretrain architectures in a task-agnostic way. We qualitatively find that multi-modal diffusion processes are able to coherently learn information across multiple data-modalities. Fine-tuning the pretrained backbone on a segmentation task outperforms a supervised baseline model and is highly label-efficient, which is not the case when fine-tuned on a classification task. Alexander Lontke, Michael Mommert, Damian Borth |
IGARSS | 3 |
| 2023 | Fine-Grained Emotional Control of Text-to-Speech: Learning to Rank Inter- and Intra-Class Emotion IntensitiesabstractState-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words or phonemes. Although still challenging, the first TTS models have been recently proposed that are able to control voice by manually assigning emotion intensity. Unfortunately, due to the neglect of intra-class distance, the intensity differences are often unrecognizable. In this paper, we propose a fine-grained controllable emotional TTS, that considers both inter- and intra-class distances and be able to synthesize speech with recognizable intensity difference. Our subjective and objective experiments demonstrate that our model exceeds two state-of-the-art controllable TTS models for controllability, emotion expressiveness and naturalness. Jón Guðnason, Damian Borth |
ICASSP | 3 |
| 2023 | Eurosat Model Zoo: A Dataset and Benchmark on Populations of Neural Networks and Its Sparsified Model TwinsabstractThe availability of large-scale labeled datasets in remote sensing and Earth observation accelerated the use of deep neural networks in this domain. In the standard workflow, data is first downloaded from satellites and then processed or analyzed locally on Earth. However, the advent of commercial satellite mega-constellations in low-earth orbit, from corporations such as SpaceX or Planet, renders the downstream link of the data for analysis as the limiting factor. Downstream capacity thus becomes a crucial bottleneck for Earth observation. Given this situation, potential future paradigms would require to run deep neural networks on the satellite itself and only download the results of data processing or analysis to Earth. In this scenario, data processing and analysis capabilities are limited by the constraints in compute and power supply of the on-board computer. This work investigates the potential of neural network sparsification techniques to reduce model size and improve efficiency in the Earth observation domain. We study neural network sparsification through the use of populations of neural networks, i.e., thousands of models trained on satellite imagery, to derive insights into the performance of sparsified neural networks for Earth observation. Dominik Honegger, Konstantin Schürholt, Linus Scheibenreif, Damian Borth |
IGARSS | 4 |
| 2023 | Dataset Distillation for EurosatabstractIn supervised learning, which is commonly used in Remote Sensing applications, the performance of a model trained on a larger dataset is generally better than or equal to a model trained on a smaller dataset. Dataset distillation is a method that extracts the discriminative features from a larger dataset to a smaller one. By doing so, the important characteristics of the original dataset that are critical for learning can be isolated. This has implications regarding computational efficiency and understanding underlying representation learning dynamics. These implications are of particular interest in the context of remote sensing, where large amounts of data are being generated and processed every day. We use dataset distillation across multiple network architectures on the RGB bands of the EuroSAT dataset to test how it would behave in a Remote Sensing scenario with real-world data. Our distilled dataset leads to a consistent out-performance of 5%-10% compared to random sampling for downstream classification tasks. Julius Lautz, Daniel Leal, Linus Scheibenreif, Damian Borth, Michael Mommert |
IGARSS | 4 |
| 2023 | Ben-Ge: Extending Bigearthnet with Geographical and Environmental DataabstractDeep learning methods have proven to be a powerful tool in the analysis of large amounts of complex Earth observation data. However, while Earth observation data are multi-modal in most cases, only single or few modalities are typically considered. In this work, we present the ben-ge dataset, which supplements the BigEarthNet-MM dataset by compiling freely and globally available geographical and environmental data. Based on this dataset, we showcase the value of combining different data modalities for the downstream tasks of patch-based land-use/land-cover classification and land-use/land-cover segmentation. ben-ge is freely available and expected to serve as a test bed for fully supervised and self-supervised Earth observation applications. Michael Mommert, Nicolas Kesseli, Joëlle Hanna, Linus Scheibenreif, Damian Borth, Begüm Demir |
IGARSS | 5 |
| 2023 | Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-SpeechabstractEffective speech emotional representations play a key role in Speech Emotion Recognition (SER) and Emotional Text-To-Speech (TTS) tasks. However, emotional speech samples are more difficult and expensive to acquire compared with Neutral style speech, which causes one issue that most related works unfortunately neglect: imbalanced datasets. Models might overfit to the majority Neutral class and fail to produce robust and effective emotional representations. In this paper, we propose an Emotion Extractor to address this issue. We use augmentation approaches to train the model and enable it to extract effective and generalizable emotional representations from imbalanced datasets. Our empirical results show that (1) for the SER task, the proposed Emotion Extractor surpasses the state-of-the-art baseline on three imbalanced datasets; (2) the produced representations from our Emotion Extractor benefit the TTS model, and enable it to synthesize more expressive speech. Jón Guðnason, Damian Borth |
INTERSPEECH | 3 |
| 2023 | Physics-Guided Multitask Learning for Estimating Power Generation and CO2 Emissions From Satellite ImageryabstractFossil fuel combustion produces large quantities of carbon dioxide (CO2), a major greenhouse gas (GHG), which is one of the main drivers of climate change. A quantitative assessment of GHG emissions is fundamental to predicting climate change effects, enforcing emission regulations, and monitoring pollution trading schemes. Unfortunately, the reporting of GHG emissions is only required in some countries, resulting in insufficient global coverage. At the same time, the transition from fossil fuels to zero carbon to limit climate change is at the heart of several ecological movements, hence the need for quantifying energy production, as well. In this work, we propose an end-to-end method to estimate power generation rates for fossil fuel power plants from satellite images, based on which we approximate GHG (CO2) emission rates. We present a physics-guided multitask deep-learning approach able to simultaneously predict from a single-satellite image of a power plant: 1) the pixel-area covered by plumes; 2) the type of fired fuel; and 3) the power generation rate. To ensure physically realistic predictions from our model we account for environmental conditions and empirical physical constraints. We then convert the predicted power generation rate into estimates for the rate at which CO2is being emitted, using a fuel-dependent conversion factor. Experimental results show that our multitask learning approach improves the power generation estimation mean absolute error (MAE) by 23% compared to a single-task network trained on the same dataset. Joëlle Hanna, Damian Borth, Michael Mommert |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Traffic Noise Estimation from Satellite Imagery with Deep LearningabstractRoad traffic noise represents a global health issue. Despite its importance, noise data are unavailable in many regions of the world. We therefore propose to approximate noise data from satellite imagery in an end-to-end Deep Learning approach. We train a U-Net segmentation model to estimate road noise based on freely available Sentinel-2 satellite imagery and existing road traffic noise estimates for Switzer-land. We are able to achieve an RMSE of 8.8 dB(A) for day-time traffic noise and 7.6 dB(A) for nighttime traffic noise with a spatial resolution of 10 m. In addition to identifying major road networks, our model succeeds to predict the spatial propagation of noise. Our results suggest that this approach provides a pathway to estimating road traffic noise for areas for which no such measures are available. Leonardo Eicher, Michael Mommert, Damian Borth |
IGARSS | 3 |
| 2022 | A Multimodal Approach for Event Detection: Study of UK Lockdowns in the Year 2020abstractSatellites allow spatially precise monitoring of the Earth, but provide only limited information on events of societal impact. Subjective societal impact, however, may be quantified at a high frequency by monitoring social media data. In this work, we propose a multi-modal data fusion framework to accurately identify periods of COVID-19-related lockdown in the United Kingdom using satellite observations (NO2measurements from Sentinel-5P) and social media (textual content of tweets from Twitter) data. We show that the data fusion of the two modalities improves the event detection accuracy on a national level and for large cities such as London. Joëlle Hanna, Linus Scheibenreif, Michael Mommert, Damian Borth |
IGARSS | 4 |
| 2022 | Zero-shot Voice Conversion via Self-supervised Prosody Representation LearningabstractVoice Conversion (VC) for unseen speakers, also known as zero-shot VC, is an attractive research topic as it enables a range of applications like voice customizing, animation production, and others. Recent work in this area made progress with disentanglement methods that separate utterance content and speaker characteristics from speech audio recordings. However, many of these methods are subject to the leakage of prosody (e.g., pitch, volume), causing the speaker voice in the synthesized speech to be different from the desired target speakers. To prevent this issue, we propose a novel self-supervised approach that effectively learns disentangled pitch and volume representations that can represent the prosody styles of different speakers. We then use the learned prosodic representations as conditional information to train and enhance our VC model for zero-shot conversion. In our experiments, we show that our prosody representations are disentangled and rich in prosody information. Moreover, we demonstrate that the addition of our prosody representations improves our VC performance and surpasses state-of-the-art zero-shot VC performances. Damian Borth |
IJCNN | 2 |
| 2022 | Generative Data Augmentation Guided by Triplet Loss for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) is crucial for humancomputer interaction but still remains a challenging problem because of two major obstacles: data scarcity and imbalance.Many datasets for SER are substantially imbalanced, where data utterances of one class (most often Neutral) are much more frequent than those of other classes.Furthermore, only a few data resources are available for many existing spoken languages.To address these problems, we exploit a GAN-based augmentation model guided by a triplet network, to improve SER performance given imbalanced and insufficient training data.We conduct experiments and demonstrate: 1) With a highly imbalanced dataset, our augmentation strategy significantly improves the SER performance (+8% recall score compared with the baseline).2) Moreover, in a cross-lingual benchmark, where we train a model with enough source language utterances but very few target language utterances (around 50 in our experiments), our augmentation strategy brings benefits for the SER performance of all three target languages. Hamed Hemati, Jón Guðnason, Damian Borth |
INTERSPEECH | 4 |
| 2022 | Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsabstractLearning representations of neural network weights given a model zoo is an emerg- ing and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we ex- tend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, per- formant and capable to outperform strong baselines as evaluated on several down- stream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions. Konstantin Schürholt, Xavier Giró-i-Nieto, Damian Borth |
NeurIPS | 4 |
| 2022 | Model Zoos: A Dataset of Diverse Populations of Neural Network ModelsabstractIn the last years, neural networks (NN) have evolved from laboratory environments to the state-of-the-art for many real-world problems. It was shown that NN models (i.e., their weights and biases) evolve on unique trajectories in weight space during training. Following, a population of such neural network models (referred to as model zoo) would form structures in weight space. We think that the geometry, curvature and smoothness of these structures contain information about the state of training and can reveal latent properties of individual models. With such model zoos, one could investigate novel approaches for (i) model analysis, (ii) discover unknown learning dynamics, (iii) learn rich representations of such populations, or (iv) exploit the model zoos for generative modelling of NN weights and biases. Unfortunately, the lack of standardized model zoos and available benchmarks significantly increases the friction for further research about populations of NNs. With this work, we publish a novel dataset of model zoos containing systematically generated and diverse populations of NN models for further research. In total the proposed model zoo dataset is based on eight image datasets, consists of 27 model zoos trained with varying hyperparameter combinations and includes 50’360 unique NN models as well as their sparsified twins, resulting in over 3’844’360 collected model states. Additionally, to the model zoo data we provide an in-depth analysis of the zoos and provide benchmarks for multiple downstream tasks. The dataset can be found at www.modelzoos.cc. Konstantin Schürholt, Diyar Taskiran, Xavier Giró-i-Nieto, Damian Borth |
NeurIPS | 5 |
| 2022 | Toward Global Estimation of Ground-Level NO2 Pollution With Deep Learning and Remote SensingabstractAir pollution is a central environmental problem in countries around the world. It contributes to climate change through the emission of greenhouse gases, and adversely impacts the health of billions of people. Despite its importance, detailed information about the spatial and temporal distribution of pollutants is complex to obtain. Ground-level monitoring stations are sparse, and approaches for modeling air pollution rely on extensive datasets which are unavailable for many locations. We introduce three techniques for the estimation of air pollution to overcome these limitations: 1) a baseline localized approach that mimics conventional land-use regression through gradient boosting; 2) an OpenStreetMap (OSM) approach with gradient boosting that is applicable beyond regions covered by detailed geographic datasets; and 3) a remote sensing-based deep learning method utilizing multiband imagery and trace-gas column density measurements from satellites. We focus on the estimation of nitrogen dioxide (NO2), a common anthropogenic air pollutant with adverse effects on the environment and human health. Our local baseline model achieves strong results with a mean absolute error (MAE) of 5.18 ±$0.16~\mu \text {g/m}^{3}$NO2. Substituting localized inputs with OSM leads to a degraded performance (MAE 7.22 ± 0.14) but enables NO2estimation at a global scale. The proposed deep learning model on remote sensing data combines high accuracy (MAE 5.5 ± 0.14) with global coverage and heteroscedastic uncertainty quantification. Our results enable the estimation of surface-level NO2pollution with high spatial resolution for any location on Earth. We illustrate this capability with an out-of-distribution test set on the US westcoast. Code1and data2are publicly available. Linus Scheibenreif, Michael Mommert, Damian Borth |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Power Plant Classification from Remote Imaging with Deep LearningabstractSatellite remote imaging enables the detailed study of land use patterns on a global scale. We investigate the possibility to improve the information content of traditional land use classification by identifying the nature of industrial sites from medium-resolution remote sensing images. In this work, we focus on classifying different types of power plants from Sentinel-2 imaging data. Using a ResNet-50 deep learning model, we are able to achieve a mean accuracy of 90.0% in distinguishing 10 different power plant types and a background class. Furthermore, we are able to identify the cooling mechanisms utilized in thermal power plants with a mean accuracy of 87.5%. Our results enable us to qualitatively investigate the energy mix from Sentinel-2 imaging data, and prove the feasibility to classify industrial sites on a global scale from freely available satellite imagery. Michael Mommert, Linus Scheibenreif, Joëlle Hanna, Damian Borth |
IGARSS | 4 |
| 2021 | A Novel Dataset and Benchmark for Surface No2 Prediction from Remote Sensing Data Including Covid Lockdown MeasuresabstractNO2is an atmospheric trace gas that contributes to global warming as a precursor of greenhouse gases and has adverse effects on human health. Surface NO2concentrations are commonly measured through strictly localized networks of air quality stations on the ground. This work presents a novel dataset of surface NO2measurements aligned with atmospheric column densities from Sentinel-5P, as well as geographic and meteorological variables and lockdown information11Available at https://github.com/HSG-AIML/NO2-dataset. The dataset provides access to data from a variety of sources through a common format and will foster data-driven research into the causes and effects of NO2pollution. We showcase the value of the new dataset on the task of surface NO2estimation with gradient boosting. The resulting models enable daily estimates and confident identification of EU NO2exposure limit breaches. Additionally, we investigate the influence of COVID-19 lockdowns on air quality in Europe and find a significant decrease in NO2levels. Linus Scheibenreif, Michael Mommert, Damian Borth |
IGARSS | 3 |
| 2021 | Learning Interpretable Concept Groups in CNNsabstractWe propose a novel training methodology---Concept Group Learning (CGL)---that encourages training of interpretable CNN filters by partitioning filters in each layer into \emph{concept groups}, each of which is trained to learn a single visual concept. We achieve this through a novel regularization strategy that forces filters in the same group to be active in similar image regions for a given layer. We additionally use a regularizer to encourage a sparse weighting of the concept groups in each layer so that a few concept groups can have greater importance than others. We quantitatively evaluate CGL's model interpretability using standard interpretability evaluation techniques and find that our method increases interpretability scores in most cases. Qualitatively we compare the image regions which are most active under filters learned using CGL versus filters learned without CGL and find that CGL activation regions more strongly concentrate around semantically relevant features. Saurabh Varshneya, Antoine Ledent, Robert A. Vandermeulen, Yunwen Lei, Matthias Enders, Damian Borth, Marius Kloft |
IJCAI | 6 |
| 2021 | Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic PredictionabstractSelf-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn hyper-representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings. Konstantin Schürholt, Dimche Kostadinov, Damian Borth |
NeurIPS | 3 |
| 2019 | Multi-Task Learning for Segmentation of Building Footprints with Deep Neural NetworksabstractThe increased availability of high-resolution satellite imagery allows to sense very detailed structures on the surface of our planet. Access to such information opens up new directions in the analysis of remote sensing imagery. While deep neural networks have achieved significant advances in semantic segmentation of high-resolution images, most of the existing approaches tend to produce predictions with poor boundaries. In this paper, we address the problem of preserving semantic segmentation boundaries in high-resolution satellite imagery by introducing a novel multi-task loss. The loss leverages multiple output representations of the segmentation mask and biases the network to focus more on pixels near boundaries. We evaluate our approach on the large-scale Inria Aerial Image Labeling Dataset which contains high-resolution images. Our results show that we are able to outperform state-of-the-art methods by 9.8% on the Intersection over Union (IoU) metric without any additional post-processing steps. Source code and all models will be available under https://github.com/bbischke/MultiTaskBuildingSegmentation. Benjamin Bischke, Patrick Helber, Joachim Folz, Damian Borth, Andreas Dengel 0001 |
ICIP | 4 |
| 2018 | Overcoming Missing and Incomplete Modalities with Generative Adversarial Networks for Building Footprint SegmentationabstractThe integration of information acquired with different modalities, spatial resolution and spectral bands has shown to improve predictive accuracies. Data fusion is therefore one of the key challenges in remote sensing. Most prior work focusing on multi-modal fusion, assumes that modalities are always available during inference. This assumption limits the applications of multi-modal models since in practice the data collection process is likely to generate data with missing, incomplete or corrupted modalities. In this paper, we show that Generative Adversarial Networks can be effectively used to overcome the problems that arise when modalities are missing or incomplete. Focusing on semantic segmentation of building footprints with missing modalities, our approach achieves an improvement of about 2% on the Intersection over Union (IoU) against the same network that relies only on the available modality. Benjamin Bischke, Patrick Helber, Florian König, Damian Borth, Andreas Dengel 0001 |
CBMI | 4 |
| 2018 | What Do Deep Networks Like to See?abstractWe propose a novel way to measure and understand convolutional neural networks by quantifying the amount of input signal they let in. To do this, an autoencoder (AE) was fine-tuned on gradients from a pre-trained classifier with fixed parameters. We compared the reconstructed samples from AEs that were fine-tuned on a set of image classifiers (AlexNet, VGG16, ResNet-50, and Inception v3) and found substantial differences. The AE learns which aspects of the input space to preserve and which ones to ignore, based on the information encoded in the backpropagated gradients. Measuring the changes in accuracy when the signal of one classifier is used by a second one, a relation of total order emerges. This order depends directly on each classifier's input signal but it does not correlate with classification accuracy or network size. Further evidence of this phenomenon is provided by measuring the normalized mutual information between original images and auto-encoded reconstructions from different fine-tuned AEs. These findings break new ground in the area of neural network understanding, opening a new way to reason, debug, and interpret their results. We present four concrete examples in the literature where observations can now be explained in terms of the input signal that a model uses. Sebastian Palacio, Joachim Folz, Jörn Hees, Federico Raue, Damian Borth, Andreas Dengel 0001 |
CVPR | 5 |
| 2018 | Saliency based Adjective Noun Pair Detection System
Marco Stricker, Syed Saqib Bukhari, Damian Borth, Andreas Dengel 0001 |
ICAART (2) | 3 |
| 2018 | Segmentation of Imbalanced Classes in Satellite Imagery using Adaptive Uncertainty Weighted Class LossabstractWe propose a novel loss function for the training of deep Convolutional Neural Networks (CNNs) focusing on land use and land cover classification in remote sensed data. In satellite imagery, object classes are often highly imbalanced leading to poor pixel-wise classification results when using standard training methods only. In this work, we introduce a loss function which leverages the per class uncertainty of the model during training together with median frequency balancing of the class pixels. We evaluate our result on aerial images of the state-of-the-art dataset Vaihingen. We obtain a significant improvement of the F1-Score and pixel accuracy against the standard cross entropy loss on the small car class. The overall Fl-Score using a single CNN achieves 89.35% resulting in an error reduction of 21.22% against the baseline. Benjamin Bischke, Patrick Helber, Damian Borth, Andreas Dengel 0001 |
IGARSS | 3 |
| 2018 | Introducing Eurosat: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover ClassificationabstractIn this paper, we address the challenge of land use and land cover classification using Sentinel-2 satellite images. The key contributions are as follows. We present a novel dataset based on Sentinel-2 satellite images covering 13 different spectral bands and consisting of 10 classes with in total 27,000 labeled images. We evaluate state-of-the-art deep Convolutional Neural Networks (CNNs) on this novel dataset with its different spectral bands. We also evaluate deep CNNs on existing remote sensing datasets and compare the obtained results. With the proposed novel dataset, we achieved an overall classification accuracy of 98.57%. The classification system resulting from the proposed research opens a gate towards various Earth observation applications. We demonstrate how the classification system can assist in improving geographical maps. Patrick Helber, Benjamin Bischke, Andreas Dengel 0001, Damian Borth |
IGARSS | 4 |
| 2017 | Which Saliency Detection Method is the Best to Estimate the Human Attention for Adjective Noun Concepts?
Marco Stricker, Syed Saqib Bukhari, Mohammad Al-Naser, Seyyed Saleh Mozafari Chanijani, Damian Borth, Andreas Dengel 0001 |
ICAART (2) | 5 |
| 2017 | Social Multimedia Sentiment AnalysisabstractSocial multimedia refers to the multimedia content (text, images, and videos) generated by social network users for social interactions. The increasing popularity of online social networks leads to a significant amount of multimedia content generated by online social network users. Researchers from both the industrial and academic have been working on a broad range of projects related to the analyzing and understanding the online multimedia content, including real world activity prediction and content recommendation. Particularly, understanding online users' opinions or sentiments is a fundamental task that can benefit many applications, such as political campaigning and commercial marketing. We present a few recent advances in social multimedia sentiment analysis. Specifically, this tutorial consists of three parts. The first part is on visual sentiment analysis. We will introduce the task of visual sentiment, its main challenges, and the state-of-the-art approaches. We will include several representative approaches to manually designing visual features for this task as well as some approaches using deep neural networks. The second part is on building multimedia sentiment analysis datasets. We will introduce the challenges, the solutions in the construction of different large-scale datasets for sentiment analysis. The final part is mainly on multimodality model for sentiment analysis. We will introduce some recent research projects on multimodality designing and learning. In addition, we will also share some applications of sentiment analysis, as well as thoughts on current challenges and future directions. Jiebo Luo 0001, Damian Borth, Quanzeng You |
ACM Multimedia | 2 |
| 2016 | An Evolutionary Algorithm to Learn SPARQL Queries for Source-Target-Pairs - Finding Patterns for Human Associations in DBpedia
Jörn Hees, Rouven Bauer, Joachim Folz, Damian Borth, Andreas Dengel 0001 |
EKAW | 4 |
| 2016 | A Tesseract-based OCR framework for historical documents lacking ground-truth textabstractComputationally transcribing historical document images to digital text often requires an initial, labor intensive recording of ground-truths by language experts to provide the OCR system with training text. This paper presents a framework for the automatic generation of training data, provided only with labeled character images and a digital font, thus removing the need for manually generated text. In contrast to sample images and their real text as ground-truth, our approach is based on the random, rule-based generation of “meaningless” text in an image file and a ground-truth text file. Furthermore, we experimentally demonstrate that there is a correlation between the similarity of the character sample images of a subset to each other and the resulting model's classification performance. This allows us to calculate upper- and lower-bound performance subsets for model generation using only the sample images themselves. We show that using more training samples does not unequivocally improve model performance, allowing us to focus on the case of one sample per character during training. Training a Tesseract model only with samples that maximize a dissimilarity metric for each character in regards to all other mean character sample images yields a character recognition error of ca. 15% on our custom benchmark of 15thcentury Latin documents, as compared to ca. 27% error rate for training a model in a traditional Tesseract style using a synthetically generated training image from (manually marked) real text. Brennan Nunamaker, Syed Saqib Bukhari, Damian Borth, Andreas Dengel 0001 |
ICIP | 3 |
| 2016 | Introducing Concept And Syntax Transition Networks for Image CaptioningabstractThe area of image captioning i.e. the automatic generation of short textual descriptions of images has experienced much progress recently. However, image captioning approaches often only focus on describing the content of the image without any emotional or sentimental dimension which is common in human captions. This paper presents an approach for image captioning designed specifically to incorporate emotions and feelings into the caption generation process. The presented approach consists of a Deep Convolutional Neural Network (CNN) for detecting Adjective Noun Pairs in the image and a novel graphical network architecture called "Concept And Syntax Transition (CAST)" network for generating sentences from these detected concepts. Philipp Blandfort, Tushar Karayil, Damian Borth, Andreas Dengel 0001 |
ICMR | 3 |
| 2016 | Contextual Enrichment of Remote-Sensed Events with Social Media StreamsabstractThe availability of satellite images for academic or commercial purpose is increasing rapidly due to efforts made by governmental agencies (NASA, ESA) to publish such data openly or commercial startups (PlanetLabs) to provide real-time satellite data. Beyond many commercial application, satellite data is helpful to create situation awareness in disaster recovery and emergency situations such as wildfires, earthquakes, or flooding. To fully utilize such data sources, we present a scalable system for the contextual enrichment of satellite images by crawling and analyzing multimedia content from social media. This information stream can provide vital information from the ground and help to complement remote sensing in situations. We use Twitter as main data source and analyze its textual, visual, temporal, geographical and social dimensions. Visualizations show different aspects of the event allowing high-level comprehension and provide deeper insights into the event as complemented by social media. Benjamin Bischke, Damian Borth, Christian Schulze 0001, Andreas Dengel 0001 |
ACM Multimedia | 2 |
| 2016 | Generating Affective Captions using Concept And Syntax Transition NetworksabstractThe area of image captioning i.e. the automatic generation of short textual descriptions of images has experienced much progress recently. However, image captioning approaches often only focus on describing the content of the image without any emotional or sentimental dimension which is common in human captions. This paper presents an approach for image captioning designed specifically to incorporate emotions and feelings into the caption generation process. The presented approach consists of a Deep Convolutional Neural Network (CNN) for detecting Adjective Noun Pairs in the image and a graphical network architecture called "Concept And Syntax Transition (CAST)" network for generating sentences from these detected concepts. Tushar Karayil, Philipp Blandfort, Damian Borth, Andreas Dengel 0001 |
ACM Multimedia | 3 |
| 2016 | Multimedia COMMONS Workshop 2016 (MMCommons 2016): Datasets, Evaluation, and ReproducibilityabstractLeveraged wisely, new datasets can inspire new multimedia methods and algorithms, as well as catalyze innovations in how their efficacy, efficiency, and generalizability can be evaluated. The availability of very large multimedia datasets like the Yahoo-Flickr Creative Commons 100 Million has offered unique opportunities for advancing the state of the art in multimedia processing, analysis, search, and visualization. The Multimedia Commons Initiative has been developing a community around the YFCC100M, including associated annotation and evaluation efforts. In addition to research in several multimedia subfields, including computer vision, image processing, and video content analysis, the YFCC100M and Multimedia Commons resources have been used in various competitions and benchmarks, such as the MediaEval Placing Task and the ACM Multimedia Grand Challenge competition. With additional annotation and curation, the data has the potential to enable major leaps forward in research. Bart Thomee, Damian Borth, Julia Bernd |
ACM Multimedia | 2 |
| 2014 | Automatic Detection of CSA Media by Multi-modal Feature Fusion for Law Enforcement SupportabstractThe growing amounts of multimedia data being made available and shared via the Internet pose an increasing problem for law enforcement to investigate the distribution and possession of child sexual abuse (CSA) media. In this paper we address the automatic detection of CSA material in image and video data by multi-modal feature description. Instead of analyzing hash sums or file names, we propose the content-based analysis on visual and, in case of videos, also audio features. To this end, we apply multiple low level features as well as SentiBank, a novel mid-level representation of visual content. In collaboration with police partners and European cyber crime units, we conducted experiments on several datasets, including real world CSA media. Our quantitative evaluation reveals the challenging nature of child pornography detection, especially in the joint presence of non-illegal pornographic data, rendering skin detection, a popular feature for detecting pornography, less discriminative. Further, the utilization of SentiBank features shows high potential for detection and explainability of such content. Overall, multi-modal feature fusion can achieve an improved detection accuracy, reducing equal error rate from 17% to 10% for images and from 16% to 8% for videos as compared to best single feature performance for the challenging task of classifying CSA content from adult media. Christian Schulze 0001, Dominik Henter, Damian Borth, Andreas Dengel 0001 |
ICMR | 3 |
| 2013 | Analysis and forecasting of trending topics in online media streamsabstractAmong the vast information available on the web, social media streams capture what people currently pay attention to and how they feel about certain topics. Awareness of such trending topics plays a crucial role in multimedia systems such as trend aware recommendation and automatic vocabulary selection for video concept detection systems. Correctly utilizing trending topics requires a better understanding of their various characteristics in different social media streams. To this end, we present the first comprehensive study across three major online and social media streams, Twitter, Google, and Wikipedia, covering thousands of trending topics during an observation period of an entire year. Our results indicate that depending on one's requirements one does not necessarily have to turn to Twitter for information about current events and that some media streams strongly emphasize content of specific categories. As our second key contribution, we further present a novel approach for the challenging task of forecasting the life cycle of trending topics in the very moment they emerge. Our fully automated approach is based on a nearest neighbor forecasting technique exploiting our assumption that semantically similar topics exhibit similar behavior. Tim Althoff, Damian Borth, Jörn Hees, Andreas Dengel 0001 |
ACM Multimedia | 2 |
| 2013 | SentiBank: large-scale ontology and classifiers for detecting sentiment and emotions in visual contentabstractA picture is worth one thousand words, but what words should be used to describe the sentiment and emotions conveyed in the increasingly popular social multimedia? We demonstrate a novel system which combines sound structures from psychology and the folksonomy extracted from social multimedia to develop a large visual sentiment ontology consisting of 1,200 concepts and associated classifiers called SentiBank. Each concept, defined as an Adjective Noun Pair (ANP), is made of an adjective strongly indicating emotions and a noun corresponding to objects or scenes that have a reasonable prospect of automatic detection. We believe such large-scale visual classifiers offer a powerful mid-level semantic representation enabling high-level sentiment analysis of social multimedia. We demonstrate novel applications made possible by SentiBank including live sentiment prediction of social media and visualization of visual content in a rich intuitive semantic space. Damian Borth, Tao Chen 0015, Rongrong Ji, Shih-Fu Chang |
ACM Multimedia | 1 |
| 2013 | Large-scale visual sentiment ontology and detectors using adjective noun pairsabstractWe address the challenge of sentiment analysis from visual content. In contrast to existing methods which infer sentiment or emotion directly from visual low-level features, we propose a novel approach based on understanding of the visual concepts that are strongly related to sentiments. Our key contribution is two-fold: first, we present a method built upon psychological theories and web mining to automatically construct a large-scale Visual Sentiment Ontology (VSO) consisting of more than 3,000 Adjective Noun Pairs (ANP). Second, we propose SentiBank, a novel visual concept detector library that can be used to detect the presence of 1,200 ANPs in an image. The VSO and SentiBank are distinct from existing work and will open a gate towards various applications enabled by automatic sentiment analysis. Experiments on detecting sentiment of image tweets demonstrate significant improvement in detection accuracy when comparing the proposed SentiBank based predictors with the text-based approaches. The effort also leads to a large publicly available resource consisting of a visual sentiment ontology, a large detector library, and the training/testing benchmark for visual sentiment analysis. Damian Borth, Rongrong Ji, Tao Chen 0015, Thomas M. Breuel, Shih-Fu Chang |
ACM Multimedia | 1 |
| 2012 | Linking visual concept detection with viewer demographicsabstractThe estimation of demographic target groups for web videos -- with applications in ad targeting -- poses a challenging problem, as the textual description and view statistics available for many clips is extremely sparse. Therefore, the goal of this paper is to link a clip's popularity across different viewer ages and genders on the one hand with the video content on the other: Employing user comments and user profiles on YouTube, we show that there is a strong correlation between demographic target groups and semantic concepts appearing in the video (like "teenage male" and "skateboarding"). Based on this observation, we suggest two approaches: First, the demographic target group of a clip is predicted automatically via a content-based concept detection. Second, should sufficient view statistics already give a good impression of a video's audience, we show that this information can serve as a valuable additional signal to disambiguate concept detection. Adrian Ulges, Markus Koch, Damian Borth |
ICMR | 3 |
| 2012 | Dynamic vocabularies for web-based concept detection by trend discoveryabstractWe present a novel approach towards automatic vocabulary selection for video concept detection. Our key idea is to expand concept vocabularies with trending topics that we mine automatically on other media like Wikipedia or Twitter. We evaluate several strategies for extending concept detection to auto-detect these topics in new videos, either by linking them to a static concept vocabulary, by a visual learning of trends on the fly, or by an expansion of the vocabulary. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ACM Multimedia | 1 |
| 2011 | Lookapp: interactive construction of web-based concept detectorsabstractWhile online platforms like YouTube and Flickr do provide massive content for training of visual concept detectors, it remains a difficult challenge to retrieve the right training content from such platforms. In this technical demonstration we present lookapp, a system for the interactive construction of web-based concept detectors. It major features are an interactive "concept-to-query" mapping for training data acquisition and an efficient detector construction based on third party cloud computing services. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ICMR | 1 |
| 2011 | Automatic concept-to-query mapping for web-based concept detector trainingabstractNowadays, online platforms like YouTube provide massive content for training of visual concept detectors. However, it remains a difficult challenge to retrieve the right training content from such platforms since the underlying query construction can be arbitrarily complex. In this paper we present an approach, which offers an automatic concept-to-query mapping for training data acquisition from such platforms. Queries are automatically constructed by a keyword selection and a category assignment using ImageNet and Google Sets as external sources. Our results demonstrate that the proposed method is able to reach retrieval results comparable to queries constructed by humans providing 76% more relevant content for detector training than a one-to-one mapping of concept names to retrieval queries would do. Damian Borth, Adrian Ulges, Thomas M. Breuel |
ACM Multimedia | 1 |
| 2009 | TubeFiler: an automatic web video categorizerabstractWhile hierarchies are powerful tools for organizing content in other application areas, current web video platforms offer only limited support for a taxonomy-based browsing. To overcome this limitation, we present a framework called TubeFiler. Its two key features are an automatic multimodal categorization of videos into a genre hierarchy, and a support of additional fine-grained hierarchy levels based on unsupervised learning. We present experimental results on real-world YouTube clips with a 2-level 46-category genre hierarchy, indicating that - though the problem is clearly challenging - good category suggestions can be achieved. For example, if TubeFiler suggests 5 categories, it hits the right one (or at least its supercategory) in 91.8% of cases. Damian Borth, Jörn Hees, Markus Koch, Adrian Ulges, Christian Schulze 0001, Thomas M. Breuel, Roberto Paredes |
ACM Multimedia | 1 |