EDBT 2026 Demo / reviewers in the wild / expert
Begüm Demir
dblp:47/622
· DBLP profile ↗
106ranked-venue papers
33as first author
46since 2021 · last 2026
0000-0003-2175-7072ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 85 · 29 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing ImageryabstractSelf-supervised learning (SSL) has become a powerful paradigm for learning from large, unlabeled datasets, particularly in computer vision (CV). However, applying SSL to multispectral remote sensing (RS) images presents unique challenges and opportunities due to the geographical and temporal variability of the data. In this paper, we introduce GeoRank, a novel regularization method for contrastive SSL that improves upon prior techniques by directly optimizing spherical distances to embed geographical relationships into the learned feature space. GeoRank outperforms or matches prior methods that integrate geographical metadata and consistently improves diverse contrastive SSL algorithms (e.g., BYOL, DINO). Beyond this, we present a systematic investigation of key adaptations of contrastive SSL for multispectral RS images, including the effectiveness of data augmentations, the impact of dataset cardinality and image size on performance, and the task dependency of temporal views. Code is available at https://github.com/tomburgert/georank. Tom Burgert, Leonard W. Hackel, Paolo Rota, Begüm Demir |
WACV | 4 |
| 2026 | Hybrid Deep Learning Models for Remote Sensing Image ProcessingabstractCore image processing tasks, such as super-resolution, denoising, deblurring, pansharpening, and atmospheric correction, underpin all optical remote sensing (RS) pipelines. Errors at this stage propagate through downstream applications, distorting land-cover maps, change detection, and climate records. Classical physics-based models capture sensor optics, radiometry, and geometry but struggle with complex noise and scene variability. In contrast, deep learning (DL) methods offer powerful data-driven solutions yet often act as closed boxes, ignoring physical constraints and overfitting to spurious patterns. Hybrid DL (HDL) approaches bridge this gap by integrating physical models with neural architectures, combining interpretability and data adaptivity. This article surveys the emerging landscape of HDL methods in RS image processing, outlining their theoretical foundations, motivations, and design philosophies. We categorize fusion strategies, from model-embedded schemes (e.g., plug-and-play (PnP) and unrolling) to model-guided learning (e.g., deep image prior (DIP) and unsupervised frameworks), and discuss how they enhance trust, robustness, and physical consistency in RS image analysis. Matthieu Muller, Daniele Picone, Begüm Demir, Gustau Camps-Valls, Mauro Dalla Mura, Magnus O. Ulfarsson, Jón Atli Benediktsson |
Proc. IEEE | 3 |
| 2026 | Radio Map Prediction From Aerial Images and Application to Coverage OptimizationabstractSeveral studies have explored deep learning algorithms to predict large-scale signal fading, or path loss, in urban communication networks. The goal is to replace costly measurement campaigns, inaccurate statistical models, or computationally expensive ray-tracing simulations with machine learning models that deliver quick and accurate predictions. We focus on predicting path loss radio maps using convolutional neural networks, leveraging aerial images alone or in combination with supplementary height information. Notably, our approach does not rely on explicit classification of environmental objects, which is often unavailable for most locations worldwide. While the prediction of radio maps using complete 3D environmental data is well-studied, the use of only aerial images remains under-explored. We address this gap by showing that state-of-the-art models developed for existing radio map datasets can be effectively adapted to this task. Additionally, we introduce a new model dubbed UNetDCN that achieves on par or better performance compared to the state-of-the-art with reduced complexity. The trained models are differentiable, and therefore they can be incorporated in various network optimization algorithms. While an extensive discussion is beyond this paper’s scope, we demonstrate this through an example optimizing the directivity of base stations in cellular networks via backpropagation to enhance coverage. Fabian Jaensch, Giuseppe Caire, Begüm Demir |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppressionabstractThe hypothesis that Convolutional Neural Networks (CNNs) are inherently texture-biased has shaped much of the discourse on feature use in deep learning. We revisit this hypothesis by examining limitations in the cue-conflict experiment by Geirhos et al. To address these limitations, we propose a domain-agnostic framework that quantifies feature reliance through systematic suppression of shape, texture, and color cues, avoiding the confounds of forced-choice conflicts. By evaluating humans and neural networks under controlled suppression conditions, we find that CNNs are not inherently texture-biased but predominantly rely on local shape features. Nonetheless, this reliance can be substantially mitigated through modern training strategies or architectures (ConvNeXt, ViTs). We further extend the analysis across computer vision, medical imaging, and remote sensing, revealing that reliance patterns differ systematically: computer vision models prioritize shape, medical imaging models emphasize color, and remote sensing models exhibit a stronger reliance on texture. Code is available at https://github.com/tomburgert/feature-reliance. Tom Burgert, Oliver Stoll, Paolo Rota, Begüm Demir |
NeurIPS | 4 |
| 2025 | Sea-Undistort: A Dataset for Through-Water Image Restoration in High-Resolution Airborne Bathymetric MappingabstractAccurate image-based bathymetric mapping in shallow waters remains challenging due to the complex optical distortions such as wave induced patterns, scattering and sunglint, introduced by the dynamic water surface, the water column properties, and solar illumination. In this work, we introduce Sea-Undistort, a comprehensive synthetic dataset of 1200 paired 512×512 through-water scenes rendered in Blender. Each pair comprises a distortion-free and a distorted view, featuring realistic water effects such as sun glint, waves, and scattering over diverse seabeds. Accompanied by per-image metadata such as camera parameters, sun position, and average depth, Sea-Undistort enables supervised training that is otherwise infeasible in real environments. We use Sea-Undistort to benchmark two state-of-the-art image restoration methods alongside an enhanced lightweight diffusion-based framework with an early-fusion sunglint mask. When applied to real aerial data, the enhanced diffusion model delivers more complete Digital Surface Models (DSMs) of the seabed, especially in deeper areas, reduces bathymetric errors, suppresses glint and scattering, and crisply restores fine seabed details. The Sea-Undistort dataset and code will be released upon acceptance. Maximilian Kromer, Panagiotis Agrafiotis, Begüm Demir |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Continual Self-Supervised Learning With Masked Autoencoders in Remote SensingabstractThe development of continual learning (CL) methods, which aim to learn new tasks in a sequential manner from the training data acquired continuously, has gained great attention in remote sensing (RS). The existing CL methods in RS, while learning new tasks, enhance robustness towards catastrophic forgetting. This is achieved by using a large number of labeled training samples, which is costly and not always feasible to gather in RS. To address this problem, we propose a novel continual self-supervised learning method in the context of masked autoencoders (denoted as CoSMAE). The proposed CoSMAE consists of two components: i) data mixup; and ii) model mixup knowledge distillation. Data mixup is associated with retaining information on previous data distributions by interpolating images from the current task with those from the previous tasks. Model mixup knowledge distillation is associated with distilling knowledge from past models and the current model simultaneously by interpolating their model weights to form a teacher for the knowledge distillation. The two components complement each other to regularize the MAE at the data and model levels to facilitate better generalization across tasks and reduce the risk of catastrophic forgetting. Experimental results show that CoSMAE achieves significant improvements of up to 4.94% over state-of-the-art CL methods applied to MAE. Upon acceptance of the letter, the code will be made publicly available at: https://git.tu-berlin.de/rsim/CoSMAE. Lars Möllenbrok, Behnood Rasti, Begüm Demir |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Exploring Masked Autoencoders for Sensor-Agnostic Image Retrieval in Remote SensingabstractSelf-supervised learning through masked autoencoders (MAEs) has recently attracted great attention for remote sensing (RS) image representation learning (IRL), and thus embodies a significant potential for content-based image retrieval (CBIR) from ever-growing RS image archives. However, the existing MAE-based CBIR studies in RS assume that the considered RS images are acquired by a single image sensor, and thus are only suitable for unimodal CBIR problems. The effectiveness of MAEs for cross-sensor CBIR, which aims to search semantically similar images across different image modalities, has not been explored yet. In this article, we take the first step to explore the effectiveness of MAEs for sensor-agnostic CBIR in RS. To this end, we present a systematic overview on the possible adaptations of the vanilla MAE to exploit masked image modeling (MIM) on multisensor RS image archives [denoted as cross-sensor masked autoencoders [(CSMAEs)] in the context of CBIR. Based on different adjustments applied to the vanilla MAE, we introduce different CSMAE models. We also provide an extensive experimental analysis of these CSMAE models. We finally derive a guideline to exploit MIM for unimodal and cross-modal CBIR problems in RS. The code of this work is publicly available athttps://github.com/jakhac/CSMAE. Jakob Hackstein, Gencer Sumbul, Kai Norman Clasen, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MAGICBATHYNET: A Multimodal Remote Sensing Dataset for Bathymetry Prediction and Pixel-Based Classification in Shallow WatersabstractAccurate, detailed, and regularly updated bathymetry, coupled with complex semantic content, is crucial for the under-mapped shallow water areas facing intense climatological and anthropogenic pressures. Current methods exploiting remote sensing imagery to derive bathymetry or pixel-based seabed classes mainly exploit non-open data. This lack of openly accessible benchmark archives prevents the wider use of deep learning methods in such applications. To address this issue, in this paper we present the MagicBathyNet, which is a benchmark dataset made up of image patches of Sentinel-2, SPOT-6 and aerial imagery, bathymetry in raster format and annotations of seabed classes. MagicBathyNet is then exploited to benchmark state-of-the-art methods in learning-based bathymetry and pixel-based classification. Dataset, pre-trained weights, and code are publicly available at www.magicbathy.eu/magicbathynet.html. Panagiotis Agrafiotis, Lukasz Janowski, Dimitrios Skarlatos 0001, Begüm Demir |
IGARSS | 4 |
| 2024 | Estimating Physical Information Consistency of Channel Data Augmentation for Remote Sensing ImagesabstractThe application of data augmentation for deep learning (DL) methods plays an important role in achieving state-of-the-art results in supervised, semi-supervised, and self-supervised image classification. In particular, channel transformations (e.g., solarize, grayscale, brightness adjustments) are integrated into data augmentation pipelines for remote sensing (RS) image classification tasks. However, contradicting beliefs exist about their proper applications to RS images. A common point of critique is that the application of channel augmentation techniques may lead to physically inconsistent spectral data (i.e., pixel signatures). To shed light on the open debate, we propose an approach to estimate whether a channel augmentation technique affects the physical information of RS images. To this end, the proposed approach estimates a score that measures the alignment of a pixel signature within a time series that can be naturally subject to deviations caused by factors such as acquisition conditions or phenological states of vegetation. We compare the scores associated with original and augmented pixel signatures to evaluate the physical consistency. Experimental results on a multi-label image classification task show that channel augmentations yielding a score that exceeds the expected deviation of original pixel signatures can not improve the performance of a baseline model trained without augmentation. Tom Burgert, Begüm Demir |
IGARSS | 2 |
| 2024 | Transformer-Based Federated Learning For Multi-Label Remote Sensing Image ClassificationabstractFederated learning (FL) aims to collaboratively learn deep learning model parameters from decentralized data archives (i.e., clients) without accessing training data on clients. However, the training data across clients might be not independent and identically distributed (non-IID), which may result in difficulty in achieving optimal model convergence. In this work, we investigate the capability of state-of-the-art transformer architectures (which are MLP-Mixer, ConvMixer, PoolFormer) to address the challenges related to non-IID training data across various clients in the context of FL for multi-label classification (MLC) problems in remote sensing (RS). The considered transformer architectures are compared among themselves and with the ResNet-50 architecture in terms of their: 1) robustness to training data heterogeneity; 2) local training complexity; and 3) aggregation complexity under different non-IID levels. The experimental results obtained on the BigEarthNet-S2 benchmark archive demonstrate that the considered architectures increase the generalization ability with the cost of higher local training and aggregation complexities. On the basis of our analysis, some guidelines are derived for a proper selection of transformer architecture in the context of FL for RS MLC. The code of this work is publicly available at https://git.tu-berlin.de/rsim/FL-Transformer. Baris Büyüktas, Kenneth Weitzel, Sebastian Völkers, Felix Zailskas, Begüm Demir |
IGARSS | 5 |
| 2024 | Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time SeriesabstractUsing images acquired by different satellite sensors has shown to improve classification performance in the framework of crop mapping from satellite image time series (SITS). Existing state-of-the-art architectures use self-attention mechanisms to process the temporal dimension and convolutions for the spatial dimension of SITS. Motivated by the success of purely attention-based architectures in crop mapping from single-modal SITS, we introduce several multi-modal multi-temporal transformer-based architectures. Specifically, we investigate the effectiveness of Early Fusion, Cross Attention Fusion and Synchronized Class Token Fusion within the Temporo-Spatial Vision Transformer (TSViT). Experimental results demonstrate significant improvements over state-of-the-art architectures with both convolutional and self-attention components. Theresa Follath, David Mickisch, Jan Hemmerling, Stefan Erasmi, Marcel Schwieder, Begüm Demir |
IGARSS | 6 |
| 2024 | Annotation Cost-Efficient Active Learning for Deep Metric Learning-Driven Remote Sensing Image RetrievalabstractDeep metric learning (DML) has shown to be effective for content-based image retrieval (CBIR) in remote sensing (RS). Most of the DML methods for CBIR rely on a high number of annotated images to accurately learn model parameters of deep neural networks (DNNs). However, gathering such data is time-consuming and costly. To address this, we propose an annotation cost-efficient active learning (ANNEAL) method tailored to DML-driven CBIR in RS. ANNEAL aims to create a small but informative training set made up of similar and dissimilar image pairs to be used for accurately learning a metric space. The informativeness of image pairs is evaluated by combining uncertainty and diversity criteria. To assess the uncertainty of image pairs, we introduce two algorithms: 1) metric-guided uncertainty estimation (MGUE) and 2) binary-classifier-guided uncertainty estimation (BCGUE). MGUE algorithm automatically estimates a threshold value that acts as a boundary between similar and dissimilar image pairs based on the distances in the metric space. The closer the similarity between image pairs is to the estimated threshold value, the higher their uncertainty. BCGUE algorithm estimates the uncertainty of the image pairs based on the confidence of the classifier in assigning correct similarity labels. The diversity criterion is assessed through a clustering-based strategy. ANNEAL combines either MGUE or BCGUE algorithm with the clustering-based strategy to select the most informative image pairs, which are then labeled by expert annotators as similar or dissimilar. This way of annotating images significantly reduces the annotation cost compared with annotating images with land-use land-cover class labels. Experimental results on two RS benchmark datasets demonstrate the effectiveness of our method. The code of this work is publicly available athttps://git.tu-berlin.de/rsim/anneal_tgrs. Genc Hoxha, Gencer Sumbul, Julia Henkel, Lars Möllenbrok, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Multi-Label Noise Robust Collaborative Learning for Remote Sensing Image ClassificationabstractThe development of accurate methods for multi-label classification (MLC) of remote sensing (RS) images is one of the most important research topics in RS. The MLC methods based on convolutional neural networks (CNNs) have shown strong performance gains in RS. However, they usually require a high number of reliable training images annotated with multiple land-cover class labels. Collecting such data is time-consuming and costly. To address this problem, the publicly available thematic products, which can include noisy labels, can be used to annotate RS images with zero-labeling cost. However, multi-label noise (which can be associated with wrong and missing label annotations) can distort the learning process of the MLC methods. To address this problem, we propose a novel multi-label noise robust collaborative learning (RCML) method to alleviate the negative effects of multi-label noise during the training phase of a CNN model. RCML identifies, ranks, and excludes noisy multi-labels in RS images based on three main modules: 1) the discrepancy module; 2) the group lasso module; and 3) the swap module. The discrepancy module ensures that the two networks learn diverse features, while producing the same predictions. The task of the group lasso module is to detect the potentially noisy labels assigned to multi-labeled training images, while the swap module is devoted to exchange the ranking information between two networks. Unlike the existing methods that make assumptions about noise distribution, our proposed RCML does not make any prior assumption about the type of noise in the training set. The experiments conducted on two multi-label RS image archives confirm the robustness of the proposed RCML under extreme multi-label noise rates. Our code is publicly available at: https://www.noisy-labels-in-rs.org. Ahmet Kerem Aksoy, Mahdyar Ravanbakhsh, Begüm Demir |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Learning Across Decentralized Multi-Modal Remote Sensing Archives with Federated LearningabstractThe development of federated learning (FL) methods, which aim to learn from distributed databases (i.e., clients) without accessing data on clients, has recently attracted great attention. Most of these methods assume that the clients are associated with the same data modality. However, remote sensing (RS) images in different clients can be associated with different data modalities that can improve the classification performance when jointly used. To address this problem, in this paper we introduce a novel multi-modal FL framework that aims to learn from decentralized multi-modal RS image archives for RS image classification problems. The proposed framework is made up of three modules: 1) multimodal fusion (MF); 2) feature whitening (FW); and 3) mutual information maximization (MIM). The MF module performs iterative model averaging to learn without accessing data on clients in the case that clients are associated with different data modalities. The FW module aligns the representations learned among the different clients. The MIM module maximizes the similarity of images from different modalities. Experimental results show the effectiveness of the proposed framework compared to iterative model averaging, which is a widely used algorithm in FL. The code of the proposed framework is publicly available at https://git.tu-berlin.de/rsim/MMFL. Baris Büyüktas, Gencer Sumbul, Begüm Demir |
IGARSS | 3 |
| 2023 | HySpecNet-11k: a Large-Scale Hyperspectral Dataset for Benchmarking Learning-Based Hyperspectral Image Compression MethodsabstractThe development of learning-based hyperspectral image compression methods has recently attracted great attention in remote sensing. Such methods require a high number of hyperspectral images to be used during training to optimize all parameters and reach a high compression performance. However, existing hyperspectral datasets are not sufficient to train and evaluate learning-based compression methods, which hinders the research in this field. To address this problem, in this paper we present HySpecNet-11k that is a large-scale hyperspectral benchmark dataset made up of 11,483 nonoverlapping image patches. Each patch is a portion of 128 × 128 pixels with 224 spectral bands and a ground sample distance of 30 m. We exploit HySpecNet-11k to benchmark the current state of the art in learning-based hyperspectral image compression by focussing our attention on various 1D, 2D and 3D convolutional autoencoder architectures. Nevertheless, HySpecNet-11k can be used for any unsupervised learning task in the framework of hyperspectral image analysis. The dataset, our code and the pre-trained weights are publicly available at https://hyspecnet.rsim.berlin. Martin Hermann Paul Fuchs, Begüm Demir |
IGARSS | 2 |
| 2023 | LIT-4-RSVQA: Lightweight Transformer-Based Visual Question Answering in Remote SensingabstractVisual question answering (VQA) methods in remote sensing (RS) aim to answer natural language questions with respect to an RS image. Most of the existing methods require a large amount of computational resources, which limits their application in operational scenarios in RS. To address this issue, in this paper we present an effective lightweight transformer-based VQA in RS (LiT-4-RSVQA) architecture for efficient and accurate VQA in RS. Our architecture consists of: i) a lightweight text encoder module; ii) a lightweight image encoder module; iii) a fusion module; and iv) a classification module. The experimental results obtained on a VQA benchmark dataset demonstrate that our proposed LiT-4-RSVQA architecture provides accurate VQA results while significantly reducing the computational requirements on the executing hardware. Leonard W. Hackel, Kai Norman Clasen, Mahdyar Ravanbakhsh, Begüm Demir |
IGARSS | 4 |
| 2023 | Annotation Cost Efficient Active Learning for Content Based Image RetrievalabstractDeep metric learning (DML) based methods have been found very effective for content-based image retrieval (CBIR) in remote sensing (RS). For accurately learning the model parameters of deep neural networks, most of the DML methods require a high number of annotated training images, which can be costly to gather. To address this problem, in this paper we present an annotation cost efficient active learning (AL) method (denoted as ANNEAL). The proposed method aims to iteratively enrich the training set by annotating the most informative image pairs as similar or dissimilar, while accurately modelling a deep metric space. This is achieved by two consecutive steps. In the first step the pairwise image similarity is modelled based on the available training set. Then, in the second step the most uncertain and diverse (i.e., informative) image pairs are selected to be annotated. Unlike the existing AL methods for CBIR, at each AL iteration of ANNEAL a human expert is asked to annotate the most informative image pairs as similar/dissimilar. This significantly reduces the annotation cost compared to annotating images with land-use/land cover class labels. Experimental results show the effectiveness of our method. The code of ANNEAL is publicly available at https://git.tu-berlin.de/rsim/ANNEAL. Julia Henkel, Genc Hoxha, Gencer Sumbul, Lars Möllenbrok, Begüm Demir |
IGARSS | 5 |
| 2023 | Transformer-Based Multi-Modal Learning for Multi-Label Remote Sensing Image ClassificationabstractIn this paper, we introduce a novel Synchronized Class Token Fusion (SCT Fusion) architecture in the framework of multi-modal multi-label classification (MLC) of remote sensing (RS) images. The proposed architecture leverages modality-specific attention-based transformer encoders to process varying input modalities, while exchanging information across modalities by synchronizing the special class tokens after each transformer encoder block. The synchronization involves fusing the class tokens with a trainable fusion transformation, resulting in a synchronized class token that contains information from all modalities. As the fusion transformation is trainable, it allows to reach an accurate representation of the shared features among different modalities. Experimental results show the effectiveness of the proposed architecture over single-modality architectures and an early fusion multi-modal architecture when evaluated on a multi-modal MLC dataset. The code of the proposed architecture is publicly available at https://git.tu-berlin.de/rsim/sct-fusion. David Hoffmann, Kai Norman Clasen, Begüm Demir |
IGARSS | 3 |
| 2023 | Active Learning Guided Fine-Tuning for Enhancing Self-Supervised based Multi-Label Classification of Remote Sensing ImagesabstractIn recent years, deep neural networks (DNNs) have been found very successful for multi-label classification (MLC) of remote sensing (RS) images. Self-supervised pre-training combined with fine-tuning on a randomly selected small training set has become a popular approach to minimize annotation efforts of data-demanding DNNs. However, finetuning on a small and biased training set may limit model performance. To address this issue, we investigate the effectiveness of the joint use of self-supervised pre-training with active learning (AL). The considered AL strategy aims at guiding the MLC fine-tuning of a self-supervised model by selecting informative training samples to annotate in an iterative manner. Experimental results show the effectiveness of applying AL-guided fine-tuning (particularly for the case where strong class-imbalance is present in MLC problems) compared to the application of fine-tuning using a randomly constructed small training set. Lars Möllenbrok, Begüm Demir |
IGARSS | 2 |
| 2023 | Ben-Ge: Extending Bigearthnet with Geographical and Environmental DataabstractDeep learning methods have proven to be a powerful tool in the analysis of large amounts of complex Earth observation data. However, while Earth observation data are multi-modal in most cases, only single or few modalities are typically considered. In this work, we present the ben-ge dataset, which supplements the BigEarthNet-MM dataset by compiling freely and globally available geographical and environmental data. Based on this dataset, we showcase the value of combining different data modalities for the downstream tasks of patch-based land-use/land-cover classification and land-use/land-cover segmentation. ben-ge is freely available and expected to serve as a test bed for fully supervised and self-supervised Earth observation applications. Michael Mommert, Nicolas Kesseli, Joëlle Hanna, Linus Scheibenreif, Damian Borth, Begüm Demir |
IGARSS | 6 |
| 2023 | Label Noise Robust Image Representation Learning Based on Supervised Variational Autoencoders in Remote SensingabstractDue to the publicly available thematic maps and crowd-sourced data, remote sensing (RS) image annotations can be gathered at zero cost for training deep neural networks (DNNs). However, such annotation sources may increase the risk of including noisy labels in training data, leading to inaccurate RS image representation learning (IRL). To address this issue, in this paper we propose a label noise robust IRL method that aims to prevent the interference of noisy labels on IRL, independently from the learning task being considered in RS. To this end, the proposed method combines a supervised variational autoencoder (SVAE) with any kind of DNN. This is achieved by defining variational generative process based on image features. This allows us to define the importance of each training sample for IRL based on the loss values acquired from the SVAE and the task head of the considered DNN. Then, the proposed method imposes lower importance to images with noisy labels, while giving higher importance to those with correct labels during IRL. Experimental results show the effectiveness of the proposed method when compared to well-known label noise robust IRL methods applied to RS images. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/RS-IRL-SVAE. Gencer Sumbul, Begüm Demir |
IGARSS | 2 |
| 2023 | Deep Active Learning for Multi-Label Classification of Remote Sensing ImagesabstractIn this letter, we introduce deep active learning (AL) for multi-label classification (MLC) problems in remote sensing (RS). In particular, we investigate the effectiveness of several AL query functions for MLC of RS images. Unlike the existing AL query functions (which are defined for single-label classification or semantic segmentation problems), each query function in this paper is based on the evaluation of two criteria: i) multi-label uncertainty; and ii) multi-label diversity. The multi-label uncertainty criterion is associated to the confidence of the deep neural networks (DNNs) in correctly assigning multi-labels to each image. To assess this criterion, we investigate three strategies: i) learning multi-label loss ordering; ii) measuring temporal discrepancy of multi-label predictions; and iii) measuring magnitude of approximated gradient embeddings. The multi-label diversity criterion is associated to the selection of a set of images that are as diverse as possible to each other that prevents redundancy among them. To assess this criterion, we exploit a clustering based strategy. We combine each of the above-mentioned uncertainty strategies with the clustering based diversity strategy, resulting in three different query functions. All the considered query functions are introduced for the first time in the framework of MLC problems in RS. Experimental results obtained on two benchmark archives show that these query functions result in the selection of a highly informative set of samples at each iteration of the AL process. Lars Möllenbrok, Gencer Sumbul, Begüm Demir |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Generative Reasoning Integrated Label Noise Robust Deep Image Representation LearningabstractThe development of deep learning based image representation learning (IRL) methods has attracted great attention for various image understanding problems. Most of these methods require the availability of a set of high quantity and quality of annotated training images, which can be time-consuming, complex and costly to gather. To reduce labeling costs, crowdsourced data, automatic labeling procedures or citizen science projects can be considered. However, such approaches increase the risk of including label noise in training data. It may result in overfitting on noisy labels when discriminative reasoning is employed as in most of the existing methods. This leads to sub-optimal learning procedures, and thus inaccurate characterization of images. To address this issue, in this paper, we introduce a generative reasoning integrated label noise robust deep representation learning (GRID) approach. The proposed GRID approach aims to model the complementary characteristics of discriminative and generative reasoning for IRL under noisy labels. To this end, we first integrate generative reasoning into discriminative reasoning through a supervised variational autoencoder. This allows the proposed GRID approach to automatically detect training samples with noisy labels. Then, through our label noise robust hybrid representation learning strategy, GRID adjusts the whole learning procedure for IRL of these samples through generative reasoning and that of the other samples through discriminative reasoning. Our approach learns discriminative image representations while preventing interference of noisy labels during training independently from the IRL method being selected. Thus, unlike the existing label noise robust methods, GRID does not depend on the type of annotation, label noise, neural network architecture, loss function or learning task, and thus can be directly utilized for various image understanding problems. Experimental results show the effectiveness of the proposed GRID approach compared to the state-of-the-art methods. The code of the proposed approach is publicly available at https://github.com/gencersumbul/GRID. Gencer Sumbul, Begüm Demir |
IEEE Trans. Image Process. | 2 |
| 2023 | A hybrid CUDA, OpenMP, and MPI parallel TCA-based domain adaptation for classification of very high-resolution remote sensing imagesabstractAbstract Domain Adaptation (DA) is a technique that aims at extracting information from a labeled remote sensing image to allow classifying a different image obtained by the same sensor but at a different geographical location. This is a very complex problem from the computational point of view, specially due to the very high-resolution of multispectral images. TCANet is a deep learning neural network for DA classification problems that has been proven as very accurate for solving them. TCANet consists of several stages based on the application of convolutional filters obtained through Transfer Component Analysis (TCA) computed over the input images. It does not require backpropagation training, in contrast to the usual CNN-based networks, as the convolutional filters are directly computed based on the TCA transform applied over the training samples. In this paper, a hybrid parallel TCA-based domain adaptation technique for solving the classification of very high-resolution multispectral images is presented. It is designed for efficient execution on a multi-node computer by using Message Passing Interface (MPI), exploiting the available Graphical Processing Units (GPUs), and making efficient use of each multicore node by using Open Multi-Processing (OpenMP). As a result, an accurate DA technique from the point of view of classification and with high speedup values over the sequential version is obtained, increasing the applicability of the technique to real problems. Alberto S. Garea, Dora Blanco Heras, Francisco Argüello, Begüm Demir |
J. Supercomput. | 4 |
| 2022 | Unsupervised Contrastive Hashing for Cross-Modal Retrieval in Remote SensingabstractThe development of cross-modal retrieval systems that can search and retrieve semantically relevant data across different modalities based on a query in any modality has attracted great attention in remote sensing (RS). In this paper, we focus our attention on cross-modal text-image retrieval, where queries from one modality (e.g., text) can be matched to archive entries from another (e.g., image). Most of the existing cross-modal text-image retrieval systems in RS require a high number of labeled training samples and also do not allow fast and memory-efficient retrieval. These issues limit the applicability of the existing cross-modal retrieval systems for large-scale applications in RS. To address this problem, in this paper we introduce a novel unsupervised cross-modal contrastive hashing (DUCH) method for text-image retrieval in RS. To this end, the proposed DUCH is made up of two main modules: 1) feature extraction module, which extracts deep representations of two modalities; 2) hashing module that learns to generate cross-modal binary hash codes from the extracted representations. We introduce a novel multi-objective loss function including: i) contrastive objectives that enable similarity preservation in intra- and inter-modal similarities; ii) an adversarial objective that is enforced across two modalities for cross-modal representation consistency; and iii) binarization objectives for generating hash codes. Experimental results show that the proposed DUCH outperforms state-of-the-art methods. Our code is publicly available at https://git.tu-berlin.de/rsim/duch. Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir |
ICASSP | 3 |
| 2022 | An Unsupervised Cross-Modal Hashing Method Robust to Noisy Training Image-Text Correspondences in Remote SensingabstractThe development of accurate and scalable cross-modal image-text retrieval methods, where queries from one modality (e.g., text) can be matched to archive entries from another (e.g., remote sensing image) has attracted great attention in remote sensing (RS). Most of the existing methods assume that a reliable multi-modal training set with accurately matched text-image pairs is existing. However, this assumption may not always hold since the multi-modal training sets may include noisy pairs (i.e., textual descriptions/captions associated to training images can be noisy), distorting the learning process of the retrieval methods. To address this problem, we propose a novel unsupervised cross-modal hashing method robust to the noisy image-text correspondences (CHNR). CHNR consists of three modules: 1) feature extraction module, which extracts feature representations of image-text pairs; 2) noise detection module, which detects potential noisy correspondences; and 3) hashing module that generates cross-modal binary hash codes. The proposed CHNR includes two training phases: i) meta-learning phase that uses a small portion of clean (i.e., reliable) data to train the noise detection module in an adversarial fashion; and ii) the main training phase for which the trained noise detection module is used to identify noisy correspondences while the hashing module is trained on the noisy multi-modal training set. Experimental results show that the proposed CHNR outperforms state-of-the-art methods. Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir |
ICIP | 3 |
| 2022 | A Novel Self-Supervised Cross-Modal Image Retrieval Method in Remote SensingabstractDue to the availability of multi-modal remote sensing (RS) image archives, one of the most important research topics is the development of cross-modal RS image retrieval (CM-RSIR) methods that search semantically similar images across different modalities. Existing CM-RSIR methods require the availability of a high quality and quantity of annotated training images. The collection of a sufficient number of reliable labeled images is time consuming, complex and costly in operational scenarios, and can significantly affect the final accuracy of CM-RSIR. In this paper, we introduce a novel self-supervised CM-RSIR method that aims to: i) model mutual-information between different modalities in a self-supervised manner; ii) retain the distributions of modal-specific feature spaces similar to each other; and iii) define the most similar images within each modality without requiring any annotated training image. To this end, we propose a novel objective including three loss functions that simultaneously: i) maximize mutual information of different modalities for inter-modal similarity preservation; ii) minimize the angular distance of multi-modal image tuples for the elimination of inter-modal discrepancies; and iii) increase cosine similarity of the most similar images within each modality for the characterization of intra-modal similarities. Experimental results show the effectiveness of the proposed method compared to state-of-the-art methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/SS-CM-RSIR. Gencer Sumbul, Begüm Demir |
ICIP | 3 |
| 2022 | Deep Metric Learning-Based Semi-Supervised Regression with Alternate LearningabstractThis paper introduces a novel deep metric learning-based semi-supervised regression (DML-S2R) method for parameter estimation problems. The proposed DML-S2R method aims to mitigate the problems of insufficient amount of labeled samples without collecting any additional sample with a target value. To this end, it is made up of two main steps: i) pairwise similarity modeling with scarce labeled data; and ii) triplet-based metric learning with abundant unlabeled data. The first step aims to model pairwise sample similarities by using a small number of labeled samples. This is achieved by estimating the target value differences of labeled samples with a Siamese neural network (SNN). The second step aims to learn a triplet-based metric space (in which similar samples are close to each other and dissimilar samples are far apart from each other) when the number of labeled samples is insufficient. This is achieved by employing the SNN of the first step for triplet-based deep metric learning that exploits not only labeled samples but also unlabeled samples. For the end-to-end training of DML-S2R, we investigate an alternate learning strategy for the two steps. Due to this strategy, the encoded information in each step becomes a guidance for learning phase of the other step. The experimental results confirm the success of DML-S2R compared to the state-of-the-art semi-supervised regression methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/DML-S2R. Adina Zell, Gencer Sumbul, Begüm Demir |
ICIP | 3 |
| 2022 | Weakly Supervised Semantic Segmentation of Remote Sensing Images for Tree Species Classification Based on Explanation MethodsabstractThe collection of a high number of pixel-based labeled training samples for tree species identification is time consuming and costly in operational forestry applications. To address this problem, in this paper we investigate the effectiveness of explanation methods for deep neural networks in performing weakly supervised semantic segmentation using only image-level labels. Specifically, we consider four methods: i) class activation maps (CAM); ii) gradient-based CAM; iii) pixel correlation module; and iv) self-enhancing maps (SEM). We compare these methods with each other using both quantitative and qualitative measures of their segmentation accuracy, as well as their computational requirements. Experimental results obtained on an aerial image archive show that: i) considered explanation techniques are highly relevant for the identification of tree species with weak supervision; and ii) the SEM outperforms the other considered methods. The code for this paper is publicly available at https://git.tu-berlin.de/rsim/rs_wsss. Steve Ahlswede, Thekke Madam Nimisha, Christian Schulz 0012, Birgit Kleinschmit, Begüm Demir |
IGARSS | 5 |
| 2022 | Causality for Remote Sensing: An Exploratory StudyabstractCausality is one of the most important topics in a Machine Learning (ML) research, and it gives insights beyond the dependency of data points. Causality is a very vital concept also for investigating the dynamic surface of our living planet. However, there are not many attempts for integrating a causal model in Remote Sensing (RS) methodologies. Hence, in this paper, we propose to use patch-based RS images and to represent each patch-based image by a single variable (e.g. entropy). Then we use a Structural Equation Model (SEM) to study their cause-effect relation. Moreover, the SEM is a simple causal model characterized by a Directed Acyclic Graph (DAG). Its nodes are causal variables, and its edges represent causal relationships among causal variables if and only if causal variables are dependent. Soronzonbold Otgonbaatar, Mihai Datcu, Begüm Demir |
IGARSS | 3 |
| 2022 | Coreset of Hyperspectral Images on a Small Quantum ComputerabstractMachine Learning (ML) techniques are employed to analyze and process big Remote Sensing (RS) data, and one well-known ML technique is a Support Vector Machine (SVM). An SVM is a quadratic programming (QP) problem, and a D-Wave quantum annealer (D-Wave QA) promises to solve this QP problem more efficiently than a conventional computer. However, the D-Wave QA cannot solve directly the SVM due to its very few input qubits. Hence, we use a coreset ("core of a dataset") of given EO data for training an SVM on this small D-Wave QA. The coreset is a small, representative weighted subset of an original dataset, and any training models generate competitive classes by using the coreset in contrast to by using its original dataset. We measured the closeness between an original dataset and its coreset by employing a Kullback-Leibler (KL) divergence measure. Moreover, we trained the SVM on the coreset data by using both a D-Wave QA and a conventional method. We conclude that the coreset characterizes the original dataset with very small KL divergence measure. In addition, we present our KL divergence results for demonstrating the closeness between our original data and its coreset. As practical RS data, we use Hyperspectral Image (HSI) of Indian Pine, USA Soronzonbold Otgonbaatar, Mihai Datcu, Begüm Demir |
IGARSS | 3 |
| 2022 | A Novel Framework to Jointly Compress and Index Remote Sensing Images for Efficient Content-Based RetrievalabstractRemote sensing (RS) images are usually stored in compressed format to reduce the storage size of the archives. Thus, existing content-based image retrieval (CBIR) systems in RS require decoding images before applying CBIR (which is computationally demanding in the case of large-scale CBIR problems). To address this problem, in this paper, we present a joint framework that simultaneously learns RS image compression and indexing. Thus, it eliminates the need for decoding RS images before applying CBIR. The proposed framework is made up of two modules. The first module compresses RS images based on an autoencoder architecture. The second module produces hash codes with a high discrimination capability by employing soft pairwise, bit-balancing and classification loss functions. We also introduce a two stage learning strategy with gradient manipulation techniques to obtain image representations that are compatible with both RS image indexing and compression. Experimental results show the efficacy of the proposed framework when compared to widely used approaches in RS. The code of the proposed framework is available at https://git.tu-berlin.de/rsim/RS-JCIF. Gencer Sumbul, Thekke Madam Nimisha, Begüm Demir |
IGARSS | 4 |
| 2022 | Deep Learning Driven Content-Based Image Time-Series Retrieval in Remote Sensing ArchivesabstractThe rapid evolution of satellite imaging systems has resulted in sharp increases of image archive volumes. Multitemporal images constitute a sizeable portion of these time-series databases. Accordingly, development of accurate content based time-series retrieval (CBTSR) methods in massive archives of RS images attracts much research interest. Given a user-defined query time series, CBTSR aims at identifying within a massive archive image time series that show characteristics similar to those of the query time series. In this paper, we focus our attention to CBTSR in pairs of RS images, aiming to search and retrieve bi-temporal image pairs containing changes similar to those modeled in the query. To this end, we introduce two deep learning-based methods in the framework of CBTSR. The first method, called deep change vector retrieval (DVCR), is based on selected deep features extracted from the change vector analysis. The second method, called autoencoder with early fusion (AEEF) uses an autoencoder architecture to recreate the time difference images and the latent codes produced by this network. Experimental results show the effectiveness of the proposed methods for CBTSR problems. The code of the proposed methods is available at: https://github.com/OnatV/ChangeRetrieval. Onat Vuran, Oguzhan Akcin, Mahdyar Ravanbakhsh, Bülent Sankur, Begüm Demir |
IGARSS | 5 |
| 2022 | Satellite Image Search in AgoraEOabstractThe growing operational capability of global Earth Observation (EO) creates new opportunities for data-driven approaches to understand and protect our planet. However, the current use of EO archives is very restricted due to the huge archive sizes and the limited exploration capabilities provided by EO platforms. To address this limitation, we have recently proposed MiLaN, a content-based image retrieval approach for fast similarity search in satellite image archives. MiLaN is a deep hashing network based on metric learning that encodes high-dimensional image features into compact binary hash codes. We use these codes as keys in a hash table to enable real-time nearest neighbor search and highly accurate retrieval. In this demonstration, we showcase the efficiency of MiLaN by integrating it with EarthQube, a browser and search engine within AgoraEO. EarthQube supports interactive visual exploration and Query-by-Example over satellite image repositories. Demo visitors will interact with EarthQube playing the role of different users that search images in a large-scale remote sensing archive by their semantic content and apply other filters. Ahmet Kerem Aksoy, Pavel Dushev, Eleni Tzirita Zacharatou, Holmer Hemsen, Marcela Charfuelan, Jorge-Arnulfo Quiané-Ruiz, Begüm Demir, Volker Markl |
Proc. VLDB Endow. | 7 |
| 2022 | On the Effects of Different Types of Label Noise in Multi-Label Remote Sensing Image ClassificationabstractThe development of accurate methods for multi-label classification (MLC) of remote sensing (RS) images is one of the most important research topics in RS. To address MLC problems, the use of deep neural networks that require a high number of reliable training images annotated by multiple land-cover class labels (multi-labels) has been found popular in RS. However, collecting such annotations is time-consuming and costly. A common procedure to obtain annotations at zero labeling cost is to rely on thematic products or crowdsourced labels. As a drawback, these procedures come with the risk of label noise that can distort the learning process of the MLC algorithms. In the literature, most label noise robust methods are designed for single-label classification (SLC) problems in computer vision (CV), where each image is annotated by a single label. Unlike SLC, label noise in MLC can be associated with: 1) subtractive label noise (a land cover class label is not assigned to an image while that class is present in the image); 2) additive label noise (a land cover class label is assigned to an image, although that class is not present in the given image); and 3) mixed label noise (a combination of both). In this paper, we investigate three different noise robust CV SLC methods (Self-Adaptive Training, Early-Learning Regularization, and Joint Co-Regularized Training) and adapt them to be robust for multi-label noise scenarios in RS. During experiments, we study the effects of different types of multi-label noise and evaluate the adapted methods rigorously. To this end, we also introduce a synthetic multi-label noise injection strategy that is more adequate to simulate operational scenarios compared to the uniform label noise injection strategy, in which the labels of absent and present classes are flipped at uniform probability. Further, we study the relevance of different evaluation metrics in MLC problems under noisy multi-labels. On the basis of the theoretical and experimental analyses, some guidelines for a proper design of label noise robust MLC methods are derived. Tom Burgert, Mahdyar Ravanbakhsh, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Plasticity-Stability Preserving Multi-Task Learning for Remote Sensing Image RetrievalabstractDeep learning-based multi-task learning (MTL) methods have recently attracted attention for content-based image retrieval (CBIR) applications in remote sensing (RS). For a given set of tasks (e.g., scene classification, semantic segmentation, and image reconstruction), existing MTL methods employ a joint optimization algorithm on the direct aggregation of task-specific loss functions. Such an approach may provide limited CBIR performance when: 1) tasks compete or even distract each other; 2) one of the tasks dominates the whole learning procedure; or 3) characterization of each task is underperformed compared to single-task learning. This is mainly due to the lack of: 1) plasticity condition (which is associated with sensitivity to new information) or 2) stability condition (which is associated with protection from radical disruptions by new information) of the whole learning procedure. To avoid this issue, as a first time, we propose a novel plasticity-stability preserving MTL (PLASTA-MTL) approach to ensure the plasticity and the stability conditions of the whole learning procedure independently of the number and type of tasks. This is achieved by defining two novel loss functions. The first loss function is the plasticity preserving loss (PPL) function that aims to enforce the global image representation space to be sensitive to new information learned with each task. This is achieved by minimizing the difference of gradient magnitudes for the global representation and task-specific embedding spaces. The second loss function is the stability preserving loss (SPL) function that aims to protect the global representation space radically disrupted by a new task. This is achieved by minimizing the angular distances between the task gradients over global representation space. To effectively employ the proposed loss functions, we also introduce a novel sequential optimization algorithm. Experimental results show the effectiveness of the proposed approach compared to the state-of-the-art MTL methods in the context of CBIR. Gencer Sumbul, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Informative and Representative Triplet Selection for Multilabel Remote Sensing Image RetrievalabstractLearning the similarity between remote sensing (RS) images forms the foundation for content-based RS image retrieval (CBIR). Recently, deep metric learning approaches that map the semantic similarity of images into an embedding (metric) space have been found very popular in RS. A common approach for learning the metric space relies on the selection of triplets of similar (positive) and dissimilar (negative) images to a reference image called as an anchor. Choosing triplets is a difficult task particularly for multi-label RS CBIR, where each training image is annotated by multiple class labels. To address this problem, in this paper we propose a novel triplet sampling method in the framework of deep neural networks (DNNs) defined for multi-label RS CBIR problems. The proposed method selects a small set of the most representative and informative triplets based on two main steps. In the first step, a set of anchors that are diverse to each other in the embedding space is selected from the current mini-batch using an iterative algorithm. In the second step, different sets of positive and negative images are chosen for each anchor by evaluating the relevancy, hardness and diversity of the images among each other based on a novel strategy. Experimental results obtained on two multi-label benchmark archives show that the selection of the most informative and representative triplets in the context of DNNs results in: i) reducing the computational complexity of the training phase of the DNNs without any significant loss on the performance; and ii) an increase in learning speed since informative triplets allow fast convergence. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/image-retrieval-from-triplets. Gencer Sumbul, Mahdyar Ravanbakhsh, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Towards Simultaneous Image Compression and Indexing for Scalable Content-Based Retrieval in Remote SensingabstractDue to the rapidly growing remote-sensing (RS) image archives, images are usually stored in a compressed format for reducing their storage sizes. Thus, most of the existing content-based RS image retrieval systems require fully decoding images (i.e., decompression) that is computationally demanding for large-scale archives. To address this issue, we introduce a novel approach devoted to simultaneous RS image compression and indexing for scalable content-based image retrieval (denoted as SCI-CBIR). The proposed SCI-CBIR prevents the requirement of decoding RS images before image search and retrieval. To this end, it includes two main steps: 1) deep-learning-based compression and 2) deep-hashing-based indexing. The first step effectively compresses RS images by employing a pair of deep encoder and decoder neural networks and an entropy model. The second step produces hash codes with a high discrimination capability for RS images by employing pairwise, bit-balancing, and classification loss functions. For the training of the SCI-CBIR approach, we also introduce a novel multistage learning procedure with automatic loss weighting techniques to characterize RS image representations that are appropriate for both RS image indexing and compression. The proposed learning procedure enables automatically weighting different loss functions considered for the proposed approach instead of computationally demanding grid search. Experimental results show the effectiveness of the proposed approach when compared to widely used approaches in RS. The code of the proposed approach is available athttps://git.tu-berlin.de/rsim/SCI-CBIR. Gencer Sumbul, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | A Consensual Collaborative Learning Method for Remote Sensing Image Classification Under Noisy Multi-LabelsabstractCollecting a large number of reliable training images annotated by multiple land-cover class labels in the framework of multi-label classification is time-consuming and costly in remote sensing (RS). To address this problem, publicly available thematic products are often used for annotating RS images with zero-labeling-cost. However, such an approach may result in constructing a training set with noisy multi-labels, distorting the learning process. To address this problem, we propose a Consensual Collaborative Multi-Label Learning (CCML) method. The proposed CCML identifies, ranks and corrects training images with noisy multi-labels through four main modules: 1) discrepancy module; 2) group lasso module; 3) flipping module; and 4) swap module. The discrepancy module ensures that the two networks learn diverse features, while obtaining the same predictions. The group lasso module detects the potentially noisy labels by estimating the label uncertainty based on the aggregation of two collaborative networks. The flipping module corrects the identified noisy labels, whereas the swap module exchanges the ranking information between the two networks. The experimental results confirm the success of the proposed CCML under high (synthetically added) multi-label noise rates. The code of the proposed method is publicly available at https://noisy-labels-in-rs.org. Ahmet Kerem Aksoy, Mahdyar Ravanbakhsh, Tristan Kreuziger, Begüm Demir |
ICIP | 4 |
| 2021 | RSVQA Meets Bigearthnet: A New, Large-Scale, Visual Question Answering Dataset for Remote SensingabstractVisual Question Answering is a new task that can facilitate the extraction of information from images through textual queries: it aims at answering an open-ended question formulated in natural language about a given image. In this work, we introduce a new dataset to tackle the task of visual question answering on remote sensing images: this large-scale, open access dataset extracts image/question/answer triplets from the BigEarthNet dataset. This new dataset contains close to 15 millions samples and is openly available. We present the dataset construction procedure, its characteristics and first results using a deep-learning based methodology. These first results show that the task of visual question answering is challenging and opens new interesting research avenues at the interface of remote sensing and natural language processing. The dataset and the code to create and process it are open and freely available on https://rsvqa.sylvainlobry.com/ Sylvain Lobry, Begüm Demir, Devis Tuia |
IGARSS | 2 |
| 2021 | A Novel Graph-Theoretic Deep Representation Learning Method for Multi-Label Remote Sensing Image RetrievalabstractThis paper presents a novel graph-theoretic deep representation learning method in the framework of multi-label remote sensing (RS) image retrieval problems. The proposed method aims to extract and exploit multi-label co-occurrence relationships associated to each RS image in the archive. To this end, each training image is initially represented with a graph structure that provides region-based image representation combining both local information and the related spatial organization. Unlike the other graph-based methods, the proposed method contains a novel learning strategy to train a deep neural network for automatically predicting a graph structure of each RS image in the archive. This strategy employs a region representation learning loss function to characterize the image content based on its multi-label co-occurrence relationship. Experimental results show the effectiveness of the proposed method for retrieval problems in RS compared to state-of-the-art deep representation learning methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/GT-DRL-CBIR. Gencer Sumbul, Begüm Demir |
IGARSS | 2 |
| 2021 | Unsupervised Remote Sensing Image Retrieval Using Probabilistic Latent Semantic HashingabstractUnsupervised hashing methods have attracted considerable attention in large-scale remote sensing (RS) image retrieval, due to their capability for massive data processing with significantly reduced storage and computation. Although existing unsupervised hashing methods are suitable for operational applications, they exhibit limitations when accurately modeling the complex semantic content present in RS images using binary codes (in an unsupervised manner). To address this problem, in this letter, we introduce a novel unsupervised hashing method that takes advantage of the generative nature of probabilistic topic models to encapsulate the hidden semantic patterns of the data into the final binary representation. Specifically, we introduce a new probabilistic latent semantic hashing (pLSH) model to effectively learn the hash codes using three main steps: 1) data grouping, where the input RS archive is clustered into several groups; 2) topic computation, where the pLSH model is used to uncover highly descriptive hidden patterns from each group; and 3) hash code generation, where the data probability distributions are thresholded to generate the final binary codes. Our experimental results, obtained on two benchmark archives, reveal that the proposed method significantly outperforms state-of-the-art unsupervised hashing methods. Rubén Fernández-Beltran, Begüm Demir, Filiberto Pla, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Metric-Learning-Based Deep Hashing Network for Content-Based Retrieval of Remote Sensing ImagesabstractHashing methods have recently been shown to be very effective in the retrieval of remote sensing (RS) images due to their computational efficiency and fast search speed. Common hashing methods in RS are based on hand-crafted features on top of which they learn a hash function, which provides the final binary codes. However, these features are not optimized for the final task (i.e., retrieval using binary codes). On the other hand, modern deep neural networks (DNNs) have shown an impressive success in learning optimized features for a specific task in an end-to-end fashion. Unfortunately, typical RS data sets are composed of only a small number of labeled samples, which make the training (or fine-tuning) of big DNNs problematic and prone to overfitting. To address this problem, in this letter, we introduce a metric-learning-based hashing network, which: 1) implicitly uses a big, pretrained DNN as an intermediate representation step without the need of retraining or fine-tuning; 2) learns a semantic-based metric space where the features are optimized for the target retrieval task; and 3) computes compact binary hash codes for fast search. Experiments carried out on two RS benchmarks highlight that the proposed network significantly improves the retrieval performance under the same retrieval time when compared to the state-of-the-art hashing methods in RS. Subhankar Roy, Enver Sangineto, Begüm Demir, Nicu Sebe |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Remote-Sensing Image Scene Classification With Deep Neural Networks in JPEG 2000 Compressed DomainabstractTo reduce the storage requirements, remote-sensing (RS) images are usually stored in compressed format. Existing scene classification approaches using deep neural networks (DNNs) require to fully decompress the images, which is a computationally demanding task in operational applications. To address this issue, in this article, we propose a novel approach to achieve scene classification in Joint Photographic Experts Group (JPEG) 2000 compressed RS images. The proposed approach consists of two main steps: 1) approximation of the finer resolution subbands of reversible biorthogonal wavelet filters used in JPEG 2000 and 2) characterization of the high-level semantic content of approximated wavelet subbands and scene classification based on the learned descriptors. This is achieved by taking codestreams associated with the coarsest resolution wavelet subband as input to approximate finer resolution subbands using a number of transposed convolutional layers. Then, a series of convolutional layers models the high-level semantic content of the approximated wavelet subband. Thus, the proposed approach models the multiresolution paradigm given in the JPEG 2000 compression algorithm in an end-to-end trainable unified neural network. In the classification stage, the proposed approach takes only the coarsest resolution wavelet subbands as input, thereby reducing the time required to apply decoding. Experimental results performed on two benchmark aerial image archives demonstrate that the proposed approach significantly reduces the computational time with similar classification accuracies when compared with traditional RS scene classification approaches (which requires full image decompression). Akshara Preethy Byju, Gencer Sumbul, Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SD-RSIC: Summarization-Driven Deep Remote Sensing Image CaptioningabstractDeep neural networks (DNNs) have been recently found popular for image captioning problems in remote sensing (RS). Existing DNN-based approaches rely on the availability of a training set made up of a high number of RS images with their captions. However, captions of training images may contain redundant information (they can be repetitive or semantically similar to each other), resulting in information deficiency while learning a mapping from the image domain to the language domain. To overcome this limitation, in this article, we present a novel summarization-driven RS image captioning (SD-RSIC) approach. The proposed approach consists of three main steps. The first step obtains the standard image captions by jointly exploiting convolutional neural networks (CNNs) with long short-term memory (LSTM) networks. The second step, unlike the existing RS image captioning methods, summarizes the ground-truth captions of each training image into a single caption by exploiting sequence to sequence neural networks and eliminates the redundancy present in the training set. The third step automatically defines the adaptive weights associated with each RS image to combine the standard captions with the summarized captions based on the semantic content of the image. This is achieved by a novel adaptive weighting strategy defined in the context of LSTM networks. Experimental results obtained on the RSCID, UCM-Captions, and Sydney-Captions data sets show the effectiveness of the proposed approach compared with the state-of-the-art RS image captioning approaches. The code of the proposed approach is publicly available athttps://gitlab.tubit.tu-berlin.de/rsim/SD-RSIC. Gencer Sumbul, Sonali Nayak, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Learning Convolutional Sparse Coding on Complex Domain for Interferometric Phase RestorationabstractInterferometric phase restoration has been investigated for decades and most of the state-of-the-art methods have achieved promising performances for InSAR phase restoration. These methods generally follow the nonlocal filtering processing chain, aiming at circumventing the staircase effect and preserving the details of phase variations. In this article, we propose an alternative approach for InSAR phase restoration, that is, Complex Convolutional Sparse Coding (ComCSC) and its gradient regularized version. To the best of the authors' knowledge, this is the first time that we solve the InSAR phase restoration problem in a deconvolutional fashion. The proposed methods can not only suppress interferometric phase noise, but also avoid the staircase effect and preserve the details. Furthermore, they provide an insight into the elementary phase components for the interferometric phases. The experimental results on synthetic and realistic high- and medium-resolution data sets from TerraSAR-X StripMap and Sentinel-1 interferometric wide swath mode, respectively, show that our method outperforms those previous state-of-the-art methods based on nonlocal InSAR filters, particularly the state-of-the-art method: InSAR-BM3D. The source code of this article will be made publicly available for reproducible research inside the community. Jian Kang 0005, Danfeng Hong, Jialin Liu 0003, Gerald Baier, Naoto Yokoya, Begüm Demir |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Band-Wise Multi-Scale CNN Architecture for Remote Sensing Image Scene ClassificationabstractMost of the existing convolutional neural network (CNN) architectures in the framework of image scene classification problems are designed for modeling RGB image bands. Direct application of these architectures to the high-dimensional remote sensing (RS) scene classification can be insufficient to accurately describe the spectral content. To address this issue, we propose a novel CNN architecture for the feature embedding of high-dimensional RS images. The proposed architecture aims at: 1) decoupling the spectral and spatial feature extraction for sufficiently describing the complex information content of images; and 2) taking advantage of multi-scale representations of different land-use and land-cover classes present in the images. To this end, the proposed architecture is mainly composed of: 1) a convolutional layer for band-wise extraction of multi-scale spatial features; 2) a convolutional layer for pixel-wise extraction of spectral features; and 3) standard 2D convolution and residual blocks for further feature learning. Experiments on BigEarthNet validate the effectiveness of the proposed method, when compared to the state-of-the-art CNN architectures. Jian Kang 0005, Begüm Demir |
IGARSS | 2 |
| 2020 | S2-CGAN: Self-Supervised Adversarial Representation Learning for Binary Change Detection in Multispectral ImagesabstractDeep Neural Networks have recently demonstrated promising performance in binary change detection (CD) problems in remote sensing (RS), requiring a large amount of labeled multitemporal training samples. Since collecting such data is time-consuming and costly, most of the existing methods rely on pre-trained networks on publicly available computer vision (CV) datasets. However, because of the differences in image characteristics in CV and RS, this approach limits the performance of the existing CD methods. To address this problem, we propose a self-supervised conditional Generative Adversarial Network (S2-cGAN). The proposed S2-cGAN is trained to generate only the distribution of unchanged samples. To this end, the proposed method consists of two main steps: 1) Generating a reconstructed version of the input image as an unchanged image 2) Learning the distribution of unchanged samples through an adversarial game. Unlike the existing GAN based methods (which only use the discriminator during the adversarial training to supervise the generator), the S2-cGAN directly exploits the discriminator likelihood to solve the binary CD task. Experimental results show the effectiveness of the proposed S2-cGAN when compared to the state of the art CD methods. Jose Luis Holgado Alvarez, Mahdyar Ravanbakhsh, Begüm Demir |
IGARSS | 3 |
| 2020 | A Comparative Study of Deep Learning Loss Functions for Multi-Label Remote Sensing Image ClassificationabstractThis paper analyzes and compares different deep learning loss functions in the framework of multi-label remote sensing (RS) image scene classification problems. We consider seven loss functions: 1) cross-entropy loss; 2) focal loss; 3) weighted cross-entropy loss; 4) Hamming loss; 5) Huber loss; 6) ranking loss; and 7) sparseMax loss. All the considered loss functions are analyzed for the first time in RS. After a theoretical analysis, an experimental analysis is carried out to compare the considered loss functions in terms of their: 1) overall accuracy; 2) class imbalance awareness (for which the number of samples associated to each class significantly varies); 3) convexibility and differentiability; and 4) learning efficiency (i.e., convergence speed). On the basis of our analysis, some guidelines are derived for a proper selection of a loss function in multi-label RS scene classification problems. Hichame Yessou, Gencer Sumbul, Begüm Demir |
IGARSS | 3 |
| 2020 | A Progressive Content-Based Image Retrieval in JPEG 2000 Compressed Remote Sensing ArchivesabstractDue to the dramatically increased volume of remote sensing (RS) image archives, images are usually stored in a compressed format to reduce the storage size. Existing content-based RS image retrieval (CBIR) systems require as input fully decoded images, thus resulting in a computationally demanding task in the case of large-scale CBIR problems. To overcome this limitation, in this article, we present a novel CBIR system that achieves a coarse-to-fine progressive RS image description and retrieval in the partially decoded Joint Photographic Experts Group (JPEG) 2000 compressed domain. The proposed system initially: 1) decodes the code blocks associated only to the coarse wavelet resolution and 2) discards the most irrelevant images to the query image based on the similarities computed on the coarse resolution wavelet features of the query and archive images. Then, the code blocks associated with the subsequent resolution of the remaining images are decoded and the most irrelevant images are discarded by computing similarities considering the image features associated with both resolutions. This is achieved by using the pyramid match kernel similarity measure that assigns higher weights to the features associated with the finer wavelet resolution than to those related to the coarse wavelet resolution. These processes are iterated until the codestreams associated with the highest wavelet resolution are decoded. Then, the final retrieval is performed on a very small set of completely decoded images. Experimental results obtained on two benchmark archives of aerial images point out that the proposed system is much faster while providing a similar retrieval accuracy than the standard CBIR systems. Akshara Preethy Byju, Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | GPU-accelerated registration of hyperspectral images using KAZE features
Álvaro Ordóñez, Francisco Argüello, Dora Blanco Heras, Begüm Demir |
J. Supercomput. | 4 |
| 2019 | Retrieving Images with Generated Textual DescriptionsabstractThis paper presents a novel remote sensing (RS) image retrieval system that is defined based on generation and exploitation of textual descriptions that model the content of RS images. The proposed RS image retrieval system is composed of three main steps. The first one generates textual descriptions of the content of the RS images combining a convolutional neural network (CNN) and a recurrent neural network (RNN) to extract the features of the images and to generate the descriptions of their content, respectively. The second step encodes the semantic content of the generated descriptions using word embedding techniques able to produce semantically rich word vectors. The third step retrieves the most similar images with respect to the query image by measuring the similarity between the encoded generated textual descriptions of the query image and those of the archive. Experimental results on RS image archive composed of RS images acquired by unmanned aerial vehicles (UAVs) are reported and discussed. Genc Hoxha, Farid Melgani, Begüm Demir |
IGARSS | 3 |
| 2019 | Weighted Support Vector Machines for Tree Species Classification Using Lidar DataabstractTree species classification at individual tree crowns (ITCs) level using remote sensing data requires the availability of a sufficient number of reliable reference samples. Two main issues that affect the classification performance are: a) an imbalanced distribution of the tree species classes; and b) the presence of unreliable samples due to field collection errors, coordinates misalignments, etc. In this study, we present a weighted Support Vector Machine (WSVM) classifier that addresses these problems by considering: 1) different weights for different classes of tree species to mitigate the effects of the class imbalance distribution; and 2) different weights for different training samples according to their importance for the considered classification problem. Experimental results obtained on a study area located in the Italian Alps showed that the proposed method increased the overall and kappa accuracies of about 2%, and the mean class accuracy of about 10% with respect to a standard SVM. Begüm Demir, Michele Dalponte |
IGARSS | 2 |
| 2019 | Bigearthnet: A Large-Scale Benchmark Archive for Remote Sensing Image UnderstandingabstractThis paper presents the BigEarthNet that is a new large-scale multi-label Sentinel-2 benchmark archive. The BigEarthNet consists of 590, 326 Sentinel-2 image patches, each of which is a section of i) 120 × 120 pixels for 10m bands; ii) 60×60 pixels for 20m bands; and iii) 20×20 pixels for 60m bands. Unlike most of the existing archives, each image patch is annotated by multiple land-cover classes (i.e., multi-labels) that are provided from the CORINE Land Cover database of the year 2018 (CLC 2018). The BigEarthNet is significantly larger than the existing archives in remote sensing (RS) and thus is much more convenient to be used as a training source in the context of deep learning. This paper first addresses the limitations of the existing archives and then describes the properties of the BigEarthNet. Experimental results obtained in the framework of RS image scene classification problems show that a shallow Convolutional Neural Network (CNN) architecture trained on the BigEarthNet provides much higher accuracy compared to a state-of-the-art CNN model pre-trained on the ImageNet (which is a very popular large-scale benchmark archive in computer vision). The BigEarthNet opens up promising directions to advance operational RS applications and research in massive Sentinel-2 image archives. Gencer Sumbul, Marcela Charfuelan, Begüm Demir, Volker Markl |
IGARSS | 3 |
| 2019 | A Novel Multi-Attention Driven System for Multi-Label Remote Sensing Image ClassificationabstractThis paper presents a novel multi-attention driven system that jointly exploits Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) in the context of multi-label remote sensing (RS) image classification. The proposed system consists of four main modules. The first module aims to extract preliminary local descriptors of RS image bands that can be associated to different spatial resolutions. To this end, we introduce a K-Branch CNN, in which each branch extracts descriptors of image bands that have the same spatial resolution. The second module aims to model spatial relationship among local descriptors. This is achieved by a bidirectional RNN architecture, in which Long Short-Term Memory nodes enrich local descriptors by considering spatial relationships of local areas (image patches). The third module aims to define multiple attention scores for local descriptors. This is achieved by a novel patch-based multi-attention mechanism that takes into account the joint occurrence of multiple land-cover classes and provides the attention-based local descriptors. The last module exploits these descriptors for multi-label RS image classification. Experimental results obtained on the BigEarth-Net that is a large-scale Sentinel-2 benchmark archive show the effectiveness of the proposed method compared to a state of the art method. Gencer Sumbul, Begüm Demir |
IGARSS | 2 |
| 2019 | An Unsupervised Multicode Hashing Method for Accurate and Scalable Remote Sensing Image RetrievalabstractHashing methods have recently attracted great attention for approximate nearest neighbor search in massive remote sensing (RS) image archives due to their computational and storage effectiveness. The existing hashing methods in RS represent each image with a single-hash code that is usually obtained by applying hash functions to global image representations. Such an approach may not optimally represent the complex information content of RS images. To overcome this problem, in this letter, we present a simple yet effective unsupervised method that represents each image with primitive-cluster sensitive multi-hash codes (each of which corresponds to a primitive present in the image). To this end, the proposed method consists of two main steps: 1) characterization of images by descriptors of primitive-sensitive clusters and 2) definition of multi-hash codes from the descriptors of the primitive-sensitive clusters. After obtaining multi-hash codes for each image, retrieval of images is achieved based on a multi-hash-code-matching scheme. Any hashing method that provides single-hash code can be embedded within the proposed method to provide primitive-sensitive multi-hash codes. Compared with state-of-the-art single-code hashing methods in RS, the proposed method achieves higher retrieval accuracy under the same retrieval time, and thus it is more efficient for operational applications. Thomas Reato, Begüm Demir, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Extended attribute profiles on GPU applied to hyperspectral image classification
Pedro G. Bascoy, Pablo Quesada-Barriuso, Dora Blanco Heras, Francisco Argüello, Begüm Demir, Lorenzo Bruzzone |
J. Supercomput. | 5 |
| 2018 | From Big Data to Big Information and Big Knowledge: the Case of Earth Observation DataabstractSome particularly important rich sources of open and free big geospatial data are the Earth observation (EO) programs of various countries such as the Landsat program of the US and the Copernicus programme of the European Union. EO data is a paradigmatic case of big data and the same is true for the big information and big knowledge extracted from it. EO data (satellite images and in-situ data), and the information and knowledge extracted from it, can be utilized in many applications with financial and environmental impact in areas such as emergency management, climate change, agriculture and security. Konstantina Bereta, Manolis Koubarakis, Stefan Manegold, George Stamoulis 0001, Begüm Demir |
CIKM | 5 |
| 2018 | Semantic-Fusion Gans for Semi-Supervised Satellite Image ClassificationabstractMost of the public satellite image datasets contain only a small number of annotated images. The lack of a sufficient quantity of labeled data for training is a bottleneck for the use of modern deep-learning based classification approaches in this domain. In this paper we propose a semi -supervised approach to deal with this problem. We use the discriminator (D) of a Generative Adversarial Network (GAN) as the final classifier, and we train D using both labeled and unlabeled data. The main novelty we introduce is the representation of the visual information fed to D by means of two different channels: the original image and its “semantic” representation, the latter being obtained by means of an external network trained on ImageNet. The two channels are fused in D and jointly used to classify fake images, real labeled and real unlabeled images. We show that using only 100 labeled images, the proposed approach achieves an accuracy close to 69% and a significant improvement with respect to other GAN-based semi-supervised methods. Although we have tested our approach only on satellite images, we do not use any domain-specific knowledge. Thus, our method can be applied to other semi-supervised domains. Subhankar Roy, Enver Sangineto, Nicu Sebe, Begüm Demir |
ICIP | 4 |
| 2018 | Integration of Remote Sensing with A Hydroclimatological Model for an Improved Monitoring of Alpine GlaciersabstractIn this work, we present a framework to integrate physically based hydroclimatological models and remote sensing products, by exploiting their advantages and overcoming their limitations, for an improved understanding and estimation of the alpine glacier accumulation and ablation processes. The capability of remote sensing to well represent the spatial variability of the snow cover over the glaciers is used to correct possible errors in the model simulations, thus obtaining a more reliable estimation of annual glacier mass balance. The proposed approach is tested on the glaciers in the Rofen Valley (Austria) by employing the AMUNDSEN model for accumulation and ablation processes simulation and Landsat-5/7/8 data for glacier zone mapping from 1998 to 2016. Mattia Callegari, Carlo Marin, Daniel Günther 0001, Philipp Rastner, Lorenzo Bruzzone, Begüm Demir, Thomas Marke, Ulrich Strasser, Marc Zebisch, Claudia Notarnicola |
IGARSS | 6 |
| 2018 | A Novel Data Fusion Technique for Snow Parameter RetrievalabstractThe main idea of this study is the development of an innovative data fusion method through which state-of-the-art remotely sensed products and hydrological modelling simulations can be integrated to improve the retrieval and the reliability of snow cover and snow water equivalent mapping. The proposed method is based on a machine learning technique, Support Vector Machine (SVM), and on exploitation of two well-instrumented test-sites in EUREGIO region for the validation. Results show an improvement of performances with respect to single products from remote sensing and model. On EUREGIO scale the accuracy of snow cover mapping obtained from fusion reaches 0.95. Ludovica De Gregorio, Mattia Callegari, Carlo Marin, Marc Zebisch, Lorenzo Bruzzone, Begüm Demir, Ulrich Strasser, Daniel Günther 0001, Thomas Marke, Claudia Notarnicola |
IGARSS | 6 |
| 2018 | Deep Metric and Hash-Code Learning for Content-Based Retrieval of Remote Sensing ImagesabstractThe growing volume of Remote Sensing (RS) image archives demands for feature learning techniques and hashing functions which can: (1) accurately represent the semantics in the RS images; and (2) have quasi real-time performance during retrieval. This paper aims to address both challenges at the same time, by learning a semantic-based metric space for content based RS image retrieval while simultaneously producing binary hash codes for an efficient archive search. This double goal is achieved by training a deep network using a combination of different loss functions which, on the one hand, aim at clustering semantically similar samples (i.e., images), and, on the other hand, encourage the network to produce final activation values (i.e., descriptors) that can be easily binarized. Moreover, since RS annotated training images are too few to train a deep network from scratch, we propose to split the image representation problem in two different phases. In the first we use a general-purpose, pre-trained network to produce an intermediate representation, and in the second we train our hashing network using a relatively small set of training images. Experiments on two aerial benchmark archives show that the proposed method outperforms previous state-of-the-art hashing approaches by up to 5.4% using the same number of hash bits per image. Subhankar Roy, Enver Sangineto, Begüm Demir, Nicu Sebe |
IGARSS | 3 |
| 2018 | Advanced Local Binary Patterns for Remote Sensing Image RetrievalabstractThe standard Local Binary Pattern (LBP) is considered among the most computationally efficient remote sensing (RS) image descriptors in the framework of large-scale content based RS image retrieval (CBIR). However, it has limited discrimination capability for characterizing high dimensional RS images with complex semantic content. There are several LBP variants introduced in computer vision that can be extended to RS CBIR to efficiently overcome the above-mentioned problem. To this end, this paper presents a comparative study in order to analyze and compare advanced LBP variants in RS CBIR domain. We initially introduce a categorization of the LBP variants based on the specific CBIR problems in RS, and analyze the most recent methodological developments associated to each category. All the considered LBP variants are introduced for the first time in the framework of RS image retrieval problems, and have been experimentally compared in terms of their: 1) discrimination capability to model high-level semantic information present in RS images (and thus the retrieval performance); and 2) computational complexities associated to retrieval and feature extraction time. Issayas Tekeste, Begüm Demir |
IGARSS | 2 |
| 2018 | Multilabel Remote Sensing Image Retrieval Using a Semisupervised Graph-Theoretic MethodabstractConventional supervised content-based remote sensing (RS) image retrieval systems require a large number of already annotated images to train a classifier for obtaining high retrieval accuracy. Most systems assume that each training image is annotated by a single label associated to the most significant semantic content of the image. However, this assumption does not fit well with the complexity of RS images, where an image might have multiple land-cover classes (i.e., multilabels). Moreover, annotating images with multilabels is costly and time consuming. To address these issues, in this paper, we introduce a semisupervised graph-theoretic method in the framework of multilabel RS image retrieval problems. The proposed method is based on four main steps. The first step segments each image in the archive and extracts the features of each region. The second step constructs an image neighborhood graph and uses a correlated label propagation algorithm to automatically assign a set of labels to each image in the archive by exploiting only a small number of training images annotated with multilabels. The third step associates class labels with image regions by a novel region labeling strategy, whereas the final step retrieves the images similar to a given query image by a subgraph matching strategy. Experiments carried out on an archive of aerial images show the effectiveness of the proposed method when compared with the state-of-the-art RS content-based image retrieval methods. Bindita Chaudhuri, Begüm Demir, Subhasis Chaudhuri, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Multiple Kernel Learning for Remote Sensing Image ClassificationabstractThis paper presents multiple kernel learning (MKL) in the context of remote sensing (RS) image classification problems by illustrating main characteristics of different MKL algorithms and analyzing their properties in RS domain. A categorization of different MKL algorithms is initially introduced, and some promising MKL algorithms for each category are presented. In particular, MKL algorithms presented only in machine learning are introduced in RS. Then, the investigated MKL algorithms are theoretically compared in terms of their: 1) computational complexities; 2) accuracy with different qualities of kernels; and 3) accuracy with different numbers of kernels. After the theoretical comparison, experimental analyses are carried out to compare different MKL algorithms in terms of: 1) model selection and 2) feature fusion problems. On the basis of the theoretical and experimental analyses of MKL algorithms, some guidelines for a proper selection of the MKL algorithms are derived. Saeid Niazmardi, Begüm Demir, Lorenzo Bruzzone, Abdolreza Safari, Saeid Homayouni |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | A novel system for content based retrieval of multi-label remote sensing imagesabstractThis paper presents a novel content based remote sensing (RS) image retrieval system that consists of: i) a spatial and spectral image description scheme; and ii) a sparsity based supervised retrieval method. Spatial image description is based on the scale invariant feature transform (SIFT), while a novel descriptor defined based on the bag of spectral values is proposed to express spectral features. With the conjunction of these two feature vectors RS image retrieval is instrumented via a sparse reconstruction-based approach. These sparse reconstructions are used to estimate the likelihood of a scene to contain a land-cover class label. Applying this method separately for each land-cover class, one achieves retrieval in the framework of multi-label remote sensing image retrieval. Experimental results obtained on an archive of hyperspectral images show the effectiveness of the proposed system. Osman Emre Dai, Begüm Demir, Bülent Sankur, Lorenzo Bruzzone |
IGARSS | 2 |
| 2017 | Primitive cluster sensitive hashing for scalable content-based image retrieval in remote sensing archivesabstractThis paper proposes a novel unsupervised method based on primitive cluster sensitive hashing for fast and accurate image retrieval in large remote sensing (RS) archives. The proposed method consists of a three-steps algorithm. In the first step, each image in the archive is characterized by primitive clusters' descriptors. These descriptors are obtained through an unsupervised approach, which automatically extracts the image regions' descriptors and then associates them with primitive clusters. In the second step the primitive clusters' descriptors are transformed into multi-hash codes to represent each image. Then, in the last step, a multi-hash-code-matching scheme is applied to retrieve the images in the archive that are very similar to a query image. Experiments carried out on an archive of aerial images show that the proposed method provides distinctive multi-hash codes associated to the primitive clusters. Thus, it is more accurate than standard hashing methods, particularly under complex RS image retrieval tasks. Thomas Reato, Begüm Demir, Lorenzo Bruzzone |
IGARSS | 2 |
| 2016 | Quad-tree based compressed histogram attribute profiles for classification of very high resolution imagesabstractThis paper presents a novel quad-tree based compressed histogram attribute profile (QT-CHAP) for classification of very high resolution remote sensing images. The QT-CHAP characterizes the marginal local distribution of attribute filter responses to model the spatial context of each sample with a very small number of image features. This is achieved based on a three steps algorithm that comprises a novel non-uniform quantization strategy to the compression of the information present in standard histogram attribute profiles. Due to the proposed quad-tree based non-uniform quantization strategy, the proposed QT-CHAP results in an optimized tradeoff between information extraction and number of considered features. Experimental results confirm the effectiveness of the proposed QT-CHAP in terms of computational complexity, storage requirements and classification accuracy when compared to the other state of the art attribute profile based methods. Romano Battiti, Begüm Demir, Lorenzo Bruzzone |
IGARSS | 2 |
| 2016 | A comparative study on Multiple Kernel Learning for remote sensing image classificationabstractThis paper analyzes and compares different Multiple Kernel Learning (MKL) algorithms for the classification of remote sensing (RS) images. The main purpose of the comparison is to identify advantages and disadvantages of different MKL algorithms in terms of their computational time and classification accuracy. Furthermore, some guidelines on the proper selection of the MKL algorithms associated with different RS image classification problems are derived. Saeid Niazmardi, Begüm Demir, Lorenzo Bruzzone, Abdolreza Safari, Saeid Homayouni |
IGARSS | 2 |
| 2016 | Region-Based Retrieval of Remote Sensing Images Using an Unsupervised Graph-Theoretic ApproachabstractThis letter introduces a novel unsupervised graph-theoretic approach in the framework of region-based retrieval of remote sensing (RS) images. The proposed approach is characterized by two main steps: (1) modeling each image by a graph, which provides region-based image representation combining both local information and related spatial organization, and (2) retrieving the images in the archive that are most similar to the query image by evaluating graph-based similarities. In the first step, each image is initially segmented into distinct regions and then modeled by an attributed relational graph, where nodes and edges represent region characteristics and their spatial relationships, respectively. In the second step, a novel inexact graph matching strategy, which jointly exploits a subgraph isomorphism algorithm and a spectral graph embedding technique, is applied to match corresponding graphs and to retrieve images in the order of graph similarity. Experiments carried out on an archive of aerial images point out that the proposed approach significantly improves the retrieval performance compared to the state-of-the-art unsupervised RS image retrieval methods.(RS) images. Bindita Chaudhuri, Begüm Demir, Lorenzo Bruzzone, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | A Novel Hybrid Method for the Correction of the Theoretical Model Inversion in Bio/Geophysical Parameter EstimationabstractThis paper presents a novel hybrid method to the estimation of bio/geophysical parameters, which models and corrects deviations from correct target values when theoretical electromagnetic models are used for the inversion process. The proposed hybrid method integrates theoretical models with empirical observations associated to a few field reference samples. This is achieved based on two steps. In the first step, deviations between estimations obtained by a theoretical model and empirical observations are initially computed. Then, deviations associated to unlabeled samples (for which reference measures are not existing) are characterized based on two different strategies: 1) the global deviation bias strategy (which assumes that the deviations of samples are constant within the input space); and 2) the local deviation bias strategy (which assumes that the deviations of samples are variable within different portions of the input space). In the second step, the theoretical model estimates of unlabeled samples are corrected based on the estimated deviations. The experimental analysis carried out in the context of soil moisture content retrieval from microwave remotely sensed data confirms the effectiveness of the proposed hybrid estimation method. Davide Castelletti, Luca Pasolli, Lorenzo Bruzzone, Claudia Notarnicola, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Hashing-Based Scalable Remote Sensing Image Search and Retrieval in Large ArchivesabstractLarge-scale remote sensing (RS) image search and retrieval have recently attracted great attention, due to the rapid evolution of satellite systems, that results in a sharp growing of image archives. An exhaustive search through linear scan from such archives is time demanding and not scalable in operational applications. To overcome such a problem, this paper introduces hashing-based approximate nearest neighbor search for fast and accurate image search and retrieval in large RS data archives. The hashing aims at mapping high-dimensional image feature vectors into compact binary hash codes, which are indexed into a hash table that enables real-time search and accurate retrieval. Such binary hash codes can also significantly reduce the amount of memory required for storing the RS images in the auxiliary archives. In particular, in this paper, we introduce in RS two kernel-based nonlinear hashing methods. The first hashing method defines hash functions in the kernel space by using only unlabeled images, while the second method leverages on the semantic similarity extracted by annotated images to describe much distinctive hash functions in the kernel space. The effectiveness of considered hashing methods is analyzed in terms of RS image retrieval accuracy and retrieval time. Experiments carried out on an archive of aerial images point out that the presented hashing methods are much faster, while keeping a similar (or even higher) retrieval accuracy, than those typically used in RS, which exploit an exact nearest neighbor search. Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Histogram-Based Attribute Profiles for Classification of Very High Resolution Remote Sensing ImagesabstractMorphological attribute profiles (APs) obtained by the sequential application of morphological attribute filters to images have been found very effective in remote sensing (RS) to characterize spatial properties of objects in a scene. However, a direct use of the APs can be insufficient to provide a complete characterization of spatial information when complex texture is present in the considered images. To overcome this problem, in this paper, we present the novel histogram-based morphological APs (HAPs). The HAPs model the marginal local distribution of attribute filter responses to better characterize the texture information, and they are obtained based on a three-step algorithm. In the first step, the standard APs are constructed by sequentially applying attribute filters to the considered image. In the second step, a local histogram is calculated for each sample of each image in the APs. Then, in the final step, the local histograms of the same pixel locations in the APs are stacked, resulting in a texture descriptor whose components represent local distributions of the filter responses for the related pattern. Finally, the very-high-dimensional HAPs are classified by a support vector machine (SVM) classifier with histogram intersection kernel. Experimental results obtained by considering two very high resolution panchromatic images show the effectiveness of the proposed HAPs, which sharply improve the accuracy of the SVM classifier with respect to standard AP-based methods. Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | A cluster-based appraoch to content based time series retrieval (CBTSR)abstractGiven a user-defined image time series (i.e., the query time series), content based time series retrieval (CBTSR) is the process of identifying other time series that show properties similar to the query. When dealing with time series, the elements of the content based retrieval process require to be redefined in order to take into account the time variable. In this perspective, the design of the query, the feature extraction, and retrieval itself have to be reformulated. Here we focus our attention to CBTSR in pairs of images. The goal is to identify bi-temporal images showing a specific kind of change (associated with changes on the ground) modeled by the query. Attention is devoted to the design of the auxiliary archive modeling the change information and on the retrieval algorithm. Experiments on an archive of Landsat images confirmed the effectiveness of the proposed approach. Francesca Bovolo, Begüm Demir, Lorenzo Bruzzone |
IGARSS | 2 |
| 2015 | Fast and accurate image classification with histogram based features and additive kernel SVMabstractKernel-based image classification methods rely on the considered kernel functions that can be chosen with respect to prior information on the adopted features. In remote sensing, histogram features have recently gained an increasing interest due to their capability to address several critical classification problems (e.g., the problem of curse of dimensionality) when appropriate kernels and classifiers are selected. In view of that, in this paper we introduce in remote sensing additive kernels in the context of support vector machine classification (AK-SVM), which are suitable kernels for histogram based feature representations. In particular, we investigate the Histogram Intersection kernel and the chi-square kernel within the AK-SVM. Moreover, we present fast implementations of the AK-SVM to significantly speed up the classification phase of the SVM. Experimental results show the effectiveness of the AK-SVM in terms of classification accuracy and computational time when compared to SVMs with standard kernels. Begüm Demir, Lorenzo Bruzzone |
IGARSS | 1 |
| 2015 | Histogram based attribute profiles for classification of very high resolution remote sensing imagesabstractThis paper presents a novel histogram based attribute profiles (HAPs) technique for classification of very high resolution remote sensing images. The HAPs characterize the marginal local distribution of attribute filter responses to model the texture information. This is achieved based on a two steps algorithm. In the first step the standard attribute profiles (AP) are built through sequential application of attribute filters to the considered image. In the second step a local histogram is initially computed for each sample of each image in the APs. Then the local histograms of the same pixel locations in the APs are concatenated. Accordingly, each sample is characterized by a texture descriptor whose components model local distributions of the filter responses. Finally the very high dimensional HAPs are classified by a Support Vector Machine classifier with histogram intersection kernel, which is very effective for high dimensional histogram-based feature representations. Experimental results confirm the effectiveness of the proposed HAPs with respect to standard APs. Begüm Demir, Lorenzo Bruzzone |
IGARSS | 1 |
| 2015 | An analysis of the capabilities of COSMO-SKYMED and RADARSAT systems for agricultural area monitoringabstractThis research aims at analyzing the integration of C and X band data collected from Radarsat2 (RS2) and COSMO-SkyMed (CSK) systems on some test areas in Italy, in order to estimate the main geophysical parameters of soil and vegetation, such as soil moisture and vegetation biomass. A check of the sensitivity of SAR signal to the soil parameters was first carried out on both test sites. Over the South-Tyrol area a retrieval approach based on the Support Vector Regression methodology, which was already tested in this area using C-band data from ENVISAT/ASAR data, was carried out. From these preliminary results it can be concluded that X-band images combined with C-band images could provide valuable information for the retrieval of SMC, even though further investigations should be carried out on a larger time-series and larger set of samples. Simonetta Paloscia, Simone Pettinato, Emanuele Santi, Claudia Notarnicola, Felix Greifeneder, Giovanni Cuozzo, Irene Nicolini, Begüm Demir, Lorenzo Bruzzone |
IGARSS | 8 |
| 2015 | A Novel Active Learning Method in Relevance Feedback for Content-Based Remote Sensing Image RetrievalabstractConventional relevance feedback (RF) schemes improve the performance of content-based image retrieval (CBIR) requiring the user to annotate a large number of images. To reduce the labeling effort of the user, this paper presents a novel active learning (AL) method to drive RF for retrieving remote sensing images from large archives in the framework of the support vector machine classifier. The proposed AL method is specifically designed for CBIR and defines an effective and as small as possible set of relevant and irrelevant images with regard to a general query image by jointly evaluating three criteria: uncertainty; diversity; and density of images in the archive. The uncertainty and diversity criteria aim at selecting the most informative images in the archive, whereas the density criterion goal is to choose the images that are representative of the underlying distribution of data in the archive. The proposed AL method assesses jointly the three criteria based on two successive steps. In the first step, the most uncertain (i.e., ambiguous) images are selected from the archive on the basis of the margin sampling strategy. In the second step, the images that are both diverse (i.e., distant) to each other and associated to the high-density regions of the image feature space in the archive are chosen from the most uncertain images. This step is achieved by a novel clustering-based strategy. The proposed AL method for driving the RF contributes to mitigate problems of unbalanced and biased set of relevant and irrelevant images. Experimental results show the effectiveness of the proposed AL method. Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Kernel-based hashing for content-based image retrval in large remote sensing data archiveabstractThis paper presents hashing based approximate nearest neighbor search algorithms that allow fast and accurate image retrieval in huge remote sensing data archives. Hashing methods aim at mapping high-dimensional image feature vectors into short binary codes based on hashing functions. Then, the image retrieval is accomplished according to Hamming distances of image hash codes. In particular, in this paper two hashing methods are adopted for RS image retrieval problems. The former aims at defining hash functions in the kernel space by using only unlabeled images. The latter leverages on the semantic similarity given in terms of annotated images to define much distinctive hash functions in the kernel space. The effectiveness of both methods is analyzed in terms of RS image retrieval accuracy as well as retrieval time. Experiments carried out on an archive of aerial images show that the presented hashing methods are one hundred times faster than those that exploit an exact nearest neighbor search while keeping a high retrieval accuracy. Begüm Demir, Lorenzo Bruzzone |
IGARSS | 1 |
| 2014 | An Effective Strategy to Reduce the Labeling Cost in the Definition of Training Sets by Active LearningabstractThis letter proposes a novel strategy for reducing the cost of in situ sample labeling for the definition of training sets by active learning (AL) in the framework of supervised classification of remote sensing images. AL methods define a training set according to an iterative procedure that at each iteration requires the labeling of a set of new samples selected by the classifier. The proposed strategy can be embedded in any AL method to identify the most informative area on the ground where focusing each AL iteration to reduce the overall cost (in terms of time) of labeling. To this end, at each iteration, the most uncertain unlabeled samples are initially identified. Then, the area on the ground (having a size predefined by the user) that has the highest spatial density of informative (i.e., uncertain and diverse) unlabeled samples is selected by the proposed strategy, and the AL technique is applied only to the samples of that area. This results in a decrease of the overall labeling cost with respect to that required by the use of a given technique in a standard way. Experimental results obtained by embedding the presented strategy in different literature active learning methods confirm its effectiveness. Begüm Demir, Luca Minello, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | A multiple criteria active learning method for support vector regression
Begüm Demir, Lorenzo Bruzzone |
Pattern Recognit. | 1 |
| 2014 | Definition of Effective Training Sets for Supervised Classification of Remote Sensing Images by a Novel Cost-Sensitive Active Learning MethodabstractThis paper proposes a novel cost-sensitive active learning (CSAL) method to the definition of reliable training sets for the classification of remote sensing images with support vector machines. Unlike standard active learning (AL) methods, the proposed CSAL method redefines AL by assuming that the labeling cost of samples during ground survey is not identical, but depends on both the samples accessibility and the traveling time to the considered locations. The proposed CSAL method selects the most informative samples on the basis of three criteria: 1) uncertainty; 2) diversity; and 3) labeling cost. The labeling cost of the samples is modeled by a novel cost function that exploits ancillary data such as the road network map and the digital elevation model of the considered area. In the proposed method, the three criteria are applied in two consecutive steps. In the first step, the most uncertain samples are selected, whereas in the second step the uncertain samples that are diverse and have low labeling cost are chosen. In order to select the uncertain samples that optimize the diversity and cost criteria, we propose two different optimization algorithms. The first algorithm is defined on the basis of a sequential forward selection optimization strategy, whereas the second one relies on a genetic algorithm. Experimental results show the effectiveness of the proposed CSAL method compared to standard AL methods that neglect the labeling cost. Begüm Demir, Luca Minello, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | An effective active learning method for interactive content-based retrieval in remote sensing imagesabstractThis paper presents a novel active learning (AL) technique to drive relevance feedback in content based image retrieval (CBIR) from earth observation data archives. The proposed AL method aims at defining an effective set of relevant and irrelevant images with respect to the query image as small as possible. This is achieved on the basis of a joint evaluation of three criteria: i) uncertainty, ii) diversity and iii) density of images. The uncertainty and diversity criteria aims at choosing the most informative images in the archive, whereas the density criterion aims at selecting those that are representative of the underlying distribution of images in the archive. In the proposed AL method, the three criteria are applied in two consecutive steps. In the first step the most uncertain images are selected based on well-known margin sampling strategy. In the second step the images that are associated to high density regions in the archive and are diverse (i.e., distant) to each other are chosen from the most uncertain ones on the basis of a novel clustering based strategy. Experimental results show the effectiveness of the proposed AL method, particularly when a poor initial set of relevant and irrelevant images is available. Begüm Demir, Lorenzo Bruzzone |
IGARSS | 1 |
| 2013 | Sequential cascade classification of image time series by exploiting multiple pairwise change detectionabstractThis paper presents a novel sequential cascade classification technique for automatically updating land-cover maps by classifying remote sensing image time series. We assume that a reliable training set is initially available only for one of the images (i.e., the source domain) in the time series, whereas it is not for an image being classified (i.e., the target domain). Unlike the standard cascade classification method, the proposed method aims at exploiting all the images in the time series acquired between the target and source domains to effectively classify the target domain. To this end, initially `pseudo' training sets of the images are defined by a multiple pairwise change detection based transfer learning strategy. Then, the target domain is classified by the proposed sequential cascade classification method, exploiting the temporal correlation between images. Experimental results obtained on a time series of Landsat multispectral images show the effectiveness of the proposed technique with respect to the standard cascade classification. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IGARSS | 1 |
| 2013 | Updating Land-Cover Maps by Classification of Image Time Series: A Novel Change-Detection-Driven Transfer Learning ApproachabstractThis paper proposes a novel change-detection-driven transfer learning (TL) approach to update land-cover maps by classifying remote-sensing images acquired on the same area at different times (i.e., image time series). The proposed approach requires that a reliable training set is available only for one of the images (i.e., the source domain) in the time series whereas it is not for another image to be classified (i.e., the target domain). Unlike other literature TL methods, no additional assumptions on either the similarity between class distributions or the presence of the same set of land-cover classes in the two domains are required. The proposed method aims at defining a reliable training set for the target domain, taking advantage of the already available knowledge on the source domain. This is done by applying an unsupervised-change-detection method to target and source domains and transferring class labels of detected unchanged training samples from the source to the target domain to initialize the target-domain training set. The training set is then optimized by a properly defined novel active learning (AL) procedure. At the early iterations of AL, priority in labeling is given to samples detected as being changed, whereas in the remaining ones, the most informative samples are selected from changed and unchanged unlabeled samples. Finally, the target image is classified. Experimental results show that transferring the class labels from the source domain to the target domain provides a reliable initial training set and that the priority rule for AL results in a fast convergence to the desired accuracy with respect to Standard AL. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Classification of Time Series of Multispectral Images With Limited Training DataabstractImage classification usually requires the availability of reliable reference data collected for the considered image to train supervised classifiers. Unfortunately when time series of images are considered, this is seldom possible because of the costs associated with reference data collection. In most of the applications it is realistic to have reference data available for one or few images of a time series acquired on the area of interest. In this paper, we present a novel system for automatically classifying image time series that takes advantage of image(s) with an associated reference information (i.e., the source domain) to classify image(s) for which reference information is not available (i.e., the target domain). The proposed system exploits the already available knowledge on the source domain and, when possible, integrates it with a minimum amount of new labeled data for the target domain. In addition, it is able to handle possible significant differences between statistical distributions of the source and target domains. Here, the method is presented in the context of classification of remote sensing image time series, where ground reference data collection is a highly critical and demanding task. Experimental results show the effectiveness of the proposed technique. The method can work on multimodal (e.g., multispectral) images. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Trans. Image Process. | 1 |
| 2012 | A novel system for classification of image time series with limited ground reference dataabstractThis paper presents a novel system for automatically updating land-cover maps by classifying remote sensing image time series. The proposed system assumes that a reliable training set is available only for one of the images (i.e., the source domain) in the time series, whereas it is not for another image to be classified (i.e., the target domain). To effectively classify the target domain the proposed system includes two steps: i) low-cost definition of the training set for the target domain; and ii) target domain classification according to the Bayesian cascade decision rule that exploits the temporal correlation between domains. In the proposed system, the low cost training set for the target domain is defined on the basis of transfer and active learning methods, which also use the temporal dependence information between the domains. Experimental results obtained on a time series of Landsat multispectral images show the effectiveness of the proposed technique. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IGARSS | 1 |
| 2012 | A cost-sensitive active learning technique for the definition of effective training sets for supervised classifiersabstractThis paper presents a novel cost-sensitive active learning technique (CSAL) to define effective training sets for the classification of remote sensing images. Unlike the standard active learning methods, the proposed technique redefines AL by assuming that the labeling cost of samples when ground survey is used is not uniform and depends both on the samples accessibility and the traveling time to the considered locations. Accordingly, the proposed CSAL technique is based on the joint evaluation of three criteria for the selection of the most informative samples that have a low labeling cost: i) uncertainty, ii) diversity and iii) cost efficiency. The labeling cost of the samples is assessed by using ancillary data like the road map and the digital elevation model of the considered area. Experimental results show the effectiveness of the proposed CSAL method compared to the standard active learning methods that neglect the labeling cost. Begüm Demir, Luca Minello, Lorenzo Bruzzone |
IGARSS | 1 |
| 2012 | Detection of Land-Cover Transitions in Multitemporal Remote Sensing Images With Active-Learning-Based Compound ClassificationabstractThis paper presents a novel iterative active learning (AL) technique aimed at defining effective multitemporal training sets to be used for the supervised detection of land-cover transitions in a pair of remote sensing images acquired on the same area at different times. The proposed AL technique is developed in the framework of the Bayes' rule for compound classification. At each iteration, it selects the pair of spatially aligned unlabeled pixels in the two images that are classified with the maximum uncertainty. These pixels are then labeled by an external supervisor and included in the training set. The uncertainty of a pair of pixels is assessed by the joint entropy defined by considering two possible different simplifying assumptions: 1) class-conditional independence and 2) temporal independence between multitemporal images. Accordingly, different algorithms are introduced. The proposed joint-entropy-based AL algorithms for compound classification are compared with each other and with a marginal-entropy-based AL technique (in which the entropy is computed separately on single-date images) applied to the postclassification comparison method. The experimental results obtained on two multispectral and multitemporal data sets show the effectiveness of the proposed technique. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | Detection of land-cover transitions in multitemporal images with active-learning based compound classificationabstractThis paper presents a novel active learning (AL) technique for the compound classification of multitemporal remote-sensing images for the detection of land-cover transitions. The proposed AL technique is based on the selection of unlabeled pairs of samples that have maximum uncertainty on their labels assigned by a classifier implemented according to the Bayes rule for compound classification. Uncertainty of a pair of samples is assessed by joint entropy defined on the basis of two different simplifying assumptions: i) class-conditional independence, and ii) temporal independence between multitemporal images. Accordingly, two algorithms for the proposed joint entropy based AL technique are introduced. The proposed joint entropy based AL algorithms are compared to each other and with a marginal entropy (entropy computed separately on single-date images) based AL technique. Experimental results obtained on two multispectral images show the effectiveness of the proposed technique. Begüm Demir, Francesca Bovolo, Lorenzo Bruzzone |
IGARSS | 1 |
| 2011 | Hyperspectral Image Classification Using Denoising of Intrinsic Mode FunctionsabstractThis letter proposes the use of denoising in conjunction with 2-D empirical mode decomposition (2D-EMD) of hyperspectral image bands for higher classification accuracy. Initially, 2D-EMD is performed to hyperspectral image bands for decomposition into intrinsic mode functions (IMFs). Then, denoising is applied to the first IMF of each band because this IMF includes local high-spatial-frequency components. Features reconstructed as the sums of lower order IMFs are then used for classification. Support vector machine classification is used as a classification approach in this letter. Experimental results show that the proposed technique can provide a higher classification accuracy. Begüm Demir, Sarp Ertürk, M. Kemal Güllü |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2011 | Batch-Mode Active-Learning Methods for the Interactive Classification of Remote Sensing ImagesabstractThis paper investigates different batch-mode active-learning (AL) techniques for the classification of remote sensing (RS) images with support vector machines. This is done by generalizing to multiclass problem techniques defined for binary classifiers. The investigated techniques exploit different query functions, which are based on the evaluation of two criteria: uncertainty and diversity. The uncertainty criterion is associated to the confidence of the supervised algorithm in correctly classifying the considered sample, while the diversity criterion aims at selecting a set of unlabeled samples that are as more diverse (distant one another) as possible, thus reducing the redundancy among the selected samples. The combination of the two criteria results in the selection of the potentially most informative set of samples at each iteration of the AL process. Moreover, we propose a novel query function that is based on a kernel-clustering technique for assessing the diversity of samples and a new strategy for selecting the most informative representative sample from each cluster. The investigated and proposed techniques are theoretically and experimentally compared with state-of-the-art methods adopted for RS applications. This is accomplished by considering very high resolution multispectral and hyperspectral images. By this comparison, we observed that the proposed method resulted in better accuracy with respect to other investigated and state-of-the art methods on both the considered data sets. Furthermore, we derived some guidelines on the design of AL systems for the classification of different types of RS images. Begüm Demir, Claudio Persello, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2010 | Empirical mode decomposition based decision fusion for higher hyperspectral image classification accuracyabstractThis paper proposes a novel Empirical Mode Decomposition (EMD) based decision fusion approach for accurate classification of hyperspectral images. The proposed method consists of three steps. In the first step, EMD, which iteratively decomposes the data into so called Intrinsic Mode Functions (IMFs) in accordance with the intrinsic characteristics of data, is applied to each hyperspectral image band for decomposition. In the second step, the IMFs are assumed as different representations of data, and original hyperspectral data as well as IMF based representations are classified by Support Vector Machine (SVM), independently from each other, to obtain independent decisions. In the final step, these independent decisions are fused by a decision fusion rule to get the final classification result. Provided experimental results demonstrate that the proposed EMD based decision approach results in improved SVM classification. Begüm Demir, Sarp Ertürk |
IGARSS | 1 |
| 2010 | Empirical Mode Decomposition of Hyperspectral Images for Support Vector Machine ClassificationabstractThis paper presents the utilization of empirical mode decomposition (EMD) of hyperspectral images to increase the classification accuracy using support vector machine (SVM)-based classification. EMD has been shown in the literature to be particularly suitable for nonlinear and nonstationary signals and is used in this paper to decompose hyperspectral image bands into several intrinsic mode functions (IMFs) and a final residue. EMD is utilized in this paper to improve hyperspectral-image-classification accuracy by effectively exploiting the feature that EMD performs a decomposition that is spatially adaptive with respect to intrinsic features. This paper presents two different approaches for improved hyperspectral image classification making use of EMD. In the first approach, IMFs corresponding to each hyperspectral image band are obtained and the sums of lower order IMFs are used as new features for classification with SVM. In the second approach, the pieces of information contained in the first and second IMFs of each hyperspectral image band are combined using composite kernels for SVM classification with higher accuracy. Begüm Demir, Sarp Ertürk |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2009 | Improving SVM classification accuracy using a hierarchical approach for hyperspectral imagesabstractThis paper proposes to combine standard SVM classification with a hierarchical approach to increase SVM classification accuracy as well as reduce computational load of SVM testing. Support vectors are obtained by applying SVM training to the entire original training data. For classification, multi-level two-dimensional wavelet decomposition is applied to each hyperspectral image band and low spatial frequency components of each level are used for hierarchical classification. Initially, conventional SVM classification is carried out in the highest hierarchical level (lowest resolution) using all support vectors and a one-to-one multiclass classification strategy, so that all pixels in the lowest resolution are classified. In the sub-sequent levels (higher resolutions) pixels are classified using the class information of the corresponding neighbor pixels of the upper level. Therefore, the classification at a lower level is carried out using only the support vectors of classes to which corresponding neighbor pixels in the higher level are assigned to. Because classification with all support vectors is only utilized at the lowest resolution and classification of higher resolutions requires a subset of the support vectors, this approach reduces the overall computational load of SVM classification and provides reduced SVM testing time compared to standard SVM. Furthermore, the proposed approach provides significantly better classification accuracy as it exploits spatial correlation thanks to hierarchical processing. Begüm Demir, Sarp Ertürk |
ICIP | 1 |
| 2009 | Improved quality multiple description 3D mesh coding with optimal filteringabstractA wavelet subdivision surfaces based multiple description 3D model coding with optimal reconstruction filtering approach is proposed in this paper to transmit 3D models over mediums with possible packet loss with higher quality. The proposed method is based on the wavelet subdivision surfaces strategy for 3D model coding and introduces optimal reconstruction filtering. Initially, remeshing is performed for any 3D model to obtain a semi-regular mesh structure which includes an irregular base mesh and regular wavelet (detail) coefficients of finer levels. Next, multiple descriptions of detail coefficients are obtained using Multiple Description Scalar Quantization (MDSQ) to be encoded using SPIHT. In the proposed approach, reconstructed multiple descriptions are applied to optimal filters, that are obtained at the encoder so as to reduce the reconstruction error between the original model and encoded model. Each transmitted description therefore includes encoded 3D data as well as optimal reconstruction filter coefficients. Experimental results show that the optimal reconstruction filtering approach reduces the distortion of standard MDSQ based 3D model coding. Begüm Demir, Sarp Ertürk, Oguzhan Urhan |
ICIP | 1 |
| 2009 | An Empirical Mode Decomposition and Composite Kernel Approach to Increase Hyperspectral Image Classification AccuracyabstractThis paper proposes to increase the classification accuracy of hyperspectral images based on Empirical Mode Decomposition (EMD) algorithm and composite kernels. EMD is a signal decomposition algorithm and decomposes signals into several Intrinsic Mode Functions (IMFs) and a final residue. In this paper, two-dimensional EMD is initially applied to each hyperspectral image band separately and IMFs of hyperspectral image bands are obtained. Composite kernels are used to combine the information contained in the first IMFs and second IMFs of all bands and kernel based Support Vector Machine (SVM) is used for classification. Experimental results confirm the usefulness of the proposed approach compared to direct SVM approach. Begüm Demir, Sarp Ertürk |
IGARSS (2) | 1 |
| 2009 | Wavelet Shrinkage Denoising of Intrinsic Mode Functions of Hyperspectral Image Bands for Classification with High AccuracyabstractThis paper proposes Empirical Mode Decomposition (EMD) followed by wavelet shrinkage denoising in hyperspectral image classification to improve classification accuracy. EMD decomposes signals into several Intrinsic Mode Functions (IMFs) and a final residue. In this paper, firstly, EMD is applied to each hyperspectral image band separately to obtain the IMFs of all image bands. Then, the first IMF of each band is applied to wavelet shrinkage denoising, as this IMF includes all local high spatial frequency components. The sums of lower order IMFs are then used to reconstruct hyperspectral image bands that are used as new features for classification. Support Vector Machine (SVM) based classification is used as classification approach in this paper. Experimental results show the effectiveness of the proposed approach. Begüm Demir, Sarp Ertürk, M. Kemal Güllü |
IGARSS (3) | 1 |
| 2009 | Clustering-Based Extraction of Border Training Patterns for Accurate SVM Classification of Hyperspectral ImagesabstractThis letter presents an accurate support vector machine (SVM)-based hyperspectral image classification algorithm, which uses border training patterns that are close to the separating hyperplane. Border training patterns are obtained in two consecutive steps. In the first step, clustering is performed to training data of each class, and cluster centers are taken as initial training data for SVM. In the second step, the reduced-size training data composed of cluster centers are used in SVM training, and cluster centers obtained as support vectors at this step are regarded to be located close to the hyperplane border. Original training samples are contained in clusters for which the cluster centers are obtained to be close to the hyperplane border and the corresponding cluster centers are then together assigned as border training patterns. These border training patterns are then used in the training of the SVM classifier. Experimental results show that it is possible to significantly increase the classification accuracy of SVM using border training patterns obtained with the proposed approach. Begüm Demir, Sarp Ertürk |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2009 | A Low-Complexity Approach for the Color Display of Hyperspectral Remote-Sensing Images Using One-Bit-Transform-Based Band SelectionabstractThis paper presents a new approach for the color display of hyperspectral images. It is proposed to use the one-bit transform (1BT) of hyperspectral image bands to select three suitable bands for red, green, and blue (RGB) display. The proposed approach has low complexity and is very suitable for hardware implementation. A dedicated hardware architecture that computes the transitions in the 1BT of hyperspectral image bands to determine bands that contain more information and the corresponding field-programmable gate array implementation of the proposed architecture are presented. In the proposed approach, less-structured bands are initially eliminated using the total number of transitions in the 1BT of hyperspectral image bands. Then, three suitable bands are selected from within this remaining set of well-structured bands for RGB color display. The proposed approach provides a new method for facilitating the color display of hyperspectral images, which has very low complexity. Begüm Demir, Anil Çelebi, Sarp Ertürk |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2008 | Spectral Magnitude and Spectral Derivative Feature Fusion for Improved Classification of Hyperspectral ImagesabstractThis paper proposes to increase the classification accuracy of hyperspectral images by fusing spectral magnitude features and spectral derivative features. Principle component analysis (PCA) is used as feature extraction method to reduce the final number of features of the hyperspectral data before feature fusion. PCA is applied separately to magnitude and derivative features to determine significant components of each. Different fusion approaches of the significant components of magnitude features and the significant components of the first as well as second spectral derivatives are evaluated to construct the desired number of final features. Support vector machine (SVM) classification is used for classification of hyperspectral images after feature fusion and it is shown that the proposed approach improves classification accuracy. Begüm Demir, Sarp Ertürk |
IGARSS (3) | 1 |
| 2008 | Reducing the Computational Load of Hyperspectral Band Selection Using the One-Bit Transform of Hyperspectral BandsabstractThis paper concentrates on reducing the computational complexity of hyperspectral image band selection algorithms via one-bit transform which can be obtained using simple filtering and comparison operations. Firstly, one-bit transform of each band is obtained and noisy and less-discriminative bands, which are decided according to the total number of vertical and horizontal transitions in their one-bit representations, are eliminated. Then remained bands are forwarded to the band selection and classification algorithms. Steepest ascents band selection and support vector machine classification are used to demonstrate the performance of the proposed approach. Experimental results show that the proposed approach not only reduces the computational load of the band selection process, but also provides similar or even higher classification accuracy. Begüm Demir, Sarp Ertürk |
IGARSS (2) | 1 |
| 2008 | Empirical Mode Decomposition Pre-Process for Higher Accuracy Hyperspectral Image ClassificationabstractThis paper proposes empirical mode decomposition (EMD) based pre-process to increase classification accuracy of hyperspectral images. EMD is an adaptive and non-linear signal decomposition approach and decomposes the data into intrinsic mode functions (IMFs) and a residue. In this paper, EMD is applied to each hyperspectral image band to obtain IMFs. After EMD is performed to each band, new bands are reconstructed as the sum of higher level IMFs and classification is executed over these new bands. Support vector machine (SVM) is used to show the classification performance of the proposed approach. Experimental results show that, utilization of the first two IMFs significantly increases the classification accuracy compared to applying SVM directly to the original data set. Begüm Demir, Sarp Ertürk |
IGARSS (2) | 1 |
| 2007 | Hyperspectral data classification using RVM with pre-segmentation and RANSACabstractRelevance vector machines (RVMs) and support vector machines (SVMs) are known to outperform classical supervised classification algorithms. RVMs have some advantages compared to SVMs, the most important being more sparsity. This paper presents hyperspectral image classification based on relevance vector machines with two different unsupervised segmentation methods as well as RANSAC (RANdom SAmple Consencus) applied before RVM classification. Compression is achieved using k-means or phase correlation based unsupervised segmentation, or using RANSAC cross-validation before the RVM classification step. Approximately the same hyperspectral data classification accuracy can be obtained with a smaller relevance vector rate and faster training time for the proposed pre-segmented RVM classification approach compared with direct RVM classification. The proposed approach can be used to improve the sparsity of RVM classification even further, and is particularly suitable for low-complexity applications. Begüm Demir, Sarp Ertürk |
IGARSS | 1 |
| 2007 | Phase correlation based supervised classification of hyperspectral images using multiple class representativesabstractIn this paper it is genuinely proposed to use a modified phase correlation (MPC) based supervised classification approach for hyperspectral images. The hyperspectral spectrum of each pixel is initially subsampled to gain robustness against noise and spatial variability, and phase correlation is applied to determine spectral similarity to class feature vectors. For this purpose it is required to obtain class feature vectors in the training phase. It is shown that the classification accuracy can be improved if multiple representative feature vectors are utilized for each class. These multiple representatives are selected from training data by finding training vectors of the same class that are less similar, so as to represent the class as good as possible with different representatives. Prediction is made according to the maximum value of the phase correlation results between new samples and the class representatives. Begüm Demir, Sarp Ertürk |
IGARSS | 1 |
| 2007 | Hyperspectral Image Classification Using Relevance Vector MachinesabstractThis letter presents a hyperspectral image classification method based on relevance vector machines (RVMs). Support vector machine (SVM)-based approaches have been recently proposed for hyperspectral image classification and have raised important interest. In this letter, it is genuinely proposed to use an RVM-based approach for the classification of hyperspectral images. It is shown that approximately the same classification accuracy is obtained using RVM-based classification, with a significantly smaller relevance vector rate and, therefore, much faster testing time, compared with SVM-based classification. This feature makes the RVM-based hyperspectral classification approach more suitable for applications that require low complexity and, possibly, real-time classification. Begüm Demir, Sarp Ertürk |
IEEE Geosci. Remote. Sens. Lett. | 1 |