VLDB 2026 Research / reviewers in the wild / expert
Gencer Sumbul
dblp:211/7067 · also Gencer Sümbül
· DBLP profile ↗
23ranked-venue papers
14as first author
18since 2021 · last 2025
0000-0003-3690-3052ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 11 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsabstractMonitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video datasets limits the development of powerful video understanding models needed to process the vast amount of fieldwork data collected. To advance research in wild animal behavior monitoring we present MammAlps, a multi-modal and multi-view dataset of wildlife behavior monitoring from 9 camera-traps in the Swiss National Park. Mam-mAlps contains over 14 hours of video with audio, 2D segmentation maps and 8.5 hours of individual tracks densely labeled for species and behavior. Based on 6‘135 single animal clips, we propose the first hierarchical and multi-modal animal behavior recognition benchmark using audio, video and reference scene segmentation maps as inputs. Furthermore, we also propose a second ecology-oriented benchmark aiming at identifying activities, species, number of individuals and meteorological conditions from 397 multi-view and long-term ecological events, including false positive triggers. We advocate that both tasks are complementary and contribute to bridging the gap between machine learning and ecology. Code and data are available at https://github.com/eceo-epfl/MammAlps. Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul, Alexander Mathis, Devis Tuia |
CVPR | 4 |
| 2025 | SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
Gencer Sumbul, Chang Xu 0027, Emanuele Dalsasso, Devis Tuia |
ICCV | 1 |
| 2025 | Exploring Masked Autoencoders for Sensor-Agnostic Image Retrieval in Remote SensingabstractSelf-supervised learning through masked autoencoders (MAEs) has recently attracted great attention for remote sensing (RS) image representation learning (IRL), and thus embodies a significant potential for content-based image retrieval (CBIR) from ever-growing RS image archives. However, the existing MAE-based CBIR studies in RS assume that the considered RS images are acquired by a single image sensor, and thus are only suitable for unimodal CBIR problems. The effectiveness of MAEs for cross-sensor CBIR, which aims to search semantically similar images across different image modalities, has not been explored yet. In this article, we take the first step to explore the effectiveness of MAEs for sensor-agnostic CBIR in RS. To this end, we present a systematic overview on the possible adaptations of the vanilla MAE to exploit masked image modeling (MIM) on multisensor RS image archives [denoted as cross-sensor masked autoencoders [(CSMAEs)] in the context of CBIR. Based on different adjustments applied to the vanilla MAE, we introduce different CSMAE models. We also provide an extensive experimental analysis of these CSMAE models. We finally derive a guideline to exploit MIM for unimodal and cross-modal CBIR problems in RS. The code of this work is publicly available athttps://github.com/jakhac/CSMAE. Jakob Hackstein, Gencer Sumbul, Kai Norman Clasen, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Annotation Cost-Efficient Active Learning for Deep Metric Learning-Driven Remote Sensing Image RetrievalabstractDeep metric learning (DML) has shown to be effective for content-based image retrieval (CBIR) in remote sensing (RS). Most of the DML methods for CBIR rely on a high number of annotated images to accurately learn model parameters of deep neural networks (DNNs). However, gathering such data is time-consuming and costly. To address this, we propose an annotation cost-efficient active learning (ANNEAL) method tailored to DML-driven CBIR in RS. ANNEAL aims to create a small but informative training set made up of similar and dissimilar image pairs to be used for accurately learning a metric space. The informativeness of image pairs is evaluated by combining uncertainty and diversity criteria. To assess the uncertainty of image pairs, we introduce two algorithms: 1) metric-guided uncertainty estimation (MGUE) and 2) binary-classifier-guided uncertainty estimation (BCGUE). MGUE algorithm automatically estimates a threshold value that acts as a boundary between similar and dissimilar image pairs based on the distances in the metric space. The closer the similarity between image pairs is to the estimated threshold value, the higher their uncertainty. BCGUE algorithm estimates the uncertainty of the image pairs based on the confidence of the classifier in assigning correct similarity labels. The diversity criterion is assessed through a clustering-based strategy. ANNEAL combines either MGUE or BCGUE algorithm with the clustering-based strategy to select the most informative image pairs, which are then labeled by expert annotators as similar or dissimilar. This way of annotating images significantly reduces the annotation cost compared with annotating images with land-use land-cover class labels. Experimental results on two RS benchmark datasets demonstrate the effectiveness of our method. The code of this work is publicly available athttps://git.tu-berlin.de/rsim/anneal_tgrs. Genc Hoxha, Gencer Sumbul, Julia Henkel, Lars Möllenbrok, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Learning Across Decentralized Multi-Modal Remote Sensing Archives with Federated LearningabstractThe development of federated learning (FL) methods, which aim to learn from distributed databases (i.e., clients) without accessing data on clients, has recently attracted great attention. Most of these methods assume that the clients are associated with the same data modality. However, remote sensing (RS) images in different clients can be associated with different data modalities that can improve the classification performance when jointly used. To address this problem, in this paper we introduce a novel multi-modal FL framework that aims to learn from decentralized multi-modal RS image archives for RS image classification problems. The proposed framework is made up of three modules: 1) multimodal fusion (MF); 2) feature whitening (FW); and 3) mutual information maximization (MIM). The MF module performs iterative model averaging to learn without accessing data on clients in the case that clients are associated with different data modalities. The FW module aligns the representations learned among the different clients. The MIM module maximizes the similarity of images from different modalities. Experimental results show the effectiveness of the proposed framework compared to iterative model averaging, which is a widely used algorithm in FL. The code of the proposed framework is publicly available at https://git.tu-berlin.de/rsim/MMFL. Baris Büyüktas, Gencer Sumbul, Begüm Demir |
IGARSS | 2 |
| 2023 | Annotation Cost Efficient Active Learning for Content Based Image RetrievalabstractDeep metric learning (DML) based methods have been found very effective for content-based image retrieval (CBIR) in remote sensing (RS). For accurately learning the model parameters of deep neural networks, most of the DML methods require a high number of annotated training images, which can be costly to gather. To address this problem, in this paper we present an annotation cost efficient active learning (AL) method (denoted as ANNEAL). The proposed method aims to iteratively enrich the training set by annotating the most informative image pairs as similar or dissimilar, while accurately modelling a deep metric space. This is achieved by two consecutive steps. In the first step the pairwise image similarity is modelled based on the available training set. Then, in the second step the most uncertain and diverse (i.e., informative) image pairs are selected to be annotated. Unlike the existing AL methods for CBIR, at each AL iteration of ANNEAL a human expert is asked to annotate the most informative image pairs as similar/dissimilar. This significantly reduces the annotation cost compared to annotating images with land-use/land cover class labels. Experimental results show the effectiveness of our method. The code of ANNEAL is publicly available at https://git.tu-berlin.de/rsim/ANNEAL. Julia Henkel, Genc Hoxha, Gencer Sumbul, Lars Möllenbrok, Begüm Demir |
IGARSS | 3 |
| 2023 | Label Noise Robust Image Representation Learning Based on Supervised Variational Autoencoders in Remote SensingabstractDue to the publicly available thematic maps and crowd-sourced data, remote sensing (RS) image annotations can be gathered at zero cost for training deep neural networks (DNNs). However, such annotation sources may increase the risk of including noisy labels in training data, leading to inaccurate RS image representation learning (IRL). To address this issue, in this paper we propose a label noise robust IRL method that aims to prevent the interference of noisy labels on IRL, independently from the learning task being considered in RS. To this end, the proposed method combines a supervised variational autoencoder (SVAE) with any kind of DNN. This is achieved by defining variational generative process based on image features. This allows us to define the importance of each training sample for IRL based on the loss values acquired from the SVAE and the task head of the considered DNN. Then, the proposed method imposes lower importance to images with noisy labels, while giving higher importance to those with correct labels during IRL. Experimental results show the effectiveness of the proposed method when compared to well-known label noise robust IRL methods applied to RS images. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/RS-IRL-SVAE. Gencer Sumbul, Begüm Demir |
IGARSS | 1 |
| 2023 | Deep Active Learning for Multi-Label Classification of Remote Sensing ImagesabstractIn this letter, we introduce deep active learning (AL) for multi-label classification (MLC) problems in remote sensing (RS). In particular, we investigate the effectiveness of several AL query functions for MLC of RS images. Unlike the existing AL query functions (which are defined for single-label classification or semantic segmentation problems), each query function in this paper is based on the evaluation of two criteria: i) multi-label uncertainty; and ii) multi-label diversity. The multi-label uncertainty criterion is associated to the confidence of the deep neural networks (DNNs) in correctly assigning multi-labels to each image. To assess this criterion, we investigate three strategies: i) learning multi-label loss ordering; ii) measuring temporal discrepancy of multi-label predictions; and iii) measuring magnitude of approximated gradient embeddings. The multi-label diversity criterion is associated to the selection of a set of images that are as diverse as possible to each other that prevents redundancy among them. To assess this criterion, we exploit a clustering based strategy. We combine each of the above-mentioned uncertainty strategies with the clustering based diversity strategy, resulting in three different query functions. All the considered query functions are introduced for the first time in the framework of MLC problems in RS. Experimental results obtained on two benchmark archives show that these query functions result in the selection of a highly informative set of samples at each iteration of the AL process. Lars Möllenbrok, Gencer Sumbul, Begüm Demir |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Generative Reasoning Integrated Label Noise Robust Deep Image Representation LearningabstractThe development of deep learning based image representation learning (IRL) methods has attracted great attention for various image understanding problems. Most of these methods require the availability of a set of high quantity and quality of annotated training images, which can be time-consuming, complex and costly to gather. To reduce labeling costs, crowdsourced data, automatic labeling procedures or citizen science projects can be considered. However, such approaches increase the risk of including label noise in training data. It may result in overfitting on noisy labels when discriminative reasoning is employed as in most of the existing methods. This leads to sub-optimal learning procedures, and thus inaccurate characterization of images. To address this issue, in this paper, we introduce a generative reasoning integrated label noise robust deep representation learning (GRID) approach. The proposed GRID approach aims to model the complementary characteristics of discriminative and generative reasoning for IRL under noisy labels. To this end, we first integrate generative reasoning into discriminative reasoning through a supervised variational autoencoder. This allows the proposed GRID approach to automatically detect training samples with noisy labels. Then, through our label noise robust hybrid representation learning strategy, GRID adjusts the whole learning procedure for IRL of these samples through generative reasoning and that of the other samples through discriminative reasoning. Our approach learns discriminative image representations while preventing interference of noisy labels during training independently from the IRL method being selected. Thus, unlike the existing label noise robust methods, GRID does not depend on the type of annotation, label noise, neural network architecture, loss function or learning task, and thus can be directly utilized for various image understanding problems. Experimental results show the effectiveness of the proposed GRID approach compared to the state-of-the-art methods. The code of the proposed approach is publicly available at https://github.com/gencersumbul/GRID. Gencer Sumbul, Begüm Demir |
IEEE Trans. Image Process. | 1 |
| 2022 | A Novel Self-Supervised Cross-Modal Image Retrieval Method in Remote SensingabstractDue to the availability of multi-modal remote sensing (RS) image archives, one of the most important research topics is the development of cross-modal RS image retrieval (CM-RSIR) methods that search semantically similar images across different modalities. Existing CM-RSIR methods require the availability of a high quality and quantity of annotated training images. The collection of a sufficient number of reliable labeled images is time consuming, complex and costly in operational scenarios, and can significantly affect the final accuracy of CM-RSIR. In this paper, we introduce a novel self-supervised CM-RSIR method that aims to: i) model mutual-information between different modalities in a self-supervised manner; ii) retain the distributions of modal-specific feature spaces similar to each other; and iii) define the most similar images within each modality without requiring any annotated training image. To this end, we propose a novel objective including three loss functions that simultaneously: i) maximize mutual information of different modalities for inter-modal similarity preservation; ii) minimize the angular distance of multi-modal image tuples for the elimination of inter-modal discrepancies; and iii) increase cosine similarity of the most similar images within each modality for the characterization of intra-modal similarities. Experimental results show the effectiveness of the proposed method compared to state-of-the-art methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/SS-CM-RSIR. Gencer Sumbul, Begüm Demir |
ICIP | 1 |
| 2022 | Deep Metric Learning-Based Semi-Supervised Regression with Alternate LearningabstractThis paper introduces a novel deep metric learning-based semi-supervised regression (DML-S2R) method for parameter estimation problems. The proposed DML-S2R method aims to mitigate the problems of insufficient amount of labeled samples without collecting any additional sample with a target value. To this end, it is made up of two main steps: i) pairwise similarity modeling with scarce labeled data; and ii) triplet-based metric learning with abundant unlabeled data. The first step aims to model pairwise sample similarities by using a small number of labeled samples. This is achieved by estimating the target value differences of labeled samples with a Siamese neural network (SNN). The second step aims to learn a triplet-based metric space (in which similar samples are close to each other and dissimilar samples are far apart from each other) when the number of labeled samples is insufficient. This is achieved by employing the SNN of the first step for triplet-based deep metric learning that exploits not only labeled samples but also unlabeled samples. For the end-to-end training of DML-S2R, we investigate an alternate learning strategy for the two steps. Due to this strategy, the encoded information in each step becomes a guidance for learning phase of the other step. The experimental results confirm the success of DML-S2R compared to the state-of-the-art semi-supervised regression methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/DML-S2R. Adina Zell, Gencer Sumbul, Begüm Demir |
ICIP | 2 |
| 2022 | A Novel Framework to Jointly Compress and Index Remote Sensing Images for Efficient Content-Based RetrievalabstractRemote sensing (RS) images are usually stored in compressed format to reduce the storage size of the archives. Thus, existing content-based image retrieval (CBIR) systems in RS require decoding images before applying CBIR (which is computationally demanding in the case of large-scale CBIR problems). To address this problem, in this paper, we present a joint framework that simultaneously learns RS image compression and indexing. Thus, it eliminates the need for decoding RS images before applying CBIR. The proposed framework is made up of two modules. The first module compresses RS images based on an autoencoder architecture. The second module produces hash codes with a high discrimination capability by employing soft pairwise, bit-balancing and classification loss functions. We also introduce a two stage learning strategy with gradient manipulation techniques to obtain image representations that are compatible with both RS image indexing and compression. Experimental results show the efficacy of the proposed framework when compared to widely used approaches in RS. The code of the proposed framework is available at https://git.tu-berlin.de/rsim/RS-JCIF. Gencer Sumbul, Thekke Madam Nimisha, Begüm Demir |
IGARSS | 1 |
| 2022 | Plasticity-Stability Preserving Multi-Task Learning for Remote Sensing Image RetrievalabstractDeep learning-based multi-task learning (MTL) methods have recently attracted attention for content-based image retrieval (CBIR) applications in remote sensing (RS). For a given set of tasks (e.g., scene classification, semantic segmentation, and image reconstruction), existing MTL methods employ a joint optimization algorithm on the direct aggregation of task-specific loss functions. Such an approach may provide limited CBIR performance when: 1) tasks compete or even distract each other; 2) one of the tasks dominates the whole learning procedure; or 3) characterization of each task is underperformed compared to single-task learning. This is mainly due to the lack of: 1) plasticity condition (which is associated with sensitivity to new information) or 2) stability condition (which is associated with protection from radical disruptions by new information) of the whole learning procedure. To avoid this issue, as a first time, we propose a novel plasticity-stability preserving MTL (PLASTA-MTL) approach to ensure the plasticity and the stability conditions of the whole learning procedure independently of the number and type of tasks. This is achieved by defining two novel loss functions. The first loss function is the plasticity preserving loss (PPL) function that aims to enforce the global image representation space to be sensitive to new information learned with each task. This is achieved by minimizing the difference of gradient magnitudes for the global representation and task-specific embedding spaces. The second loss function is the stability preserving loss (SPL) function that aims to protect the global representation space radically disrupted by a new task. This is achieved by minimizing the angular distances between the task gradients over global representation space. To effectively employ the proposed loss functions, we also introduce a novel sequential optimization algorithm. Experimental results show the effectiveness of the proposed approach compared to the state-of-the-art MTL methods in the context of CBIR. Gencer Sumbul, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Informative and Representative Triplet Selection for Multilabel Remote Sensing Image RetrievalabstractLearning the similarity between remote sensing (RS) images forms the foundation for content-based RS image retrieval (CBIR). Recently, deep metric learning approaches that map the semantic similarity of images into an embedding (metric) space have been found very popular in RS. A common approach for learning the metric space relies on the selection of triplets of similar (positive) and dissimilar (negative) images to a reference image called as an anchor. Choosing triplets is a difficult task particularly for multi-label RS CBIR, where each training image is annotated by multiple class labels. To address this problem, in this paper we propose a novel triplet sampling method in the framework of deep neural networks (DNNs) defined for multi-label RS CBIR problems. The proposed method selects a small set of the most representative and informative triplets based on two main steps. In the first step, a set of anchors that are diverse to each other in the embedding space is selected from the current mini-batch using an iterative algorithm. In the second step, different sets of positive and negative images are chosen for each anchor by evaluating the relevancy, hardness and diversity of the images among each other based on a novel strategy. Experimental results obtained on two multi-label benchmark archives show that the selection of the most informative and representative triplets in the context of DNNs results in: i) reducing the computational complexity of the training phase of the DNNs without any significant loss on the performance; and ii) an increase in learning speed since informative triplets allow fast convergence. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/image-retrieval-from-triplets. Gencer Sumbul, Mahdyar Ravanbakhsh, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Towards Simultaneous Image Compression and Indexing for Scalable Content-Based Retrieval in Remote SensingabstractDue to the rapidly growing remote-sensing (RS) image archives, images are usually stored in a compressed format for reducing their storage sizes. Thus, most of the existing content-based RS image retrieval systems require fully decoding images (i.e., decompression) that is computationally demanding for large-scale archives. To address this issue, we introduce a novel approach devoted to simultaneous RS image compression and indexing for scalable content-based image retrieval (denoted as SCI-CBIR). The proposed SCI-CBIR prevents the requirement of decoding RS images before image search and retrieval. To this end, it includes two main steps: 1) deep-learning-based compression and 2) deep-hashing-based indexing. The first step effectively compresses RS images by employing a pair of deep encoder and decoder neural networks and an entropy model. The second step produces hash codes with a high discrimination capability for RS images by employing pairwise, bit-balancing, and classification loss functions. For the training of the SCI-CBIR approach, we also introduce a novel multistage learning procedure with automatic loss weighting techniques to characterize RS image representations that are appropriate for both RS image indexing and compression. The proposed learning procedure enables automatically weighting different loss functions considered for the proposed approach instead of computationally demanding grid search. Experimental results show the effectiveness of the proposed approach when compared to widely used approaches in RS. The code of the proposed approach is available athttps://git.tu-berlin.de/rsim/SCI-CBIR. Gencer Sumbul, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A Novel Graph-Theoretic Deep Representation Learning Method for Multi-Label Remote Sensing Image RetrievalabstractThis paper presents a novel graph-theoretic deep representation learning method in the framework of multi-label remote sensing (RS) image retrieval problems. The proposed method aims to extract and exploit multi-label co-occurrence relationships associated to each RS image in the archive. To this end, each training image is initially represented with a graph structure that provides region-based image representation combining both local information and the related spatial organization. Unlike the other graph-based methods, the proposed method contains a novel learning strategy to train a deep neural network for automatically predicting a graph structure of each RS image in the archive. This strategy employs a region representation learning loss function to characterize the image content based on its multi-label co-occurrence relationship. Experimental results show the effectiveness of the proposed method for retrieval problems in RS compared to state-of-the-art deep representation learning methods. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/GT-DRL-CBIR. Gencer Sumbul, Begüm Demir |
IGARSS | 1 |
| 2021 | Remote-Sensing Image Scene Classification With Deep Neural Networks in JPEG 2000 Compressed DomainabstractTo reduce the storage requirements, remote-sensing (RS) images are usually stored in compressed format. Existing scene classification approaches using deep neural networks (DNNs) require to fully decompress the images, which is a computationally demanding task in operational applications. To address this issue, in this article, we propose a novel approach to achieve scene classification in Joint Photographic Experts Group (JPEG) 2000 compressed RS images. The proposed approach consists of two main steps: 1) approximation of the finer resolution subbands of reversible biorthogonal wavelet filters used in JPEG 2000 and 2) characterization of the high-level semantic content of approximated wavelet subbands and scene classification based on the learned descriptors. This is achieved by taking codestreams associated with the coarsest resolution wavelet subband as input to approximate finer resolution subbands using a number of transposed convolutional layers. Then, a series of convolutional layers models the high-level semantic content of the approximated wavelet subband. Thus, the proposed approach models the multiresolution paradigm given in the JPEG 2000 compression algorithm in an end-to-end trainable unified neural network. In the classification stage, the proposed approach takes only the coarsest resolution wavelet subbands as input, thereby reducing the time required to apply decoding. Experimental results performed on two benchmark aerial image archives demonstrate that the proposed approach significantly reduces the computational time with similar classification accuracies when compared with traditional RS scene classification approaches (which requires full image decompression). Akshara Preethy Byju, Gencer Sumbul, Begüm Demir, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | SD-RSIC: Summarization-Driven Deep Remote Sensing Image CaptioningabstractDeep neural networks (DNNs) have been recently found popular for image captioning problems in remote sensing (RS). Existing DNN-based approaches rely on the availability of a training set made up of a high number of RS images with their captions. However, captions of training images may contain redundant information (they can be repetitive or semantically similar to each other), resulting in information deficiency while learning a mapping from the image domain to the language domain. To overcome this limitation, in this article, we present a novel summarization-driven RS image captioning (SD-RSIC) approach. The proposed approach consists of three main steps. The first step obtains the standard image captions by jointly exploiting convolutional neural networks (CNNs) with long short-term memory (LSTM) networks. The second step, unlike the existing RS image captioning methods, summarizes the ground-truth captions of each training image into a single caption by exploiting sequence to sequence neural networks and eliminates the redundancy present in the training set. The third step automatically defines the adaptive weights associated with each RS image to combine the standard captions with the summarized captions based on the semantic content of the image. This is achieved by a novel adaptive weighting strategy defined in the context of LSTM networks. Experimental results obtained on the RSCID, UCM-Captions, and Sydney-Captions data sets show the effectiveness of the proposed approach compared with the state-of-the-art RS image captioning approaches. The code of the proposed approach is publicly available athttps://gitlab.tubit.tu-berlin.de/rsim/SD-RSIC. Gencer Sumbul, Sonali Nayak, Begüm Demir |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | A Comparative Study of Deep Learning Loss Functions for Multi-Label Remote Sensing Image ClassificationabstractThis paper analyzes and compares different deep learning loss functions in the framework of multi-label remote sensing (RS) image scene classification problems. We consider seven loss functions: 1) cross-entropy loss; 2) focal loss; 3) weighted cross-entropy loss; 4) Hamming loss; 5) Huber loss; 6) ranking loss; and 7) sparseMax loss. All the considered loss functions are analyzed for the first time in RS. After a theoretical analysis, an experimental analysis is carried out to compare the considered loss functions in terms of their: 1) overall accuracy; 2) class imbalance awareness (for which the number of samples associated to each class significantly varies); 3) convexibility and differentiability; and 4) learning efficiency (i.e., convergence speed). On the basis of our analysis, some guidelines are derived for a proper selection of a loss function in multi-label RS scene classification problems. Hichame Yessou, Gencer Sumbul, Begüm Demir |
IGARSS | 2 |
| 2019 | Bigearthnet: A Large-Scale Benchmark Archive for Remote Sensing Image UnderstandingabstractThis paper presents the BigEarthNet that is a new large-scale multi-label Sentinel-2 benchmark archive. The BigEarthNet consists of 590, 326 Sentinel-2 image patches, each of which is a section of i) 120 × 120 pixels for 10m bands; ii) 60×60 pixels for 20m bands; and iii) 20×20 pixels for 60m bands. Unlike most of the existing archives, each image patch is annotated by multiple land-cover classes (i.e., multi-labels) that are provided from the CORINE Land Cover database of the year 2018 (CLC 2018). The BigEarthNet is significantly larger than the existing archives in remote sensing (RS) and thus is much more convenient to be used as a training source in the context of deep learning. This paper first addresses the limitations of the existing archives and then describes the properties of the BigEarthNet. Experimental results obtained in the framework of RS image scene classification problems show that a shallow Convolutional Neural Network (CNN) architecture trained on the BigEarthNet provides much higher accuracy compared to a state-of-the-art CNN model pre-trained on the ImageNet (which is a very popular large-scale benchmark archive in computer vision). The BigEarthNet opens up promising directions to advance operational RS applications and research in massive Sentinel-2 image archives. Gencer Sumbul, Marcela Charfuelan, Begüm Demir, Volker Markl |
IGARSS | 1 |
| 2019 | A Novel Multi-Attention Driven System for Multi-Label Remote Sensing Image ClassificationabstractThis paper presents a novel multi-attention driven system that jointly exploits Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) in the context of multi-label remote sensing (RS) image classification. The proposed system consists of four main modules. The first module aims to extract preliminary local descriptors of RS image bands that can be associated to different spatial resolutions. To this end, we introduce a K-Branch CNN, in which each branch extracts descriptors of image bands that have the same spatial resolution. The second module aims to model spatial relationship among local descriptors. This is achieved by a bidirectional RNN architecture, in which Long Short-Term Memory nodes enrich local descriptors by considering spatial relationships of local areas (image patches). The third module aims to define multiple attention scores for local descriptors. This is achieved by a novel patch-based multi-attention mechanism that takes into account the joint occurrence of multiple land-cover classes and provides the attention-based local descriptors. The last module exploits these descriptors for multi-label RS image classification. Experimental results obtained on the BigEarth-Net that is a large-scale Sentinel-2 benchmark archive show the effectiveness of the proposed method compared to a state of the art method. Gencer Sumbul, Begüm Demir |
IGARSS | 1 |
| 2019 | Multisource Region Attention Network for Fine-Grained Object Recognition in Remote Sensing ImageryabstractFine-grained object recognition concerns the identification of the type of an object among a large number of closely related subcategories. Multisource data analysis that aims to leverage the complementary spectral, spatial, and structural information embedded in different sources is a promising direction toward solving the fine-grained recognition problem that involves low between-class variance, small training set sizes for rare classes, and class imbalance. However, the common assumption of coregistered sources may not hold at the pixel level for small objects of interest. We present a novel methodology that aims to simultaneously learn the alignment of multisource data and the classification model in a unified framework. The proposed method involves a multisource region attention network that computes per-source feature representations, assigns attention scores to candidate regions sampled around the expected object locations by using these representations, and classifies the objects by using an attention-driven multisource representation that combines the feature representations and the attention scores from all sources. All components of the model are realized using deep neural networks and are learned in an end-to-end fashion. Experiments using RGB, multispectral, and LiDAR elevation data for classification of street trees showed that our approach achieved 64.2% and 47.3% accuracies for the 18-class and 40-class settings, respectively, which correspond to 13% and 14.3% improvement relative to the commonly used feature concatenation approach from multiple sources. Gencer Sumbul, Ramazan Gokberk Cinbis, Selim Aksoy |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Fine-Grained Object Recognition and Zero-Shot Learning in Remote Sensing ImageryabstractFine-grained object recognition that aims to identify the type of an object among a large number of subcategories is an emerging application with the increasing resolution that exposes new details in image data. Traditional fully supervised algorithms fail to handle this problem where there is low between-class variance and high within-class variance for the classes of interest with small sample sizes. We study an even more extreme scenario named zero-shot learning (ZSL) in which no training example exists for some of the classes. ZSL aims to build a recognition model for new unseen categories by relating them to seen classes that were previously learned. We establish this relation by learning a compatibility function between image features extracted via a convolutional neural network and auxiliary information that describes the semantics of the classes of interest by using training samples from the seen classes. Then, we show how knowledge transfer can be performed for the unseen classes by maximizing this function during inference. We introduce a new data set that contains 40 different types of street trees in 1-ft spatial resolution aerial data, and evaluate the performance of this model with manually annotated attributes, a natural language model, and a scientific taxonomy as auxiliary information. The experiments show that the proposed model achieves 14.3% recognition accuracy for the classes with no training examples, which is significantly better than a random guess accuracy of 6.3% for 16 test classes, and three other ZSL algorithms. Gencer Sumbul, Ramazan Gokberk Cinbis, Selim Aksoy |
IEEE Trans. Geosci. Remote. Sens. | 1 |