Genc Hoxha

dblp:253/3663 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-2552-5615ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Annotation Cost-Efficient Active Learning for Deep Metric Learning-Driven Remote Sensing Image Retrieval
abstract
Deep metric learning (DML) has shown to be effective for content-based image retrieval (CBIR) in remote sensing (RS). Most of the DML methods for CBIR rely on a high number of annotated images to accurately learn model parameters of deep neural networks (DNNs). However, gathering such data is time-consuming and costly. To address this, we propose an annotation cost-efficient active learning (ANNEAL) method tailored to DML-driven CBIR in RS. ANNEAL aims to create a small but informative training set made up of similar and dissimilar image pairs to be used for accurately learning a metric space. The informativeness of image pairs is evaluated by combining uncertainty and diversity criteria. To assess the uncertainty of image pairs, we introduce two algorithms: 1) metric-guided uncertainty estimation (MGUE) and 2) binary-classifier-guided uncertainty estimation (BCGUE). MGUE algorithm automatically estimates a threshold value that acts as a boundary between similar and dissimilar image pairs based on the distances in the metric space. The closer the similarity between image pairs is to the estimated threshold value, the higher their uncertainty. BCGUE algorithm estimates the uncertainty of the image pairs based on the confidence of the classifier in assigning correct similarity labels. The diversity criterion is assessed through a clustering-based strategy. ANNEAL combines either MGUE or BCGUE algorithm with the clustering-based strategy to select the most informative image pairs, which are then labeled by expert annotators as similar or dissimilar. This way of annotating images significantly reduces the annotation cost compared with annotating images with land-use land-cover class labels. Experimental results on two RS benchmark datasets demonstrate the effectiveness of our method. The code of this work is publicly available athttps://git.tu-berlin.de/rsim/anneal_tgrs.
Genc Hoxha, Gencer Sumbul, Julia Henkel, Lars Möllenbrok, Begüm Demir
IEEE Trans. Geosci. Remote. Sens.1
2023 Annotation Cost Efficient Active Learning for Content Based Image Retrieval
abstract
Deep metric learning (DML) based methods have been found very effective for content-based image retrieval (CBIR) in remote sensing (RS). For accurately learning the model parameters of deep neural networks, most of the DML methods require a high number of annotated training images, which can be costly to gather. To address this problem, in this paper we present an annotation cost efficient active learning (AL) method (denoted as ANNEAL). The proposed method aims to iteratively enrich the training set by annotating the most informative image pairs as similar or dissimilar, while accurately modelling a deep metric space. This is achieved by two consecutive steps. In the first step the pairwise image similarity is modelled based on the available training set. Then, in the second step the most uncertain and diverse (i.e., informative) image pairs are selected to be annotated. Unlike the existing AL methods for CBIR, at each AL iteration of ANNEAL a human expert is asked to annotate the most informative image pairs as similar/dissimilar. This significantly reduces the annotation cost compared to annotating images with land-use/land cover class labels. Experimental results show the effectiveness of our method. The code of ANNEAL is publicly available at https://git.tu-berlin.de/rsim/ANNEAL.
Julia Henkel, Genc Hoxha, Gencer Sumbul, Lars Möllenbrok, Begüm Demir
IGARSS2
2023 Improving Image Captioning Systems With Postprocessing Strategies
abstract
Image captioning (IC) systems are generally based on encoder–decoder architecture where convolutional neural networks (CNNs) are employed to represent an image with discriminative features and recurrent neural networks (RNNs) sequentially generate a sentence description. Even though a lot of effort has been devoted lately to designing reliable IC systems, the task is far from being solved. The generated descriptions can be affected by different errors related to the attributes and the objects present in the scene. Moreover, once an error occurs, it can be propagated in the recurrent layers of the decoder generating non-accurate descriptions. To solve this problem, we propose two postprocessing strategies applied to the generated descriptions to rectify the errors and improve their quality. The proposed postprocessing strategies are based on hidden Markov models (HMMs) and Viterbi algorithm. The proposed postprocessing strategies can be applied to any encoder–decoder IC system. They are applied at test time once the IC system is trained. In particular, we propose to rectify a sentence once it is fully generated (post-generation strategy) or at each time instant of the generation process (in-generation strategy). Experiments conducted on four different IC datasets confirm the promising capabilities of the proposed postprocessing strategies to rectify the output of a simple encoder–decoder by generating more coherent descriptions. The achieved results are competitive and sometimes better than complex IC systems.
Genc Hoxha, Giacomo Scuccato, Farid Melgani
IEEE Trans. Geosci. Remote. Sens.1
2022 Change Captioning: A New Paradigm for Multitemporal Remote Sensing Image Analysis
abstract
Change detection (CD) is among the most important applications in remote sensing that allows identifying the changes that occurred in a given geographical area across different times. Even though CD systems have seen a lot of progress in RS, their output is either a binary map highlighting the changing area or a semantic change map that indicates the type of change for each pixel. The change maps are often difficult to interpret by end-users and they omit important information such are relationships and attributes of the changed areas. Motivated by the recent advancement of image captioning in the RS community, in this article we propose to describe the changes over bi-temporal images through change sentence descriptions. The aim of this article is to provide a user-friendly interpretation of the occurred changes. To this end, we propose two change captioning (CC) systems that take as input bi-temporal images and generate coherent sentence descriptions of the occurred changes. Convolutional neural networks (CNNs) are used to extract discriminative features from the bi-temporal images and recurrent neural networks (RNNs) or support vector machines (SVMs) are exploited to generate coherent change descriptions. Furthermore, in absence of a CC dataset to test our systems, we propose two new datasets. One is based on very high-resolution RGB images and the other one is based on multispectral RS images. The obtained experimental results show promising capabilities of the proposed systems to generate coherent change descriptions from the bi-temporal images. The datasets are available at the following link: https://disi.unitn.it/~melgani/datasets.html.
Genc Hoxha, Seloua Chouaf, Farid Melgani, Youcef Smara
IEEE Trans. Geosci. Remote. Sens.1
2022 A Novel SVM-Based Decoder for Remote Sensing Image Captioning
abstract
Most of the remote sensing image captioning (IC) models are based on encoder–decoder frameworks where a convolutional neural network (CNN) encodes the image information and a recurrent neural network (RNN) decodes the image information into a sentence description. In order to achieve good accuracies, encoder–decoder frameworks relying on RNNs typically require a huge amount of annotated samples. Furthermore, they demand high and expensive computational power in order to have reasonable training and testing time. In this article, we aim to address these issues by introducing a novel decoder that is based on support vector machines (SVMs). In particular, instead of RNNs, we propose a novel network of SVMs to decode the image information into a sentence description. The proposed IC system is particularly interesting when just a limited amount of training samples is available. Experiments conducted on four different IC datasets confirm the promising capability of the proposed IC system to generate descriptions that are highly correlated with the image content. The proposed IC system is characterized by short training and inference times compared to other state-of-the-art models.
Genc Hoxha, Farid Melgani
IEEE Trans. Geosci. Remote. Sens.1
2021 Captioning Changes in Bi-Temporal Remote Sensing Images
abstract
Motivated by the good performance recorded for image captioning (IC) techniques in different remote sensing (RS) applications, we propose in this paper a change detection (CD) system based on IC. It aims to the creation of a user-friendly solution that describes, via human-like sentences, the changes detected comparing bi-temporal images acquired over the same geographical area. The model we propose, is based on a convolutional neural network (CNN), and a multimodal recurrent neural network (RNN). Our experiments have been performed combining a set of aerial images with semantic information that we generated to describe the changes observed for different types of objects. This work yielded to encouraging results, evaluated using the BLEU metric.
Seloua Chouaf, Genc Hoxha, Youcef Smara, Farid Melgani
IGARSS2
2021 An Active Learning Strategy for SVM-Based Captioning
abstract
Remote sensing (RS) image captioning (IC) is a novel technique introduced recently in the RS community to enrich the description of very high resolution (VHR) images. The goal of RSIC is to generate a sentence that summarizes the content of an image. In general, RSIC is developed in a supervised way where annotated samples are needed to train the system. However, obtaining annotated samples is costly and time consuming in particular when the labels are sentence descriptions that are very subjective. In order to cope with the problem of having large amounts of training samples in this work we propose an active learning solution for RSIC that selects the most important samples to label and include in the training. Experimental results performed on UCM caption dataset show the promising effectiveness of the proposed active learning strategy.
Genc Hoxha, Farid Melgani
IGARSS1
2021 Improving Text Encoding for Retro-Remote Sensing
abstract
A recent work on retro-remote sensing (converting ancient text descriptions into images) was proposed using a multilabel encoding scheme in which an input text description is represented by a binary vector indicating the presence or absence of specific objects. However, this kind of encoding disregards information such as object attributes and spatial relationship between multiple objects in a description, resulting in images that do not semantically (fully) conform to the input description. In this letter, we propose an improved text-encoding mechanism that takes into account different levels of information available from an input text. The encoded text is then used as conditional information to guide the image synthesis process using generative adversarial networks (GANs). Besides, we present a modified GAN architecture intending to improve the semantic content of the generated images. Both the qualitative and quantitative results obtained indicate that the proposed method is particularly promising.
Mesay Belete Bejiga, Genc Hoxha, Farid Melgani
IEEE Geosci. Remote. Sens. Lett.2
2020 Remote Sensing Image Captioning with SVM-Based Decoding
abstract
With the fast development of remote sensing (RS) technology we are now able to acquire high resolution images. To cope with the new challenges of analyzing such images, a recently introduced tool is RS image captioning (IC). With respect to conventional techniques such as scene classification, RSIC provides more information about an image. It aims to generate a description that summarizes the content of an image. Most of RSIC systems are based on deep learning frameworks (encoder-decoder). The performance of these frameworks strongly depend on the number of annotated samples used during training. In this paper, we propose an alternative RSIC system for a relatively small dataset based on support vector machines (SVMs). A pre-trained CNN is used to extract the image visual features and a network of SVMs is used to generate the descriptions. Experimental results on a RS image archive composed of images acquired by unmanned aerial vehicles (UAV), show that the proposed IC system could be an interesting alternative to deep learning frameworks when only small training samples are available.
Genc Hoxha, Farid Melgani
IGARSS1
2019 Retrieving Images with Generated Textual Descriptions
abstract
This paper presents a novel remote sensing (RS) image retrieval system that is defined based on generation and exploitation of textual descriptions that model the content of RS images. The proposed RS image retrieval system is composed of three main steps. The first one generates textual descriptions of the content of the RS images combining a convolutional neural network (CNN) and a recurrent neural network (RNN) to extract the features of the images and to generate the descriptions of their content, respectively. The second step encodes the semantic content of the generated descriptions using word embedding techniques able to produce semantically rich word vectors. The third step retrieves the most similar images with respect to the query image by measuring the similarity between the encoded generated textual descriptions of the query image and those of the archive. Experimental results on RS image archive composed of RS images acquired by unmanned aerial vehicles (UAVs) are reported and discussed.
Genc Hoxha, Farid Melgani, Begüm Demir
IGARSS1