Mahdyar Ravanbakhsh

dblp:173/5349 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
9since 2021 · last 2024
0000-0002-6456-867XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 Multi-Label Noise Robust Collaborative Learning for Remote Sensing Image Classification
abstract
The development of accurate methods for multi-label classification (MLC) of remote sensing (RS) images is one of the most important research topics in RS. The MLC methods based on convolutional neural networks (CNNs) have shown strong performance gains in RS. However, they usually require a high number of reliable training images annotated with multiple land-cover class labels. Collecting such data is time-consuming and costly. To address this problem, the publicly available thematic products, which can include noisy labels, can be used to annotate RS images with zero-labeling cost. However, multi-label noise (which can be associated with wrong and missing label annotations) can distort the learning process of the MLC methods. To address this problem, we propose a novel multi-label noise robust collaborative learning (RCML) method to alleviate the negative effects of multi-label noise during the training phase of a CNN model. RCML identifies, ranks, and excludes noisy multi-labels in RS images based on three main modules: 1) the discrepancy module; 2) the group lasso module; and 3) the swap module. The discrepancy module ensures that the two networks learn diverse features, while producing the same predictions. The task of the group lasso module is to detect the potentially noisy labels assigned to multi-labeled training images, while the swap module is devoted to exchange the ranking information between two networks. Unlike the existing methods that make assumptions about noise distribution, our proposed RCML does not make any prior assumption about the type of noise in the training set. The experiments conducted on two multi-label RS image archives confirm the robustness of the proposed RCML under extreme multi-label noise rates. Our code is publicly available at: https://www.noisy-labels-in-rs.org.
Ahmet Kerem Aksoy, Mahdyar Ravanbakhsh, Begüm Demir
IEEE Trans. Neural Networks Learn. Syst.2
2023 LIT-4-RSVQA: Lightweight Transformer-Based Visual Question Answering in Remote Sensing
abstract
Visual question answering (VQA) methods in remote sensing (RS) aim to answer natural language questions with respect to an RS image. Most of the existing methods require a large amount of computational resources, which limits their application in operational scenarios in RS. To address this issue, in this paper we present an effective lightweight transformer-based VQA in RS (LiT-4-RSVQA) architecture for efficient and accurate VQA in RS. Our architecture consists of: i) a lightweight text encoder module; ii) a lightweight image encoder module; iii) a fusion module; and iv) a classification module. The experimental results obtained on a VQA benchmark dataset demonstrate that our proposed LiT-4-RSVQA architecture provides accurate VQA results while significantly reducing the computational requirements on the executing hardware.
Leonard W. Hackel, Kai Norman Clasen, Mahdyar Ravanbakhsh, Begüm Demir
IGARSS3
2022 Unsupervised Contrastive Hashing for Cross-Modal Retrieval in Remote Sensing
abstract
The development of cross-modal retrieval systems that can search and retrieve semantically relevant data across different modalities based on a query in any modality has attracted great attention in remote sensing (RS). In this paper, we focus our attention on cross-modal text-image retrieval, where queries from one modality (e.g., text) can be matched to archive entries from another (e.g., image). Most of the existing cross-modal text-image retrieval systems in RS require a high number of labeled training samples and also do not allow fast and memory-efficient retrieval. These issues limit the applicability of the existing cross-modal retrieval systems for large-scale applications in RS. To address this problem, in this paper we introduce a novel unsupervised cross-modal contrastive hashing (DUCH) method for text-image retrieval in RS. To this end, the proposed DUCH is made up of two main modules: 1) feature extraction module, which extracts deep representations of two modalities; 2) hashing module that learns to generate cross-modal binary hash codes from the extracted representations. We introduce a novel multi-objective loss function including: i) contrastive objectives that enable similarity preservation in intra- and inter-modal similarities; ii) an adversarial objective that is enforced across two modalities for cross-modal representation consistency; and iii) binarization objectives for generating hash codes. Experimental results show that the proposed DUCH outperforms state-of-the-art methods. Our code is publicly available at https://git.tu-berlin.de/rsim/duch.
Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir
ICASSP2
2022 An Unsupervised Cross-Modal Hashing Method Robust to Noisy Training Image-Text Correspondences in Remote Sensing
abstract
The development of accurate and scalable cross-modal image-text retrieval methods, where queries from one modality (e.g., text) can be matched to archive entries from another (e.g., remote sensing image) has attracted great attention in remote sensing (RS). Most of the existing methods assume that a reliable multi-modal training set with accurately matched text-image pairs is existing. However, this assumption may not always hold since the multi-modal training sets may include noisy pairs (i.e., textual descriptions/captions associated to training images can be noisy), distorting the learning process of the retrieval methods. To address this problem, we propose a novel unsupervised cross-modal hashing method robust to the noisy image-text correspondences (CHNR). CHNR consists of three modules: 1) feature extraction module, which extracts feature representations of image-text pairs; 2) noise detection module, which detects potential noisy correspondences; and 3) hashing module that generates cross-modal binary hash codes. The proposed CHNR includes two training phases: i) meta-learning phase that uses a small portion of clean (i.e., reliable) data to train the noise detection module in an adversarial fashion; and ii) the main training phase for which the trained noise detection module is used to identify noisy correspondences while the hashing module is trained on the noisy multi-modal training set. Experimental results show that the proposed CHNR outperforms state-of-the-art methods.
Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir
ICIP2
2022 Deep Learning Driven Content-Based Image Time-Series Retrieval in Remote Sensing Archives
abstract
The rapid evolution of satellite imaging systems has resulted in sharp increases of image archive volumes. Multitemporal images constitute a sizeable portion of these time-series databases. Accordingly, development of accurate content based time-series retrieval (CBTSR) methods in massive archives of RS images attracts much research interest. Given a user-defined query time series, CBTSR aims at identifying within a massive archive image time series that show characteristics similar to those of the query time series. In this paper, we focus our attention to CBTSR in pairs of RS images, aiming to search and retrieve bi-temporal image pairs containing changes similar to those modeled in the query. To this end, we introduce two deep learning-based methods in the framework of CBTSR. The first method, called deep change vector retrieval (DVCR), is based on selected deep features extracted from the change vector analysis. The second method, called autoencoder with early fusion (AEEF) uses an autoencoder architecture to recreate the time difference images and the latent codes produced by this network. Experimental results show the effectiveness of the proposed methods for CBTSR problems. The code of the proposed methods is available at: https://github.com/OnatV/ChangeRetrieval.
Onat Vuran, Oguzhan Akcin, Mahdyar Ravanbakhsh, Bülent Sankur, Begüm Demir
IGARSS3
2022 On the Effects of Different Types of Label Noise in Multi-Label Remote Sensing Image Classification
abstract
The development of accurate methods for multi-label classification (MLC) of remote sensing (RS) images is one of the most important research topics in RS. To address MLC problems, the use of deep neural networks that require a high number of reliable training images annotated by multiple land-cover class labels (multi-labels) has been found popular in RS. However, collecting such annotations is time-consuming and costly. A common procedure to obtain annotations at zero labeling cost is to rely on thematic products or crowdsourced labels. As a drawback, these procedures come with the risk of label noise that can distort the learning process of the MLC algorithms. In the literature, most label noise robust methods are designed for single-label classification (SLC) problems in computer vision (CV), where each image is annotated by a single label. Unlike SLC, label noise in MLC can be associated with: 1) subtractive label noise (a land cover class label is not assigned to an image while that class is present in the image); 2) additive label noise (a land cover class label is assigned to an image, although that class is not present in the given image); and 3) mixed label noise (a combination of both). In this paper, we investigate three different noise robust CV SLC methods (Self-Adaptive Training, Early-Learning Regularization, and Joint Co-Regularized Training) and adapt them to be robust for multi-label noise scenarios in RS. During experiments, we study the effects of different types of multi-label noise and evaluate the adapted methods rigorously. To this end, we also introduce a synthetic multi-label noise injection strategy that is more adequate to simulate operational scenarios compared to the uniform label noise injection strategy, in which the labels of absent and present classes are flipped at uniform probability. Further, we study the relevance of different evaluation metrics in MLC problems under noisy multi-labels. On the basis of the theoretical and experimental analyses, some guidelines for a proper design of label noise robust MLC methods are derived.
Tom Burgert, Mahdyar Ravanbakhsh, Begüm Demir
IEEE Trans. Geosci. Remote. Sens.2
2022 Informative and Representative Triplet Selection for Multilabel Remote Sensing Image Retrieval
abstract
Learning the similarity between remote sensing (RS) images forms the foundation for content-based RS image retrieval (CBIR). Recently, deep metric learning approaches that map the semantic similarity of images into an embedding (metric) space have been found very popular in RS. A common approach for learning the metric space relies on the selection of triplets of similar (positive) and dissimilar (negative) images to a reference image called as an anchor. Choosing triplets is a difficult task particularly for multi-label RS CBIR, where each training image is annotated by multiple class labels. To address this problem, in this paper we propose a novel triplet sampling method in the framework of deep neural networks (DNNs) defined for multi-label RS CBIR problems. The proposed method selects a small set of the most representative and informative triplets based on two main steps. In the first step, a set of anchors that are diverse to each other in the embedding space is selected from the current mini-batch using an iterative algorithm. In the second step, different sets of positive and negative images are chosen for each anchor by evaluating the relevancy, hardness and diversity of the images among each other based on a novel strategy. Experimental results obtained on two multi-label benchmark archives show that the selection of the most informative and representative triplets in the context of DNNs results in: i) reducing the computational complexity of the training phase of the DNNs without any significant loss on the performance; and ii) an increase in learning speed since informative triplets allow fast convergence. The code of the proposed method is publicly available at https://git.tu-berlin.de/rsim/image-retrieval-from-triplets.
Gencer Sumbul, Mahdyar Ravanbakhsh, Begüm Demir
IEEE Trans. Geosci. Remote. Sens.2
2021 A Consensual Collaborative Learning Method for Remote Sensing Image Classification Under Noisy Multi-Labels
abstract
Collecting a large number of reliable training images annotated by multiple land-cover class labels in the framework of multi-label classification is time-consuming and costly in remote sensing (RS). To address this problem, publicly available thematic products are often used for annotating RS images with zero-labeling-cost. However, such an approach may result in constructing a training set with noisy multi-labels, distorting the learning process. To address this problem, we propose a Consensual Collaborative Multi-Label Learning (CCML) method. The proposed CCML identifies, ranks and corrects training images with noisy multi-labels through four main modules: 1) discrepancy module; 2) group lasso module; 3) flipping module; and 4) swap module. The discrepancy module ensures that the two networks learn diverse features, while obtaining the same predictions. The group lasso module detects the potentially noisy labels by estimating the label uncertainty based on the aggregation of two collaborative networks. The flipping module corrects the identified noisy labels, whereas the swap module exchanges the ranking information between the two networks. The experimental results confirm the success of the proposed CCML under high (synthetically added) multi-label noise rates. The code of the proposed method is publicly available at https://noisy-labels-in-rs.org.
Ahmet Kerem Aksoy, Mahdyar Ravanbakhsh, Tristan Kreuziger, Begüm Demir
ICIP2
2021 Learning Self-Awareness for Autonomous Vehicles: Exploring Multisensory Incremental Models
abstract
The technology for autonomous vehicles is close to replacing human drivers by artificial systems endowed with high-level decision-making capabilities. In this regard, systems must learn about the usual vehicle's behavior to predict imminent difficulties before they happen. An autonomous agent should be capable of continuously interacting with multi-modal dynamic environments while learning unseen novel concepts. Such environments are not often available to train the agent on it, so the agent should have an understanding of its own capacities and limitations. This understanding is usually called self-awareness. This paper proposes a multi-modal self-awareness modeling of signals coming from different sources. This paper shows how different machine learning techniques can be used under a generic framework to learn single modality models by using Dynamic Bayesian Networks. In the presented case, a probabilistic switching model and a bank of generative adversarial networks are employed to model a vehicle's positional and visual information respectively. Our results include experiments performed on a real vehicle, highlighting the potentiality of the proposed approach at detecting abnormalities in real scenarios.
Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marín-Plaza, David Martín 0001, Lucio Marcenaro, Carlo S. Regazzoni
IEEE Trans. Intell. Transp. Syst.1
2020 Human-Machine Collaboration for Medical Image Segmentation
abstract
Image segmentation is a ubiquitous step in almost any medical image study. Deep learning-based approaches achieve state-of-the-art in the majority of image segmentation benchmarks. However, end-to-end training of such models requires sufficient annotation. In this paper, we propose a method based on conditional Generative Adversarial Network (cGAN) to address segmentation in semi-supervised setup and in a human-in-the-loop fashion. More specifically, we use the generator in the GAN to synthesize segmentations on unlabeled data and use the discriminator to identify unreliable slices for which expert annotation is required. The quantitative results on a conventional standard benchmark show that our method is comparable with the state-of-the-art fully supervised methods in slice-level evaluation, despite of requiring far less annotated data.
Mahdyar Ravanbakhsh, Vadim Tschernezki, Felix Last, Tassilo Klein, Kayhan Batmanghelich, Volker Tresp, Moin Nabi
ICASSP1
2020 S2-CGAN: Self-Supervised Adversarial Representation Learning for Binary Change Detection in Multispectral Images
abstract
Deep Neural Networks have recently demonstrated promising performance in binary change detection (CD) problems in remote sensing (RS), requiring a large amount of labeled multitemporal training samples. Since collecting such data is time-consuming and costly, most of the existing methods rely on pre-trained networks on publicly available computer vision (CV) datasets. However, because of the differences in image characteristics in CV and RS, this approach limits the performance of the existing CD methods. To address this problem, we propose a self-supervised conditional Generative Adversarial Network (S2-cGAN). The proposed S2-cGAN is trained to generate only the distribution of unchanged samples. To this end, the proposed method consists of two main steps: 1) Generating a reconstructed version of the input image as an unchanged image 2) Learning the distribution of unchanged samples through an adversarial game. Unlike the existing GAN based methods (which only use the discriminator during the adversarial training to supervise the generator), the S2-cGAN directly exploits the discriminator likelihood to solve the binary CD task. Experimental results show the effectiveness of the proposed S2-cGAN when compared to the state of the art CD methods.
Jose Luis Holgado Alvarez, Mahdyar Ravanbakhsh, Begüm Demir
IGARSS2
2019 Training Adversarial Discriminators for Cross-Channel Abnormal Event Detection in Crowds
abstract
Abnormal crowd behaviour detection attracts a large interest due to its importance in video surveillance scenarios. However, the ambiguity and the lack of sufficient abnormal ground truth data makes end-to-end training of large deep networks hard in this domain. In this paper we propose to use Generative Adversarial Nets (GANs), which are trained to generate only the normal distribution of the data. During the adversarial GAN training, a discriminator (D) is used as a supervisor for the generator network (G) and vice versa. At testing time we use D to solve our discriminative task (abnormality detection), where D has been trained without the need of manually-annotated abnormal data. Moreover, in order to prevent G learn a trivial identity function, we use a cross-channel approach, forcing G to transform raw-pixel data in motion information and vice versa. The quantitative results on standard benchmarks show that our method outperforms previous state-of-the-art methods in both the frame-level and the pixel-level evaluation.
Mahdyar Ravanbakhsh, Enver Sangineto, Moin Nabi, Nicu Sebe
WACV1
2018 Fast but Not Deep: Efficient Crowd Abnormality Detection with Local Binary Tracklets
abstract
In this paper, an efficient method for crowd abnormal behavior detection and localization is introduced. Despite the significant improvements of deep-learning-based methods in this field, but still, they are not fully applicable for the real-time applications. We propose a simple yet effective descriptor based on binary tracklets, containing both orientation and magnitude information in a single feature. The results of the proposed method are comparable with deep-based methods while it performs more efficiently. The evaluation of our descriptors on three different datasets yields a promising result in abnormality detection, which is competitive with the state-of-the-art methods.
Mahdyar Ravanbakhsh, Hossein Mousavi, Moin Nabi, Lucio Marcenaro, Carlo S. Regazzoni
AVSS1
2018 Learning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving
abstract
This paper presents a novel approach for learning self-awareness models for autonomous vehicles. Proposed technique is based on the availability of synchronized multi-sensor dynamic data related to different maneuvering tasks performed by a human operator. It is shown that different machine learning approaches can be used to first learn single modality models using coupled Dynamic Bayesian Networks; such models are then correlated at event level to discover contextual multimodal concepts. In the presented case, visual perception and localization are used as modalities. Cross-correlations among modalities in time is discovered from data and are described as probabilistic links connecting shared and private multi-modal DBNs at the event (discrete) level. Results are presented on experiments performed on an autonomous vehicle, highlighting potentiality of the proposed approach to allow anomaly detection and autonomous decision making based on learned self-awareness models.
Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marín-Plaza, David Martín 0001, Lucio Marcenaro, Carlo S. Regazzoni
FUSION1
2018 a Multi-Perspective Approach to Anomaly Detection for Self -Aware Embodied Agents
abstract
This paper focuses on multi-sensor anomaly detection for moving cognitive agents using both external and private first-person visual observations. Both observation types are used to characterize agents motion in a given environment. The proposed method generates locally uniform motion models by dividing a Gaussian process that approximates agents displacements on the scene and provides a Shared Level (SL) self-awareness based on Environment Centered (EC) models. Such models are then used to train in a semi-unsupervised way a set of Generative Adversarial Networks (GANs) that produce an estimation of external and internal parameters of moving agents. Obtained results exemplify the feasibility of using multi-perspective data for predicting and analyzing trajectory information.
Mohamad Baydoun, Mahdyar Ravanbakhsh, Damian Campo, Pablo Marín-Plaza, David Martín 0001, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni
ICASSP2
2018 Hierarchy of Gans for Learning Embodied Self-Awareness Model
abstract
In recent years several architectures have been proposed to learn embodied agents complex self-awareness models. In this paper, dynamic incremental self-awareness (SA) models are proposed that allow experiences done by an agent to be modeled in a hierarchical fashion, starting from more simple situations to more structured ones. Each situation is learned from subsets of private agent perception data as a model capable to predict normal behaviors and detect abnormalities. Hierarchical SA models have been already proposed using low dimensional sensorial inputs. In this work, a hierarchical model is introduced by means of a cross-modal Generative Adversarial Networks (GANs) processing high dimensional visual data. Different levels of the GANs are detected in a self-supervised manner using GANs discriminators decision boundaries. Real experiments on semi-autonomous ground vehicles are presented.
Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marín-Plaza, David Martín 0001, Lucio Marcenaro, Carlo S. Regazzoni
ICIP1
2018 Plug-and-Play CNN for Crowd Motion Analysis: An Application in Abnormal Event Detection
abstract
Most of the crowd abnormal event detection methods rely on complex hand-crafted features to represent the crowd motion and appearance. Convolutional Neural Networks (CNN) have shown to be a powerful instrument with excellent representational capacities, which can leverage the need for hand-crafted features. In this paper, we show that keeping track of the changes in the CNN feature across time can be used to effectively detect local anomalies. Specifically, we propose to measure local abnormality by combining semantic information (inherited from existing CNN models) with low-level optical-flow. One of the advantages of this method is that it can be used without the fine-tuning phase. The proposed method is validated on challenging abnormality detection datasets and the results show the superiority of our approach compared with the state-of-theart methods.
Mahdyar Ravanbakhsh, Moin Nabi, Hossein Mousavi, Enver Sangineto, Nicu Sebe
WACV1
2017 Abnormal event detection in videos using generative adversarial nets
abstract
In this paper we address the abnormality detection problem in crowded scenes. We propose to use Generative Adversarial Nets (GANs), which are trained using normal frames and corresponding optical-flow images in order to learn an internal representation of the scene normality. Since our GANs are trained with only normal data, they are not able to generate abnormal events. At testing time the real data are compared with both the appearance and the motion representations reconstructed by our GANs and abnormal areas are detected by computing local differences. Experimental results on challenging abnormality detection datasets show the superiority of the proposed method compared to the state of the art in both frame-level and pixel-level abnormality detection tasks.
Mahdyar Ravanbakhsh, Moin Nabi, Enver Sangineto, Lucio Marcenaro, Carlo S. Regazzoni, Nicu Sebe
ICIP1
2016 CNN-aware binary MAP for general semantic segmentation
abstract
In this paper we introduce a novel method for general semantic segmentation that can benefit from general semantics of Convolutional Neural Network (CNN). Our segmentation proposes visually and semantically coherent image segments. We use binary encoding of CNN features to overcome the difficulty of the clustering on the high-dimensional CNN feature space. These binary codes are very robust against noise and non-semantic changes in the image. These binary encoding can be embedded into the CNN as an extra layer at the end of the network. This results in real-time segmentation. To the best of our knowledge our method is the first attempt on general semantic image segmentation using CNN. All the previous papers were limited to few number of category of the images (e.g. PASCAL VOC). Experiments show that our segmentation algorithm outperform the state-of-the-art non-semantic segmentation methods by large margin.
Mahdyar Ravanbakhsh, Hossein Mousavi, Moin Nabi, Mohammad Rastegari, Carlo S. Regazzoni
ICIP1