Moin Nabi

dblp:167/0748 · DBLP profile ↗
← Back
34ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0001-7559-9888ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 since 2021Artificial intelligence and machine learning · 19 · 10 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
abstract
The generation of toxic content by large language models (LLMs) remains a critical challenge for the safe deployment of language technology.We propose a novel framework for implicit knowledge editing and controlled text generation by fine-tuning LLMs with a prototype-based contrastive perplexity objective.Central to our method is the construction of hard negatives-toxic outputs that are generated through adversarial paraphrasing to be semantically similar and model probability to their non-toxic counterparts.By training on these challenging and realistic pairs, our approach ensures robust and stable contrastive optimization.Experimental results in the domain of detoxification demonstrate that our method significantly reduces toxic generation while maintaining strong performance on downstream tasks such as commonsense reasoning and reading comprehension.Our findings highlight the effectiveness of exploiting hard negatives for attribute-aware fine-tuning. 1
Tassilo Klein, Moin Nabi
ACL (1)2
2025 Multimodal Autoregressive Pre-training of Large Vision Encoders
abstract
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G. T. da Costa, Louis Béthune, Zhe Gan, Alexander Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby
CVPR13
2023 miCSE: Mutual Information Contrastive Learning for Low-shot Sentence Embeddings
abstract
This paper presents miCSE, a mutual information-based contrastive learning framework that significantly advances the stateof-the-art in few-shot sentence embedding.The proposed approach imposes alignment between the attention pattern of different views during contrastive learning.Learning sentence embeddings with miCSE entails enforcing the structural consistency across augmented views for every sentence, making contrastive self-supervised learning more sample efficient.As a result, the proposed approach shows strong performance in the few-shot learning domain.While it achieves superior results compared to state-of-the-art methods on multiple benchmarks in few-shot learning, it is comparable in the full-shot scenario.This study opens up avenues for efficient self-supervised learning methods that are more robust than current contrastive methods for sentence embedding.1
Tassilo Klein, Moin Nabi
ACL (1)2
2023 Semi-supervised learning made simple with self-supervised clustering
abstract
Self-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partially available, motivating a recent line of work on semi-supervised methods inspired by self-supervised principles. In this paper, we propose a conceptually simple yet empirically powerful approach to turn clustering-based self-supervised methods such as SwAV or DINO into semi-supervised learners. More precisely, we introduce a multi-task framework merging a supervised objective using ground-truth labels and a self-supervised objective relying on clustering assignments with a single cross-entropy loss. This approach may be interpreted as imposing the cluster centroids to be class prototypes. Despite its simplicity, we provide empirical evidence that our approach is highly effective and achieves state-of-the-art performance on CI-FAR100 and ImageNet.
Enrico Fini, Pietro Astolfi, Karteek Alahari, Xavier Alameda-Pineda, Julien Mairal, Moin Nabi, Elisa Ricci 0001
CVPR6
2023 A soft nearest-neighbor framework for continual semi-supervised learning
abstract
Despite significant advances, the performance of state-of-the-art continual learning approaches hinges on the unrealistic scenario of fully labeled data. In this paper, we tackle this challenge and propose an approach for continual semi-supervised learning—a setting where not all the data samples are labeled. A primary issue in this scenario is the model forgetting representations of unlabeled data and overfitting the labeled samples. We leverage the power of nearest-neighbor classifiers to nonlinearly partition the feature space and flexibly model the underlying data distribution thanks to its non-parametric nature. This enables the model to learn a strong representation for the current task, and distill relevant information from previous tasks. We perform a thorough experimental evaluation and show that our method outperforms all the existing approaches by large margins, setting a solid state of the art on the continual semi-supervised learning paradigm. For example, on CIFAR-100 we surpass several others even when using at least 30 times less supervision (0.8% vs. 25% of annotations). Finally, our method works well on both low and high resolution images and scales seamlessly to more complex datasets such as ImageNet-100. Our source code is publicly available at https://github.com/kangzhiq/NNCSL.
Zhiqi Kang, Enrico Fini, Moin Nabi, Elisa Ricci 0001, Karteek Alahari
ICCV3
2023 Uncertainty-Aware Contrastive Distillation for Incremental Semantic Segmentation
abstract
A fundamental and challenging problem in deep learning is catastrophic forgetting, i.e., the tendency of neural networks to fail to preserve the knowledge acquired from old tasks when learning new tasks. This problem has been widely investigated in the research community and several Incremental Learning (IL) approaches have been proposed in the past years. While earlier works in computer vision have mostly focused on image classification and object detection, more recently some IL approaches for semantic segmentation have been introduced. These previous works showed that, despite its simplicity, knowledge distillation can be effectively employed to alleviate catastrophic forgetting. In this paper, we follow this research direction and, inspired by recent literature on contrastive learning, we propose a novel distillation framework, Uncertainty-aware Contrastive Distillation (UCD). In a nutshell, UCDis operated by introducing a novel distillation loss that takes into account all the images in a mini-batch, enforcing similarity between features associated to all the pixels from the same classes, and pulling apart those corresponding to pixels from different classes. In order to mitigate catastrophic forgetting, we contrast features of the new model with features extracted by a frozen model learned at the previous incremental step. Our experimental results demonstrate the advantage of the proposed distillation technique, which can be used in synergy with previous IL approaches, and leads to state-of-art performance on three commonly adopted benchmarks for incremental semantic segmentation.
Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Moin Nabi, Xavier Alameda-Pineda, Elisa Ricci 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 solo-learn: A Library of Self-supervised Methods for Visual Representation Learning
abstract
This paper presents solo-learn, a library of self-supervised methods for visual representation learning. Implemented in Python, using Pytorch and Pytorch lightning, the library fits both research and industry needs by featuring distributed training pipelines with mixed-precision, faster data loading via Nvidia DALI, online linear evaluation for better prototyping, and many additional training tricks. Our goal is to provide an easy-to-use library comprising a large amount of Self-supervised Learning (SSL) methods, that can be easily extended and fine-tuned by the community. solo-learn opens up avenues for exploiting large-budget SSL solutions on inexpensive smaller infrastructures and seeks to democratize SSL by making it accessible to all. The source code is available at https://github.com/vturrisi/solo-learn.
Victor G. T. da Costa, Enrico Fini, Moin Nabi, Nicu Sebe, Elisa Ricci 0001
J. Mach. Learn. Res.3
2021 Towards Zero-shot Commonsense Reasoning with Self-supervised Refinement of Language Models
abstract
Can we get existing language models and refine them for zero-shot commonsense reasoning?This paper presents an initial study exploring the feasibility of zero-shot commonsense reasoning for the Winograd Schema Challenge by formulating the task as selfsupervised refinement of a pre-trained language model.In contrast to previous studies that rely on fine-tuning annotated datasets, we seek to boost conceptualization via loss landscape refinement.To this end, we propose a novel self-supervised learning approach that refines the language model utilizing a set of linguistic perturbations of similar concept relationships.Empirical analysis of our conceptually simple framework demonstrates the viability of zero-shot commonsense reasoning on multiple benchmarks. 1
Tassilo Klein, Moin Nabi
EMNLP (1)2
2021 A Unified Objective for Novel Class Discovery
abstract
In this paper, we study the problem of Novel Class Discovery (NCD). NCD aims at inferring novel object categories in an unlabeled set by leveraging from prior knowledge of a labeled set containing different, but related classes. Existing approaches tackle this problem by considering multiple objective functions, usually involving specialized loss terms for the labeled and the unlabeled samples respectively, and often requiring auxiliary regularization terms. In this paper we depart from this traditional scheme and introduce a UNified Objective function (UNO) for discovering novel classes, with the explicit purpose of favoring synergy between supervised and unsupervised learning. Using a multi-view self-labeling strategy, we generate pseudo-labels that can be treated homogeneously with ground truth labels. This leads to a single classification objective operating on both known and unknown classes. De-spite its simplicity, UNO outperforms the state of the art by a significant margin on several benchmarks (≈ +10% on CIFAR-100 and +8% on ImageNet). The project page is available at : https://ncd-uno.github.io.
Enrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong, Moin Nabi, Elisa Ricci 0001
ICCV5
2021 EaSe: A Diagnostic Tool for VQA based on Answer Diversity
abstract
We propose EASE, a simple diagnostic tool for Visual Question Answering (VQA) which quantifies the difficulty of an image, question sample.EASE is based on the pattern of answers provided by multiple annotators to a given question.In particular, it considers two aspects of the answers: (i) their Entropy; (ii) their Semantic content.First, we prove the validity of our diagnostic to identify samples that are easy/hard for state-of-art VQA models.Second, we show that EASE can be successfully used to select the most-informative samples for training/fine-tuning. Crucially, only information that is readily available in any VQA dataset is used to compute its scores. 1
Shailza Jolly, Sandro Pezzelle, Moin Nabi
NAACL-HLT3
2021 Multimodal Prototypical Networks for Few-shot Learning
abstract
Although providing exceptional results for many computer vision tasks, state-of-the-art deep learning algorithms catastrophically struggle in low data scenarios. However, if data in additional modalities exist (e.g. text) this can compensate for the lack of data and improve the classification results. To overcome this data scarcity, we design a cross-modal feature generation framework capable of enriching the low populated embedding space in few-shot scenarios, leveraging data from the auxiliary modality. Specifically, we train a generative model that maps text data into the visual feature space to obtain more reliable prototypes. This allows to exploit data from additional modalities (e.g. text) during training while the ultimate task at test time remains classification with exclusively visual data. We show that in such cases nearest neighbor classification is a viable approach and outperform state-of-the-art single-modal and multimodal few-shot learning methods on the CUB-200 and Oxford-102 datasets.
Frederik Pahde, Mihai Marian Puscas, Tassilo Klein, Moin Nabi
WACV4
2020 Contrastive Self-Supervised Learning for Commonsense Reasoning
abstract
We propose a self-supervised method to solve Pronoun Disambiguation and Winograd Schema Challenge problems.Our approach exploits the characteristic structure of training corpora related to so-called "trigger" words, which are responsible for flipping the answer in pronoun disambiguation.We achieve such commonsense reasoning by constructing pairwise contrastive auxiliary predictions.To this end, we leverage a mutual exclusive loss regularized by a contrastive margin.Our architecture is based on the recently introduced transformer networks, BERT, that exhibits strong performance on many NLP benchmarks.Empirical results show that our method alleviates the limitation of current supervised approaches for commonsense reasoning.This study opens up avenues for exploiting inexpensive self-supervision to achieve performance gain in commonsense reasoning tasks. 1
Tassilo Klein, Moin Nabi
ACL2
2020 Online Continual Learning Under Extreme Memory Constraints
Enrico Fini, Stéphane Lathuilière, Enver Sangineto, Moin Nabi, Elisa Ricci 0001
ECCV (28)4
2020 Human-Machine Collaboration for Medical Image Segmentation
abstract
Image segmentation is a ubiquitous step in almost any medical image study. Deep learning-based approaches achieve state-of-the-art in the majority of image segmentation benchmarks. However, end-to-end training of such models requires sufficient annotation. In this paper, we propose a method based on conditional Generative Adversarial Network (cGAN) to address segmentation in semi-supervised setup and in a human-in-the-loop fashion. More specifically, we use the generator in the GAN to synthesize segmentations on unlabeled data and use the discriminator to identify unreliable slices for which expert annotation is required. The quantitative results on a conventional standard benchmark show that our method is comparable with the state-of-the-art fully supervised methods in slice-level evaluation, despite of requiring far less annotated data.
Mahdyar Ravanbakhsh, Vadim Tschernezki, Felix Last, Tassilo Klein, Kayhan Batmanghelich, Volker Tresp, Moin Nabi
ICASSP7
2019 Attention Is (not) All You Need for Commonsense Reasoning
abstract
The recently introduced BERT model exhibits strong performance on several language understanding benchmarks.In this paper, we describe a simple re-implementation of BERT for commonsense reasoning.We show that the attentions produced by BERT can be directly utilized for tasks such as the Pronoun Disambiguation Problem and Winograd Schema Challenge.Our proposed attention-guided commonsense reasoning method is conceptually simple yet empirically powerful.Experimental analysis on multiple datasets demonstrates that our proposed system performs remarkably well on all cases while outperforming the previously reported state of the art by a margin.While results suggest that BERT seems to implicitly learn to establish complex relationships between entities, solving commonsense reasoning tasks might require more than unsupervised models learned from huge text corpora.
Tassilo Klein, Moin Nabi
ACL (1)2
2019 Learning to Remember: A Synaptic Plasticity Driven Framework for Continual Learning
abstract
Models trained in the context of continual learning (CL) should be able to learn from a stream of data over an undefined period of time. The main challenges herein are: 1) maintaining old knowledge while simultaneously benefiting from it when learning new tasks, and 2) guaranteeing model scalability with a growing amount of data to learn from. In order to tackle these challenges, we introduce Dynamic Generative Memory (DGM) - synaptic plasticity driven framework for continual learning. DGM relies on conditional generative adversarial networks with learnable connection plasticity realized with neural masking. Specifically, we evaluate two variants of neural masking: applied to (i) layer activations and (ii) to connection weights directly. Furthermore, we propose a dynamic network expansion mechanism that ensures sufficient model capacity to accommodate for continually incoming tasks. The amount of added capacity is determined dynamically from the learned binary mask. We evaluate DGM in the continual class-incremental setup on visual classification tasks.
Oleksiy Ostapenko, Mihai Marian Puscas, Tassilo Klein, Patrick Jähnichen, Moin Nabi
CVPR5
2019 Prune Your Neurons Blindly: Neural Network Compression through Structured Class-blind Pruning
abstract
High performance of deep learning models typically comes at cost of considerable model size and computation time. These factors limit applicability for deployment on memory and battery constrained devices such as mobile phones or embedded systems. In this work, we propose a novel pruning technique that eliminates entire filters and neurons according to their relative L1-norm as compared to the rest of the network, yielding more compression and decreased parameters' redundancy. The resulting network is non-sparse, however, much more compact and requires no special infrastructure for deployment. We prove the viability of our method by achieving 97.4%, 86.1%, 47.8% and 53% compression of LeNet-5, VGG-16, ResNet-56 and ResNet-110 respectively, exceeding state-of-the-art compression results reported on VGG-16 and ResNet without losing any performance compared to the baseline. Our approach does not only exhibit good performance but is also easy to implement.
Abdullah Salama, Oleksiy Ostapenko, Tassilo Klein, Moin Nabi
ICASSP4
2019 Budget-Aware Adapters for Multi-Domain Learning
abstract
Multi-Domain Learning (MDL) refers to the problem of learning a set of models derived from a common deep architecture, each one specialized to perform a task in a certain domain (e.g., photos, sketches, paintings). This paper tackles MDL with a particular interest in obtaining domain-specific models with an adjustable budget in terms of the number of network parameters and computational complexity. Our intuition is that, as in real applications the number of domains and tasks can be very large, an effective MDL approach should not only focus on accuracy but also on having as few parameters as possible. To implement this idea we derive specialized deep models for each domain by adapting a pre-trained architecture but, differently from other methods, we propose a novel strategy to automatically adjust the computational complexity of the network. To this aim, we introduce Budget-Aware Adapters that select the most relevant feature channels to better handle data from a novel domain. Some constraints on the number of active switches are imposed in order to obtain a network respecting the desired complexity budget. Experimentally, we show that our approach leads to recognition accuracy competitive with state-of-the-art approaches but with much lighter networks both in terms of storage and computation.
Rodrigo Ferreira Berriel, Stéphane Lathuilière, Moin Nabi, Tassilo Klein, Thiago Oliveira-Santos, Nicu Sebe, Elisa Ricci 0001
ICCV3
2019 Self-Paced Adversarial Training for Multimodal Few-Shot Learning
abstract
State-of-the-art deep learning algorithms yield remark-able results in many visual recognition tasks. However, they still fail to provide satisfactory results in scarce data regimes. To a certain extent this lack of data can be compensated by multimodal information. Missing information in one modality of a single data point (e.g. an image) can be made up for in another modality (e.g. a textual description). Therefore, we design a few-shot learning task that is multimodal during training (i.e. image and text) and single-modal during test time (i.e. image). In this regard, we pro-pose a self-paced class-discriminative generative adversarial network incorporating multimodality in the context off ew-shot learning. The proposed approach builds upon the idea of cross-modal data generation in order to alleviate the data sparsity problem. We improve few-shot learning accuracies on the fine grained CUB and Oxford-102 datasets.
Frederik Pahde, Oleksiy Ostapenko, Patrick Jähnichen, Tassilo Klein, Moin Nabi
WACV5
2019 Low-Shot Learning From Imaginary 3D Model
abstract
Since the advent of deep learning, neural networks have demonstrated remarkable results in many visual recognition tasks, constantly pushing the limits. However, the state-of-the-art approaches are largely unsuitable in scarce data regimes. To address this shortcoming, this paper proposes employing a 3D model, which is derived from training images. Such a model can then be used to hallucinate novel viewpoints and poses for the scarce samples of the few-shot learning scenario. A self-paced learning approach allows for the selection of a diverse set of high-quality images, which facilitates the training of a classifier. The performance of the proposed approach is showcased on the fine-grained CUB-200-2011 dataset in a few-shot setting and significantly improves our baseline accuracy.
Frederik Pahde, Mihai Marian Puscas, Jannik Wolff, Tassilo Klein, Nicu Sebe, Moin Nabi
WACV6
2019 Training Adversarial Discriminators for Cross-Channel Abnormal Event Detection in Crowds
abstract
Abnormal crowd behaviour detection attracts a large interest due to its importance in video surveillance scenarios. However, the ambiguity and the lack of sufficient abnormal ground truth data makes end-to-end training of large deep networks hard in this domain. In this paper we propose to use Generative Adversarial Nets (GANs), which are trained to generate only the normal distribution of the data. During the adversarial GAN training, a discriminator (D) is used as a supervisor for the generator network (G) and vice versa. At testing time we use D to solve our discriminative task (abnormality detection), where D has been trained without the need of manually-annotated abnormal data. Moreover, in order to prevent G learn a trivial identity function, we use a cross-channel approach, forcing G to transform raw-pixel data in motion information and vice versa. The quantitative results on standard benchmarks show that our method outperforms previous state-of-the-art methods in both the frame-level and the pixel-level evaluation.
Mahdyar Ravanbakhsh, Enver Sangineto, Moin Nabi, Nicu Sebe
WACV3
2019 Self Paced Deep Learning for Weakly Supervised Object Detection
abstract
In a weakly-supervised scenario object detectors need to be trained using image-level annotation alone. Since bounding-box-level ground truth is not available, most of the solutions proposed so far are based on an iterative, Multiple Instance Learning framework in which the current classifier is used to select the highest-confidence boxes in each image, which are treated as pseudo-ground truth in the next training iteration. However, the errors of an immature classifier can make the process drift, usually introducing many of false positives in the training dataset. To alleviate this problem, we propose in this paper a training protocol based on the self-paced learning paradigm. The main idea is to iteratively select a subset of images and boxes that are the most reliable, and use them for training. While in the past few years similar strategies have been adopted for SVMs and other classifiers, we are the first showing that a self-paced approach can be used with deep-network-based classifiers in an end-to-end training pipeline. The method we propose is built on the fully-supervised Fast-RCNN architecture and can be applied to similar architectures which represent the input image as a bag of boxes. We show state-of-the-art results on Pascal VOC 2007, Pascal VOC 2010 and ILSVRC 2013. On ILSVRC 2013 our results based on a low-capacity AlexNet network outperform even those weakly-supervised approaches which are based on much higher-capacity networks.
Enver Sangineto, Moin Nabi, Dubravko Culibrk, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Fast but Not Deep: Efficient Crowd Abnormality Detection with Local Binary Tracklets
abstract
In this paper, an efficient method for crowd abnormal behavior detection and localization is introduced. Despite the significant improvements of deep-learning-based methods in this field, but still, they are not fully applicable for the real-time applications. We propose a simple yet effective descriptor based on binary tracklets, containing both orientation and magnitude information in a single feature. The results of the proposed method are comparable with deep-based methods while it performs more efficiently. The evaluation of our descriptors on three different datasets yields a promising result in abnormality detection, which is competitive with the state-of-the-art methods.
Mahdyar Ravanbakhsh, Hossein Mousavi, Moin Nabi, Lucio Marcenaro, Carlo S. Regazzoni
AVSS3
2018 Discriminative Hallucination for Multi-Modal Few-Shot Learning
abstract
State-of-the-art deep learning algorithms yield remarkable results in many visual recognition tasks. However, they still catastrophically struggle in low data scenarios. To a certain extent, the lack of visual data can be compensated by multimodal information. Information missing in one modality (e.g. image) due to the limited data can be included in the other modality (e.g. text). In this paper, we propose a benchmark for few-shot learning with multi-modal data which can be used by other researchers. We also introduced a method for few-shot fine-grained recognition, utilizing textual descriptions of the visual data. We developed a two-stage framework built upon the idea of cross-modal data hallucination. For each visual category, we first generate a set of images by conditioning on the textual description of the category using StackGANs. Next, we rank the generated images based on their class-discriminativeness and only pick the most discriminative images to extend the dataset. Lastly, a classifier invariant to our framework can be trained using an extended training set. We show the results of our proposed discriminative hallucinated method for 1-, 2-, and 5-shot learning on the CUB dataset, where the accuracy is improved by employing the multi-modal data.
Frederik Pahde, Moin Nabi, Tassilo Klein, Patrick Jähnichen
ICIP2
2018 Plug-and-Play CNN for Crowd Motion Analysis: An Application in Abnormal Event Detection
abstract
Most of the crowd abnormal event detection methods rely on complex hand-crafted features to represent the crowd motion and appearance. Convolutional Neural Networks (CNN) have shown to be a powerful instrument with excellent representational capacities, which can leverage the need for hand-crafted features. In this paper, we show that keeping track of the changes in the CNN feature across time can be used to effectively detect local anomalies. Specifically, we propose to measure local abnormality by combining semantic information (inherited from existing CNN models) with low-level optical-flow. One of the advantages of this method is that it can be used without the fine-tuning phase. The proposed method is validated on challenging abnormality detection datasets and the results show the superiority of our approach compared with the state-of-theart methods.
Mahdyar Ravanbakhsh, Moin Nabi, Hossein Mousavi, Enver Sangineto, Nicu Sebe
WACV2
2017 FOIL it! Find One mismatch between Image and Language caption
abstract
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi
ACL (1)5
2017 A cross-modal adaptation approach for brain decoding
abstract
Brain decoding has become a hot topic in many recent brain studies. In a typical neuroimaging experiment, participants are presented with different categories of stimuli while their concurrent brain activity is recorded. Then a classifier is trained on the features extracted from the recorded brain data to discriminate different target stimuli classes. It is a common practice to hypothesize that the stimulus-related information exists in the brain data if the decoder can accurately predict the target stimulus category. However, most of the neuroimaging studies suffer from few and noisy samples. These constraints affects the performance of such decoding systems. In order to cope with this limitation, a dictionary learning approach is used in this paper to transfer knowledge from the multimedia domain to the brain domain. We show that such cross-modal domain adaptation yields better performance of the learning algorithm in the brain domain. This is the first study in the direction of cross-modal adaptation by joint dictionary learning on multimedia and brain modality.
Pouya Ghaemmaghami, Moin Nabi, Yan Yan 0002, Giuseppe Riccardi, Nicu Sebe
ICASSP2
2017 Abnormal event detection in videos using generative adversarial nets
abstract
In this paper we address the abnormality detection problem in crowded scenes. We propose to use Generative Adversarial Nets (GANs), which are trained using normal frames and corresponding optical-flow images in order to learn an internal representation of the scene normality. Since our GANs are trained with only normal data, they are not able to generate abnormal events. At testing time the real data are compared with both the appearance and the motion representations reconstructed by our GANs and abnormal areas are detected by computing local differences. Experimental results on challenging abnormality detection datasets show the superiority of the proposed method compared to the state of the art in both frame-level and pixel-level abnormality detection tasks.
Mahdyar Ravanbakhsh, Moin Nabi, Enver Sangineto, Lucio Marcenaro, Carlo S. Regazzoni, Nicu Sebe
ICIP2
2017 Autonomous Crowdsourcing through Human-Machine Collaborative Learning
abstract
In this paper, we introduce a general iterative human-machine collaborative method for training crowdsource workers: the classifier (i.e., the machine) selects the highest quality examples for training the crowdsource workers (i.e., the humans). Then, the latter annotate the lower quality examples such that the classifier can be re-trained with more accurate examples. This process can be iterated several times. We tested our approach on two different tasks, Relation Extraction and Community Question Answering, which are also in two different languages, English and Arabic, respectively. Our experimental results show a significant improvement for creating Gold Standard data over distant supervision or just crowdsourcing without worker training. At the same time, our method approach the performance than state-of-the-art methods using expensive Gold Standard for training workers
Azad Abad, Moin Nabi, Alessandro Moschitti
SIGIR2
2016 Novel dataset for fine-grained abnormal behavior understanding in crowd
abstract
Despite the huge research on crowd on behavior understanding in visual surveillance community, lack of publicly available realistic datasets for evaluating crowd behavioral interaction led not to have a fair common test bed for researchers to compare the strength of their methods in the real scenarios. This work presents a novel crowd dataset contains around 45,000 video clips which annotated by one of the five different fine-grained abnormal behavior categories. We also evaluated two state-of-the-art methods on our dataset, showing that our dataset can be effectively used as a benchmark for fine-grained abnormality detection. The details of the dataset and the results of the baseline methods are presented in the paper.
Javad Haddadnia, Hossein Mousavi, Maziyar Kalantarzadeh, Moin Nabi, Vittorio Murino
AVSS5
2016 CNN-aware binary MAP for general semantic segmentation
abstract
In this paper we introduce a novel method for general semantic segmentation that can benefit from general semantics of Convolutional Neural Network (CNN). Our segmentation proposes visually and semantically coherent image segments. We use binary encoding of CNN features to overcome the difficulty of the clustering on the high-dimensional CNN feature space. These binary codes are very robust against noise and non-semantic changes in the image. These binary encoding can be embedded into the CNN as an extra layer at the end of the network. This results in real-time segmentation. To the best of our knowledge our method is the first attempt on general semantic image segmentation using CNN. All the previous papers were limited to few number of category of the images (e.g. PASCAL VOC). Experiments show that our segmentation algorithm outperform the state-of-the-art non-semantic segmentation methods by large margin.
Mahdyar Ravanbakhsh, Hossein Mousavi, Moin Nabi, Mohammad Rastegari, Carlo S. Regazzoni
ICIP3
2016 Sparse-coded cross-domain adaptation from the visual to the brain domain
abstract
Brain decoding (i.e., retrieving information from brain signals by employing machine learning algorithms) has recently received considerable attention across many communities. In a typical brain decoding paradigm, different types of stimuli are shown to the participant of the neuroimaging experiment, while his/her concurrent brain activity is captured using neuroimaging techniques. Then a machine learning algorithm is employed to categorize the measured brain signal into the target stimuli classes. Accurate prediction of the stimulus category by the algorithm is considered a positive evidence of the hypothesis of the existence of stimulus-related information in brain data. However, most of the brain decoding studies suffer from the constraint of having few and noisy samples. In order to overcome this limitation, in this paper, an adaptation paradigm is employed in order to transfer knowledge from visual domain to brain domain. We experimentally show that such adaptation procedure leads to improved results for the object recognition task in the brain domain, outperforming significantly the results achieved by the brain features alone. This is the first study in the direction of transferring knowledge by adapting representations learned on visual domain to the brain modality. We believe this paper opens up avenues for exploiting large-scale visual datasets to achieve performance gain in brain decoding.
Pouya Ghaemmaghami, Moin Nabi, Yan Yan 0002, Nicu Sebe
ICPR2
2015 Learning with dataset bias in latent subcategory models
abstract
Latent subcategory models (LSMs) offer significant improvements over training linear support vector machines (SVMs). Training LSMs is a challenging task due to the potentially large number of local optima in the objective function and the increased model complexity which requires large training set sizes. Often, larger datasets are available as a collection of heterogeneous datasets. However, previous work has highlighted the possible danger of simply training a model from the combined datasets, due to the presence of bias. In this paper, we present a model which jointly learns an LSM for each dataset as well as a compound LSM. The method provides a means to borrow statistical strength from the datasets while reducing their inherent bias. In experiments we demonstrate that the compound LSM, when tested on PASCAL, LabelMe, Caltech101 and SUN09 in a leave-one-dataset-out fashion, achieves an average improvement of over 6.5% over a previous SVM-based undoing bias approach and an average improvement of over 8.5% over a standard LSM trained on the concatenation of the datasets.
Dimitris Stamos, Samuele Martelli, Moin Nabi, Andrew M. McDonald, Vittorio Murino, Massimiliano Pontil
CVPR3
2015 Crowd motion monitoring using tracklet-based commotion measure
abstract
Abnormal detection in crowd is a challenging vision task due to the scarcity of real-world training examples and the lack of a clear definition of abnormality. To tackle these challenges, we propose a novel measure to capture the commotion of a crowd motion for the task of abnormality detection in crowd. The unsupervised nature of the proposed measure allows to detect abnormality adaptively (i.e. context dependent) with no training cost. The extensive experiments on three different levels (e.g. pixel, frame and video) show the superiority of the proposed approach compared to the state of the arts.
Hossein Mousavi, Moin Nabi, Hamed Kiani Galoogahi, Alessandro Perina, Vittorio Murino
ICIP2