Nikhil Rasiwasia

dblp:82/335 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
1since 2021 · last 2021
0000-0002-7046-851XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Image recognition and object detection · 35% Representation and self-supervised learning · 29% Information extraction and text analysis · 14%
Databases, data mining, and information retrieval
6 papers
Information retrieval · 52% Data mining · 40% Data models and query languages · 7%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
scene recognition
0.752015
Scene classification with semantic Fisher vectors · CVPR 2015
Holistic Context Models for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Scene Recognition on the Semantic Manifold · ECCV (4) 2012
Computer vision › Image recognition and object detection
image classification
0.532013
Latent Dirichlet Allocation Models for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Class-Specific Simplex-Latent Dirichlet Allocation for Image Classification · ICCV 2013
Adapted Gaussian models for image classification · CVPR 2011
Multimedia analysis and retrieval
cross-modal retrieval
0.422015
Multi-label Cross-Modal Retrieval · ICCV 2015
On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Natural language and speech › Question answering and dialogue systems › community question answering
answer selection
0.412019
Improving Answer Selection and Answer Triggering using Hard Negatives · EMNLP/IJCNLP (1) 2019
Information retrieval
retrieval models
0.322014
On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Bridging the Gap: Query by Semantic Example · IEEE Trans. Multim. 2007
Machine learning › Representation and self-supervised learning › multi-view learning
canonical correlation analysis
0.212015
Multi-label Cross-Modal Retrieval · ICCV 2015
Machine learning › Representation and self-supervised learning
fisher vector
0.212015
Scene classification with semantic Fisher vectors · CVPR 2015
Machine learning › Representation and self-supervised learning › visual representation
image representation
0.212015
Scene classification with semantic Fisher vectors · CVPR 2015
Multimedia analysis and retrieval › cross-modal retrieval
multi-label cross-modal retrieval
0.212015
Multi-label Cross-Modal Retrieval · ICCV 2015
Information retrieval
image retrieval
0.222013
Holistic Context Models for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Latent Dirichlet Allocation Models for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation
0.212013
Latent Dirichlet Allocation Models for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Natural language and speech › Information extraction and text analysis › topic model
supervised topic model
0.212013
Latent Dirichlet Allocation Models for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Natural language and speech › Information extraction and text analysis
topic model
0.212013
Class-Specific Simplex-Latent Dirichlet Allocation for Image Classification · ICCV 2013
Computer vision › Segmentation and scene understanding
context modeling
0.112012
Holistic Context Models for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Machine learning › Representation and self-supervised learning
semantic manifold
0.112012
Scene Recognition on the Semantic Manifold · ECCV (4) 2012
Data mining › predictive modeling
classification
0.112011
Improving Product Classification Using Images · ICDM 2011
Data mining › predictive modeling › classification
ensemble learning
0.112011
Improving Product Classification Using Images · ICDM 2011
Data mining › text mining › text classification
product classification
0.112011
Improving Product Classification Using Images · ICDM 2011
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
answer triggering
0.112019
Improving Answer Selection and Answer Triggering using Hard Negatives · EMNLP/IJCNLP (1) 2019
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.112010
A new approach to cross-modal multimedia retrieval · ACM Multimedia 2010
Computer vision › Image recognition and object detection
image annotation
0.112009
Holistic context modeling using semantic co-occurrences · CVPR 2009
Computer vision › Vision and language
visual context modeling
0.112009
Holistic context modeling using semantic co-occurrences · CVPR 2009
Data models and query languages › query interface
query by example
0.112007
Bridging the Gap: Query by Semantic Example · IEEE Trans. Multim. 2007
Information retrieval › search engines
semantic search
0.112007
Bridging the Gap: Query by Semantic Example · IEEE Trans. Multim. 2007
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
bag-of-features
0.012012
Holistic Context Models for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Data mining › text mining › topic model
latent topic models
0.012010
A new approach to cross-modal multimedia retrieval · ACM Multimedia 2010
Machine learning › Learning paradigms
weakly supervised learning
0.012008
Scene classification with low-dimensional semantic spaces and weak supervision · CVPR 2008

Methods — techniques the papers use, named apart from their topics

latent dirichlet allocation · 0.8canonical correlation analysis · 0.7subspace learning · 0.4hard negative mining · 0.4semantic matching · 0.4correlation matching · 0.4gaussian mixture model · 0.3class-specific simplex · 0.3support vector machine · 0.3fisher vector · 0.2dirichlet mixture · 0.2convolutional neural network · 0.2topic supervision · 0.2semantic space representation · 0.1probabilistic co-occurrence modeling · 0.1appearance classifiers · 0.1probabilistic fusion · 0.1confusion matrix · 0.1
YearPublicationVenuePosition
2021 Distantly Supervised Transformers For E-Commerce Product QA
abstract
Happy Mittal, Aniket Chakrabarti, Belhassen Bayar, Animesh Anant Sharma, Nikhil Rasiwasia. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Happy Mittal, Aniket Chakrabarti, Belhassen Bayar, Animesh Anant Sharma, Nikhil Rasiwasia
NAACL-HLT5
2019 Improving Answer Selection and Answer Triggering using Hard Negatives
abstract
Sawan Kumar, Shweta Garg, Kartik Mehta, Nikhil Rasiwasia. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sawan Kumar, Shweta Garg 0001, Kartik Mehta, Nikhil Rasiwasia
EMNLP/IJCNLP (1)4
2015 Scene classification with semantic Fisher vectors
abstract
With the help of a convolutional neural network (CNN) trained to recognize objects, a scene image is represented as a bag of semantics (BoS). This involves classifying image patches using the network and considering the class posterior probability vectors as locally extracted semantic descriptors. The image BoS is summarized using a Fisher vector (FV) embedding that exploits the properties of the space of these descriptors. The resulting representation is referred to as a semantic Fisher vector. Two implementations of a semantic FV are investigated. First involves modeling the BoS with a Dirichlet Mixture and computing the Fisher gradients for this model. Due to the difficulty of mixture modeling on a non-Euclidean probability simplex, this approach is shown to be unsuccessful. A second implementation is derived using the interpretation of semantic descriptors as parameters of a multinomial distribution. Like the parameters of any exponential family, these can be projected into their natural parameter space. For a CNN, this is shown equivalent to using inputs of its soft-max layer as patch descriptors. A semantic FV is then computed as a Gaussian Mixture FV in the space of these natural parameters. This representation is shown to outperform other alternatives such as FVs of features from the intermediate CNN layers or a classifier obtained by adapting (fine-tuning) the CNN. The proposed FV represents an embedding for object classification probabilities. As an image representation, therefore, it is complementary to the features obtained from a scene classification CNN. A combination of the two representations is shown to achieve state-of-the-art results on MIT Indoor scenes and SUN datasets.
Mandar Dixit, Dashan Gao 0001, Nikhil Rasiwasia, Nuno Vasconcelos
CVPR4
2015 Multi-label Cross-Modal Retrieval
abstract
In this work, we address the problem of cross-modal retrieval in presence of multi-label annotations. In particular, we introduce multi-label Canonical Correlation Analysis (ml-CCA), an extension of CCA, for learning shared subspaces taking into account high level semantic information in the form of multi-label annotations. Unlike CCA, ml-CCA does not rely on explicit pairing between modalities, instead it uses the multi-label information to establish correspondences. This results in a discriminative subspace which is better suited for cross-modal retrieval tasks. We also present Fast ml-CCA, a computationally efficient version of ml-CCA, which is able to handle large scale datasets. We show the efficacy of our approach by conducting extensive cross-modal retrieval experiments on three standard benchmark datasets. The results show that the proposed approach achieves state of the art retrieval performance on the three datasets.
Viresh Ranjan, Nikhil Rasiwasia, C. V. Jawahar
ICCV2
2015 Combining deep learning and unsupervised clustering to improve scene recognition performance
abstract
Deep Neural Networks (DNN) are now the state-of-the-art for many image and object recognition tasks, as illustrated by their performance on standard benchmarks. The success of DNNs is attributed to their ability to learn rich mid-level image representations, as opposed to hand-designed low-level features used in other image analysis methods. Typically a large dataset of unlabeled images is used for unsupervised feature learning, and then standard classifiers are trained on the features extracted from the images in a labeled set. In this paper, we show that clustering the images using the features from the DNN allows more accurate per-cluster classifiers to be learned, which improves the overall classification accuracy. We demonstrate the effectiveness of our approach on a scene recognition task.
Armin Kappeler, Robin D. Morris, Amar Ramesh Kamat, Nikhil Rasiwasia, Gaurav Aggarwal
MMSP4
2015 Construction and evaluation of ontological tag trees
Chetan Kumar Verma, Vijay Mahadevan, Nikhil Rasiwasia, Gaurav Aggarwal, Alejandro Jaimes, Sujit Dey
Expert Syst. Appl.3
2014 Cluster Canonical Correlation Analysis
abstract
In this paper we present cluster canonical correlation analysis (cluster-CCA) for joint dimensionality reduction of two sets of data points. Unlike the standard pairwise correspondence between the data points, in our problem each set is partitioned into multiple clusters or classes, where the class labels define correspondences between the sets. Cluster-CCA is able to learn discriminant low dimensional representations that maximize the correlation between the two sets while segregating the different classes on the learned space. Furthermore, we present a kernel extension, kernel cluster canonical correlation analysis (cluster-KCCA) that extends cluster-CCA to account for non-linear relationships. Cluster-(K)CCA is shown to be computationally efficient, the complexity being similar to standard (K)CCA. By means of experimental evaluation on benchmark datasets, cluster-(K)CCA is shown to achieve state of the art performance for cross-modal retrieval tasks.
Nikhil Rasiwasia, Dhruv Mahajan 0001, Vijay Mahadevan, Gaurav Aggarwal
AISTATS1
2014 Do We Need Annotation Experts? A Case Study in Celiac Disease Classification
Roland Kwitt, Sebastian Hegenbart, Nikhil Rasiwasia, Andreas Vécsei, Andreas Uhl
MICCAI (2)3
2014 On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval
abstract
The problem of cross-modal retrieval from multimedia repositories is considered. This problem addresses the design of retrieval systems that support queries across content modalities, for example, using an image to search for texts. A mathematical formulation is proposed, equating the design of cross-modal retrieval systems to that of isomorphic feature spaces for different content modalities. Two hypotheses are then investigated regarding the fundamental attributes of these spaces. The first is that low-level cross-modal correlations should be accounted for. The second is that the space should enable semantic abstraction. Three new solutions to the cross-modal retrieval problem are then derived from these hypotheses: correlation matching (CM), an unsupervised method which models cross-modal correlations, semantic matching (SM), a supervised technique that relies on semantic representation, and semantic correlation matching (SCM), which combines both. An extensive evaluation of retrieval performance is conducted to test the validity of the hypotheses. All approaches are shown successful for text retrieval in response to image queries and vice versa. It is concluded that both hypotheses hold, in a complementary form, although evidence in favor of the abstraction hypothesis is stronger than that for correlation.
José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Nikhil Rasiwasia, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Class-Specific Simplex-Latent Dirichlet Allocation for Image Classification
abstract
An extension of the latent Dirichlet allocation (LDA), denoted class-specific-simplex LDA (css-LDA), is proposed for image classification. An analysis of the supervised LDA models currently used for this task shows that the impact of class information on the topics discovered by these models is very weak in general. This implies that the discovered topics are driven by general image regularities, rather than the semantic regularities of interest for classification. To address this, we introduce a model that induces supervision in topic discovery, while retaining the original flexibility of LDA to account for unanticipated structures of interest. The proposed css-LDA is an LDA model with class supervision at the level of image features. In css-LDA topics are discovered per class, i.e. a single set of topics shared across classes is replaced by multiple class-specific topic sets. This model can be used for generative classification using the Bayes decision rule or even extended to discriminative classification with support vector machines (SVMs). A css-LDA model can endow an image with a vector of class and topic specific count statistics that are similar to the Bag-of-words (BoW) histogram. SVM-based discriminants can be learned for classes in the space of these histograms. The effectiveness of css-LDA model in both generative and discriminative classification frameworks is demonstrated through an extensive experimental evaluation, involving multiple benchmark datasets, where it is shown to outperform all existing LDA based image classification approaches.
Mandar Dixit, Nikhil Rasiwasia, Nuno Vasconcelos
ICCV2
2013 Latent Dirichlet Allocation Models for Image Classification
abstract
Two new extensions of latent Dirichlet allocation (LDA), denoted topic-supervised LDA (ts-LDA) and class-specific-simplex LDA (css-LDA), are proposed for image classification. An analysis of the supervised LDA models currently used for this task shows that the impact of class information on the topics discovered by these models is very weak in general. This implies that the discovered topics are driven by general image regularities, rather than the semantic regularities of interest for classification. To address this, ts-LDA models are introduced which replace the automated topic discovery of LDA with specified topics, identical to the classes of interest for classification. While this results in improvements in classification accuracy over existing LDA models, it compromises the ability of LDA to discover unanticipated structure of interest. This limitation is addressed by the introduction of css-LDA, an LDA model with class supervision at the level of image features. In css-LDA topics are discovered per class, i.e., a single set of topics shared across classes is replaced by multiple class-specific topic sets. The css-LDA model is shown to combine the labeling strength of topic-supervision with the flexibility of topic-discovery. Its effectiveness is demonstrated through an extensive experimental evaluation, involving multiple benchmark datasets, where it is shown to outperform existing LDA-based image classification approaches.
Nikhil Rasiwasia, Nuno Vasconcelos
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Scene Recognition on the Semantic Manifold
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia
ECCV (4)3
2012 Endoscopic image analysis in semantic space
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia, Andreas Uhl, Bradley C. Davis, Michael Häfner, Friedrich Wrba
Medical Image Anal.3
2012 Holistic Context Models for Visual Recognition
abstract
A novel framework to context modeling based on the probability of co-occurrence of objects and scenes is proposed. The modeling is quite simple, and builds upon the availability of robust appearance classifiers. Images are represented by their posterior probabilities with respect to a set of contextual models, built upon the bag-of-features image representation, through two layers of probabilistic modeling. The first layer represents the image in a semantic space, where each dimension encodes an appearance-based posterior probability with respect to a concept. Due to the inherent ambiguity of classifying image patches, this representation suffers from a certain amount of contextual noise. The second layer enables robust inference in the presence of this noise by modeling the distribution of each concept in the semantic space. A thorough and systematic experimental evaluation of the proposed context modeling is presented. It is shown that it captures the contextual “gist” of natural images. Scene classification experiments show that contextual classifiers outperform their appearance-based counterparts, irrespective of the precise choice and accuracy of the latter. The effectiveness of the proposed approach to context modeling is further demonstrated through a comparison to existing approaches on scene classification and image retrieval, on benchmark data sets. In all cases, the proposed approach achieves superior results.
Nikhil Rasiwasia, Nuno Vasconcelos
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Adapted Gaussian models for image classification
abstract
A general formulation of “Bayesian Adaptation” for generative and discriminative classification in the topic model framework is proposed. A generic topic-independent Gaussian mixture model, known as the background GMM, is learned using all available training data and adapted to the individual topics. In the generative framework, a Gaussian variant of the spatial pyramid model is used with a Bayes classifier. For the discriminative case, a novel predictive histogram representation for an image is presented. This builds upon the adapted topic model structure, using the individual class dictionaries and Bayesian weighting. The resulting histogram representation is evaluated for classification using a Support Vector Machine (SVM). A comparative evaluation of the proposed image models with the standard ones in the image classification literature is provided on three benchmark datasets.
Mandar Dixit, Nikhil Rasiwasia, Nuno Vasconcelos
CVPR2
2011 Improving Product Classification Using Images
abstract
Product classification in Commerce search (e.g., Google Product Search, Bing Shopping) involves associating categories to offers of products from a large number of merchants. The categorized offers are used in many tasks including product taxonomy browsing and matching merchant offers to products in the catalog. Hence, learning a product classifier with high precision and recall is of fundamental importance in order to provide high quality shopping experience. A product offer typically consists of a short textual description and an image depicting the product. Traditional approaches to this classification task is to learn a classifier using only the textual descriptions of the products. In this paper, we show that the use of images, a weaker signal in our setting, in conjunction with the textual descriptions, a more discriminative signal, can considerably improve the precision of the classification task, irrespective of the type of classifier being used. We present a novel classification approach, Confusion Driven Probabilistic Fusion++ (CDPF++), that is cognizant of the disparity in the discriminative power of different types of signals and hence makes use of the confusion matrix of dominant signal (text in our setting) to prudently leverage the weaker signal (image), for an improved performance. Our evaluation performed on data from a major Commerce search engine's catalog shows a 12% (absolute) improvement in precision at 100% coverage, and a 16% (absolute) improvement in recall at 90% precision compared to classifiers that only use textual description of products. In addition, CDPF++ also provides a more accurate classifier based only on the dominant signal (text) that can be used in situations in which only the dominant signal is available during application time.
Anitha Kannan, Partha P. Talukdar, Nikhil Rasiwasia, Qifa Ke
ICDM3
2011 Learning Pit Pattern Concepts for Gastroenterological Training
Roland Kwitt, Nikhil Rasiwasia, Nuno Vasconcelos, Andreas Uhl, Michael Häfner, Friedrich Wrba
MICCAI (3)2
2010 A new approach to cross-modal multimedia retrieval
abstract
The problem of joint modeling the text and image components of multimedia documents is studied. The text component is represented as a sample from a hidden topic model, learned with latent Dirichlet allocation, and images are represented as bags of visual (SIFT) features. Two hypotheses are investigated: that 1) there is a benefit to explicitly modeling correlations between the two components, and 2) this modeling is more effective in feature spaces with higher levels of abstraction. Correlations between the two components are learned with canonical correlation analysis. Abstraction is achieved by representing text and images at a more general, semantic level. The two hypotheses are studied in the context of the task of cross-modal document retrieval. This includes retrieving the text that most closely matches a query image, or retrieving the images that most closely match a query text. It is shown that accounting for cross-modal correlations and semantic abstraction both improve retrieval accuracy. The cross-modal model is also shown to outperform state-of-the-art image retrieval systems on a unimodal retrieval task.
Nikhil Rasiwasia, José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos
ACM Multimedia1
2009 Holistic context modeling using semantic co-occurrences
abstract
We present a simple framework to model contextual relationships between visual concepts. The new framework combines ideas from previous object-centric methods (which model contextual relationships between objects in an image, such as their co-occurrence patterns) and scene-centric methods (which learn a holistic context model from the entire image, known as its “gist”). This is accomplished without demarcating individual concepts or regions in the image. First, using the output of a generic appearance based concept detection system, a semantic space is formulated, where each axis represents a semantic feature. Next, context models are learned for each of the concepts in the semantic space, using mixtures of Dirichlet distributions. Finally, an image is represented as a vector of posterior concept probabilities under these contextual concept models. It is shown that these posterior probabilities are remarkably noise-free, and an effective model of the contextual relationships between semantic concepts in natural images. This is further demonstrated through an experimental evaluation with respect to two vision tasks, viz. scene classification and image annotation, on benchmark datasets. The results show that, besides quite simple to compute, the proposed context models attain superior performance than state of the art systems in both tasks.
Nikhil Rasiwasia, Nuno Vasconcelos
CVPR1
2008 Scene classification with low-dimensional semantic spaces and weak supervision
abstract
A novel approach to scene categorization is proposed. Similar to previous works of [11, 15, 3, 12], we introduce an intermediate space, based on a low dimensional semantic ldquothemerdquo image representation. However, instead of learning the themes in an unsupervised manner, they are learned with weak supervision, from casual image annotations. Each theme induces a probability density on the space of low-level features, and images are represented as vectors of posterior theme probabilities. This enables an image to be associated with multiple themes, even when there are no multiple associations in the training labels. An implementation is presented and compared to various existing algorithms, on benchmark datasets. It is shown that the proposed low dimensional representation correlates well with human scene understanding, and is able to learn theme co-occurrences without explicit training. It is also shown to outperform unsupervised latent-space methods, with much smaller training complexity, and to achieve performance close to the state of the art methods, which rely on much higher-dimensional image representations. Finally a study of the effect of dimensionality on the classification performance is presented, indicating that the dimensionality of theme space grows sub-linearly with the number of scene categories.
Nikhil Rasiwasia, Nuno Vasconcelos
CVPR1
2008 A systematic study of the role of context on image classification
abstract
We present the results of a systematic study of the contextual gain hypothesis for image classification. This hypothesis relates the traditional strategy of direct visual classification (DVC), and an alternative strategy based on indirect contextual classification (ICC). DVC is composed of classifiers that operate directly on pixel or feature based image representations. ICC relies on DVC to label images with respect to a pre-defined set of contextual semantic features. Image classification is then performed by a classifier that operates on the semantic space of these classifier outputs. The contextual gain hypothesis states that, in this semantic space, it is possible to design classifiers with better accuracy than those achievable with DVC. A framework for the systematic comparison of the DVC and ICC strategies is introduced, and an extensive comparison of the performance of the two strategies is carried out. Its results strongly suggest that the contextual gain hypothesis holds.
Nikhil Rasiwasia, Nuno Vasconcelos
ICIP1
2007 Bridging the Gap: Query by Semantic Example
abstract
A combination of query-by-visual-example (QBVE) and semantic retrieval (SR), denoted as query-by-semantic-example (QBSE), is proposed. Images are labeled with respect to a vocabulary of visual concepts, as is usual in SR. Each image is then represented by a vector, referred to as a semantic multinomial, of posterior concept probabilities. Retrieval is based on the query-by-example paradigm: the user provides a query image, for which 1) a semantic multinomial is computed and 2) matched to those in the database. QBSE is shown to have two main properties of interest, one mostly practical and the other philosophical. From a practical standpoint, because it inherits the generalization ability of SR inside the space of known visual concepts (referred to as the semantic space) but performs much better outside of it, QBSE produces retrieval systems that are more accurate than what was previously possible. Philosophically, because it allows a direct comparison of visual and semantic representations under a common query paradigm, QBSE enables the design of experiments that explicitly test the value of semantic representations for image retrieval. An implementation of QBSE under the minimum probability of error (MPE) retrieval framework, previously applied with success to both QBVE and SR, is proposed, and used to demonstrate the two properties. In particular, an extensive objective comparison of QBSE with QBVE is presented, showing that the former significantly outperforms the latter both inside and outside the semantic space. By carefully controlling the structure of the semantic space, it is also shown that this improvement can only be attributed to the semantic nature of the representation on which QBSE is based.
Nikhil Rasiwasia, Pedro J. Moreno 0001, Nuno Vasconcelos
IEEE Trans. Multim.1