Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Florent Perronnin

dblp:51/1537 · DBLP profile ↗
← Back
70ranked-venue papers
16as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 13 first-authorGraphics, computer vision, multimedia, augmented reality and games · 44 · 13 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
28 papers
Image recognition and object detection · 44% Representation and self-supervised learning · 14% Transfer learning and domain adaptation · 12%
Databases, data mining, and information retrieval
11 papers
Information retrieval · 100%
Computer graphics and multimedia
9 papers
Multimedia analysis and retrieval · 70% Image and video coding · 12% Computational photography and imaging · 10%
Theoretical computer science
3 papers
Algorithms and data structures · 100%

Topics — the 30 heaviest of 67, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
image classification
1.9142016
Label-Embedding for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Fisher vectors meet Neural Networks: A hybrid classification architecture · CVPR 2015
Good Practice in Large-Scale Learning for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Information retrieval
image retrieval
1.272017
Interferences in Match Kernels · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Convolutional Patch Representations for Image Retrieval: An Unsupervised Approach · Int. J. Comput. Vis. 2017
Local Convolutional Features with Unsupervised Training for Image Retrieval · ICCV 2015
Computer vision › Image recognition and object detection › image classification
large-scale image classification
0.862013
Distance-Based Image Classification: Generalizing to New Classes at Near-Zero Cost · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost · ECCV (2) 2012
Towards good practice in large-scale learning for image classification · CVPR 2012
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.632016
Label-Embedding for Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Distance-Based Image Classification: Generalizing to New Classes at Near-Zero Cost · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Label-Embedding for Attribute-Based Classification · CVPR 2013
Information retrieval › image retrieval
large-scale image retrieval
0.332014
Asymmetric Distances for Binary Embeddings · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Large-scale image retrieval with compressed Fisher vectors · CVPR 2010
Aggregating Local Image Descriptors into Compact Codes · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Natural language and speech › Language models and text generation
synthetic data
0.312018
The Reasonable Effectiveness of Synthetic Visual Data · Int. J. Comput. Vis. 2018
Information retrieval
similarity search
0.322016
Polysemous Codes · ECCV (2) 2016
Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Information retrieval › image retrieval
image representation
0.312017
Interferences in Match Kernels · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Computer vision › Image recognition and object detection
text recognition
0.212015
Label Embedding: A Frugal Baseline for Text Recognition · Int. J. Comput. Vis. 2015
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning
0.212014
Transformation Pursuit for Image Classification · CVPR 2014
Information retrieval › hashing
binary embedding
0.212014
Asymmetric Distances for Binary Embeddings · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Multimedia analysis and retrieval
image classification
0.212014
Generalized Max Pooling · CVPR 2014
Image and video processing
image representation
0.212014
Generalized Max Pooling · CVPR 2014
Multimedia analysis and retrieval › multimedia feature representation
patch-based representation
0.212014
Generalized Max Pooling · CVPR 2014
Algorithms and data structures › data structure design › search structures › hashing
locality-sensitive hashing
0.222014
Asymmetric distances for binary embeddings · CVPR 2011
Asymmetric Distances for Binary Embeddings · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Image recognition and object detection
attribute recognition
0.212013
Label-Embedding for Attribute-Based Classification · CVPR 2013
Machine learning › Representation and self-supervised learning
fisher vector
0.212013
Image Classification with the Fisher Vector: Theory and Practice · Int. J. Comput. Vis. 2013
Machine learning › Representation and self-supervised learning › visual representation
image representation
0.212013
Image Classification with the Fisher Vector: Theory and Practice · Int. J. Comput. Vis. 2013
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › semantic embedding
label embedding
0.212013
Label-Embedding for Attribute-Based Classification · CVPR 2013
Multimedia analysis and retrieval › image retrieval
hashing-based image retrieval
0.212013
Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Multimedia analysis and retrieval
image retrieval
0.212013
Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Machine learning › Kernel, tree and ensemble methods › kernel function
fisher kernel
0.232010
Improving the Fisher Kernel for Large-Scale Image Classification · ECCV (4) 2010
Large-scale image retrieval with compressed Fisher vectors · CVPR 2010
Fisher Kernels on Visual Vocabularies for Image Categorization · CVPR 2007
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot generalization
0.112012
Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost · ECCV (2) 2012
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.112012
Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost · ECCV (2) 2012
Multimedia analysis and retrieval › image retrieval
instance-level image retrieval
0.112012
Leveraging category-level labels for instance-level image retrieval · CVPR 2012
Machine learning › Efficient and distributed learning
model compression
0.112011
High-dimensional signature compression for large-scale image classification · CVPR 2011
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112011
An Efficient Approach to Semantic Segmentation · Int. J. Comput. Vis. 2011
Image and video coding › image quality assessment › image aesthetics assessment
aesthetic quality prediction
0.112011
Assessing the aesthetic quality of photographs using generic image descriptors · ICCV 2011
Multimedia analysis and retrieval
binary embedding
0.112011
Asymmetric distances for binary embeddings · CVPR 2011
Image and video coding › image quality assessment
image aesthetics assessment
0.112011
Assessing the aesthetic quality of photographs using generic image descriptors · ICCV 2011

Methods — techniques the papers use, named apart from their topics

fisher vector · 1.3convolutional neural network · 1.2iterative quantization · 0.7label embedding · 0.6metric learning · 0.6bag-of-visual-words · 0.5locality-sensitive hashing · 0.4support vector machine · 0.4stochastic gradient descent · 0.3early stopping · 0.3cross-validation · 0.3spectral hashing · 0.3nearest class mean classifier · 0.3VLAD · 0.3product quantization · 0.2polysemous codes · 0.2compatibility function · 0.2weighted ranking loss · 0.2
YearPublicationVenuePosition
2018 The Reasonable Effectiveness of Synthetic Visual Data
Adrien Gaidon, Antonio M. López 0001, Florent Perronnin
Int. J. Comput. Vis.3
2017 Convolutional Patch Representations for Image Retrieval: An Unsupervised Approach
Mattis Paulin, Julien Mairal, Matthijs Douze, Zaïd Harchaoui, Florent Perronnin, Cordelia Schmid
Int. J. Comput. Vis.5
2017 Interferences in Match Kernels
abstract
We consider the design of an image representation that embeds and aggregates a set of local descriptors into a single vector. Popular representations of this kind include the bag-of-visual-words, the Fisher vector and the VLAD. When two such image representations are compared with the dot-product, the image-to-image similarity can be interpreted as a match kernel. In match kernels, one has to deal with interference, i.e., with the fact that even if two descriptors are unrelated, their matching score may contribute to the overall similarity. We formalise this problem and propose two related solutions, both aimed at equalising the individual contributions of the local descriptors in the final representation. These methods modify the aggregation stage by including a set of per-descriptor weights. They differ by the objective function that is optimised to compute those weights. The first is a "democratisation" strategy that aims at equalising the relative importance of each descriptor in the set comparison metric. The second one involves equalising the match of a single descriptor to the aggregated vector. These concurrent methods give a substantial performance boost over the state of the art in image search with short or mid-size vectors, as demonstrated by our experiments on standard public image retrieval benchmarks.
Naila Murray, Hervé Jégou, Florent Perronnin, Andrew Zisserman
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Polysemous Codes
Matthijs Douze, Hervé Jégou, Florent Perronnin
ECCV (2)3
2016 Label-Embedding for Image Classification
abstract
Attributes act as intermediate representations that enable parameter sharing between classes, a must when training data is scarce. We propose to view attribute-based image classification as a label-embedding problem: each class is embedded in the space of attribute vectors. We introduce a function that measures the compatibility between an image and a label embedding. The parameters of this function are learned on a training set of labeled samples to ensure that, given an image, the correct classes rank higher than the incorrect ones. Results on the Animals With Attributes and Caltech-UCSD-Birds datasets show that the proposed framework outperforms the standard Direct Attribute Prediction baseline in a zero-shot learning scenario. Label embedding enjoys a built-in ability to leverage alternative sources of information instead of or in addition to attributes, such as, e.g., class hierarchies or textual descriptions. Moreover, label embedding encompasses the whole range of learning settings from zero-shot learning to regular learning with a large number of labeled examples.
Zeynep Akata, Florent Perronnin, Zaïd Harchaoui, Cordelia Schmid
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Deep Fishing: Gradient Features from Deep Nets
abstract
Convolutional Networks (ConvNets) have recently improved image recognition performance thanks to end-to-end learning of deep feed-forward models from raw pixels. Deep learning is a marked departure from the previous state of the art, the Fisher Vector (FV), which relied on gradient-based encoding of local hand-crafted features. In this paper, we discuss a novel connection between these two approaches. First, we show that one can derive gradient representations from ConvNets in a similar fashion to the FV. Second, we show that this gradient representation actually corresponds to a structured matrix that allows for efficient similarity computation. We experimentally study the benefits of transferring this representation over the outputs of ConvNet layers, and find consistent improvements on the Pascal VOC 2007 and 2012 datasets.
Albert Gordo, Adrien Gaidon, Florent Perronnin
BMVC3
2015 Fisher vectors meet Neural Networks: A hybrid classification architecture
abstract
Fisher Vectors (FV) and Convolutional Neural Networks (CNN) are two image classification pipelines with different strengths. While CNNs have shown superior accuracy on a number of classification tasks, FV classifiers are typically less costly to train and evaluate. We propose a hybrid architecture that combines their strengths: the first unsupervised layers rely on the FV while the subsequent fully-connected supervised layers are trained with back-propagation. We show experimentally that this hybrid architecture significantly outperforms standard FV systems without incurring the high cost that comes with CNNs. We also derive competitive mid-level features from our architecture that are readily applicable to other class sets and even to new tasks.
Florent Perronnin, Diane Larlus
CVPR1
2015 LEWIS: Latent Embeddings for Word Images and Their Semantics
abstract
The goal of this work is to bring semantics into the tasks of text recognition and retrieval in natural images. Although text recognition and retrieval have received a lot of attention in recent years, previous works have focused on recognizing or retrieving exactly the same word used as a query, without taking the semantics into consideration. In this paper, we ask the following question: can we predict semantic concepts directly from a word image, without explicitly trying to transcribe the word image or its characters at any point? For this goal we propose a convolutional neural network (CNN) with a weighted ranking loss objective that ensures that the concepts relevant to the query image are ranked ahead of those that are not relevant. This can also be interpreted as learning a Euclidean space where word images and concepts are jointly embedded. This model is learned in an end-to-end manner, from image pixels to semantic concepts, using a dataset of synthetically generated word images and concepts mined from a lexical database (WordNet). Our results show that, despite the complexity of the task, word images and concepts can indeed be associated with a high degree of accuracy.
Albert Gordo, Jon Almazán, Naila Murray, Florent Perronnin
ICCV4
2015 Local Convolutional Features with Unsupervised Training for Image Retrieval
abstract
Patch-level descriptors underlie several important computer vision tasks, such as stereo-matching or content-based image retrieval. We introduce a deep convolutional architecture that yields patch-level descriptors, as an alternative to the popular SIFT descriptor for image retrieval. The proposed family of descriptors, called Patch-CKN, adapt the recently introduced Convolutional Kernel Network (CKN), an unsupervised framework to learn convolutional architectures. We present a comparison framework to benchmark current deep convolutional approaches along with Patch-CKN for both patch and image retrieval, including our novel "RomePatches" dataset. Patch-CKN descriptors yield competitive results compared to supervised CNN alternatives on patch and image retrieval.
Mattis Paulin, Matthijs Douze, Zaïd Harchaoui, Julien Mairal, Florent Perronnin, Cordelia Schmid
ICCV5
2015 Discovering Beautiful Attributes for Aesthetic Image Analysis
Luca Marchesotti, Naila Murray, Florent Perronnin
Int. J. Comput. Vis.3
2015 Label Embedding: A Frugal Baseline for Text Recognition
José A. Rodríguez-Serrano, Albert Gordo, Florent Perronnin
Int. J. Comput. Vis.3
2014 Generalized Max Pooling
abstract
State-of-the-art patch-based image representations involve a pooling operation that aggregates statistics computed from local descriptors. Standard pooling operations include sum- and max-pooling. Sum-pooling lacks discriminability because the resulting representation is strongly influenced by frequent yet often uninformative descriptors, but only weakly influenced by rare yet potentially highly-informative ones. Max-pooling equalizes the influence of frequent and rare descriptors but is only applicable to representations that rely on count statistics, such as the bag-of-visual-words (BOV)and its soft- and sparse-coding extensions. We propose a novel pooling mechanism that achieves the same effect as max-pooling but is applicable beyond the BOV and especially to the state-of-the-art Fisher Vector -- hence the name Generalized Max Pooling (GMP). It involves equalizing the similarity between each patch and the pooled representation, which is shown to be equivalent to re-weighting the per-patch statistics. We show on five public image classification benchmarks that the proposed GMP can lead to significant performance gains with respect to heuristic alternatives.
Naila Murray, Florent Perronnin
CVPR2
2014 Transformation Pursuit for Image Classification
abstract
A simple approach to learning invariances in image classification consists in augmenting the training set with transformed versions of the original images. However, given a large set of possible transformations, selecting a compact subset is challenging. Indeed, all transformations are not equally informative and adding uninformative transformations increases training time with no gain in accuracy. We propose a principled algorithm -- Image Transformation Pursuit (ITP) -- for the automatic selection of a compact set of transformations. ITP works in a greedy fashion, by selecting at each iteration the one that yields the highest accuracy gain. ITP also allows to efficiently explore complex transformations, that combine basic transformations. We report results on two public benchmarks: the CUB dataset of bird images and the ImageNet 2010 challenge. Using Fisher Vector representations, we achieve an improvement from 28.2% to 45.2% in top-1 accuracy on CUB, and an improvement from 70.1% to 74.9% in top-5 accuracy on ImageNet. We also show significant improvements for deep convnet features: from 47.3% to 55.4% on CUB and from 77.9% to 81.4% on ImageNet.
Mattis Paulin, Jérôme Revaud, Zaïd Harchaoui, Florent Perronnin, Cordelia Schmid
CVPR4
2014 Instance classification with prototype selection
abstract
We address the problem of instance classification: our goal is to annotate images with tags corresponding to objects classes which exhibit small intra-class variations such as logos, products or landmarks. We propose a novel algorithm for the selection of class-specific prototypes which are used in a voting-based classification scheme. We show significant improvements over two state-of-the-art methods, namely the Fisher vector and Hamming Embedding, on two challenging methods of logos and vehicles.
Josip Krapac, Florent Perronnin, Teddy Furon, Hervé Jégou
ICMR2
2014 Comparison of face detection and image classification for detecting front seat passengers in vehicles
abstract
Due to the high volume of traffic on modern roadways, transportation agencies have proposed High Occupancy Vehicle (HOV) lanes and High Occupancy Tolling (HOT) lanes to promote car pooling. However, enforcement of the rules of these lanes is currently performed by roadside enforcement officers using visual observation. Manual roadside enforcement is known to be inefficient, costly, potentially dangerous, and ultimately ineffective. Violation rates up to 50%–80% have been reported, while manual enforcement rates of less than 10% are typical. Therefore, there is a need for automated vehicle occupancy detection to support HOV/HOT lane enforcement. A key component of determining vehicle occupancy is to determine whether or not the vehicle's front passenger seat is occupied. In this paper, we examine two methods of determining vehicle front seat occupancy using a near infrared (NIR) camera system pointed at the vehicle's front windshield. The first method examines a state-of-the-art deformable part model (DPM) based face detection system that is robust to facial pose. The second method examines state-of-the-art local aggregation based image classification using bag-of-visual-words (BOW) and Fisher vectors (FV). A dataset of 3000 images was collected on a public roadway and is used to perform the comparison. From these experiments it is clear that the image classification approach is superior for this problem.
Yusuf Artan, Peter Paul, Florent Perronnin, Aaron M. Burry
WACV3
2014 Good Practice in Large-Scale Learning for Image Classification
abstract
We benchmark several SVM objective functions for large-scale image classification. We consider one-versus-rest, multiclass, ranking, and weighted approximate ranking SVMs. A comparison of online and batch methods for optimizing the objectives shows that online methods perform as well as batch methods in terms of classification accuracy, but with a significant gain in training speed. Using stochastic gradient descent, we can scale the training to millions of images and thousands of classes. Our experimental evaluation shows that ranking-based algorithms do not outperform the one-versus-rest strategy when a large number of training examples are used. Furthermore, the gap in accuracy between the different algorithms shrinks as the dimension of the features increases. We also show that learning through cross-validation the optimal rebalancing of positive and negative examples can result in a significant improvement for the one-versus-rest strategy. Finally, early stopping can be used as an effective regularization strategy when training with online algorithms. Following these "good practices," we were able to improve the state of the art on a large subset of 10K classes and 9M images of ImageNet from 16.7 percent Top-1 accuracy to 19.1 percent.
Zeynep Akata, Florent Perronnin, Zaïd Harchaoui, Cordelia Schmid
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Asymmetric Distances for Binary Embeddings
abstract
In large-scale query-by-example retrieval, embedding image signatures in a binary space offers two benefits: data compression and search efficiency. While most embedding algorithms binarize both query and database signatures, it has been noted that this is not strictly a requirement. Indeed, asymmetric schemes that binarize the database signatures but not the query still enjoy the same two benefits but may provide superior accuracy. In this work, we propose two general asymmetric distances that are applicable to a wide variety of embedding techniques including locality sensitive hashing (LSH), locality sensitive binary codes (LSBC), spectral hashing (SH), PCA embedding (PCAE), PCAE with random rotations (PCAE-RR), and PCAE with iterative quantization (PCAE-ITQ). We experiment on four public benchmarks containing up to 1M images and show that the proposed asymmetric distances consistently lead to large improvements over the symmetric Hamming distance for all binary embedding techniques.
Albert Gordo, Florent Perronnin, Yunchao Gong, Svetlana Lazebnik
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Revisiting the Fisher vector for fine-grained classification
Philippe Henri Gosselin, Naila Murray, Hervé Jégou, Florent Perronnin
Pattern Recognit. Lett.4
2013 What is a good evaluation measure for semantic segmentation?
abstract
In this work, we consider the evaluation of the semantic segmentation task. We discuss the strengths and limitations of the few existing measures, and propose new ways to evaluate semantic segmentation. First, we argue that a per-image score instead of one computed over the entire dataset brings a lot more insight. Second, we propose to take contours more carefully into account. Based on the conducted experiments, we suggest best practices for the evaluation. Finally, we present a user study we conducted to better understand how the quality of image segmentations is perceived by humans.
Gabriela Csurka, Diane Larlus, Florent Perronnin
BMVC3
2013 Learning beautiful (and ugly) attributes
Luca Marchesotti, Florent Perronnin
BMVC2
2013 Label embedding for text recognition
José A. Rodríguez-Serrano, Florent Perronnin
BMVC2
2013 Label-Embedding for Attribute-Based Classification
abstract
Attributes are an intermediate representation, which enables parameter sharing between classes, a must when training data is scarce. We propose to view attribute-based image classification as a label-embedding problem: each class is embedded in the space of attribute vectors. We introduce a function which measures the compatibility between an image and a label embedding. The parameters of this function are learned on a training set of labeled samples to ensure that, given an image, the correct classes rank higher than the incorrect ones. Results on the Animals With Attributes and Caltech-UCSD-Birds datasets show that the proposed framework outperforms the standard Direct Attribute Prediction baseline in a zero-shot learning scenario. The label embedding framework offers other advantages such as the ability to leverage alternative sources of information in addition to attributes (e.g. class hierarchies) or to transition smoothly from zero-shot learning to learning with large quantities of data.
Zeynep Akata, Florent Perronnin, Zaïd Harchaoui, Cordelia Schmid
CVPR2
2013 Image Classification with the Fisher Vector: Theory and Practice
Jorge Sánchez 0002, Florent Perronnin, Thomas Mensink, Jakob Verbeek
Int. J. Comput. Vis.2
2013 Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image Retrieval
abstract
This paper addresses the problem of learning similarity-preserving binary codes for efficient similarity search in large-scale image collections. We formulate this problem in terms of finding a rotation of zero-centered data so as to minimize the quantization error of mapping this data to the vertices of a zero-centered binary hypercube, and propose a simple and efficient alternating minimization algorithm to accomplish this task. This algorithm, dubbed iterative quantization (ITQ), has connections to multiclass spectral clustering and to the orthogonal Procrustes problem, and it can be used both with unsupervised data embeddings such as PCA and supervised embeddings such as canonical correlation analysis (CCA). The resulting binary codes significantly outperform several other state-of-the-art methods. We also show that further performance improvements can result from transforming the data with a nonlinear kernel mapping prior to PCA or CCA. Finally, we demonstrate an application of ITQ to learning binary attributes or "classemes" on the ImageNet data set.
Yunchao Gong, Svetlana Lazebnik, Albert Gordo, Florent Perronnin
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Distance-Based Image Classification: Generalizing to New Classes at Near-Zero Cost
abstract
We study large-scale image classification methods that can incorporate new classes and training images continuously over time at negligible cost. To this end, we consider two distance-based classifiers, the k-nearest neighbor (k-NN) and nearest class mean (NCM) classifiers, and introduce a new metric learning approach for the latter. We also introduce an extension of the NCM classifier to allow for richer class representations. Experiments on the ImageNet 2010 challenge dataset, which contains over 10(6) training images of 1,000 classes, show that, surprisingly, the NCM classifier compares favorably to the more flexible k-NN classifier. Moreover, the NCM performance is comparable to that of linear SVMs which obtain current state-of-the-art performance. Experimentally, we study the generalization performance to classes that were not used to learn the metrics. Using a metric learned on 1,000 classes, we show results for the ImageNet-10K dataset which contains 10,000 classes, and obtain performance that is competitive with the current state-of-the-art while being orders of magnitude faster. Furthermore, we show how a zero-shot class prior based on the ImageNet hierarchy can improve performance when few training images are available.
Thomas Mensink, Jakob Verbeek, Florent Perronnin, Gabriela Csurka
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Large-scale document image retrieval and classification with runlength histograms and binary embeddings
Albert Gordo, Florent Perronnin, Ernest Valveny
Pattern Recognit.2
2012 Learning to rank images using semantic and aesthetic labels
abstract
Most works on image retrieval from text queries have addressed the problem of retrieving semantically relevant images. However, the ability to assess the aesthetic quality of an image is an increasingly important differentiating factor for search engines. In this work, given a semantic query, we are interested in retrieving images which are semantically relevant and score highly in terms of aesthetics/visual quality. We use large-margin classifiers and rankers to learn statistical models capable of ordering images based on the aesthetic and semantic information. In particular, we compare two families of approaches: while the first one attempts to learn a single ranker which takes into account both semantic and aesthetic information, the second one learns separate semantic and aesthetic models. We carry out a quantitative and qualitative evaluation on a recentlypublished large-scale dataset and we show that the second family of techniques significantly outperforms the first one.
Naila Murray, Luca Marchesotti, Florent Perronnin
BMVC3
2012 Leveraging category-level labels for instance-level image retrieval
abstract
In this article, we focus on the problem of large-scale instance-level image retrieval. For efficiency reasons, it is common to represent an image by a fixed-length descriptor which is subsequently encoded into a small number of bits. We note that most encoding techniques include an unsupervised dimensionality reduction step. Our goal in this work is to learn a better subspace in a supervised manner. We especially raise the following question: "can category-level labels be used to learn such a subspace?" To answer this question, we experiment with four learning techniques: the first one is based on a metric learning framework, the second one on attribute representations, the third one on Canonical Correlation Analysis (CCA) and the fourth one on Joint Subspace and Classifier Learning (JSCL). While the first three approaches have been applied in the past to the image retrieval problem, we believe we are the first to show the usefulness of JSCL in this context. In our experiments, we use ImageNet as a source of category-level labels and report retrieval results on two standard dataseis: INRIA Holidays and the University of Kentucky benchmark. Our experimental study shows that metric learning and attributes do not lead to any significant improvement in retrieval accuracy, as opposed to CCA and JSCL. As an example, we report on Holidays an increase in accuracy from 39.3% to 48.6% with 32-dimensional representations. Overall JSCL is shown to yield the best results.
Albert Gordo, José A. Rodríguez-Serrano, Florent Perronnin, Ernest Valveny
CVPR3
2012 AVA: A large-scale database for aesthetic visual analysis
abstract
With the ever-expanding volume of visual content available, the ability to organize and navigate such content by aesthetic preference is becoming increasingly important. While still in its nascent stage, research into computational models of aesthetic preference already shows great potential. However, to advance research, realistic, diverse and challenging databases are needed. To this end, we introduce a new large-scale database for conducting Aesthetic Visual Analysis: AVA. It contains over 250,000 images along with a rich variety of meta-data including a large number of aesthetic scores for each image, semantic labels for over 60 categories as well as labels related to photographic style. We show the advantages of AVA with respect to existing databases in terms of scale, diversity, and heterogeneity of annotations. We then describe several key insights into aesthetic preference afforded by AVA. Finally, we demonstrate, through three applications, how the large scale of AVA can be leveraged to improve performance on existing preference tasks.
Naila Murray, Luca Marchesotti, Florent Perronnin
CVPR3
2012 Towards good practice in large-scale learning for image classification
abstract
We propose a benchmark of several objective functions for large-scale image classification: we compare the one-vs-rest, multiclass, ranking and weighted average ranking SVMs. Using stochastic gradient descent optimization, we can scale the learning to millions of images and thousands of classes. Our experimental evaluation shows that ranking based algorithms do not outperform a one-vs-rest strategy and that the gap between the different algorithms reduces in case of high-dimensional data. We also show that for one-vs-rest, learning through cross-validation the optimal degree of imbalance between the positive and the negative samples can have a significant impact. Furthermore, early stopping can be used as an effective regularization strategy when training with stochastic gradient algorithms. Following these "good practices", we were able to improve the state-of-the-art on a large subset of 10K classes and 9M of images of lmageNet from 16.7% accuracy to 19.1%.
Florent Perronnin, Zeynep Akata, Zaïd Harchaoui, Cordelia Schmid
CVPR1
2012 Document Classification Using Multiple Views
abstract
The combination of multiple features or views when representing documents or other kinds of objects usually leads to improved results in classification (and retrieval) tasks. Most systems assume that those views will be available both at training and test time. However, some views may be too `expensive' to be available at test time. In this paper, we consider the use of Canonical Correlation Analysis to leverage `expensive' views that are available only at training time. Experimental results show that this information may significantly improve the results in a classification task.
Albert Gordo, Florent Perronnin, Ernest Valveny
Document Analysis Systems2
2012 Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost
Thomas Mensink, Jakob Verbeek, Florent Perronnin, Gabriela Csurka
ECCV (2)3
2012 Toward automatic and flexible concept transfer
Naila Murray, Sandra Skaff, Luca Marchesotti, Florent Perronnin
Comput. Graph.4
2012 Images as sets of locally weighted features
Teófilo Emídio de Campos, Gabriela Csurka, Florent Perronnin
Comput. Vis. Image Underst.3
2012 Aggregating Local Image Descriptors into Compact Codes
abstract
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image data set takes about 250 ms on one processor core.
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez 0002, Patrick Pérez, Cordelia Schmid
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 A Model-Based Sequence Similarity with Application to Handwritten Word Spotting
abstract
This paper proposes a novel similarity measure between vector sequences. We work in the framework of model-based approaches, where each sequence is first mapped to a Hidden Markov Model (HMM) and then a measure of similarity is computed between the HMMs. We propose to model sequences with semicontinuous HMMs (SC-HMMs). This is a particular type of HMM whose emission probabilities in each state are mixtures of shared Gaussians. This crucial constraint provides two major benefits. First, the a priori information contained in the common set of Gaussians leads to a more accurate estimate of the HMM parameters. Second, the computation of a similarity between two SC-HMMs can be simplified to a Dynamic Time Warping (DTW) between their mixture weight vectors, which significantly reduces the computational cost. Experiments are carried out on a handwritten word retrieval task in three different datasets-an in-house dataset of real handwritten letters, the George Washington dataset, and the IFN/ENIT dataset of Arabic handwritten words. These experiments show that the proposed similarity outperforms the traditional DTW between the original sequences, and the model-based approach which uses ordinary continuous HMMs. We also show that this increase in accuracy can be traded against a significant reduction of the computational cost.
José A. Rodríguez-Serrano, Florent Perronnin
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Synthesizing queries for handwritten word image retrieval
José A. Rodríguez-Serrano, Florent Perronnin
Pattern Recognit.2
2012 Modeling the spatial layout of images beyond spatial pyramids
Jorge Sánchez 0002, Florent Perronnin, Teófilo Emídio de Campos
Pattern Recognit. Lett.2
2011 Asymmetric distances for binary embeddings
abstract
In large-scale query-by-example retrieval, embedding image signatures in a binary space offers two benefits: data compression and search efficiency. While most embedding algorithms binarize both query and database signatures, it has been noted that this is not strictly a requirement. Indeed, asymmetric schemes which binarize the database signatures but not the query still enjoy the same two benefits but may provide superior accuracy. In this work, we propose two general asymmetric distances which are applicable to a wide variety of embedding techniques including Locality Sensitive Hashing (LSH), Locality Sensitive Binary Codes (LSBC), Spectral Hashing (SH) and Semi-Supervised Hashing (SSH). We experiment on four public benchmarks containing up to 1M images and show that the proposed asymmetric distances consistently lead to large improvements over the symmetric Hamming distance for all binary embedding techniques. We also propose a novel simple binary embedding technique - PCA Embedding (PCAE) - which is shown to yield competitive results with respect to more complex algorithms such as SH and SSH.
Albert Gordo, Florent Perronnin
CVPR2
2011 High-dimensional signature compression for large-scale image classification
abstract
We address image classification on a large-scale, i.e. when a large number of images and classes are involved. First, we study classification accuracy as a function of the image signature dimensionality and the training set size. We show experimentally that the larger the training set, the higher the impact of the dimensionality on the accuracy. In other words, high-dimensional signatures are important to obtain state-of-the-art results on large datasets. Second, we tackle the problem of data compression on very large signatures (on the order of 105dimensions) using two lossy compression strategies: a dimensionality reduction technique known as the hash kernel and an encoding technique based on product quantizers. We explain how the gain in storage can be traded against a loss in accuracy and/or an increase in CPU cost. We report results on two large databases - ImageNet and a dataset of lM Flickr images - showing that we can reduce the storage of our signatures by a factor 64 to 128 with little loss in accuracy. Integrating the decompression in the classifier learning yields an efficient and scalable training algorithm. On ILSVRC2010 we report a 74.3% accuracy at top-5, which corresponds to a 2.5% absolute improvement with respect to the state-of-the-art. On a subset of 10K classes of ImageNet we report a top-1 accuracy of 16.7%, a relative improvement of 160% with respect to the state-of-the-art.
Jorge Sánchez 0002, Florent Perronnin
CVPR2
2011 Assessing the aesthetic quality of photographs using generic image descriptors
abstract
In this paper, we automatically assess the aesthetic properties of images. In the past, this problem has been addressed by hand-crafting features which would correlate with best photographic practices (e.g. “Does this image respect the rule of thirds?”) or with photographic techniques (e.g. “Is this image a macro?”). We depart from this line of research and propose to use generic image descriptors to assess aesthetic quality. We experimentally show that the descriptors we use, which aggregate statistics computed from low-level local features, implicitly encode the aesthetic properties explicitly used by state-of-the-art methods and outperform them by a significant margin.
Luca Marchesotti, Florent Perronnin, Diane Larlus, Gabriela Csurka
ICCV2
2011 An Efficient Approach to Semantic Segmentation
Gabriela Csurka, Florent Perronnin
Int. J. Comput. Vis.2
2010 Large-scale image retrieval with compressed Fisher vectors
abstract
The problem of large-scale image search has been traditionally addressed with the bag-of-visual-words (BOV). In this article, we propose to use as an alternative the Fisher kernel framework. We first show why the Fisher representation is well-suited to the retrieval problem: it describes an image by what makes it different from other images. One drawback of the Fisher vector is that it is high-dimensional and, as opposed to the BOV, it is dense. The resulting memory and computational costs do not make Fisher vectors directly amenable to large-scale retrieval. Therefore, we compress Fisher vectors to reduce their memory footprint and speed-up the retrieval. We compare three binarization approaches: a simple approach devised for this representation and two standard compression techniques. We show on two publicly available datasets that compressed Fisher vectors perform very well using as little as a few hundreds of bits per image, and significantly better than a very recent compressed BOV approach.
Florent Perronnin, Yan Liu 0017, Jorge Sánchez 0002, Hervé Poirier
CVPR1
2010 Large-scale image categorization with explicit data embedding
abstract
Kernel machines rely on an implicit mapping of the data such that non-linear classification in the original space corresponds to linear classification in the new space. As kernel machines are difficult to scale to large training sets, it has been proposed to perform an explicit mapping of the data and to learn directly linear classifiers in the new space. In this paper, we consider the problem of learning image categorizers on large image sets (e.g. > 100k images) using bag-of-visual-words (BOV) image representations and Support Vector Machine classifiers. We experiment with three approaches to BOV embedding: 1) kernel PCA (kPCA), 2) a modified kPCA we propose for additive kernels and 3) random projections for shift-invariant kernels. We report experiments on 3 datasets: Cal-tech101, VOC07 and ImageNet. An important conclusion is that simply square-rooting BOV vectors - which corresponds to an exact mapping for the Bhattacharyya kernel - already leads to large improvements, often quite close to the best results obtained with additive kernels. Another conclusion is that, although it is possible to go beyond additive kernels, the embedding comes at a much higher cost.
Florent Perronnin, Jorge Sánchez 0002, Yan Liu 0017
CVPR1
2010 Improving the Fisher Kernel for Large-Scale Image Classification
Florent Perronnin, Jorge Sánchez 0002, Thomas Mensink
ECCV (4)1
2010 Font retrieval on a large scale: An experimental study
abstract
This paper addresses the problem of font retrieval using a query-by-example paradigm: given a font, retrieve the the most visually similar fonts. We describe a font by (a) rendering a set of reference characters, (b) extracting a feature vector for each reference character and (c) concatenating the-level character descriptors. The similarity between two fonts is simply the similarity between the vectorial representations. Our contribution is an experimental comparison of character-level descriptors of step (b) on a large dataset of 9,000 fonts. The descriptors we chose to evaluate were drawn from the literature on typed and handwritten text analysis. An important conclusion is that the SIFT descriptor, which was shown to be state-of-the-art for object recognition in photographs and for handwriting recognition, yields the best results for font retrieval.
Saurabh Kataria 0003, Luca Marchesotti, Florent Perronnin
ICIP3
2010 A Bag-of-Pages Approach to Unordered Multi-page Document Classification
abstract
We consider the problem of classifying documents containing multiple unordered pages. For this purpose, we propose a novel bag-of-pages document representation. To represent a document, one assigns every page to a prototype in a codebook of pages. This leads to a histogram representation which can then be fed to any discriminative classifier. We also consider several refinements over this initial approach. We show on two challenging datasets that the proposed approach significantly outperforms a baseline system.
Albert Gordo, Florent Perronnin
ICPR2
2010 Unsupervised writer adaptation of whole-word HMMs with application to word-spotting
José A. Rodríguez-Serrano, Florent Perronnin, Gemma Sánchez, Josep Lladós 0001
Pattern Recognit. Lett.2
2009 Hierarchical Image-Region Labeling via Structured Learning
abstract
We present a graphical model that encodes hierarchical constraints for classifying image regions at multiple scales. We show that inference can be performed efficiently and exactly, rendering it amenable to structured learning. Our model is parametrised using the outputs of a series of first-order classifiers, meaning that it learns which classifiers are useful at different scales, as well as the relationships between classifiers across scales. Example results Correct labeling, using bounding-boxes from VOC2007 (1 − ∆ = 1): Baseline, using no second-order features (1 − ∆ = 0.566): Our model The ‘nodes ’ of our graphical model correspond to overlapping image regions: Second-order features without learning (1 − ∆ = 0.551): Learning of all features (1 − ∆ = 0.770): Colour-code for labels: Edges are formed by connecting nodes at different scales: we connect two nodes precisely when the corresponding image regions overlap at adjacent scales, so that our graphical model forms a quad-tree. First-order (node) features Our image features are based on those from [2], in which image-level, region-level, and patch-level classifiers are proposed. We use all classifiers at all scales (Pr,label is the probability that the region r is labeled label): Φ nodes (r, label) = (0,..., P 1 r,label,..., 0)... (0,..., P features for first classifier n r,label,..., 0) features for nth classifier Thus we learn which classifiers are useful at which scales. Hierarchical constraints We want to ban inconsistent assignments at different scales: aeroplane
Julian J. McAuley, Teófilo Emídio de Campos, Gabriela Csurka, Florent Perronnin
BMVC4
2009 Modeling images as mixtures of reference images
abstract
A state-of-the-art approach to measure the similarity of two images is to model each image by a continuous distribution, generally a Gaussian mixture model (GMM), and to compute a probabilistic similarity between the GMMs. One limitation of traditional measures such as the Kullback-Leibler (KL) divergence and the probability product kernel (PPK) is that they measure a global match of distributions. This paper introduces a novel image representation. We propose to approximate an image, modeled by a GMM, as a convex combination of K reference image GMMs, and then to describe the image as the K-dimensional vector of mixture weights. The computed weights encode a similarity that favors local matches (i.e. matches of individual Gaussians) and is therefore fundamentally different from the KL or PPK. Although the computation of the mixture weights is a convex optimization problem, its direct optimization is difficult. We propose two approximate optimization algorithms: the first one based on traditional sampling methods, the second one based on a variational bound approximation of the true objective function. We apply this novel representation to the image categorization problem and compare its performance to traditional kernel-based methods. We demonstrate on the PASCAL VOC 2007 dataset a consistent increase in classification accuracy.
Florent Perronnin, Yan Liu 0017
CVPR1
2009 A family of contextual measures of similarity between distributions with application to image retrieval
abstract
We introduce a novel family of contextual measures of similarity between distributions: the similarity between two distributions q and p is measured in the context of a third distribution u. In our framework any traditional measure of similarity / dissimilarity has its contextual counterpart. We show that for two important families of divergences (Bregman and Csisz'ar), the contextual similarity computation consists in solving a convex optimization problem. We focus on the case of multinomials and explain how to compute in practice the similarity for several well-known measures. These contextual measures are then applied to the image retrieval problem. In such a case, the context u is estimated from the neighbors of a query q. One of the main benefits of our approach lies in the fact that using different contexts, and especially contexts at multiple scales (i.e. broad and narrow contexts), provides different views on the same problem. Combining the different views can improve retrieval accuracy. We will show on two very different datasets (one of photographs, the other of document images) that the proposed measures have a relatively small positive impact on macro Average Precision (which measures purely ranking) and a large positive impact on micro Average Precision (which measures both ranking and consistency of the scores across multiple queries).
Florent Perronnin, Yan Liu 0017, Jean-Michel Renders
CVPR1
2009 A similarity measure between vector sequences with application to handwritten word image retrieval
abstract
This article proposes a novel similarity measure between vector sequences. Recently, a model-based approach was introduced to address this issue. It consists in modeling each sequence with a continuous Hidden Markov Model (CHMM) and computing a probabilistic measure of similarity between C-HMMs. In this paper we propose to model sequences with semi-continuous HMMs (SC-HMMs): the Gaussians of the SC-HMMs are constrained to belong to a shared pool of Gaussians. This constraint provides two major benefits. First, the a priori information contained in the common set of Gaussians leads to a more accurate estimate of the HMM parameters. Second, the computation of a probabilistic similarity between two SC-HMMs can be simplified to a Dynamic Time Warping (DTW) between their mixture weight vectors, which reduces significantly the computational cost. Experimental results on a handwritten word retrieval task show that the proposed similarity outperforms the traditional DTW between the original sequences, and the model-based approach which uses C-HMMs. We also show that this increase in accuracy can be traded against a significant reduction of the computational cost (up to 100 times).
José A. Rodríguez-Serrano, Florent Perronnin, Josep Lladós 0001, Gemma Sánchez
CVPR2
2009 Fisher Kernels for Handwritten Word-spotting
abstract
The Fisher kernel is a generic framework which combines the benefits of generative and discriminative approaches to pattern classification. In this contribution, we propose to apply this framework to handwritten word-spotting. Given a word image and a keyword generative model, the idea is to generate a vector which describes how the parameters of the keyword model should be modified to best fit the word image.This vector can then be used as the input of a discriminative classifier. We compare the performance of the proposed approach with that of a generative baseline on a challenging real-world dataset of customer letters. When the kernel used by the classifier is linear, the performance improvement is marginal but the proposed system is approximately 15 times faster than the baseline. If we use a non-linear kernel devised for this task, we obtain a 15% relative reduction of the error but the detector is approximately 15 times slower.
Florent Perronnin, José A. Rodríguez-Serrano
ICDAR1
2009 Handwritten Word Image Retrieval with Synthesized Typed Queries
abstract
We propose a new method for handwritten word-spotting which does not require prior training or gathering examples for querying. More precisely, a model is trained “on the fly” with images rendered from the searched words in one or multiple computer fonts. To reduce the mismatch between the typed-text prototypes and the candidate handwritten images, we make use of: (i) local gradient histogram(LGH) features, which were shown to model word shapes robustly, and (ii) semi-continuous hidden Markov models(SC-HMM), in which the typed-text models are constrained to a “vocabulary” of handwritten shapes, thus learning a link between both types of data. Experiments show that the proposed method is effective in retrieving handwritten words, and the comparison to alternative methods reveals that the contribution of both the LGH features and the SCHMM is crucial. To the best of the authors’ knowledge, this is the first work to address this issue in a non-trivial manner.
José A. Rodríguez-Serrano, Florent Perronnin
ICDAR2
2009 Handwritten word-spotting using hidden Markov models and universal vocabularies
José A. Rodríguez-Serrano, Florent Perronnin
Pattern Recognit.2
2008 A Simple High Performance Approach to Semantic Segmentation
abstract
We propose a simple approach to semantic image segmentation. Our system scores low-level patches according to their class relevance, propagates these posterior probabilities to pixels and uses low-level segmentation to guide the semantic segmentation. The two main contributions of this paper are as follows. First, for the patch scoring, we describe each patch with a high-level descriptor based on the Fisher kernel and use a set of linear classifiers. While the Fisher kernel methodology was shown to lead to high accuracy for image classification, it has not been applied to the segmentation problem. Second, we use global image classifiers to take into account the context of the objects to be segmented. If an image as a whole is unlikely to contain an object class, then the corresponding class is not considered in the segmentation pipeline. This increases the classification accuracy and reduces the computational cost. We will show that despite its apparent simplicity, this system provides above state-of-the-art performance on the PASCAL VOC 2007 dataset and state-of-the-art performance on the MSRC 21 dataset. 1
Gabriela Csurka, Florent Perronnin
BMVC2
2008 A similarity measure between unordered vector sets with application to image categorization
abstract
We present a novel approach to compute the similarity between two unordered variable-sized vector sets. To solve this problem, several authors have proposed to model each vector set with a Gaussian mixture model (GMM) and to compute a probabilistic measure of similarity between the GMMs. The main contribution of this paper is to model each vector set with a GMM adapted from a common ldquouniversalrdquo GMM using the maximum a posteriori (MAP) criterion. The advantages of this approach are twofold. MAP provides a more accurate estimate of the GMM parameters compared to standard maximum likelihood estimation (MLE) in the challenging case where the cardinality of the vector set is small. Moreover, there is a correspondence between the Gaussians of two GMMs adapted from a common distribution and one can take advantage of this fact to compute efficiently the probabilistic similarity. This work is applied to the image categorization problem: images are modeled as bags of low-level features and classification is performed using a kernel classifier based on the proposed similarity measure. Experimental results on the PASCAL VOC 2006 and VOC 2007 databases show the excellent performance of our approach.
Yan Liu 0017, Florent Perronnin
CVPR2
2008 An analysis of the relationship between painters based on their work
abstract
We propose a methodology to analyze and visualize the relationships and influences between painters. We build a graph where each painter is a node and an edge between two nodes is weighted by the painters' similarity. The similarity of two painters is measured as a function of the similarity between their paintings. Although the image representation we use was initially developed for the detection of objects in natural images, we will show how our system can discover artistic periods and non-obvious connections between painters belonging to different artistic movements, showing its usefulness for the computer-assisted study of art.
Marco Bressan 0003, Claudio Cifarelli, Florent Perronnin
ICIP3
2008 Unsupervised writer style adaptation for handwritten word spotting
abstract
We propose a novel approach for writer adaptation in a word spotting task. The method exploits the fact that a semi-continuous hidden Markov model separates the word model parameters into (i) a shared codebook of shapes and (ii) a set of word-specific parameters. Our main contribution is to derive writer-specific word models by statistically adapting an initial universal codebook to each document. This process is unsupervised and does not even require the appearance of the keyword(s) in the searched document. Experimental results show an increase in performance when this adaptation technique is applied. To the best knowledge of the authors, this is the first work dealing with adaptation for word spotting.
José A. Rodríguez 0001, Florent Perronnin, Gemma Sánchez, Josep Lladós 0001
ICPR2
2008 Universal and Adapted Vocabularies for Generic Visual Categorization
abstract
Generic Visual Categorization (GVC) is the pattern classification problem which consists in assigning labels to an image based on its semantic content. This is a challenging task as one has to deal with inherent object/scene variations as well as changes in viewpoint, lighting and occlusion. Several state-of-the-art GVC systems use a vocabulary of visual terms to characterize images with a histogram of visual word counts. We propose a novel practical approach to GVC based on a universal vocabulary, which describes the content of all the considered classes of images, and class vocabularies obtained through the adaptation of the universal vocabulary using class-specific data. The main novelty is that an image is characterized by a set of histograms - one per class - where each histogram describes whether the image content is best modeled by the universal vocabulary or the corresponding class vocabulary. This framework is applied to two types of local image features: low-level descriptors such as the popular SIFT and high-level histograms of word co-occurrences in a spatial neighborhood. It is shown experimentally on two challenging datasets (an in-house database of 19 categories and the PASCAL VOC 2006 dataset) that the proposed approach exhibits state-of-the-art performance at a modest computational cost.
Florent Perronnin
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Fisher Kernels on Visual Vocabularies for Image Categorization
abstract
Within the field of pattern classification, the Fisher kernel is a powerful framework which combines the strengths of generative and discriminative approaches. The idea is to characterize a signal with a gradient vector derived from a generative probability model and to subsequently feed this representation to a discriminative classifier. We propose to apply this framework to image categorization where the input signals are images and where the underlying generative model is a visual vocabulary: a Gaussian mixture model which approximates the distribution of low-level features in images. We show that Fisher kernels can actually be understood as an extension of the popular bag-of-visterms. Our approach demonstrates excellent performance on two challenging databases: an in-house database of 19 object/scene categories and the recently released VOC 2006 database. It is also very practical: it has low computational needs both at training and test time and vocabularies trained on one set of categories can be applied to another set without any significant loss in performance.
Florent Perronnin, Christopher R. Dance
CVPR1
2006 Adapted Vocabularies for Generic Visual Categorization
Florent Perronnin, Christopher R. Dance, Gabriela Csurka, Marco Bressan 0003
ECCV (4)1
2005 Online face detection and user authentication
abstract
The ability to verify automatically and with great accuracy the identity of a person has become crucial in everyday life. Biometrics is an emerging topic in the field of signal processing. Our research on biometrics aims at developing a complete framework useful to control access. This technical demo shows the latest image processing techniques for face detection developed at France Telecom and for face recognition developed at Eurécom. Using only one computer and one standard webcam, our biometric system detects the user face and the recognition algorithm uses this image to enable the access to a resource, a service or a location.
Caroline Mallauran, Jean-Luc Dugelay, Florent Perronnin, Christophe Garcia
ACM Multimedia3
2005 A Probabilistic Model of Face Mapping with Local Transformations and Its Application to Person Recognition
abstract
This paper proposes a new measure of "distance" between faces. This measure involves the estimation of the set of possible transformations between face images of the same person. The global transformation, which is assumed to be too complex for direct modeling, is approximated by a patchwork of local transformations, under a constraint imposing consistency between neighboring local transformations. The proposed system of local transformations and neighboring constraints is embedded within the probabilistic framework of a two-dimensional hidden Markov model. More specifically, we model two types of intraclass variabilities involving variations in facial expressions and illumination, respectively. The performance of the resulting method is assessed on a large data set consisting of four face databases. In particular, it is shown to outperform a leading approach to face recognition, namely, the Bayesian intra/extrapersonal classifier.
Florent Perronnin, Jean-Luc Dugelay, Kenneth Rose
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 From turbo hidden Markov models to turbo state-space models [face recognition applications]
abstract
We recently introduced a novel approximation of the intractable two-dimensional hidden Markov model (2D HMM), the turbo-HMM (T-HMM), which consists of a set of interconnected horizontal and vertical 1D HMMs. In this paper, we consider the extension of this framework to the continuous state HMM, generally referred to as the state-space model (SSM). We provide efficient approximate answers to the three following problems: (1) how to compute the likelihood of a set of observations; (2) how to find the sequence of states that best "explains" a set of observations; and (3) how to estimate the model parameters given a set of observations. The application of this work to the challenging problem of face recognition, in the presence of large illumination variations, illustrates the potential of our approach.
Florent Perronnin, Jean-Luc Dugelay
ICASSP (3)1
2003 Iterative decoding of two-dimensional hidden Markov models
abstract
While the hidden Markov model (HMM) has been extensively applied to one-dimensional problems, the complexity of its extension to two-dimensions grows exponentially with the data size and is intractable in most cases of interest. We introduce an efficient algorithm for approximate decoding of 2D HMMs, i.e., searching for the most likely state sequence. The basic idea is to approximate a 2D HMM with a turbo-HMM (T-HMM), which consists of horizontal and vertical 1D HMMs that "communicate", and allow iterated decoding (ID) of rows and columns by a modified version of the forward-backward algorithm. We derive the approach and its re-estimation equations. We then compare its performance to another algorithm designed for decoding 2D HMMs: the path constrained variable state Viterbi (PCVSV) algorithm (Li, J. et al., IEEE Trans. on Sig. Processing, vol.48, no.2, 2000). Finally, we combine our approach with PCVSV and show that the combination outperforms each algorithm taken separately.
Florent Perronnin, Jean-Luc Dugelay, Kenneth Rose
ICASSP (3)1
2003 Deformable face mapping for person identification
abstract
This paper introduces a novel deformable model for face mapping and its application to automatic person identification. While most face recognition techniques directly model the face, our goal is to model the transformation between face images of the same person. As a global face transformation may be too complex to be modeled in its entirety, it is approximated by a set of local transformations with the constraint that neighboring transformations must be consistent with each other. Local transformations and neighboring constraints are embedded within the probabilistic framework of a two-dimensional hidden Markov model (2-D HMM). Experimental results on a face identification task show that the new approach compares favorably to the popular Fisherfaces algorithm.
Florent Perronnin, Jean-Luc Dugelay, Kenneth Rose
ICIP (1)1
2002 Recent advances in biometric person authentication
abstract
Biometrics is an emerging topic in the field of signal processing. While technologies (e.g. audio, video) for biometrics have mostly been studied separately, ultimately, biometric technologies could find their strongest role as interwined and complementary pieces of a multi-modal authentication system. In this paper, a short overview of voice, fingerprint, and face authentication algorithms is provided.
Jean-Luc Dugelay, Jean-Claude Junqua, Constantine Kotropoulos, Roland Kuhn 0001, Florent Perronnin, Ioannis Pitas
ICASSP5
2001 Very fast adaptation with a compact context-dependent eigenvoice model
abstract
The "eigenvoice" technique achieves rapid speaker adaptation by employing prior knowledge of speaker space obtained from reference speakers to place strong constraints on the initial model for each new speaker. It has previously been shown to yield very fast adaptation for a large-vocabulary system. In this paper, we describe a new way of applying the eigenvoice technique to context-dependent acoustic modeling, called the "eigencentroid plus delta trees" (EDT) model. Here, the context-dependent model is defined so that it consists of a speaker-dependent component with a small number of parameters linked to a speaker-independent component with far more parameters. The eigenvoice technique can then be applied to the speaker-dependent component alone to attain very fast adaptation of the entire context-dependent model (e.g., 10% relative reduction in error rate after 3 sentences). EDT requires only a small number of parameters to represent speaker space and works even if only a small amount of data is available per reference speaker.
Roland Kuhn 0001, Florent Perronnin, Patrick Nguyen, Jean-Claude Junqua, Luca Rigazio
ICASSP2
2001 Maximum-likelihood training of a bipartite acoustic model for speech recognition
Florent Perronnin, Roland Kuhn 0001, Patrick Nguyen, Jean-Claude Junqua
INTERSPEECH1