EDBT 2026 Demo / reviewers in the wild / expert
Hervé Jégou
dblp:19/2115
· DBLP profile ↗
103ranked-venue papers
22as first author
22since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 12 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 13 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Computer networks · 3 · 2 first-authorTheory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Faiss LibraryabstractVector databases typically manage large collections of embedding vectors. As AI applications are growing rapidly, the number of embeddings that need to be stored and indexed is increasing. The Faiss library is dedicated to vector similarity search, a core functionality of vector databases. Faiss is a toolkit of indexing methods and related primitives used to search, cluster, compress and transform vectors. This paper describes the trade-offs in vector search and the design principles of Faiss in terms of structure, approach to optimization and interfacing. We benchmark key features of the library and discuss a few selected use cases to highlight its broad applicability. Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson 0004, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, Hervé Jégou |
IEEE Trans. Big Data | 9 |
| 2025 | Neutral residues: revisiting adapters for model extensionabstractWe address the problem of extending a pre-trained large language model to a new domain that was not seen during training. Standard techniques, such as fine-tuning or low-rank adaptation (LoRA) are successful at domain adaptation, but do not formally add capacity to the model. This often leads to a trade-off, between performing well on the new domain vs. degrading performance on the original domain.
Here, we propose to revisit and improve adapters to extend LLMs. Our paper analyzes this extension problem from three angles: data, architecture and training procedure, which are advantageously considered jointly. The resulting method, called neutral residues, modifies adapters in a way that leads to each new residual block to output near-zeros on the original domain. This solution leads to strong results when adapting a state-of-the-art model originally trained on English to a new language. Neutral residues significantly outperforms competing approaches such as fine-tuning, LoRA or vanilla adapters in terms of the trade-off between learning the new language and not forgetting English. Franck Signe Talla, Edouard Grave, Hervé Jégou |
ICML | 3 |
| 2023 | Co-training 2L Submodels for Visual RecognitionabstractWe introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, “submodels”, with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed “co-sub”, uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including RegNet, ViT, PiT, XCiT, Swin and ConvNext. Our training strategy improves their results in comparable settings. For instance, a ViT-B pretrained with cosub on ImageNet-21k obtains 87.4% top1 acc. @448 on ImageNet-val. Hugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski, Jakob Verbeek, Hervé Jégou |
CVPR | 6 |
| 2023 | Variable Rate Allocation for Vector-Quantized AutoencodersabstractVector-quantized autoencoders have recently gained interest in image compression, generation and self-supervised learning. However, as a neural compression method, they lack the possibility to allocate a variable number of bits to each image location, e.g. according to the semantic content or local saliency. In this paper, we address this limitation in a simple yet effective way. We adopt a product quantizer (PQ) that produces a set of discrete codes for each image patch rather than a single index. This PQ-autoencoder is trained end-to-end with a structured dropout that selectively masks a variable number of codes at each location. These mechanisms force the decoder to reconstruct the original image based on partial information and allow us to control the local rate. The resulting model can compress images on a wide range of operating points of the rate-distortion curve and can be paired with any external method for saliency estimation to control the compression rate at a local level. We demonstrate the effectiveness of our approach on the popular Kodak and ImageNet datasets by measuring both distortion and perceptual quality metrics. Federico Baldassarre, Alaaeldin El-Nouby, Hervé Jégou |
ICASSP | 3 |
| 2023 | The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsabstractGenerative image modeling enables a wide range of applications but raises ethical concerns about responsible deployment. We introduce an active content tracing method combining image watermarking and Latent Diffusion Models. The goal is for all generated images to conceal an invisible watermark allowing for future detection and/or identification. The method quickly fine-tunes the latent decoder of the image generator, conditioned on a binary signature. A pre-trained watermark extractor recovers the hidden signature from any generated image and a statistical test then determines whether it comes from the generative model. We evaluate the invisibility and robustness of the watermarks on a variety of generation tasks, showing that the Stable Signature is robust to image modifications. For instance, it detects the origin of an image generated from a text prompt, then cropped to keep 10% of the content, with 90+% accuracy at a false positive rate below 10−6. Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, Teddy Furon |
ICCV | 3 |
| 2023 | Active Image Indexing
Pierre Fernandez, Matthijs Douze, Hervé Jégou, Teddy Furon |
ICLR | 3 |
| 2023 | Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood ModelsabstractLossy image compression aims to represent images in as few bits as possible while maintaining fidelity to the original. Theoretical results indicate that optimizing distortion metrics such as PSNR or MS-SSIM necessarily leads to a discrepancy in the statistics of original images from those of reconstructions, in particular at low bitrates, often manifested by the blurring of the compressed images. Previous work has leveraged adversarial discriminators to improve statistical fidelity. Yet these binary discriminators adopted from generative modeling tasks may not be ideal for image compression. In this paper, we introduce a non-binary discriminator that is conditioned on quantized local image representations obtained via VQ-VAE autoencoders. Our evaluations on the CLIC2020, DIV2K and Kodak datasets show that our discriminator is more effective for jointly optimizing distortion (e.g., PSNR) and statistical fidelity (e.g., FID) than the PatchGAN of the state-of-the-art HiFiC model. On CLIC2020, we obtain the same FID as HiFiC with 30-40% fewer bits. Matthew J. Muckley, Alaaeldin El-Nouby, Karen Ullrich, Hervé Jégou, Jakob Verbeek |
ICML | 4 |
| 2023 | Birth of a Transformer: A Memory ViewpointabstractLarge language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms in order to make them more reliable. These models appear to store vast amounts of knowledge from their training data, and to adapt quickly to new information provided in their context or prompt. We study how transformers balance these two types of knowledge by considering a synthetic setup where tokens are generated from either global or context-specific bigram distributions. By a careful empirical analysis of the training process on a simplified two-layer transformer, we illustrate the fast learning of global bigrams and the slower development of an "induction head" mechanism for the in-context bigrams. We highlight the role of weight matrices as associative memories, provide theoretical insights on how gradients enable their learning during training, and study the role of data-distributional properties. Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou, Léon Bottou |
NeurIPS | 4 |
| 2023 | ResMLP: Feedforward Networks for Image Classification With Data-Efficient TrainingabstractWe present ResMLP, an architecture built entirely upon multi-layer perceptrons for image classification. It is a simple residual network that alternates (i) a linear layer in which image patches interact, independently and identically across channels, and (ii) a two-layer feed-forward network in which channels interact independently per patch. When trained with a modern training strategy using heavy data-augmentation and optionally distillation, it attains surprisingly good accuracy/complexity trade-offs on ImageNet. We also train ResMLP models in a self-supervised setup, to further remove priors from employing a labelled dataset. Finally, by adapting our model to machine translation we achieve surprisingly good results. We share pre-trained models and our code based on the Timm library. Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, Hervé Jégou |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2022 | Three Things Everyone Should Know About Vision Transformers
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, Hervé Jégou |
ECCV (24) | 5 |
| 2022 | DeiT III: Revenge of the ViT
Hugo Touvron, Matthieu Cord, Hervé Jégou |
ECCV (24) | 3 |
| 2022 | Watermarking Images in Self-Supervised Latent SpacesabstractWe revisit watermarking techniques based on pre-trained deep networks, in the light of self-supervised approaches. We present a way to embed both marks and binary messages into their latent spaces, leveraging data augmentation at marking time. Our method can operate at any resolution and creates watermarks robust to a broad range of transformations (rotations, crops, JPEG, contrast, etc). It significantly outperforms the previous zero-bit methods, and its performance on multi-bit watermarking is on par with state-of-the-art encoder-decoder architectures trained end-to-end for watermarking. The code is available at github.com/facebookresearch/ssl_watermarking. Pierre Fernandez, Alexandre Sablayrolles, Teddy Furon, Hervé Jégou, Matthijs Douze |
ICASSP | 4 |
| 2022 | Nearest Neighbor Search with Compact Codes: A Decoder PerspectiveabstractModern approaches for fast retrieval of similar vectors on billion-scaled datasets rely on compressed-domain approaches such as binary sketches or product quantization. These methods minimize a certain loss, typically the Mean Squared Error or other objective functions tailored to the retrieval problem. In this paper, we re-interpret popular methods such as binary hashing or product quantizers as auto-encoders, and point out that they implicitly make suboptimal assumptions on the form of the decoder. We design backward-compatible decoders that improve the reconstruction of the vectors from the same codes, which translates to a better performance in nearest neighbor search. Our method significantly improves over binary hashing methods and product quantization on popular benchmarks. Kenza Amara, Matthijs Douze, Alexandre Sablayrolles, Hervé Jégou |
ICMR | 4 |
| 2021 | Gradient-based Adversarial Attacks against Text TransformersabstractWe propose the first general-purpose gradientbased adversarial attack against transformer models.Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization.We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks, outperforming prior work in terms of adversarial success rate with matching imperceptibility as per automated and human evaluation.Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs. Chuan Guo 0001, Alexandre Sablayrolles, Hervé Jégou, Douwe Kiela |
EMNLP (1) | 3 |
| 2021 | Emerging Properties in Self-Supervised Vision TransformersabstractIn this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [16] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [26], multi-crop training [9], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base. Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, Armand Joulin |
ICCV | 4 |
| 2021 | LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceabstractWe design a family of image classification architectures that optimize the trade-off between accuracy and efficiency in a high-speed regime. Our work exploits recent findings in attention-based architectures, which are competitive on highly parallel processing hardware. We revisit principles from the extensive literature on convolutional neural networks to apply them to transformers, in particular activation maps with decreasing resolutions. We also introduce the attention bias, a new way to integrate positional information in vision transformers.As a result, we propose LeViT: a hybrid neural network for fast inference image classification. We consider different measures of efficiency on different hardware platforms, so as to best reflect a wide range of application scenarios. Our extensive experiments empirically validate our technical choices and show they are suitable to most architectures. Overall, LeViT significantly outperforms existing convnets and vision transformers with respect to the speed/accuracy tradeoff. For example, at 80% ImageNet top-1 accuracy, LeViT is 5 times faster than EfficientNet on CPU. We release the code at https://github.com/facebookresearch/LeViT. Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, Matthijs Douze |
ICCV | 6 |
| 2021 | Going deeper with Image TransformersabstractTransformers have been recently adapted for large scale image classification, achieving high scores shaking up the long supremacy of convolutional neural networks. However the optimization of vision transformers has been little studied so far. In this work, we build and optimize deeper transformer networks for image classification. In particular, we investigate the interplay of architecture and optimization of such dedicated transformers. We make two architecture changes that significantly improve the accuracy of deep transformers. This leads us to produce models whose performance does not saturate early with more depth, for in-stance we obtain 86.5% top-1 accuracy on Imagenet when training with no external data, we thus attain the current sate of the art with less floating-point operations and parameters. Our best model establishes the new state of the art on Imagenet with Reassessed labels and Imagenet-V2 / match frequency, in the setting with no additional training data. We share our code and models1. Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, Hervé Jégou |
ICCV | 5 |
| 2021 | Grafit: Learning fine-grained image representations with coarse labelsabstractThis paper tackles the problem of learning a finer representation than the one provided by training labels. This enables fine-grained category retrieval of images in a collection annotated with coarse labels only.Our network is learned with a nearest-neighbor classifier objective, and an instance loss inspired by self-supervised learning. By jointly leveraging the coarse labels and the underlying fine-grained latent space, it significantly improves the accuracy of category-level retrieval methods.Our strategy outperforms all competing methods for retrieving or classifying images at a finer granularity than that available at train time. It also improves the accuracy for transfer learning tasks to fine-grained datasets. Hugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord, Hervé Jégou |
ICCV | 5 |
| 2021 | Training with Quantization Noise for Extreme Model Compression
Pierre Stock, Angela Fan, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, Armand Joulin |
ICLR | 6 |
| 2021 | Training data-efficient image transformers & distillation through attentionabstractRecently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. These high-performing vision transformers are pre-trained with hundreds of millions of images using a large infrastructure, thereby limiting their adoption. In this work, we produce competitive convolution-free transformers trained on ImageNet only using a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop) on ImageNet with no external data. We also introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention, typically from a convnet teacher. The learned transformers are competitive (85.2% top-1 acc.) with the state of the art on ImageNet, and similarly when transferred to other tasks. We will share our code and models. Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, Hervé Jégou |
ICML | 6 |
| 2021 | XCiT: Cross-Covariance Image TransformersabstractFollowing their success in natural language processing, transformers have recently shown much promise for computer vision. The self-attention operation underlying transformers yields global interactions between all tokens ,i.e. words or image patches, and enables flexible modelling of image data beyond the local interactions of convolutions. This flexibility, however, comes with a quadratic complexity in time and memory, hindering application to long sequences and high-resolution images. We propose a “transposed” version of self-attention that operates across feature channels rather than tokens, where the interactions are based on the cross-covariance matrix between keys and queries. The resulting cross-covariance attention (XCA) has linear complexity in the number of tokens, and allows efficient processing of high-resolution images.Our cross-covariance image transformer (XCiT) is built upon XCA. It combines the accuracy of conventional transformers with the scalability of convolutional architectures. We validate the effectiveness and generality of XCiT by reporting excellent results on multiple vision benchmarks, including image classification and self-supervised feature learning on ImageNet-1k, object detection and instance segmentation on COCO, and semantic segmentation on ADE20k.We will opensource our code and trained models to reproduce the reported results. Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, Hervé Jégou |
NeurIPS | 11 |
| 2021 | Billion-Scale Similarity Search with GPUsabstractSimilarity search finds application in database systems handling complex data such as images or videos, which are typically represented by high-dimensional features and require specific indexing structures. This paper tackles the problem of better utilizing GPUs for this task. While GPUs excel at data parallel tasks such as distance computation, prior approaches in this domain are bottlenecked by algorithms that expose less parallelism, such ask-min selection, or make poor use of the memory hierarchy. We propose a novel design fork-selection. We apply it in different similarity search scenarios, by optimizing brute-force, approximate and compressed-domain search based on product quantization. In all these setups, we outperform the state of the art by large margins. Our implementation operates at up to 55 percent of theoretical peak performance, enabling a nearest neighbor implementation that is 8.5 × faster than prior GPU state of the art. It enables the construction of a high accuracyk-NN graph on 95 million images from the Yfcc100M dataset in 35 minutes, and of a graph connecting 1 billion vectors in less than 12 hours on 4 Maxwell Titan X GPUs. We have open-sourced our approach for the sake of comparison and reproducibility. Jeff Johnson 0004, Matthijs Douze, Hervé Jégou |
IEEE Trans. Big Data | 3 |
| 2020 | And the Bit Goes Down: Revisiting the Quantization of Neural Networks
Pierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham, Hervé Jégou |
ICLR | 5 |
| 2020 | Radioactive data: tracing through trainingabstractData tracing determines whether particular data samples have been used to train a model. We propose a new technique, radioactive data, that makes imperceptible changes to these samples such that any model trained on them will bear an identifiable mark. Given a trained model, our technique detects the use of radioactive data and provides a level of confidence (p-value). Experiments on large-scale benchmarks (Imagenet), with standard architectures (Resnet-18, VGG-16, Densenet-121) and training procedures, show that we detect radioactive data with high confidence (p<0.0001) when only 1% of the data used to train a model is radioactive. Our radioactive mark is resilient to strong data augmentations and variations of the model architecture. As a result, it offers a much higher signal-to-noise ratio than data poisoning and backdoor methods. Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, Hervé Jégou |
ICML | 4 |
| 2019 | Spreading vectors for similarity search
Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, Hervé Jégou |
ICLR (Poster) | 4 |
| 2019 | Equi-normalization of Neural Networks
Pierre Stock, Benjamin Graham, Rémi Gribonval, Hervé Jégou |
ICLR (Poster) | 4 |
| 2019 | White-box vs Black-box: Bayes Optimal Strategies for Membership InferenceabstractMembership inference determines, given a sample and trained parameters of a machine learning model, whether the sample was part of the training set. In this paper, we derive the optimal strategy for membership inference with a few assumptions on the distribution of the parameters. We show that optimal attacks only depend on the loss function, and thus black-box attacks are as good as white-box attacks. As the optimal strategy is not tractable, we provide approximations of it leading to several inference methods, and show that existing membership inference methods are coarser approximations of this optimal strategy. Our membership attacks outperform the state of the art in various settings, ranging from a simple logistic regression to more complex architectures and datasets, such as ResNet-101 and Imagenet. Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, Yann Ollivier, Hervé Jégou |
ICML | 5 |
| 2019 | Large Memory Layers with Product KeysabstractThis paper introduces a structured memory which can be easily integrated into a neural network. The memory is very large by design and significantly increases the capacity of the architecture, by up to a billion parameters with a negligible computational overhead. Its design and access pattern is based on product keys, which enable fast and exact nearest neighbor search. The ability to increase the number of parameters while keeping the same computational budget lets the overall system strike a better trade-off between prediction accuracy and computation efficiency both at training and test time. This memory layer allows us to tackle very large scale language modeling tasks. In our experiments we consider a dataset with up to 30 billion words, and we plug our memory layer in a state-of-the-art transformer-based architecture. In particular, we found that a memory augmented model with only 12 layers outperforms a baseline transformer model with 24 layers, while being twice faster at inference time. We release our code for reproducibility purposes. Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou |
NeurIPS | 5 |
| 2019 | Fixing the train-test resolution discrepancyabstractData-augmentation is key to the training of neural networks for image classification. This paper first shows that existing augmentations induce a significant discrepancy between the size of the objects seen by the classifier at train and test time: in fact, a lower train resolution improves the classification at test time! We then propose a simple strategy to optimize the classifier performance, that employs different train and test resolutions. It relies on a computationally cheap fine-tuning of the network at the test resolution. This enables training strong classifiers using small training images, and therefore significantly reduce the training time. For instance, we obtain 77.1% top-1 accuracy on ImageNet with a ResNet-50 trained on 128x128 images, and 79.8% with one trained at 224x224. A ResNeXt-101 32x48d pre-trained with weak supervision on 940 million 224x224 images and further optimized with our technique for test resolution 320x320 achieves 86.4% top-1 accuracy (top-5: 98.0%). To the best of our knowledge this is the highest ImageNet single-crop accuracy to date. Hugo Touvron, Andrea Vedaldi, Matthijs Douze, Hervé Jégou |
NeurIPS | 4 |
| 2019 | Understanding and Improving Kernel Local Descriptors
Arun Mukundan, Giorgos Tolias, Andrei Bursuc, Hervé Jégou, Ondrej Chum |
Int. J. Comput. Vis. | 4 |
| 2018 | LAMV: Learning to Align and Match Videos With Kernelized Temporal LayersabstractThis paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal alignments by maximizing the scores between two sequences of vectors, according to a time-sensitive similarity metric parametrized in the Fourier domain. We learn this layer with a temporal proposal strategy, in which we minimize a triplet loss that takes into account both the localization accuracy and the recognition rate. We evaluate our approach on video alignment, copy detection and event retrieval. Our approach outperforms the state on the art on temporal video alignment and video copy detection datasets in comparable setups. It also attains the best reported results for particular event search, while precisely aligning videos. Lorenzo Baraldi 0001, Matthijs Douze, Rita Cucchiara, Hervé Jégou |
CVPR | 4 |
| 2018 | Low-Shot Learning With Large-Scale DiffusionabstractThis paper considers the problem of inferring image labels from images when only a few annotated examples are available at training time. This setup is often referred to as low-shot learning, where a standard approach is to retrain the last few layers of a convolutional neural network learned on separate classes for which training examples are abundant. We consider a semi-supervised setting based on a large collection of images to support label propagation. This is possible by leveraging the recent advances on large-scale similarity graph construction. We show that despite its conceptual simplicity, scaling label propagation up to hundred millions of images leads to state of the art accuracy in the low-shot learning regime. Matthijs Douze, Arthur Szlam, Bharath Hariharan, Hervé Jégou |
CVPR | 4 |
| 2018 | Link and Code: Fast Indexing With Graphs and Compact Regression CodesabstractSimilarity search approaches based on graph walks have recently attained outstanding speed-accuracy trade-offs, taking aside the memory requirements. In this paper, we revisit these approaches by considering, additionally, the memory constraint required to index billions of images on a single server. This leads us to propose a method based both on graph traversal and compact representations. We encode the indexed vectors using quantization and exploit the graph structure to refine the similarity estimation. In essence, our method takes the best of these two worlds: the search strategy is based on nested graphs, thereby providing high precision with a relatively small set of comparisons. At the same time it offers a significant memory compression. As a result, our approach outperforms the state of the art on operating points considering 64-128 bytes per vector, as demonstrated by our results on two billion-scale public benchmarks. Matthijs Douze, Alexandre Sablayrolles, Hervé Jégou |
CVPR | 3 |
| 2018 | Loss in Translation: Learning Bilingual Word Mapping with a Retrieval CriterionabstractContinuous word representations learned separately on distinct languages can be aligned so that their words become comparable in a common space.Existing works typically solve a quadratic problem to learn a orthogonal matrix aligning a bilingual lexicon, and use a retrieval criterion for inference.In this paper, we propose an unified formulation that directly optimizes a retrieval criterion in an end-to-end fashion.Our experiments on standard benchmarks show that our approach outperforms the state of the art on word translation, with the biggest improvements observed for distant language pairs such as English-Chinese. Armand Joulin, Piotr Bojanowski, Tomás Mikolov, Hervé Jégou, Edouard Grave |
EMNLP | 4 |
| 2018 | Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou |
ICLR (Poster) | 5 |
| 2018 | Memory Vectors for Similarity Search in High-Dimensional SpacesabstractWe study an indexing architecture to store and search in a database of high-dimensional vectors from the perspective of statistical signal processing and decision theory. This architecture is composed of several memory units, each of which summarizes a fraction of the database by a single representative vector. The potential similarity of the query to one of the vectors stored in the memory unit is gauged by a simple correlation with the memory unit's representative vector. This representative optimizes the test of the following hypothesis: the query is independent from any vector in the memory unit versus the query is a simple perturbation of one of the stored vectors. Compared to exhaustive search, our approach finds the most similar database vectors significantly faster without a noticeable reduction in search quality. Interestingly, the reduction of complexity is provably better in high-dimensional spaces. We empirically demonstrate its practical interest in a large-scale image search scenario with off-the-shelf state-of-the-art descriptors. Ahmet Iscen, Teddy Furon, Vincent Gripon, Michael G. Rabbat, Hervé Jégou |
IEEE Trans. Big Data | 5 |
| 2017 | How should we evaluate supervised hashing?abstractHashing produces compact representations for documents, to perform tasks like classification or retrieval based on these short codes. When hashing is supervised, the codes are trained using labels on the training data. This paper first shows that the evaluation protocols used in the literature for supervised hashing are not satisfactory: we show that a trivial solution that encodes the output of a classifier significantly outperforms existing supervised or semi-supervised methods, while using much shorter codes. We then propose two alternative protocols for supervised hashing: one based on retrieval on a disjoint set of classes, and another based on transfer learning to new classes. We provide two baseline methods for image-related tasks to assess the performance of (semi-)supervised hashing: without coding and with unsupervised codes. These baselines give a lower- and upper-bound on the performance of a supervised hashing scheme. Alexandre Sablayrolles, Matthijs Douze, Nicolas Usunier, Hervé Jégou |
ICASSP | 4 |
| 2017 | Efficient softmax approximation for GPUsabstractWe propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word distribution to form clusters that explicitly minimize the expectation of computation time. Our approach further reduces the computational cost by exploiting the specificities of modern architectures and matrix-matrix vector operations, making it particularly suited for graphical processing units. Our experiments carried out on standard benchmarks, such as EuroParl and One Billion Word, show that our approach brings a large gain in efficiency over standard approximations while achieving an accuracy close to that of the full softmax. The code of our method is available at https://github.com/facebookresearch/adaptive-softmax. Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, Hervé Jégou |
ICML | 5 |
| 2017 | Tubelets: Unsupervised Action Proposals from Spatiotemporal Super-VoxelsabstractThis paper considers the problem of localizing actions in videos as sequences of bounding boxes. The objective is to generate action proposals that are likely to include the action of interest, ideally achieving high recall with few proposals. Our contributions are threefold. First, inspired by selective search for object proposals, we introduce an approach to generate action proposals from spatiotemporal super-voxels in an unsupervised manner, we call them Tubelets . Second, along with the static features from individual frames our approach advantageously exploits motion. We introduce independent motion evidence as a feature to characterize how the action deviates from the background and explicitly incorporate such motion information in various stages of the proposal generation. Finally, we introduce spatiotemporal refinement of Tubelets, for more precise localization of actions, and pruning to keep the number of Tubelets limited. We demonstrate the suitability of our approach by extensive experiments for action proposal quality and action localization on three public datasets: UCF Sports, MSR-II and UCF101. For action proposal quality, our unsupervised proposals beat all other existing approaches on the three datasets. For action localization, we show top performance on both the trimmed videos of UCF Sports and UCF101 as well as the untrimmed videos of MSR-II. Mihir Jain, Jan C. van Gemert, Hervé Jégou, Patrick Bouthemy, Cees Snoek |
Int. J. Comput. Vis. | 3 |
| 2017 | Interferences in Match KernelsabstractWe consider the design of an image representation that embeds and aggregates a set of local descriptors into a single vector. Popular representations of this kind include the bag-of-visual-words, the Fisher vector and the VLAD. When two such image representations are compared with the dot-product, the image-to-image similarity can be interpreted as a match kernel. In match kernels, one has to deal with interference, i.e., with the fact that even if two descriptors are unrelated, their matching score may contribute to the overall similarity. We formalise this problem and propose two related solutions, both aimed at equalising the individual contributions of the local descriptors in the final representation. These methods modify the aggregation stage by including a set of per-descriptor weights. They differ by the objective function that is optimised to compute those weights. The first is a "democratisation" strategy that aims at equalising the relative importance of each descriptor in the set comparison metric. The second one involves equalising the match of a single descriptor to the aggregated vector. These concurrent methods give a substantial performance boost over the state of the art in image search with short or mid-size vectors, as demonstrated by our experiments on standard public image retrieval benchmarks. Naila Murray, Hervé Jégou, Florent Perronnin, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Guest Editorial: Large-Scale Multimedia Data Retrieval, Classification, and UnderstandingabstractThe papers in this special section focus on multimedia data retrieval and classification via large-scale systems. Today, large collections of multimedia data are explosively created in different fields and have attracted increasing interest in the multimedia research area. Large-scale multimedia data provide great unprecedented opportunities to address many challenging research problems, e.g., enabling generic visual classification to bridge the well-known semantic gap by exploring large-scale data, offering a promising possibility for in-depth multimedia understanding, as well as discerning patterns and making better decisions by analyzing the large pool of data. Therefore, the techniques for large-scale multimedia retrieval, classification, and understanding are highly desired. Simultaneously, the explosion of multimedia data puts urgent needs for more sophisticated and robust models and algorithms to retrieve, classify, and understand these data. Another interesting challenge is, how can the traditional machine learning algorithms be scaled up to millions and even billions of items with thousands of dimensionalities? This motivated the community to design parallel and distributed machine learning platforms, exploiting GPUs as well as developing practical algorithms. Besides, it is also important to exploit the commonalities and differences between different tasks, e.g., image retrieval and classification have much in common while different indexing methods evolve in a mutually supporting way. Jingkuan Song, Hervé Jégou, Cees Snoek, Qi Tian 0001, Nicu Sebe |
IEEE Trans. Multim. | 2 |
| 2016 | Polysemous Codes
Matthijs Douze, Hervé Jégou, Florent Perronnin |
ECCV (2) | 2 |
| 2016 | Approximate Search with Quantized Sparse Representations
Himalaya Jain, Patrick Pérez, Rémi Gribonval, Joaquin Zepeda, Hervé Jégou |
ECCV (7) | 5 |
| 2016 | Circulant Temporal Encoding for Video Retrieval and Temporal Alignment
Matthijs Douze, Jérôme Revaud, Jakob Verbeek, Hervé Jégou, Cordelia Schmid |
Int. J. Comput. Vis. | 4 |
| 2016 | Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images
Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
Int. J. Comput. Vis. | 3 |
| 2016 | Erratum to: Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images
Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
Int. J. Comput. Vis. | 3 |
| 2015 | Early burst detection for memory-efficient image retrievalabstractRecent works show that image comparison based on local descriptors is corrupted by visual bursts, which tend to dominate the image similarity. The existing strategies, like power-law normalization, improve the results by discounting the contribution of visual bursts to the image similarity. Miaojing Shi, Yannis Avrithis, Hervé Jégou |
CVPR | 3 |
| 2015 | Kernel Local Descriptors with Implicit Rotation MatchingabstractIn this work we design a kernelized local feature descriptor and propose a matching scheme for aligning patches quickly and automatically. We analyze the SIFT descriptor from a kernel view and identify and reproduce some of its underlying benefits. We overcome the quantization artifacts of SIFT by encoding pixel attributes in a continous manner via explicit feature maps. Experiments performed on the patch dataset of Brown et al. [3] show the superiority of our descriptor over methods based on supervised learning. Andrei Bursuc, Giorgos Tolias, Hervé Jégou |
ICMR | 3 |
| 2015 | Multiple Measurements and Joint Dimensionality Reduction for Large Scale Image Search with Short VectorsabstractThis paper addresses the construction of a short-vector (128D) image representation for large-scale image and particular object retrieval. In particular, the method of joint dimensionality reduction of multiple vocabularies is considered. We study a variety of vocabulary generation techniques: different k-means initializations, different descriptor transformations, different measurement regions for descriptor extraction. Our extensive evaluation shows that different combinations of vocabularies, each partitioning the descriptor space in a different yet complementary manner, results in a significant performance improvement, which exceeds the state-of-the-art. Filip Radenovic, Hervé Jégou, Ondrej Chum |
ICMR | 2 |
| 2015 | Memory Vectors for Particular Object Retrieval with Multiple QueriesabstractWe address the problem of retrieving all the images containing a specific object in a large image collection, where the input query is given as a set of representative images of the object. This problem is referred to as multiple queries in the literature. For images described with bag-of-visual-words (BOW), one of the best performing approach amounts to simply averaging the query descriptors. Ronan Sicre, Hervé Jégou |
ICMR | 2 |
| 2015 | Temporal Matching Kernel with Explicit Feature MapsabstractThis paper proposes a framework for content-based video retrieval that addresses various tasks as particular event retrieval, copy detection or video synchronization. Given a video query, the method is able to efficiently retrieve, from a large collection, similar video events or near-duplicates with temporarily consistent excerpts. As a byproduct of the representation, it provides a precise temporal alignment of the query and the detected video excerpts. Sébastien Poullot, Shunsuke Tsukatani, Phuong Anh Nguyen 0002, Hervé Jégou, Shin'ichi Satoh 0001 |
ACM Multimedia | 4 |
| 2015 | Rotation and translation covariant match kernels for image retrieval
Giorgos Tolias, Andrei Bursuc, Teddy Furon, Hervé Jégou |
Comput. Vis. Image Underst. | 4 |
| 2015 | A Comparison of Dense Region Detectors for Image Search and Fine-Grained ClassificationabstractWe consider a pipeline for image classification or search based on coding approaches like bag of words or Fisher vectors. In this context, the most common approach is to extract the image patches regularly in a dense manner on several scales. This paper proposes and evaluates alternative choices to extract patches densely. Beyond simple strategies derived from regular interest region detectors, we propose approaches based on superpixels, edges, and a bank of Zernike filters used as detectors. The different approaches are evaluated on recent image retrieval and fine-grained classification benchmarks. Our results show that the regular dense detector is outperformed by other methods in most situations, leading us to improve the state-of-the-art in comparable setups on standard retrieval and fined-grained benchmarks. As a byproduct of our study, we show that existing methods for blob and superpixel extraction achieve high accuracy if the patches are extracted along the edges and not around the detected regions. Ahmet Iscen, Giorgos Tolias, Philippe Henri Gosselin, Hervé Jégou |
IEEE Trans. Image Process. | 4 |
| 2014 | Action Localization with Tubelets from MotionabstractThis paper considers the problem of action localization, where the objective is to determine when and where certain actions appear. We introduce a sampling strategy to produce 2D+t sequences of bounding boxes, called tubelets. Compared to state-of-the-art alternatives, this drastically reduces the number of hypotheses that are likely to include the action of interest. Our method is inspired by a recent technique introduced in the context of image localization. Beyond considering this technique for the first time for videos, we revisit this strategy for 2D+t sequences obtained from super-voxels. Our sampling strategy advantageously exploits a criterion that reflects how action related motion deviates from background motion. We demonstrate the interest of our approach by extensive experiments on two public datasets: UCF Sports and MSR-II. Our approach significantly outperforms the state-of-the-art on both datasets, while restricting the search of actions to a fraction of possible bounding box sequences. Mihir Jain, Jan C. van Gemert, Hervé Jégou, Patrick Bouthemy, Cees Snoek |
CVPR | 3 |
| 2014 | Triangulation Embedding and Democratic Aggregation for Image SearchabstractWe consider the design of a single vector representation for an image that embeds and aggregates a set of local patch descriptors such as SIFT. More specifically we aim to construct a dense representation, like the Fisher Vector or VLAD, though of small or intermediate size. We make two contributions, both aimed at regularizing the individual contributions of the local descriptors in the final representation. The first is a novel embedding method that avoids the dependency on absolute distances by encoding directions. The second contribution is a "democratization" strategy that further limits the interaction of unrelated descriptors in the aggregation stage. These methods are complementary and give a substantial performance boost over the state of the art in image search with short or mid-size vectors, as demonstrated by our experiments on standard public image retrieval benchmarks. Hervé Jégou, Andrew Zisserman |
CVPR | 1 |
| 2014 | Orientation Covariant Aggregation of Local Descriptors with Embeddings
Giorgos Tolias, Teddy Furon, Hervé Jégou |
ECCV (6) | 3 |
| 2014 | Beyond "project and sign" for cosine estimation with binary codesabstractMany nearest neighbor search algorithms rely on encoding real vectors into binary vectors. The most common strategy projects the vectors onto random directions and takes the sign to produce so-called sketches. This paper discusses the sub-optimality of this choice, and proposes a better encoding strategy based on the quantization and reconstruction points of view. Our second contribution is a novel asymmetric estimator for the cosine similarity. Similar to previous asymmetric schemes, the query is not quantized and the similarity is computed in the compressed domain. Both our contribution leads to improve the quality of nearest neighbor search with binary codes. Its efficiency compares favorably against a recent encoding technique. Raghavendran Balu, Teddy Furon, Hervé Jégou |
ICASSP | 3 |
| 2014 | Instance classification with prototype selectionabstractWe address the problem of instance classification: our goal is to annotate images with tags corresponding to objects classes which exhibit small intra-class variations such as logos, products or landmarks. We propose a novel algorithm for the selection of class-specific prototypes which are used in a voting-based classification scheme. We show significant improvements over two state-of-the-art methods, namely the Fisher vector and Hamming Embedding, on two challenging methods of logos and vehicles. Josip Krapac, Florent Perronnin, Teddy Furon, Hervé Jégou |
ICMR | 4 |
| 2014 | The Yael LibraryabstractThis paper introduces Yael, a library implementing computationally intensive functions used in large scale image retrieval, such as neighbor search, clustering and inverted files. The library offers interfaces for C, Python and Matlab. Along with a brief tutorial, we analyze and discuss some of our implementation choices, and their impact on efficiency. Matthijs Douze, Hervé Jégou |
ACM Multimedia | 2 |
| 2014 | A Group Testing Framework for Similarity Search in High-dimensional SpacesabstractThis paper introduces a group testing framework for detecting large similarities between high-dimensional vectors, such as descriptors used in state-of-the-art description of multimedia documents.At the crossroad of multimedia information retrieval and signal processing, we produce a set of group representations that jointly encode several vectors into a single one, in the spirit of group testing approaches. By comparing a query vector to several of these intermediate representations, we screen the large values taken by the similarities between the query and all the vectors, at a fraction of the cost of exhaustive similarity calculation. Unlike concurrent indexing methods that suffer from the curse of dimensionality, our method exploits the properties of high-dimensional spaces. It therefore complements other strategies for approximate nearest neighbor search. Our preliminary experiments demonstrate the potential of group testing for searching large databases of multimedia objects represented by vectors. We obtain a large improvement in terms of the theoretical complexity, at the cost of a small or negligible decrease of accuracy.We hope that this preliminary work will pave the way to subsequent works for multimedia retrieval with limited resources. Miaojing Shi, Teddy Furon, Hervé Jégou |
ACM Multimedia | 3 |
| 2014 | Visual query expansion with or without geometry: Refining local descriptors by feature aggregation
Giorgos Tolias, Hervé Jégou |
Pattern Recognit. | 2 |
| 2014 | Revisiting the Fisher vector for fine-grained classification
Philippe Henri Gosselin, Naila Murray, Hervé Jégou, Florent Perronnin |
Pattern Recognit. Lett. | 3 |
| 2013 | Oriented pooling for dense and non-dense rotation-invariant featuresabstractInternational audience Wanlei Zhao, Guillaume Gravier, Hervé Jégou |
BMVC | 3 |
| 2013 | Better Exploiting Motion for Better Action RecognitionabstractSeveral recent works on action recognition have attested the importance of explicitly integrating motion characteristics in the video description. This paper establishes that adequately decomposing visual motion into dominant and residual motions, both in the extraction of the space-time trajectories and for the computation of descriptors, significantly improves action recognition algorithms. Then, we design a new motion descriptor, the DCS descriptor, based on differential motion scalar quantities, divergence, curl and shear features. It captures additional information on the local motion patterns enhancing results. Finally, applying the recent VLAD coding technique proposed in image retrieval provides a substantial improvement for action recognition. Our three contributions are complementary and lead to outperform all reported results by a significant margin on three challenging datasets, namely Hollywood 2, HMDB51 and Olympic Sports. Mihir Jain, Hervé Jégou, Patrick Bouthemy |
CVPR | 2 |
| 2013 | Event Retrieval in Large Video Collections with Circulant Temporal EncodingabstractThis paper presents an approach for large-scale event retrieval. Given a video clip of a specific event, eg, the wedding of Prince William and Kate Middleton, the goal is to retrieve other videos representing the same event from a dataset of over 100k videos. Our approach encodes the frame descriptors of a video to jointly represent their appearance and temporal order. It exploits the properties of circulant matrices to compare the videos in the frequency domain. This offers a significant gain in complexity and accurately localizes the matching parts of videos. Furthermore, we extend product quantization to complex vectors in order to compress our descriptors, and to compare them in the compressed domain. Our method outperforms the state of the art both in search quality and query time on two large-scale video benchmarks for copy detection, Trecvid and CCWeb. Finally, we introduce a challenging dataset for event retrieval, EVVE, and report the performance on this dataset. Jérôme Revaud, Matthijs Douze, Cordelia Schmid, Hervé Jégou |
CVPR | 4 |
| 2013 | Stable Hyper-pooling and Query Expansion for Event DetectionabstractThis paper makes two complementary contributions to event retrieval in large collections of videos. First, we propose hyper-pooling strategies that encode the frame descriptors into a representation of the video sequence in a stable manner. Our best choices compare favorably with regular pooling techniques based on k-means quantization. Second, we introduce a technique to improve the ranking. It can be interpreted either as a query expansion method or as a similarity adaptation based on the local context of the query video descriptor. Experiments on public benchmarks show that our methods are complementary and improve event retrieval results, without sacrificing efficiency. Matthijs Douze, Jérôme Revaud, Cordelia Schmid, Hervé Jégou |
ICCV | 4 |
| 2013 | To Aggregate or Not to aggregate: Selective Match Kernels for Image SearchabstractThis paper considers a family of metrics to compare images based on their local descriptors. It encompasses the VLAD descriptor and matching techniques such as Hamming Embedding. Making the bridge between these approaches leads us to propose a match kernel that takes the best of existing techniques by combining an aggregation procedure with a selective match kernel. Finally, the representation underpinning this kernel is approximated, providing a large scale image search both precise and scalable, as shown by our experiments on several benchmarks. Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
ICCV | 3 |
| 2013 | Query-Adaptive Asymmetrical Dissimilarities for Visual Object RetrievalabstractVisual object retrieval aims at retrieving, from a collection of images, all those in which a given query object appears. It is inherently asymmetric: the query object is mostly included in the database image, while the converse is not necessarily true. However, existing approaches mostly compare the images with symmetrical measures, without considering the different roles of query and database. This paper first measure the extent of asymmetry on large-scale public datasets reflecting this task. Considering the standard bag-of-words representation, we then propose new asymmetrical dissimilarities accounting for the different inlier ratios associated with query and database images. These asymmetrical measures depend on the query, yet they are compatible with an inverted file structure, without noticeably impacting search efficiency. Our experiments show the benefit of our approach, and show that the visual object retrieval task is better treated asymmetrically, in the spirit of state-of-the-art text retrieval. Cai-Zhi Zhu, Hervé Jégou, Shin'ichi Satoh 0001 |
ICCV | 2 |
| 2013 | Retrieving geo-location of videos with a divide & conquer hierarchical multimodal approachabstractThis paper presents a strategy to identify the geographic location of videos. First, it relies on a multi-modal cascade pipeline that exploits the available sources of information, namely the user's upload history, his social network and a visual-based matching technique. Second, we present a novel divide & conquer strategy to better exploit the tags associated with the input video. It pre-selects one or several geographic area of interest of higher expected relevance and performs a deeper analysis inside the selected area(s) to return the coordinates most likely to be related to the input tags. The experiments were conducted as part of the MediaEval 2012 Placing Task. Our approach, which differs significantly from the other submitted techniques, achieves the best results on this benchmark when considering the same amount of external information, i.e. when not using any gazetteers nor any other kind of external information. Michele Trevisiol, Hervé Jégou, Jonathan Delhumeau, Guillaume Gravier |
ICMR | 2 |
| 2013 | Revisiting the VLAD image representationabstractRecent works on image retrieval have proposed to index images by compact representations encoding powerful local descriptors, such as the closely related VLAD and Fisher vector. By combining such a representation with a suitable coding technique, it is possible to encode an image in a few dozen bytes while achieving excellent retrieval results. This paper revisits some assumptions proposed in this context regarding the handling of "visual burstiness", and shows that ad-hoc choices are implicitly done which are not desirable. Focusing on VLAD without loss of generality, we propose to modify several steps of the original design. Albeit simple, these modifications significantly improve VLAD and make it compare favorably against the state of the art. Jonathan Delhumeau, Philippe Henri Gosselin, Hervé Jégou, Patrick Pérez |
ACM Multimedia | 3 |
| 2013 | Sim-min-hash: an efficient matching technique for linking large image collectionsabstractOne of the most successful method to link all similar images within a large collection is min-Hash, which is a way to significantly speed-up the comparison of images when the underlying image representation is bag-of-words. However, the quantization step of min-Hash introduces important information loss. In this paper, we propose a generalization of min-Hash, called Sim-min-Hash, to compare sets of real-valued vectors. We demonstrate the effectiveness of our approach when combined with the Hamming embedding similarity. Experiments on large-scale popular benchmarks demonstrate that Sim-min-Hash is more accurate and faster than min-Hash for similar image search. Linking a collection of one million images described by 2 billion local descriptors is done in 7 minutes on a single core machine. Wanlei Zhao, Hervé Jégou, Guillaume Gravier |
ACM Multimedia | 2 |
| 2012 | Large-scale and larger-scale image search
Hervé Jégou |
BMVC | 1 |
| 2012 | Negative Evidences and Co-occurences in Image Retrieval: The Benefit of PCA and Whitening
Hervé Jégou, Ondrej Chum |
ECCV (2) | 1 |
| 2012 | BABAZ: A large scale audio search system for video copy detectionabstractThis paper presents BABAZ, an audio search system to search modified segments in large databases of music or video tracks. It is based on an efficient audio feature matching system which exploits the reciprocal nearest neighbors to produce a per-match similarity score. Temporal consistency is taken into account based on the audio matches, and boundary estimation allows the precise localization of the matching segments. The method is mainly intended for video retrieval based on their audio track, as typically evaluated in the copy detection task of TRECVID evaluation campaigns. The evaluation conducted on music retrieval shows that our system is comparable to a reference audio fingerprinting system for music retrieval, and significantly outperforms it on audio-based video retrieval, as shown by our experiments conducted on the dataset used in the copy detection task of TRECVID'2010 campaign. Hervé Jégou, Jonathan Delhumeau, Jiangbo Yuan, Guillaume Gravier, Patrick Gros |
ICASSP | 1 |
| 2012 | Anti-sparse coding for approximate nearest neighbor searchabstractThis paper proposes a binarization scheme for vectors of high dimension based on the recent concept of anti-sparse coding, and shows its excellent performance for approximate nearest neighbor search. Unlike other binarization schemes, this framework allows, up to a scaling factor, the explicit reconstruction from the binary representation of the original vector. The paper also shows that random projections which are used in Locality Sensitive Hashing algorithms, are significantly outperformed by regular frames for both synthetic and real data if the number of bits exceeds the vector dimensionality, i.e., when high precision is required. Hervé Jégou, Teddy Furon, Jean-Jacques Fuchs |
ICASSP | 1 |
| 2012 | Hamming embedding similarity-based image classificationabstractIn this paper, we propose a novel image classification framework based on patch matching. More precisely, we adapt the Hamming Embedding technique, first introduced for image search to improve the bag-of-words representation. This matching technique allows the fast comparison of descriptors based on their binary signatures, which refines the matching rule based on visual words and thereby limits the quantization error. Then, in order to allow the use of efficient and suitable linear kernel-based SVM classification, we propose a mapping method to cast the scores output by the Hamming Embedding matching technique into a proper similarity space. Comparative experiments of our proposed approach and other existing encoding methods on two challenging datasets PASCAL VOC 2007 and Caltech-256, report the interest of the proposed scheme, which outperforms all methods based on patch matching and even provide competitive results compared with the state-of-the-art coding techniques. Mihir Jain, Rachid Benmokhtar, Hervé Jégou, Patrick Gros |
ICMR | 3 |
| 2012 | Aggregating Local Image Descriptors into Compact CodesabstractThis paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image data set takes about 250 ms on one processor core. Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez 0002, Patrick Pérez, Cordelia Schmid |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Reconstructing an image from its local descriptorsabstractThis paper shows that an image can be approximately reconstructed based on the output of a blackbox local description software such as those classically used for image indexing. Our approach consists first in using an off-the-shelf image database to find patches that are visually similar to each region of interest of the unknown input image, according to associated local descriptors. These patches are then warped into input image domain according to interest region geometry and seamlessly stitched together. Final completion of still missing texture-free regions is obtained by smooth interpolation. As demonstrated in our experiments, visually meaningful reconstructions are obtained just based on image local descriptors like SIFT, provided the geometry of regions of interest is known. The reconstruction most often allows the clear interpretation of the semantic image content. As a result, this work raises critical issues of privacy and rights when local descriptors of photos or videos are given away for indexing and search purpose. Philippe Weinzaepfel, Hervé Jégou, Patrick Pérez |
CVPR | 2 |
| 2011 | Searching in one billion vectors: Re-rank with source codingabstractRecent indexing techniques inspired by source coding have been shown successful to index billions of high-dimensional vectors in memory. In this paper, we propose an approach that re-ranks the neighbor hypotheses obtained by these compressed-domain indexing methods. In contrast to the usual post-verification scheme, which performs exact distance calculation on the short-list of hypotheses, the estimated distances are refined based on short quantization codes, to avoid reading the full vectors from disk. We have released a new public dataset of one billion 128 dimensional vectors and proposed an experimental setup to evaluate high dimensional indexing algorithms on a realistic scale. Experiments show that our method accurately and efficiently re-ranks the neighbor hypotheses using little memory compared to the full vectors representation. Hervé Jégou, Romain Tavenard, Matthijs Douze, Laurent Amsaleg |
ICASSP | 1 |
| 2011 | Asymmetric hamming embedding: taking the best of our bits for large scale image searchabstractThis paper proposes an asymmetric Hamming Embedding scheme for large scale image search based on local descriptors. The comparison of two descriptors relies on an vector-to-binary code comparison, which limits the quantization error associated with the query compared with the original Hamming Embedding method. The approach is used in combination with an inverted file structure that offers high efficiency, comparable to that of a regular bag-of-features retrieval system. The comparison is performed on two popular datasets. Our method consistently improves the search quality over the symmetric version. The trade-off between memory usage and precision is evaluated, showing that the method is especially useful for short binary signatures. Mihir Jain, Hervé Jégou, Patrick Gros |
ACM Multimedia | 2 |
| 2011 | Bag-of-colors for improved image searchabstractThis paper investigates the use of color information when used within a state-of-the-art large scale image search system. We introduce a simple yet effective and efficient color signature generation procedure. It is used either to produce global or local descriptors. As a global descriptor, it outperforms several state-of-the-art color description methods, in particular the bag-of-words method based on color SIFT. As a local descriptor, our signature is used jointly with SIFT descriptors (no color) to provide complementary information. This significantly improves the recognition rate, outperforming the state of the art on two image search benchmarks. We provide an open source package of our signature (http://www.kooaba.com/en/learnmore/labs/). Christian Wengert, Matthijs Douze, Hervé Jégou |
ACM Multimedia | 3 |
| 2011 | Product Quantization for Nearest Neighbor SearchabstractThis paper introduces a product quantization-based approach for approximate nearest neighbor search. The idea is to decompose the space into a Cartesian product of low-dimensional subspaces and to quantize each subspace separately. A vector is represented by a short code composed of its subspace quantization indices. The euclidean distance between two vectors can be efficiently estimated from their codes. An asymmetric version increases precision, as it computes the approximate distance between a vector and a code. Experimental results show that our approach searches for nearest neighbors efficiently, in particular in combination with an inverted file system. Results for SIFT and GIST image descriptors show excellent search accuracy, outperforming three state-of-the-art approaches. The scalability of our approach is validated on a data set of two billion vectors. Hervé Jégou, Matthijs Douze, Cordelia Schmid |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Aggregating local descriptors into a compact image representationabstractWe address the problem of image search on a very large scale, where three constraints have to be considered jointly: the accuracy of the search, its efficiency, and the memory usage of the representation. We first propose a simple yet efficient way of aggregating local image descriptors into a vector of limited dimension, which can be viewed as a simplification of the Fisher kernel representation. We then show how to jointly optimize the dimension reduction and the indexing algorithm, so that it best preserves the quality of vector comparison. The evaluation shows that our approach significantly outperforms the state of the art: the search accuracy is comparable to the bag-of-features approach for an image representation that fits in 20 bytes. Searching a 10 million image dataset takes about 50ms. Hervé Jégou, Matthijs Douze, Cordelia Schmid, Patrick Pérez |
CVPR | 1 |
| 2010 | Compact Video Description for Copy Detection with Precise Temporal Alignment
Matthijs Douze, Hervé Jégou, Cordelia Schmid, Patrick Pérez |
ECCV (1) | 2 |
| 2010 | Searching with expectationsabstractHandling large amounts of data, such as large image databases, requires the use of approximate nearest neighbor search techniques. Recently, Hamming embedding methods such as spectral hashing have addressed the problem of obtaining compact binary codes optimizing the trade-off between the memory usage and the probability of retrieving the true nearest neighbors. In this paper, we formulate the problem of generating compact signatures as a rate-distortion problem. In the spirit of source coding algorithms, we aim at minimizing the reconstruction error on the squared distances with a constraint on the memory usage. The vectors are ranked based on the distance estimates to the query vector. Experiments on image descriptors show a significant improvement over spectral hashing. Harsimrat Sandhawalia, Hervé Jégou |
ICASSP | 2 |
| 2010 | Improving Bag-of-Features for Large Scale Image Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid |
Int. J. Comput. Vis. | 1 |
| 2010 | Accurate Image Search Using the Contextual Dissimilarity MeasureabstractThis paper introduces the contextual dissimilarity measure, which significantly improves the accuracy of bag-of-features-based image search. Our measure takes into account the local distribution of the vectors and iteratively estimates distance update terms in the spirit of Sinkhorn's scaling algorithm, thereby modifying the neighborhood structure. Experimental results show that our approach gives significantly better results than a standard distance and outperforms the state of the art in terms of accuracy on the Nistér-Stewénius and Lola data sets. This paper also evaluates the impact of a large number of parameters, including the number of descriptors, the clustering method, the visual vocabulary size, and the distance measure. The optimal parameter choice is shown to be quite context-dependent. In particular, using a large number of descriptors is interesting only when using our dissimilarity measure. We have also evaluated two novel variants: multiple assignment and rank aggregation. They are shown to further improve accuracy at the cost of higher memory usage and lower efficiency. Hervé Jégou, Cordelia Schmid, Hedi Harzallah, Jakob Verbeek |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Locality sensitive hashing: A comparison of hash function types and querying mechanisms
Loïc Paulevé, Hervé Jégou, Laurent Amsaleg |
Pattern Recognit. Lett. | 2 |
| 2010 | An Image-Based Approach to Video Copy Detection With Spatio-Temporal Post-FilteringabstractThis paper introduces a video copy detection system which efficiently matches individual frames and then verifies their spatio-temporal consistency. The approach for matching frames relies on a recent local feature indexing method, which is at the same time robust to significant video transformations and efficient in terms of memory usage and computation time. We match either keyframes or uniformly sampled frames. To further improve the results, a verification step robustly estimates a spatio-temporal model between the query video and the potentially corresponding video segments. Experimental results evaluate the different parameters of our system and measure the trade-off between accuracy and efficiency. We show that our system obtains excellent results for the TRECVID 2008 copy detection task. Matthijs Douze, Hervé Jégou, Cordelia Schmid |
IEEE Trans. Multim. | 2 |
| 2009 | On the burstiness of visual elementsabstractBurstiness, a phenomenon initially observed in text retrieval, is the property that a given visual element appears more times in an image than a statistically independent model would predict. In the context of image search, burstiness corrupts the visual similarity measure, i.e., the scores used to rank the images. In this paper, we propose a strategy to handle visual bursts for bag-of-features based image search systems. Experimental results on three reference datasets show that our method significantly and consistently outperforms the state of the art. Hervé Jégou, Matthijs Douze, Cordelia Schmid |
CVPR | 1 |
| 2009 | Packing bag-of-featuresabstractOne of the main limitations of image search based on bag-of-features is the memory usage per image. Only a few million images can be handled on a single machine in reasonable response time. In this paper, we first evaluate how the memory usage is reduced by using lossless index compression. We then propose an approximate representation of bag-of-features obtained by projecting the corresponding histogram onto a set of pre-defined sparse projection functions, producing several image descriptors. Coupled with a proper indexing structure, an image is represented by a few hundred bytes. A distance expectation criterion is then used to rank the images. Our method is at least one order of magnitude faster than standard bag-of-features while providing excellent search quality. Hervé Jégou, Matthijs Douze, Cordelia Schmid |
ICCV | 1 |
| 2009 | Computation of posterior marginals on aggregated state models for soft source decodingabstractOptimum soft decoding of sources compressed with variable length codes and quasi-arithmetic codes, transmitted over noisy channels, can be performed on a bit/symbol trellis. However, the number of states of the trellis is a quadratic function of the sequence length leading to a decoding complexity which is not tractable for practical applications. The decoding complexity can be significantly reduced by using an aggregated state model, while still achieving close to optimum performance in terms of bit error rate and frame error rate. However, symbol a posteriori probabilities can not be directly derived on these models and the symbol error rate (SER) may not be minimized. This paper describes a two-step decoding algorithm that achieves close to optimal decoding performance in terms of SER on aggregated state models. A performance and complexity analysis of the proposed algorithm is given. Simon Malinowski, Hervé Jégou, Christine Guillemot |
IEEE Trans. Commun. | 2 |
| 2008 | Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid |
ECCV (1) | 1 |
| 2008 | Query adaptative locality sensitive hashingabstractIt is well known that high-dimensional nearest-neighbor retrieval is very expensive. Many signal processing methods suffer from this computing cost. Dramatic performance gains can be obtained by using approximate search, such as the popular Locality-Sensitive Hashing. This paper improves LSH by performing an on-line selection of the most appropriate hash functions from a pool of functions. An additional improvement originates from the use of E& lattices for geometric hashing instead of one-dimensional random projections. A performance study based on state-of-the-art high-dimensional descriptors computed on real images shows that our improvements to LSH greatly reduce the search complexity for a given level of accuracy. Hervé Jégou, Laurent Amsaleg, Cordelia Schmid, Patrick Gros |
ICASSP | 1 |
| 2007 | A contextual dissimilarity measure for accurate and efficient image searchabstractIn this paper we present two contributions to improve accuracy and speed of an image search system based on bag-of-features: a contextual dissimilarity measure (CDM) and an efficient search structure for visual word vectors. Our measure (CDM) takes into account the local distribution of the vectors and iteratively estimates distance correcting terms. These terms are subsequently used to update an existing distance, thereby modifying the neighborhood structure. Experimental results on the Nister-Stewenius dataset show that our approach significantly outperforms the state-of-the-art in terms of accuracy. Our efficient search structure for visual word vectors is a two-level scheme using inverted files. The first level partitions the image set into clusters of images. At query time, only a subset of clusters of the second level has to be searched. This method allows fast querying in large sets of images. We evaluate the gain in speed and the loss in accuracy on large datasets (up to 500k images). Hervé Jégou, Hedi Harzallah, Cordelia Schmid |
CVPR | 1 |
| 2007 | Entropy Coding With Variable-Length Rewriting SystemsabstractThis paper describes a family of codes for entropy coding of memoryless sources. These codes are defined by sets of production rules of the form $a\bar{l} \rightarrow\bar{b}$ , where $a$ is a source symbol, and $\bar{l},\bar{b}$ are sequences of bits. The coding process can be modeled as a finite-state machine (FSM). A method to construct codes which preserve the lexicographic order in the binary-coded representation is described. For a given constraint on the number of states for the coding process, this method allows the construction of codes with a better compression efficiency than the HuTucker codes. A second method is proposed to construct codes such that the marginal bit probability of the compressed bitstream converges to 0.5 as the sequence length increases. This property is achieved even if the probability distribution function is not known by the encoder. Hervé Jégou, Christine Guillemot |
IEEE Trans. Commun. | 1 |
| 2007 | Synchronization Recovery and State Model Reduction for Soft Decoding of Variable Length CodesabstractVariable length codes (VLCs) exhibit loss of synchronization problems when transmitted over noisy channels. Trellis decoding techniques based on Maximum A Posteriori (MAP) estimators are often used to minimize the error rate on the estimated sequence. If the number of symbols and/or bits transmitted is known by the decoder, termination constraints can be incorporated in the decoding process. All the paths in the trellis which do not lead to a valid sequence length are suppressed. This correspondence presents an analytic method to assess the expected error resilience of a VLC when trellis decoding with a sequence length constraint is used. The approach is based on the computation, for a given code, of the amount of information brought by the constraint. It is then shown that this quantity is not significantly altered by appropriate trellis states aggregation. This proves that the performance obtained by running a length-constrained Viterbi decoder on aggregated state models approaches the one obtained with the bit/symbol trellis, with a significantly reduced complexity. It is then shown that the complexity can be further decreased by projecting the state model on two state models of reduced size. Simon Malinowski, Hervé Jégou, Christine Guillemot |
IEEE Trans. Inf. Theory | 2 |
| 2006 | Error recovery properties of quasi-arithmetic codes and soft decoding with length constraintabstractIn this paper, we propose a method to analyse the error recovery properties of quasi-arithmetic codes. This method is adapted from the one proposed in J. Maxted and J. Robinson (1985) for variable length codes. The expected number of symbols affected by a single bit error and the probability mass function of the gain/loss (P.F. Swaszek and P. DiCicco, 1995) of symbols following a single bit error can be computed with this method. A method to estimate this probability mass function when a bitstream is sent over a binary symmetrical channel is then proposed. The aggregated state model for soft decoding of variable length codes proposed in H. Jegou et al. (2005) is then extended to quasi-arithmetic codes, as the synchronisation recovery properties of both kind of codes are similar. The soft decoding results of this scheme reveal high performance with a reasonable computing cost Simon Malinowski, Hervé Jégou, Christine Guillemot |
ISIT | 2 |
| 2005 | Entropy coding with variable length re-writing systemsabstractThis paper describes a new set of block source codes well suited for data compression. These codes are defined by sets of productions rules of the form allowbar rarr blowbar where a isin A represents a value from the source alphabet A and llowbar, blowbar are small sequences of bits. These codes naturally encompass other variable length codes (VLCs) such as Huffman codes. It is shown that these codes may have a similar or even a shorter mean description length than Huffman codes for the same encoding and decoding complexity. A first code design method allowing to preserve the lexicographic order in the bit domain is described. The corresponding codes have the same mean description length (mdl) as Huffman codes from which they are constructed. Therefore, they outperform from a compression point of view the Hu-Tucker codes designed to offer the lexicographic property in the bit domain. A second construction method allows to obtain codes such that the marginal bit probability converges to 0.5 as the sequence length increases and this is achieved even if the probability distribution function is not known by the encoder Hervé Jégou, Christine Guillemot |
ISIT | 1 |
| 2005 | Robust multiplexed codes for compression of heterogeneous dataabstractCompression systems of real signals (images, video, audio) generate sources of information with different levels of priority which are then encoded with variable-length codes (VLCs). This paper addresses the issue of robust transmission of such VLC encoded heterogeneous sources over error-prone channels. VLCs are very sensitive to channel noise: when some bits are altered, synchronization losses can occur at the receiver. This paper describes a new family of codes, called multiplexed codes, that confine the de-synchronization phenomenon to low-priority data while reaching asymptotically the entropy bound for both (low- and high-priority) sources. The idea consists of creating fixed-length codes for high-priority information and of using the inherent redundancy to describe low-priority data, hence the name multiplexed codes. Theoretical and simulation results reveal a very high error resilience at almost no cost in compression efficiency. Hervé Jégou, Christine Guillemot |
IEEE Trans. Inf. Theory | 1 |
| 2004 | First-order multiplexed source codes for error-resilient entropy codingabstractThis paper describes a new class of variable length codes (VLCs) that allow to exploit first-order source statistics while still being resilient to transmission errors. This paper extends the work of [H. Jegou et al., (2003)] to take into account the source conditional probabilities. Theoretical performances in terms of compression efficiency and error resilience are analyzed. Hervé Jégou, Christine Guillemot |
ISIT | 1 |
| 2003 | Error-resilient binary multiplexed source codesabstractThe paper addresses the issue of robust transmission of VLC encoded sources over error-prone channels. We have recently introduced a new family of codes, called multiplexed codes. They exploit the fact that real signal compression systems generate sources of information with different levels of priority. Multiplexed codes allow the desynchronization phenomenon to be confined to low priority data while allowing the entropy bound to be reached asymptotically for both (low and high priority) sources. A multiplexing procedure based on an iterative Euclidian decomposition has been proposed. This paper introduces a variant of multiplexed codes, called binary multiplexed codes, together with a very simple multiplexing algorithm that exploits the structure of variable length codetrees. It is shown analytically and experimentally that this family of codes is more error resilient than fixed length codes while reaching the compression efficiency of classical variable length codes. Hervé Jégou, Christine Guillemot |
ICASSP (4) | 1 |
| 2003 | Source multiplexed codes for error-prone channelsabstractCompression systems of real signals (images, video, audio) generate sources of information with different levels of priority, which are then encoded with variable length codes (VLC). This paper addresses the issue of robust transmission of such VLC encoded heterogeneous sources over error-prone channels. VLCs are very sensitive to channel noise: when some bits are altered, synchronization losses can occur at the receiver. This paper describes a new family of codes, called multiplexed codes, that allow to confine the de-synchronization phenomenon to low priority data while allowing to reach asymptotically the entropy bound for both (low and high priority) sources. The idea consists in creating fixed length codes for high priority information and in using the inherent redundancy to describe low priority data, hence the name "multiplexed codes". Simulation results reveal very high error resilience at almost no cost in compression efficiency. Hervé Jégou, Christine Guillemot |
ICC | 1 |