David Picard

dblp:87/1166 · DBLP profile ↗
← Back
61ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0002-6296-4222ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 33 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
abstract
Global visual geolocation consists in predicting where an image was captured anywhere on Earth. Since not all images can be localized with the same precision, this task inherently involves a degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we propose the first generative approach for visual geolocation based on diffusion and flow matching, and an extension to Riemannian flow matching, where the denoising process operates directly on the Earth's surface. Our model achieves state-of-the-art performance on three visual geolocation benchmarks: OpenStreetView-5M, YFCC100M, and iNat21. In addition, we introduce the task of probabilistic visual geolocation, where the model predicts a probability distribution over all possible locations instead of a single point. We implement new metrics and baselines for this task, demonstrating the advantages of our generative approach. Codes and models are available here.
Nicolas Dufour, Vicky Kalogeiton, David Picard, Loïc Landrieu
CVPR3
2025 A deep learning-based pipeline for the conservation assessment of bindings in archives and libraries
Valérie Lee-Gouet, Lahcen Yamoun, Zacharie Rodière, Camille Simon 0001, Michel Jordan, Julien Longhi, David Picard
Multim. Tools Appl.7
2024 Don't Drop Your Samples! Coherence-Aware Training Benefits Conditional Diffusion
abstract
Conditional diffusion models are powerful generative models that can leverage various types of conditional information, such as class labels, segmentation masks, or text captions. However, in many real-world scenarios, conditional infor-mation may be noisy or unreliable due to human annotation errors or weak alignment. In this paper, we propose the Coherence-Aware Diffusion (CAD), a novel method that in-tegrates coherence in conditional information into diffusion models, allowing them to learn from noisy annotations with-out discarding data. We assume that each data point has an associated coherence score that reflects the quality of the conditional information. We then condition the diffusion model on both the conditional information and the coherence score. In this way, the model learns to ignore or discount the conditioning when the coherence is low. We show that CAD is theoretically sound and empirically effective on various conditional generation tasks. Moreover, we show that lever-aging coherence generates realistic and diverse samples that respect conditional information better than models trained on cleaned datasets where samples with low coherence have been discarded. Code and weights here.
Nicolas Dufour, Victor Besnier, Vicky Kalogeiton, David Picard
CVPR4
2024 An Analysis of Initial Training Strategies for Exemplar-Free Class-Incremental Learning
abstract
Class-Incremental Learning (CIL) aims to build classification models from data streams. At each step of the CIL process, new classes must be integrated into the model. Due to catastrophic forgetting, CIL is particularly challenging when examples from past classes cannot be stored, the case on which we focus here. To date, most approaches are based exclusively on the target dataset of the CIL process. However, the use of models pre-trained in a self-supervised way on large amounts of data has recently gained momentum. The initial model of the CIL process may only use the first batch of the target dataset, or also use pre-trained weights obtained on an auxiliary dataset. The choice between these two initial learning strategies can significantly influence the performance of the incremental learning model, but has not yet been studied in depth. Performance is also influenced by the choice of the CIL algorithm, the neural architecture, the nature of the target task, the distribution of classes in the stream and the number of examples available for learning. We conduct a comprehensive experimental study to assess the roles of these factors. We present a statistical analysis framework that quantifies the relative contribution of each factor to incremental performance. Our main finding is that the initial training strategy is the dominant factor influencing the average incremental accuracy, but that the choice of CIL algorithm is more important in preventing forgetting. Based on this analysis, we propose practical recommendations for choosing the right initial training strategy for a given incremental learning use case. These recommendations are intended to facilitate the practical deployment of incremental learning.
Grégoire Petit, Michaël Soumm, Eva Feillet, Adrian Popescu 0001, Bertrand Delezoide, David Picard, Céline Hudelot
WACV6
2023 H3WB: Human3.6M 3D WholeBody Dataset and Benchmark
abstract
We present a benchmark for 3D human whole-body pose estimation, which involves identifying accurate 3D keypoints on the entire human body, including face, hands, body, and feet. Currently, the lack of a fully annotated and accurate 3D whole-body dataset results in deep networks being trained separately on specific body parts, which are combined during inference. Or they rely on pseudo-groundtruth provided by parametric body models which are not as accurate as detection based methods. To overcome these issues, we introduce the Human3.6M 3D WholeBody (H3WB) dataset, which provides whole-body annotations for the Human3.6M dataset using the COCO Wholebody layout. H3WB comprises 133 whole-body keypoint annotations on 100K images, made possible by our new multi-view pipeline. We also propose three tasks: i) 3D whole-body pose lifting from 2D complete whole-body pose, ii) 3D whole-body pose lifting from 2D incomplete whole-body pose, and iii) 3D whole-body pose estimation from a single RGB image. Additionally, we report several baselines from popular methods for these tasks. Furthermore, we also provide automated 3D whole-body annotations of TotalCapture and experimentally show that when used with H3WB it helps to improve the performance.
Nermin Samet, David Picard
ICCV3
2023 Unveiling the Latent Space Geometry of Push-Forward Generative Models
abstract
Many deep generative models are defined as a push-forward of a Gaussian measure by a continuous generator, such as Generative Adversarial Networks (GANs) or Variational Auto-Encoders (VAEs). This work explores the latent space of such deep generative models. A key issue with these models is their tendency to output samples outside of the support of the target distribution when learning disconnected distributions. We investigate the relationship between the performance of these models and the geometry of their latent space. Building on recent developments in geometric measure theory, we prove a sufficient condition for optimality in the case where the dimension of the latent space is larger than the number of modes. Through experiments on GANs, we demonstrate the validity of our theoretical results and gain new insights into the latent space geometry of these models. Additionally, we propose a truncation method that enforces a simplicial cluster structure in the latent space and improves the performance of GANs.
Thibaut Issenhuth, Ugo Tanielian, Jérémie Mary, David Picard
ICML4
2023 FeTrIL: Feature Translation for Exemplar-Free Class-Incremental Learning
abstract
Exemplar-free class-incremental learning is very challenging due to the negative effect of catastrophic forgetting. A balance between stability and plasticity of the incremental process is needed in order to obtain good accuracy for past as well as new classes. Existing exemplar-free class-incremental methods focus either on successive fine tuning of the model, thus favoring plasticity, or on using a feature extractor fixed after the initial incremental state, thus favoring stability. We introduce a method which combines a fixed feature extractor and a pseudo-features generator to improve the stability-plasticity balance. The generator uses a simple yet effective geometric translation of new class features to create representations of past classes, made of pseudo-features. The translation of features only requires the storage of the centroid representations of past classes to produce their pseudo-features. Actual features of new classes and pseudo-features of past classes are fed into a linear classifier which is trained incrementally to discriminate between all classes. The incremental process is much faster with the proposed method compared to mainstream ones which update the entire deep model. Experiments are performed with three challenging datasets, and different incremental settings. A comparison with ten existing methods shows that our method outperforms the others in most cases. FeTrIL code is available at https://github.com/GregoirePetit/FeTrIL.
Grégoire Petit, Adrian Popescu 0001, Hugo Schindler, David Picard, Bertrand Delezoide
WACV4
2023 SSP-Net: Scalable sequential pyramid networks for real-Time 3D human pose regression
Diogo C. Luvizon, Hedi Tabia, David Picard
Pattern Recognit.3
2022 Decanus to Legatus: Synthetic Training for 2D-3D Human Pose Lifting
David Picard
ACCV (4)2
2022 SCAM! Transferring Humans Between Images with Semantic Cross Attention Modulation
Nicolas Dufour, David Picard, Vicky Kalogeiton
ECCV (14)2
2022 Improving Deep Metric Learning with Virtual Classes and Examples Mining
abstract
In deep metric learning, the training procedure relies on sampling informative tuples. However, as the training procedure progresses, it becomes nearly impossible to sample relevant hard negative examples without proper mining strategies or generation-based methods. Recent work on hard negative generation have shown great promises to solve the mining problem. However, this generation process is difficult to tune and often leads to incorrectly labeled examples. To tackle this issue, we introduce MIRAGE, a generation-based method that relies on virtual classes entirely composed of generated examples that act as buffer areas between the training classes. We empirically show that virtual classes significantly improve the results on popular datasets (Cub-200-2011 and Cars-196) compared to other generation methods.
Pierre Jacob, David Picard, Aymeric Histace
ICIP2
2022 Latent reweighting, an almost free improvement for GANs
abstract
Standard formulations of GANs, where a continuous function deforms a connected latent space, have been shown to be misspecified when fitting different classes of images. In particular, the generator will necessarily sample some low-quality images in between the classes. Rather than modifying the architecture, a line of works aims at improving the sampling quality from pre-trained generators at the expense of increased computational cost. Building on this, we introduce an additional network to predict latent importance weights and two associated sampling methods to avoid the poorest samples. This idea has several advantages: 1) it provides a way to inject disconnectedness into any GAN architecture, 2) since the rejection happens in the latent space, it avoids going through both the generator and the discriminator, saving computation time, 3) this importance weights formulation provides a principled way to reduce the Wasserstein’s distance to the target distribution. We demonstrate the effectiveness of our method on several datasets, both synthetic and high-dimensional.
Thibaut Issenhuth, Ugo Tanielian, David Picard, Jérémie Mary
WACV3
2022 Consensus-Based Optimization for 3D Human Pose Estimation in Camera Coordinates
Diogo C. Luvizon, David Picard, Hedi Tabia
Int. J. Comput. Vis.2
2021 Unsupervised Word Representations Learning with Bilinear Convolutional Network on Characters
abstract
In this paper, we propose a new unsupervised method for learning word embedding with raw characters as input representations, bypassing the problems arising from the use of a dictionary.To achieve this purpose, we translate the distributional hypothesis into a unsupervised metric learning objective, which allows to consider only an encoder instead of an encoder-decoder architecture.We propose to use a convolutional neural network with bilinear product blocks and residual connections to encode co-occurrences patterns.We show the effectiveness of our approach by comparing it with classical word embedding methods such as fastText and GloVe on several benchmarks.
Thomas Luka, Laure Soulier, David Picard
ESANN3
2021 Triggering Failures: Out-Of-Distribution detection by learning from local adversarial attacks in Semantic Segmentation
abstract
In this paper, we tackle the detection of out-of-distribution (OOD) objects in semantic segmentation. By analyzing the literature, we found that current methods are either accurate or fast but not both which limits their usability in real world applications. To get the best of both aspects, we propose to mitigate the common shortcomings by following four design principles: decoupling the OOD detection from the segmentation task, observing the entire segmentation network instead of just its output, generating training data for the OOD detector by leveraging blind spots in the segmentation network and focusing the generated data on localized regions in the image to simulate OOD objects. Our main contribution is a new OOD detection architecture called ObsNet associated with a dedicated training scheme based on Local Adversarial Attacks (LAA). We validate the soundness of our approach across numerous ablation studies. We also show it obtains top performances both in speed and accuracy when compared to ten recent methods of the literature on three different datasets.
Victor Besnier, Andrei Bursuc, David Picard, Alexandre Briot
ICCV3
2021 Image Collation: Matching Illustrations in Manuscripts
Ryad Kaoua, Xi Shen 0001, Alexandra Durr, Stavros Lazaris, David Picard, Mathieu Aubry
ICDAR (4)5
2021 Learning Uncertainty for Safety-Oriented Semantic Segmentation in Autonomous Driving
abstract
In this paper, we show how uncertainty estimation can be leveraged to enable safety critical image segmentation in autonomous driving, by triggering a fallback behavior if a target accuracy cannot be guaranteed. We introduce a new uncertainty measure based on disagreeing predictions as measured by a dissimilarity function. We propose to estimate this dissimilarity by training a deep neural architecture in parallel to the task-specific network. It allows this observer to be dedicated to the uncertainty estimation, and let the task-specific network make predictions. We propose to use self-supervision to train the observer, which implies that our method does not require additional training data. We show experimentally that our proposed approach is much less computationally intensive at inference time than competing methods (e.g., MC Dropout), while delivering better results on safety-oriented evaluation metrics on the CamVid dataset, especially in the case of glare artifacts.
Victor Besnier, David Picard, Alexandre Briot
ICIP2
2021 Multi-Task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
abstract
Human pose estimation and action recognition are related tasks since both problems are strongly dependent on the human body representation and analysis. Nonetheless, most recent methods in the literature handle the two problems separately. In this article, we propose a multi-task framework for jointly estimating 2D or 3D human poses from monocular color images and classifying human actions from video sequences. We show that a single architecture can be used to solve both problems in an efficient way and still achieves state-of-the-art or comparable results at each task while running with a throughput of more than 100 frames per second. The proposed method benefits from high parameters sharing between the two tasks by unifying still images and video clips processing in a single pipeline, allowing the model to be trained with data from different categories simultaneously and in a seamlessly way. Additionally, we provide important insights for end-to-end training the proposed multi-task model by decoupling key prediction parts, which consistently leads to better accuracy on both tasks. The reported results on four datasets (MPII, Human3.6M, Penn Action and NTU RGB+D) demonstrate the effectiveness of our method on the targeted tasks. Our source code and trained weights are publicly available at https://github.com/dluvizon/deephar.
Diogo C. Luvizon, David Picard, Hedi Tabia
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 DIABLO: Dictionary-based attention block for deep metric learning
Pierre Jacob, David Picard, Aymeric Histace, Edouard Klein
Pattern Recognit. Lett.2
2020 Deepzzle: Solving Visual Jigsaw Puzzles With Deep Learning and Shortest Path Optimization
abstract
We tackle the image reassembly problem with wide space between the fragments, in such a way that the patterns and colors continuity is mostly unusable. The spacing emulates the erosion of which the archaeological fragments suffer. We crop-square the fragments borders to compel our algorithm to learn from the content of the fragments. We also complicate the image reassembly by removing fragments and adding pieces from other sources. We use a two-step method to obtain the reassemblies: 1) a neural network predicts the positions of the fragments despite the gaps between them; 2) a graph that leads to the best reassemblies is made from these predictions. In this paper, we notably investigate the effect of branch-cut in the graph of reassemblies. We also provide a comparison with the literature, solve complex images reassemblies, explore at length the dataset, and propose a new metric that suits its specificities.
Marie-Morgane Paumard, David Picard, Hedi Tabia
IEEE Trans. Image Process.2
2019 Metric Learning With HORDE: High-Order Regularizer for Deep Embeddings
abstract
Learning an effective similarity measure between image representations is key to the success of recent advances in visual search tasks (e.g. verification or zero-shot learning). Although the metric learning part is well addressed, this metric is usually computed over the average of the extracted deep features. This representation is then trained to be discriminative. However, these deep features tend to be scattered across the feature space. Consequently, the representations are not robust to outliers, object occlusions, background variations, etc. In this paper, we tackle this scattering problem with a distribution-aware regularization named HORDE. This regularizer enforces visually-close images to have deep features with the same distribution which are well localized in the feature space. We provide a theoretical analysis supporting this regularization effect. We also show the effectiveness of our approach by obtaining state-of-the-art results on 4 well-known datasets (Cub-200-2011, Cars-196, Stanford Online Products and Inshop Clothes Retrieval).
Pierre Jacob, David Picard, Aymeric Histace, Edouard Klein
ICCV2
2019 Efficient Codebook and Factorization for Second Order Representation Learning
abstract
Learning rich and compact representations is an open topic in many fields such as object recognition or image retrieval. Deep neural networks have made a major breakthrough during the last few years for these tasks but their representations are not necessary as rich as needed nor as compact as expected. To build richer representations, high order statistics have been exploited and have shown excellent performances, but they produce higher dimensional features. While this drawback has been partially addressed with factorization schemes, the original compactness of first order models has never been retrieved, or at the cost of a strong performance decrease. Our method, by jointly integrating codebook strategy to factorization scheme, is able to produce compact representations while keeping the second order performances with few additional parameters. This formulation leads to state-of-the-art results on three image retrieval datasets.
Pierre Jacob, David Picard, Aymeric Histace, Edouard Klein
ICIP2
2019 Human pose regression by combining indirect part detection and contextual information
Diogo C. Luvizon, Hedi Tabia, David Picard
Comput. Graph.3
2019 Distributed optimization for deep learning with gossip exchange
Michaël Blot, David Picard, Nicolas Thome, Matthieu Cord
Neurocomputing2
2018 2D/3D Pose Estimation and Action Recognition Using Multitask Deep Learning
abstract
Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still images and human action recognition from video sequences. We show that a single architecture can be used to solve the two problems in an efficient way and still achieves state-of-the-art results. Additionally, we demonstrate that optimization from end-to-end leads to significantly higher accuracy than separated learning. The proposed architecture can be trained with data from different categories simultaneously in a seamlessly way. The reported results on four datasets (MPII, Human3.6M, Penn Action and NTU) demonstrate the effectiveness of our method on the targeted tasks.
Diogo C. Luvizon, David Picard, Hedi Tabia
CVPR2
2018 Image Reassembly Combining Deep Learning and Shortest Path Problem
Marie-Morgane Paumard, David Picard, Hedi Tabia
ECCV (6)2
2018 Leveraging Implicit Spatial Information in Global Features for Image Retrieval
abstract
Most image retrieval methods use global features that aggregate local distinctive patterns into a single representation. However, the aggregation process destroys the relative spatial information by considering orderless sets of local descriptors. We propose to integrate relative spatial information into the aggregation process by taking into account co-occurrences of local patterns in a tensor framework. The resulting signature called Improved Spatial Tensor Aggregation (ISTA) is able to reach state of the art performances on well known datasets such as Holidays, Oxford5k and Paris6k.
Pierre Jacob, David Picard, Aymeric Histace, Edouard Klein
ICIP2
2018 Jigsaw Puzzle Solving Using Local Feature Co-Occurrences in Deep Neural Networks
abstract
Archaeologists are in dire need of automated object reconstruction methods. Fragments reassembly is close to puzzle problems, which may be solved by computer vision algorithms. As they are often beaten on most image related tasks by deep learning algorithms, we study a classification method that can solve jigsaw puzzles. In this paper, we focus on classifying the relative position: given a couple of fragments, we compute their local relation (e.g. on top). We propose several enhancements over the state of the art in this domain, which is outperformed by our method by 25%. We propose an original dataset composed of pictures from the Metropolitan Museum of Art. We propose a greedy reconstruction method based on the predicted relative positions.
Marie-Morgane Paumard, David Picard, Hedi Tabia
ICIP2
2018 Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image Embeddings
abstract
Designing powerful tools that support cooking activities has rapidly gained popularity due to the massive amounts of available data, as well as recent advances in machine learning that are capable of analyzing them. In this paper, we propose a cross-modal retrieval model aligning visual and textual data (like pictures of dishes and their recipes) in a shared representation space. We describe an effective learning scheme, capable of tackling large-scale problems, and validate it on the Recipe1M dataset containing nearly 1 million picture-recipe pairs. We show the effectiveness of our approach regarding previous state-of-the-art models and present qualitative results over computational cooking use cases.
Micael Carvalho, Rémi Cadène, David Picard, Laure Soulier, Nicolas Thome, Matthieu Cord
SIGIR3
2017 Learning features combination for human action recognition from skeleton sequences
Diogo C. Luvizon, Hedi Tabia, David Picard
Pattern Recognit. Lett.3
2016 Preserving local spatial information in image similarity using tensor aggregation of local features
abstract
In this paper, we propose an aggregation scheme of local descriptors that preserves local spatial information. Our method is based on the binary product of similarities of nearby matching pairs of descriptors. The similarities are linearized using a tensor framework. We show our approach can be used with any local descriptors, handcrafted like SIFT, or learned like the outputs of convolutional layers in deep neural networks. We perform experiments on the Holidays dataset that show the soundness of the approach.
David Picard
ICIP1
2016 Local polynomial space-time descriptors for action classification
Olivier Kihl, David Picard, Philippe Henri Gosselin
Mach. Vis. Appl.2
2015 Asynchronous decentralized convex optimization through short-term gradient averaging
Jérôme Fellus, David Picard, Philippe Henri Gosselin
ESANN2
2015 Asynchronous gossip principal components analysis
Jérôme Fellus, David Picard, Philippe Henri Gosselin
Neurocomputing2
2015 A unified framework for local visual descriptors evaluation
Olivier Kihl, David Picard, Philippe Henri Gosselin
Pattern Recognit.2
2014 Covariance Descriptors for 3D Shape Matching and Retrieval
abstract
Several descriptors have been proposed in the past for 3D shape analysis, yet none of them achieves best performance on all shape classes. In this paper we propose a novel method for 3D shape analysis using the covariance matrices of the descriptors rather than the descriptors themselves. Covariance matrices enable efficient fusion of different types of features and modalities. They capture, using the same representation, not only the geometric and the spatial properties of a shape region but also the correlation of these properties within the region. Covariance matrices, however, lie on the manifold of Symmetric Positive Definite (SPD) tensors, a special type of Riemannian manifolds, which makes comparison and clustering of such matrices challenging. In this paper we study covariance matrices in their native space and make use of geodesic distances on the manifold as a dissimilarity measure. We demonstrate the performance of this metric on 3D face matching and recognition tasks. We then generalize the Bag of Features paradigm, originally designed in Euclidean spaces, to the Riemannian manifold of SPD matrices. We propose a new clustering procedure that takes into account the geometry of the Riemannian manifold. We evaluate the performance of the proposed Bag of Covariance Matrices framework on 3D shape matching and retrieval applications and demonstrate its superiority compared to descriptor-based techniques.
Hedi Tabia, Hamid Laga, David Picard, Philippe Henri Gosselin
CVPR3
2014 Dimensionality reduction in decentralized networks by Gossip aggregation of principal components analyzers
Jérôme Fellus, David Picard, Philippe Henri Gosselin
ESANN2
2014 Semantic pooling for image categorization using multiple kernel learning
abstract
In this paper, we propose a new method for taking into account the spatial information in image categorization. More specifically, we remove the loss of spatial information in Bag of Words related methods by computing the image signature over specific regions selected by object detectors. We propose to select the detectors using Multiple Kernel Learning techniques. We carry out experiments on the well known VOC 2007 dataset, and show our semantic pooling obtains promising results.
Thibaut Durand, David Picard, Nicolas Thome, Matthieu Cord
ICIP2
2014 Incremental learning of latent structural SVM for weakly supervised image classification
abstract
Visual learning with weak supervision is a promising research area, since it offers the possibility to build large image datasets at reasonable cost. In this paper, we address the problem of weakly supervised object detection, where the goal is to predict the label of the image using object position as latent variable. We propose a new method that builds upon the Latent Structural SVM (LSSVM) formalism. Specifically, we introduce an original coarse-to-fine approach that limits the evolution of the latent parameter subspace. This incremental strategy drives the learning towards better solutions, providing a model with increased predictive accuracy. In addition, this leads to a significant speed up during learning and inference compared to standard sliding window methods. Experiments carried out on Mammal dataset validate the good performances and fast training of the method compared to state-of-the-art works.
Thibaut Durand, Nicolas Thome, Matthieu Cord, David Picard
ICIP4
2014 Dimensionality reduction of visual features using sparse projectors for content-based image retrieval
abstract
In web-scale image retrieval, the most effective strategy is to aggregate local descriptors into a high dimensionality signature and then reduce it to a small dimensionality. Thanks to this strategy, web-scale image databases can be represented with small index and explored using fast visual similarities. However, the computation of this index has a very high complexity, because of the high dimensionality of signature projectors. In this work, we propose a new efficient method to greatly reduce the signature dimensionality with low computational and storage costs. Our method is based on the linear projection of the signature onto a small subspace using a sparse projection matrix. We report several experimental results on two standard datasets (Inria Holidays and Oxford) and with 100k image distractors. We show that our method reduces both the projectors storage cost and the computational cost of projection step while incurring a very slight loss in mAP (mean Average Precision) performance of these computed signatures.
Romain Negrel, David Picard, Philippe Henri Gosselin
ICIP2
2014 Photographic paper texture classification using model deviation of local visual descriptors
abstract
This paper investigates the classification of photographic paper textures using visual descriptors. Such classification is called fine grain due to the very low inter-class variability. We propose a novel image representation for photographic paper texture categorization, relying on the incorporation of a powerful local descriptor into an efficient higher-order model deviation where texture is represented by computing statistics on the occurrences of specific local visual patterns. We perform an evaluation on two different challenging datasets of photographic paper textures and show such advanced methods indeed outperforms existing descriptors.
David Picard, Ngoc-Son Vu, Inbar Fijalkow
ICIP1
2014 Efficient Metric Learning Based Dimension Reduction Using Sparse Projectors for Image Near Duplicate Retrieval
abstract
In this paper, we tackle the storage and computational cost of linear projections used in dimensionality reduction for near duplicate image retrieval. We propose a new method based on metric learning with a lower training cost than existing methods. Moreover, by adding a sparsity constraint, we obtain a projection matrix with a low storage and projection cost. We carry out experiments on a well known near duplicate image dataset and show our algorithm behaves correctly. Retrieval performances are shown to be promising when compared to the memory footprint and the projection cost of the obtained sparse matrix.
Romain Negrel, David Picard, Philippe Henri Gosselin
ICPR2
2013 Fast Approximation of Distance Between Elastic Curves using Kernels
abstract
Elastic shape analysis on non-linear Riemannian manifolds provides an efficient and elegant way for simultaneous comparison and registration of non-rigid shapes. In such formulation, shapes become points on some high dimensional shape space. A geodesic between two points corresponds to the optimal deformation needed to register one shape onto another. The length of the geodesic provides a proper metric for shape comparison. However, the computation of geodesics, and therefore the metric, is computationally very expensive as it involves a search over the space of all possible rotations and re- parameterization. This problem is even more important in shape retrieval scenarios where the query shape is compared to every element in the collection to search. In this paper, we propose a new procedure for metric approximation using the framework of kernel functions. We will demonstrate that this provides a fast approximation of the metric while preserving its invariance properties.
Hedi Tabia, David Picard, Hamid Laga, Philippe Henri Gosselin
BMVC2
2013 Machine Learning and Content-Based Multimedia Retrieval
Philippe Henri Gosselin, David Picard
ESANN2
2013 Multi-criteria search algorithm: An efficient approximate k-NN algorithm for image retrieval
abstract
We propose a new method for approximate k-NN search in large scale image databases, based on top-k multi-criteria search techniques. The method defines a simple index structure based on sorted lists, which provides a good compromise between fast retrieval, storage requirements and update cost. The search algorithm delivers approximate results with guarantees about false negatives, with fast emergence of good approximations, monotonically improved and leading if necessary to an exact result. Experiments with the on-disk implementation show that our method produces very good approximate results several times faster than the Baseline method.
Mehdi Badr, Dan Vodislav, David Picard, Shaoyi Yin, Philippe Henri Gosselin
ICIP3
2013 A unified formalism for video descriptors
abstract
In this paper, we propose a unified formalism for video descriptors. This formalism is based on the descriptors decomposition in three levels: primitive, scattering and projection. With this framework, we are able to rewrite easily all the usual descriptors in the literature such as HOG, HOF, SURF. Then, we propose a new projection method based on approximation with a finite expansion of orthogonal polynomials. Using our framework, we extend all usual descriptors by switching the projection step. The experiments are carried out on the well known KTH dataset and on the more challenging Hollywood 2 action classification dataset and show state of the art results.
Olivier Kihl, David Picard, Philippe Henri Gosselin
ICIP2
2013 3D shape similarity using vectors of locally aggregated tensors
abstract
In this paper, we present an efficient 3D object retrieval method invariant to scale, orientation and pose. Our approach is based on the dense extraction of discriminative local descriptors extracted from 2D views. We aggregate the descriptors into a single vector signature using tensor products. The similarity between 3D models can then be efficiently computed with a simple dot product. Experiments on the SHREC12 commonly-used benchmark demonstrate that our approach obtains superior performance in searching for generic shapes.
Hedi Tabia, David Picard, Hamid Laga, Philippe Henri Gosselin
ICIP2
2013 Efficient image signatures and similarities using tensor products of local descriptors
David Picard, Philippe Henri Gosselin
Comput. Vis. Image Underst.1
2013 JKernelMachines: a simple framework for kernel machine
David Picard, Nicolas Thome, Matthieu Cord
J. Mach. Learn. Res.1
2012 Learning geometric combinations of Gaussian kernels with alternating Quasi-Newton algorithm
David Picard, Nicolas Thome, Matthieu Cord, Alain Rakotomamonjy
ESANN1
2012 Compact tensor based image representation for similarity search
abstract
Within the Content Based Image Retrieval (CBIR) framework, one of the main challenges is to tackle the scalability issues. We propose a new compact signature for similarity search. We use an original method to perform a high compression of signatures while retraining their effectiveness. We propose an embedding method that maps large signatures into a low-dimensional Hilbert space. We evaluated the method on Holidays database and compared the results with methods of state-of-the-art.
Romain Negrel, David Picard, Philippe Henri Gosselin
ICIP2
2012 Classification of Urban Scenes from Geo-referenced Images in Urban Street-View Context
abstract
This paper addresses the challenging problem of scene classification in street-view georeferenced images of urban environments. More precisely, the goal of this task is semantic image classification, consisting in predicting in a given image, the presence or absence of a pre-defined class (e.g. shops, vegetation, etc.). The approach is based on the BOSSA representation, which enriches the Bag of Words (BoW) model, in conjunction with the Spatial Pyramid Matching scheme and kernel-based machine learning techniques. The proposed method handles problems that arise in large scale urban environments due to acquisition conditions (static and dynamic objects/pedestrians) combined with the continuous acquisition of data along the vehicle's direction, the varying light conditions and strong occlusions (due to the presence of trees, traffic signs, cars, etc.) giving rise to high intra-class variability. Experiments were conducted on a large dataset of high resolution images collected from two main avenues from the 12th district in Paris and the approach shows promising results.
Corina Iovan, David Picard, Nicolas Thome, Matthieu Cord
ICMLA (2)2
2012 Using spatial pyramids with compacted VLAT for image categorization
Romain Negrel, David Picard, Philippe Henri Gosselin
ICPR2
2012 An application of swarm intelligence to distributed image retrieval
David Picard, Arnaud Revel, Matthieu Cord
Inf. Sci.1
2011 Improving image similarity with vectors of locally aggregated tensors
abstract
Within the Content Based Image Retrieval (CBIR) framework, three main points can be highlighted: visual descriptors extraction, image signatures and their associated similarity measures, and machine learning based relevance functions. While the first and the last points have vastly improved in re- cent years, this paper addresses the second point. We propose a novel approach to compute vector representations extending state of the art methods in the field. Furthermore, our method can be viewed as a linearization of efficient well known kernel methods. The evaluation shows that our representation significantly improve state of the art results on the difficult VOC2007 database by a fair margin.
David Picard, Philippe Henri Gosselin
ICIP1
2010 An efficient system for combining complementary kernels in complex visual categorization tasks
abstract
Recently, increasing interest has been brought to improve image categorization performances by combining multiple descriptors. However, very few approaches have been proposed for combining features based on complementary aspects, and evaluating the performances in realistic databases. In this paper, we tackle the problem of combining different feature types (edge and color), and evaluate the performance gain in the very challenging VOC 2009 benchmark. Our contribution is three-fold. First, we propose new local color descriptors, unifying edge and color feature extraction into the “Bag Of Word” model. Second, we improve the Spatial Pyramid Matching (SPM) scheme for better incorporating spatial information into the similarity measurement. Last but not least, we propose a new combination strategy based on ℓ1Multiple Kernel Learning (MKL) that simultaneously learns individual kernel parameters and the kernel combination. Experiments prove the relevance of the proposed approach, which outperforms baseline combination methods while being computationally effective.
David Picard, Nicolas Thome, Matthieu Cord
ICIP1
2009 Geometric consistency checking for local-descriptor based document retrieval
abstract
International audience
Eduardo Valle, David Picard, Matthieu Cord
ACM Symposium on Document Engineering2
2008 Long term learning for image retrieval over networks
abstract
In this paper, we present a long term learning system for content based image retrieval over a network. Relevant feedback is used among different sessions to learn both the similarity function and the best routing for the searched category. Our system is based on mobile agents crawling the network in search of relevant images. An ant-behavior algorithm is used to learn the category dependent routing. With experiments on trecvid'05 key-frame dataset, we show that the smart association of category dependent routing and active learning leads to an improvement of the quality of the retrieval over time.
David Picard, Arnaud Revel, Matthieu Cord
ICIP1
2008 Image Retrieval Over Networks: Active Learning Using Ant Algorithm
abstract
In this article, we present a framework for distributed content based image retrieval with online learning based on ant-like mobile agents. Mobile agents crawl the network to find images matching a given example query. The images retrieved are shown to the user who labels them, following the classical relevant feedback scheme. The labels are used both to improve the similarity measure used for the retrieval and to learn paths leading to sites containing relevant images. The relevant paths are learned in an ethologically inspired way. We made experiments on the trecvid 2005 keyframe dataset showing that learning both the similarity function and the localization of the relevant images leads to a significant improvement. We also present an extension with the reuse of learned paths for later sessions leading to a further improvement.
David Picard, Matthieu Cord, Arnaud Revel
IEEE Trans. Multim.1
2006 CBIR in Distributed Databases using a Multi-Agent System
abstract
Information retrieval techniques have to face both the growing amount of data to be processed and the "natural"' distribution of these data over the network. Hence, we introduce in this paper a new architecture for image retrieval in distributed image databases, based on multi-agent, systems. Our system, inspired by "ant-agents", uses labels provided by the user for learning both the searched category of images and die path to the most relevant databases. We then show how effective can be our architecture on a generalist image database network.
David Picard, Matthieu Cord, Arnaud Revel
ICIP1
2006 Performances of Mobile-Agents for Interactive Image Retrieval
abstract
In this paper, we present a system for image retrieval over a network of computer based on "ant-like" mobile-agents. Image databases are hosted on the network, and the user wants to find all the images matching a specific concept (cars, flower, Italy, etc...). Usually, content based image retrieval systems (CBIR) do not consider the dispertion of the data among the network. We train a SVM classifier with examples annotated by the user and then launch mobile agents which explore the network in order to retrieve the most relevant images. Several interactive session (launching of agents then annotation of the results) are made to improve the classifier. Experiments are made both to see the influence of localization of the search concept on the quality of the learning, and to focus on the quality of the agent based solution compared to a centralizing system within a fixed amount of time for the interaction
David Picard, Matthieu Cord
Web Intelligence1