Matthieu Guillaumin

dblp:99/1687 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 9 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Image recognition and object detection · 27% 3D vision · 22% Learning paradigms · 12%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 42% Machine learning and data management · 36% Data mining · 22%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
large-scale image classification
0.732016
Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
From categories to subcategories: Large-scale image classification with partial class label refinement · CVPR 2015
Incremental Learning of NCM Forests for Large-Scale Image Classification · CVPR 2014
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.612022
ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.612022
ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022
Machine learning › Learning paradigms
incremental learning
0.422016
Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Incremental Learning of NCM Forests for Large-Scale Image Classification · CVPR 2014
Machine learning › Efficient and distributed learning
auto labeling
0.322014
ImageNet Auto-Annotation with Segmentation Propagation · Int. J. Comput. Vis. 2014
Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012
Computer vision › Segmentation and scene understanding › video segmentation
segmentation propagation
0.322014
ImageNet Auto-Annotation with Segmentation Propagation · Int. J. Comput. Vis. 2014
Segmentation Propagation in ImageNet · ECCV (7) 2012
Computer vision › Face, body and person analysis
face recognition
0.332012
Face Recognition from Caption-Based Supervision · Int. J. Comput. Vis. 2012
Multiple Instance Metric Learning from Automatically Labeled Bags of Faces · ECCV (1) 2010
Is that you? Metric learning approaches for face identification · ICCV 2009
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.212016
Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Vision and language › cross-modal supervision
caption-based supervision
0.222012
Face Recognition from Caption-Based Supervision · Int. J. Comput. Vis. 2012
Automatic face naming with caption-based supervision · CVPR 2008
Computer vision › 3D vision
feature matching
0.212014
Quantized Kernel Learning for Feature Matching · NIPS 2014
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212014
Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations · ECCV (5) 2014
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.212014
Food-101 - Mining Discriminative Components with Random Forests · ECCV (6) 2014
Computer vision › Segmentation and scene understanding
image segmentation
0.212014
Closed-Form Approximate CRF Training for Scalable Image Segmentation · ECCV (3) 2014
Computer vision › Video understanding and tracking
object tracking
0.212014
Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations · ECCV (5) 2014
Computer vision › 3D vision › inverse rendering
material estimation
0.212022
ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022
Machine learning › Optimization for machine learning
energy minimization
0.212013
Fast Energy Minimization Using Learned State Filters · CVPR 2013
Computer vision › Image recognition and object detection
object detection
0.212013
Prime Object Proposals with Randomized Prim's Algorithm · ICCV 2013
Computer vision › Image recognition and object detection › object detection
object proposal generation
0.212013
Prime Object Proposals with Randomized Prim's Algorithm · ICCV 2013
Multimedia analysis and retrieval › event understanding
event recognition
0.212013
Event Recognition in Photo Collections with a Stopwatch HMM · ICCV 2013
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis
image collection analysis
0.212013
Event Recognition in Photo Collections with a Stopwatch HMM · ICCV 2013
Computer vision › Image recognition and object detection
image classification
0.222012
Multimodal semi-supervised learning for image classification · CVPR 2010
Segmentation Propagation in ImageNet · ECCV (7) 2012
Computer vision › Image recognition and object detection › object detection
bounding box annotation
0.112012
Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.112012
Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012
Computer vision › Image recognition and object detection
object localization
0.112012
Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112012
Segmentation Propagation in ImageNet · ECCV (7) 2012
Machine learning › Learning paradigms
semi-supervised learning
0.112010
Multimodal semi-supervised learning for image classification · CVPR 2010
Computer vision › Face, body and person analysis › face recognition
face identification
0.112009
Is that you? Metric learning approaches for face identification · ICCV 2009
Information retrieval
similarity learning
0.112009
TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009
Information retrieval
similarity search
0.112009
TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009
Multimedia analysis and retrieval
image annotation
0.112009
TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009

Methods — techniques the papers use, named apart from their topics

random forest · 0.8support vector machine · 0.4learned filtering · 0.3latent sub-event modeling · 0.3hidden markov model · 0.3nearest class mean classifier · 0.2regularized objective · 0.2sigmoidal modulation · 0.2nearest neighbor model · 0.2metric learning · 0.2trajectory annotations · 0.2quantization · 0.2nearest class mean forests · 0.2kernel learning · 0.2closed-form approximate CRF training · 0.2superpixel graph · 0.2randomized prim's algorithm · 0.2TRW-S · 0.2
YearPublicationVenuePosition
2025 LATTECLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
abstract
Large-scale vision-language pre-trained (VLP) models (e.g., CLIP [46]) are renowned for their versatility, as they can be applied to diverse applications in a zero-shot setup. However, when these models are used in specific domains, their performance often falls short due to domain gaps or the under-representation of these domains in the training data. While fine-tuning VLP models on custom datasets with human-annotated labels can address this issue, annotating even a small-scale dataset (e.g., 100k samples) can be an expensive endeavor, often requiring expert annotators if the task is complex. To address these challenges, we propose LATTECLIP, an unsupervised method for fine-tuning CLIP models on classification with known class names in custom domains, without relying on human annotations. Our method leverages Large Multimodal Models (LMMs) to generate expressive textual descriptions for both individual images and groups of images. These provide additional contextual information to guide the fine-tuning process in the custom domains. Since LMM-generated descriptions are prone to hallucination or missing details, we introduce a novel strategy to distill only the useful information and stabilise the training. Specifically, we learn rich per-class prototype representations from noisy generated texts and dual pseudo-labels. Our experiments on 10 domain-specific datasets show that LATTECLIP outperforms pre-trained zero-shot methods by an average improvement of +4.74 points in top-1 accuracy and other state-of-the-art unsupervised methods by +3.45 points.
Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin, Raoul de Charette, Loris Bazzani
WACV3
2022 ABO: Dataset and Benchmarks for Real-World 3D Object Understanding
abstract
We introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with com-plex geometries and physically-based materials that cor-respond to real, household objects. We derive challenging benchmarks that exploit the unique properties of ABO and measure the current limits of the state-of-the-art on three open problems for real-world 3D object understanding: single-view 3D reconstruction, material estimation, and cross-domain multi-view object retrieval.
Jasmine Collins, Shubham Goel 0001, Kenan Deng 0001, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Tomás F. Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, Jitendra Malik
CVPR11
2016 Incremental Learning of Random Forests for Large-Scale Image Classification
abstract
Large image datasets such as ImageNet or open-ended photo websites like Flickr are revealing new challenges to image classification that were not apparent in smaller, fixed sets. In particular, the efficient handling of dynamically growing datasets, where not only the amount of training data but also the number of classes increases over time, is a relatively unexplored problem. In this challenging setting, we study how two variants of Random Forests (RF) perform under four strategies to incorporate new classes while avoiding to retrain the RFs from scratch. The various strategies account for different trade-offs between classification accuracy and computational efficiency. In our extensive experiments, we show that both RF variants, one based on Nearest Class Mean classifiers and the other on SVMs, outperform conventional RFs and are well suited for incrementally learning new classes. In particular, we show that RFs initially trained with just 10 classes can be extended to 1,000 classes with an acceptable loss of accuracy compared to training from the full data and with great computational savings compared to retraining for each new batch of classes.
Marko Ristin, Matthieu Guillaumin, Juergen Gall, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 From categories to subcategories: Large-scale image classification with partial class label refinement
abstract
The number of digital images is growing extremely rapidly, and so is the need for their classification. But, as more images of pre-defined categories become available, they also become more diverse and cover finer semantic differences. Ultimately, the categories themselves need to be divided into subcategories to account for that semantic refinement. Image classification in general has improved significantly over the last few years, but it still requires a massive amount of manually annotated data. Subdividing categories into subcategories multiples the number of labels, aggravating the annotation problem. Hence, we can expect the annotations to be refined only for a subset of the already labeled data, and exploit coarser labeled data to improve classification. In this work, we investigate how coarse category labels can be used to improve the classification of subcategories. To this end, we adopt the framework of Random Forests and propose a regularized objective function that takes into account relations between categories and subcategories. Compared to approaches that disregard the extra coarse labeled data, we achieve a relative improvement in subcategory classification accuracy of up to 22% in our large-scale image classification experiments.
Marko Ristin, Juergen Gall, Matthieu Guillaumin, Luc Van Gool
CVPR3
2014 Non-maximum Suppression for Object Detection by Passing Messages Between Windows
Rasmus Rothe, Matthieu Guillaumin, Luc Van Gool
ACCV (1)2
2014 Learning to Rank Histograms for Object Retrieval
Danfeng Qin, Matthieu Guillaumin, Luc Van Gool
BMVC3
2014 Incremental Learning of NCM Forests for Large-Scale Image Classification
abstract
In recent years, large image data sets such as "ImageNet", "TinyImages" or ever-growing social networks like "Flickr" have emerged, posing new challenges to image classification that were not apparent in smaller image sets. In particular, the efficient handling of dynamically growing data sets, where not only the amount of training images, but also the number of classes increases over time, is a relatively unexplored problem. To remedy this, we introduce Nearest Class Mean Forests (NCMF), a variant of Random Forests where the decision nodes are based on nearest class mean (NCM) classification. NCMFs not only outperform conventional random forests, but are also well suited for integrating new classes. To this end, we propose and compare several approaches to incorporate data from new classes, so as to seamlessly extend the previously trained forest instead of re-training them from scratch. In our experiments, we show that NCMFs trained on small data sets with 10 classes can be extended to large data sets with 1000 classes without significant loss of accuracy compared to training from scratch on the full data.
Marko Ristin, Matthieu Guillaumin, Juergen Gall, Luc Van Gool
CVPR2
2014 Food-101 - Mining Discriminative Components with Random Forests
Lukas Bossard, Matthieu Guillaumin, Luc Van Gool
ECCV (6)2
2014 Closed-Form Approximate CRF Training for Scalable Image Segmentation
Alexander Kolesnikov 0003, Matthieu Guillaumin, Vittorio Ferrari, Christoph H. Lampert
ECCV (3)2
2014 Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations
Santiago Manen, Junseok Kwon, Matthieu Guillaumin, Luc Van Gool
ECCV (5)3
2014 Quantized Kernel Learning for Feature Matching
Danfeng Qin, Xuanli Chen, Matthieu Guillaumin, Luc Van Gool
NIPS3
2014 ImageNet Auto-Annotation with Segmentation Propagation
Matthieu Guillaumin, Daniel Küttel, Vittorio Ferrari
Int. J. Comput. Vis.1
2013 Fast Energy Minimization Using Learned State Filters
abstract
Pairwise discrete energies defined over graphs are ubiquitous in computer vision. Many algorithms have been proposed to minimize such energies, often concentrating on sparse graph topologies or specialized classes of pairwise potentials. However, when the graph is fully connected and the pairwise potentials are arbitrary, the complexity of even approximate minimization algorithms such as TRW-S grows quadratically both in the number of nodes and in the number of states a node can take. Moreover, recent applications are using more and more computationally expensive pairwise potentials. These factors make it very hard to employ fully connected models. In this paper we propose a novel, generic algorithm to approximately minimize any discrete pairwise energy function. Our method exploits tractable sub-energies to filter the domain of the function. The parameters of the filter are learnt from instances of the same class of energies with good candidate solutions. Compared to existing methods, it efficiently handles fully connected graphs, with many states per node, and arbitrary pairwise potentials, which might be expensive to compute. We demonstrate experimentally on two applications that our algorithm is much more efficient than other generic minimization algorithms such as TRW-S, while returning essentially identical solutions.
Matthieu Guillaumin, Luc Van Gool, Vittorio Ferrari
CVPR1
2013 Event Recognition in Photo Collections with a Stopwatch HMM
abstract
The task of recognizing events in photo collections is central for automatically organizing images. It is also very challenging, because of the ambiguity of photos across different event classes and because many photos do not convey enough relevant information. Unfortunately, the field still lacks standard evaluation data sets to allow comparison of different approaches. In this paper, we introduce and release a novel data set of personal photo collections containing more than 61,000 images in 807 collections, annotated with 14 diverse social event classes. Casting collections as sequential data, we build upon recent and state-of-the-art work in event recognition in videos to propose a latent sub-event approach for event recognition in photo collections. However, photos in collections are sparsely sampled over time and come in bursts from which transpires the importance of specific moments for the photographers. Thus, we adapt a discriminative hidden Markov model to allow the transitions between states to be a function of the time gap between consecutive images, which we coin as Stopwatch Hidden Markov model (SHMM). In our experiments, we show that our proposed model outperforms approaches based only on feature pooling or a classical hidden Markov model. With an average accuracy of 56%, we also highlight the difficulty of the data set and the need for future advances in event recognition in photo collections.
Lukas Bossard, Matthieu Guillaumin, Luc Van Gool
ICCV2
2013 Prime Object Proposals with Randomized Prim's Algorithm
abstract
Generic object detection is the challenging task of proposing windows that localize all the objects in an image, regardless of their classes. Such detectors have recently been shown to benefit many applications such as speeding-up class-specific object detection, weakly supervised learning of object detectors and object discovery. In this paper, we introduce a novel and very efficient method for generic object detection based on a randomized version of Prim's algorithm. Using the connectivity graph of an image's super pixels, with weights modelling the probability that neighbouring super pixels belong to the same object, the algorithm generates random partial spanning trees with large expected sum of edge weights. Object localizations are proposed as bounding-boxes of those partial trees. Our method has several benefits compared to the state-of-the-art. Thanks to the efficiency of Prim's algorithm, it samples proposals very quickly: 1000 proposals are obtained in about 0.7s. With proposals bound to super pixel boundaries yet diversified by randomization, it yields very high detection rates and windows that tightly fit objects. In extensive experiments on the challenging PASCAL VOC 2007 and 2012 and SUN2012 benchmark datasets, we show that our method improves over state-of-the-art competitors for a wide range of evaluation scenarios.
Santiago Manen, Matthieu Guillaumin, Luc Van Gool
ICCV2
2012 Large-scale knowledge transfer for object localization in ImageNet
abstract
ImageNet is a large-scale database of object classes with millions of images. Unfortunately only a small fraction of them is manually annotated with bounding-boxes. This prevents useful developments, such as learning reliable object detectors for thousands of classes. In this paper we propose to automatically populate ImageNet with many more bounding-boxes, by leveraging existing manual annotations. The key idea is to localize objects of a target class for which annotations are not available, by transferring knowledge from related source classes with available annotations. We distinguish two kinds of source classes: ancestors and siblings. Each source provides knowledge about the plausible location, appearance and context of the target objects, which induces a probability distribution over windows in images of the target class. We learn to combine these distributions so as to maximize the location accuracy of the most probable window. Finally, we employ the combined distribution in a procedure to jointly localize objects in all images of the target class. Through experiments on 0.5 million images from 219 classes we show that our technique (i) annotates a wide range of classes with bounding-boxes; (ii) effectively exploits the hierarchical structure of ImageNet, since all sources and types of knowledge we propose contribute to the results; (iii) scales efficiently.
Matthieu Guillaumin, Vittorio Ferrari
CVPR1
2012 Segmentation Propagation in ImageNet
Daniel Küttel, Matthieu Guillaumin, Vittorio Ferrari
ECCV (7)2
2012 Combining Image-Level and Segment-Level Models for Automatic Annotation
Daniel Küttel, Matthieu Guillaumin, Vittorio Ferrari
MMM2
2012 Face Recognition from Caption-Based Supervision
Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid
Int. J. Comput. Vis.1
2010 Multimodal semi-supervised learning for image classification
abstract
In image categorization the goal is to decide if an image belongs to a certain category or not. A binary classifier can be learned from manually labeled images; while using more labeled examples improves performance, obtaining the image labels is a time consuming process. We are interested in how other sources of information can aid the learning process given a fixed amount of labeled images. In particular, we consider a scenario where keywords are associated with the training images, e.g. as found on photo sharing websites. The goal is to learn a classifier for images alone, but we will use the keywords associated with labeled and unlabeled images to improve the classifier using semi-supervised learning. We first learn a strong Multiple Kernel Learning (MKL) classifier using both the image content and keywords, and use it to score unlabeled images. We then learn classifiers on visual features only, either support vector machines (SVM) or least-squares regression (LSR), from the MKL output values on both the labeled and unlabeled images. In our experiments on 20 classes from the PASCAL VOC'07 set and 38 from the MIR Flickr set, we demonstrate the benefit of our semi-supervised approach over only using the labeled images. We also present results for a scenario where we do not use any manual labeling but directly learn classifiers from the image tags. The semi-supervised approach also improves classification accuracy in this case.
Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid
CVPR1
2010 Multiple Instance Metric Learning from Automatically Labeled Bags of Faces
Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid
ECCV (1)1
2009 TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation
abstract
Image auto-annotation is an important open problem in computer vision. For this task we propose TagProp, a discriminatively trained nearest neighbor model. Tags of test images are predicted using a weighted nearest-neighbor model to exploit labeled training images. Neighbor weights are based on neighbor rank or distance. TagProp allows the integration of metric learning by directly maximizing the log-likelihood of the tag predictions in the training set. In this manner, we can optimally combine a collection of image similarity metrics that cover different aspects of image content, such as local shape descriptors, or global color histograms. We also introduce a word specific sigmoidal modulation of the weighted neighbor tag predictions to boost the recall of rare words. We investigate the performance of different variants of our model and compare to existing work. We present experimental results for three challenging data sets. On all three, TagProp makes a marked improvement as compared to the current state-of-the-art.
Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid
ICCV1
2009 Is that you? Metric learning approaches for face identification
abstract
Face identification is the problem of determining whether two face images depict the same person or not. This is difficult due to variations in scale, pose, lighting, background, expression, hairstyle, and glasses. In this paper we present two methods for learning robust distance measures: (a) a logistic discriminant approach which learns the metric from a set of labelled image pairs (LDML) and (b) a nearest neighbour approach which computes the probability for two images to belong to the same class (MkNN). We evaluate our approaches on the Labeled Faces in the Wild data set, a large and very challenging data set of faces from Yahoo! News. The evaluation protocol for this data set defines a restricted setting, where a fixed set of positive and negative image pairs is given, as well as an unrestricted one, where faces are labelled by their identity. We are the first to present results for the unrestricted setting, and show that our methods benefit from this richer training data, much more so than the current state-of-the-art method. Our results of 79.3% and 87.5% correct for the restricted and unrestricted setting respectively, significantly improve over the current state-of-the-art result of 78.5%. Confidence scores obtained for face identification can be used for many applications e.g. clustering or recognition from a single training example. We show that our learned metrics also improve performance for these tasks.
Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid
ICCV1
2008 Automatic face naming with caption-based supervision
abstract
We consider two scenarios of naming people in databases of news photos with captions: (i) finding faces of a single person, and (ii) assigning names to all faces. We combine an initial text-based step, that restricts the name assigned to a face to the set of names appearing in the caption, with a second step that analyzes visual features of faces. By searching for groups of highly similar faces that can be associated with a name, the results of purely text-based search can be greatly ameliorated. We improve a recent graph-based approach, in which nodes correspond to faces and edges connect highly similar faces. We introduce constraints when optimizing the objective function, and propose improvements in the low-level methods used to construct the graphs. Furthermore, we generalize the graph-based approach to face naming in the full data set. In this multi-person naming case the optimization quickly becomes computationally demanding, and we present an important speed-up using graph-flows to compute the optimal name assignments in documents. Generative models have previously been proposed to solve the multi-person naming task. We compare the generative and graph-based methods in both scenarios, and find significantly better performance using the graph-based methods in both cases.
Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid
CVPR1