EDBT 2026 Demo / reviewers in the wild / expert
Matthieu Guillaumin
dblp:99/1687
· DBLP profile ↗
24ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 9 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Image recognition and object detection · 27% 3D vision · 22% Learning paradigms · 12% | |
| Computer graphics and multimedia
4 papers |
Multimedia analysis and retrieval · 100% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 42% Machine learning and data management · 36% Data mining · 22% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › image classification
large-scale image classification |
0.7 | 3 | 2016 | Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016 From categories to subcategories: Large-scale image classification with partial class label refinement · CVPR 2015 Incremental Learning of NCM Forests for Large-Scale Image Classification · CVPR 2014 |
Computer vision › 3D vision › 3d shape analysis
3d shape understanding |
0.6 | 1 | 2022 | ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.6 | 1 | 2022 | ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022 |
Machine learning › Learning paradigms
incremental learning |
0.4 | 2 | 2016 | Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016 Incremental Learning of NCM Forests for Large-Scale Image Classification · CVPR 2014 |
Machine learning › Efficient and distributed learning
auto labeling |
0.3 | 2 | 2014 | ImageNet Auto-Annotation with Segmentation Propagation · Int. J. Comput. Vis. 2014 Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012 |
Computer vision › Segmentation and scene understanding › video segmentation
segmentation propagation |
0.3 | 2 | 2014 | ImageNet Auto-Annotation with Segmentation Propagation · Int. J. Comput. Vis. 2014 Segmentation Propagation in ImageNet · ECCV (7) 2012 |
Computer vision › Face, body and person analysis
face recognition |
0.3 | 3 | 2012 | Face Recognition from Caption-Based Supervision · Int. J. Comput. Vis. 2012 Multiple Instance Metric Learning from Automatically Labeled Bags of Faces · ECCV (1) 2010 Is that you? Metric learning approaches for face identification · ICCV 2009 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.2 | 1 | 2016 | Incremental Learning of Random Forests for Large-Scale Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Computer vision › Vision and language › cross-modal supervision
caption-based supervision |
0.2 | 2 | 2012 | Face Recognition from Caption-Based Supervision · Int. J. Comput. Vis. 2012 Automatic face naming with caption-based supervision · CVPR 2008 |
Computer vision › 3D vision
feature matching |
0.2 | 1 | 2014 | Quantized Kernel Learning for Feature Matching · NIPS 2014 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.2 | 1 | 2014 | Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations · ECCV (5) 2014 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.2 | 1 | 2014 | Food-101 - Mining Discriminative Components with Random Forests · ECCV (6) 2014 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.2 | 1 | 2014 | Closed-Form Approximate CRF Training for Scalable Image Segmentation · ECCV (3) 2014 |
Computer vision › Video understanding and tracking
object tracking |
0.2 | 1 | 2014 | Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations · ECCV (5) 2014 |
Computer vision › 3D vision › inverse rendering
material estimation |
0.2 | 1 | 2022 | ABO: Dataset and Benchmarks for Real-World 3D Object Understanding · CVPR 2022 |
Machine learning › Optimization for machine learning
energy minimization |
0.2 | 1 | 2013 | Fast Energy Minimization Using Learned State Filters · CVPR 2013 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2013 | Prime Object Proposals with Randomized Prim's Algorithm · ICCV 2013 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.2 | 1 | 2013 | Prime Object Proposals with Randomized Prim's Algorithm · ICCV 2013 |
Multimedia analysis and retrieval › event understanding
event recognition |
0.2 | 1 | 2013 | Event Recognition in Photo Collections with a Stopwatch HMM · ICCV 2013 |
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis
image collection analysis |
0.2 | 1 | 2013 | Event Recognition in Photo Collections with a Stopwatch HMM · ICCV 2013 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 2 | 2012 | Multimodal semi-supervised learning for image classification · CVPR 2010 Segmentation Propagation in ImageNet · ECCV (7) 2012 |
Computer vision › Image recognition and object detection › object detection
bounding box annotation |
0.1 | 1 | 2012 | Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012 |
Machine learning › Transfer learning and domain adaptation
knowledge transfer |
0.1 | 1 | 2012 | Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012 |
Computer vision › Image recognition and object detection
object localization |
0.1 | 1 | 2012 | Large-scale knowledge transfer for object localization in ImageNet · CVPR 2012 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.1 | 1 | 2012 | Segmentation Propagation in ImageNet · ECCV (7) 2012 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2010 | Multimodal semi-supervised learning for image classification · CVPR 2010 |
Computer vision › Face, body and person analysis › face recognition
face identification |
0.1 | 1 | 2009 | Is that you? Metric learning approaches for face identification · ICCV 2009 |
Information retrieval
similarity learning |
0.1 | 1 | 2009 | TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009 |
Information retrieval
similarity search |
0.1 | 1 | 2009 | TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009 |
Multimedia analysis and retrieval
image annotation |
0.1 | 1 | 2009 | TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation · ICCV 2009 |
Methods — techniques the papers use, named apart from their topics
random forest · 0.8support vector machine · 0.4learned filtering · 0.3latent sub-event modeling · 0.3hidden markov model · 0.3nearest class mean classifier · 0.2regularized objective · 0.2sigmoidal modulation · 0.2nearest neighbor model · 0.2metric learning · 0.2trajectory annotations · 0.2quantization · 0.2nearest class mean forests · 0.2kernel learning · 0.2closed-form approximate CRF training · 0.2superpixel graph · 0.2randomized prim's algorithm · 0.2TRW-S · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LATTECLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic TextsabstractLarge-scale vision-language pre-trained (VLP) models (e.g., CLIP [46]) are renowned for their versatility, as they can be applied to diverse applications in a zero-shot setup. However, when these models are used in specific domains, their performance often falls short due to domain gaps or the under-representation of these domains in the training data. While fine-tuning VLP models on custom datasets with human-annotated labels can address this issue, annotating even a small-scale dataset (e.g., 100k samples) can be an expensive endeavor, often requiring expert annotators if the task is complex. To address these challenges, we propose LATTECLIP, an unsupervised method for fine-tuning CLIP models on classification with known class names in custom domains, without relying on human annotations. Our method leverages Large Multimodal Models (LMMs) to generate expressive textual descriptions for both individual images and groups of images. These provide additional contextual information to guide the fine-tuning process in the custom domains. Since LMM-generated descriptions are prone to hallucination or missing details, we introduce a novel strategy to distill only the useful information and stabilise the training. Specifically, we learn rich per-class prototype representations from noisy generated texts and dual pseudo-labels. Our experiments on 10 domain-specific datasets show that LATTECLIP outperforms pre-trained zero-shot methods by an average improvement of +4.74 points in top-1 accuracy and other state-of-the-art unsupervised methods by +3.45 points. Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin, Raoul de Charette, Loris Bazzani |
WACV | 3 |
| 2022 | ABO: Dataset and Benchmarks for Real-World 3D Object UnderstandingabstractWe introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with com-plex geometries and physically-based materials that cor-respond to real, household objects. We derive challenging benchmarks that exploit the unique properties of ABO and measure the current limits of the state-of-the-art on three open problems for real-world 3D object understanding: single-view 3D reconstruction, material estimation, and cross-domain multi-view object retrieval. Jasmine Collins, Shubham Goel 0001, Kenan Deng 0001, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Tomás F. Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, Jitendra Malik |
CVPR | 11 |
| 2016 | Incremental Learning of Random Forests for Large-Scale Image ClassificationabstractLarge image datasets such as ImageNet or open-ended photo websites like Flickr are revealing new challenges to image classification that were not apparent in smaller, fixed sets. In particular, the efficient handling of dynamically growing datasets, where not only the amount of training data but also the number of classes increases over time, is a relatively unexplored problem. In this challenging setting, we study how two variants of Random Forests (RF) perform under four strategies to incorporate new classes while avoiding to retrain the RFs from scratch. The various strategies account for different trade-offs between classification accuracy and computational efficiency. In our extensive experiments, we show that both RF variants, one based on Nearest Class Mean classifiers and the other on SVMs, outperform conventional RFs and are well suited for incrementally learning new classes. In particular, we show that RFs initially trained with just 10 classes can be extended to 1,000 classes with an acceptable loss of accuracy compared to training from the full data and with great computational savings compared to retraining for each new batch of classes. Marko Ristin, Matthieu Guillaumin, Juergen Gall, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | From categories to subcategories: Large-scale image classification with partial class label refinementabstractThe number of digital images is growing extremely rapidly, and so is the need for their classification. But, as more images of pre-defined categories become available, they also become more diverse and cover finer semantic differences. Ultimately, the categories themselves need to be divided into subcategories to account for that semantic refinement. Image classification in general has improved significantly over the last few years, but it still requires a massive amount of manually annotated data. Subdividing categories into subcategories multiples the number of labels, aggravating the annotation problem. Hence, we can expect the annotations to be refined only for a subset of the already labeled data, and exploit coarser labeled data to improve classification. In this work, we investigate how coarse category labels can be used to improve the classification of subcategories. To this end, we adopt the framework of Random Forests and propose a regularized objective function that takes into account relations between categories and subcategories. Compared to approaches that disregard the extra coarse labeled data, we achieve a relative improvement in subcategory classification accuracy of up to 22% in our large-scale image classification experiments. Marko Ristin, Juergen Gall, Matthieu Guillaumin, Luc Van Gool |
CVPR | 3 |
| 2014 | Non-maximum Suppression for Object Detection by Passing Messages Between Windows
Rasmus Rothe, Matthieu Guillaumin, Luc Van Gool |
ACCV (1) | 2 |
| 2014 | Learning to Rank Histograms for Object Retrieval
Danfeng Qin, Matthieu Guillaumin, Luc Van Gool |
BMVC | 3 |
| 2014 | Incremental Learning of NCM Forests for Large-Scale Image ClassificationabstractIn recent years, large image data sets such as "ImageNet", "TinyImages" or ever-growing social networks like "Flickr" have emerged, posing new challenges to image classification that were not apparent in smaller image sets. In particular, the efficient handling of dynamically growing data sets, where not only the amount of training images, but also the number of classes increases over time, is a relatively unexplored problem. To remedy this, we introduce Nearest Class Mean Forests (NCMF), a variant of Random Forests where the decision nodes are based on nearest class mean (NCM) classification. NCMFs not only outperform conventional random forests, but are also well suited for integrating new classes. To this end, we propose and compare several approaches to incorporate data from new classes, so as to seamlessly extend the previously trained forest instead of re-training them from scratch. In our experiments, we show that NCMFs trained on small data sets with 10 classes can be extended to large data sets with 1000 classes without significant loss of accuracy compared to training from scratch on the full data. Marko Ristin, Matthieu Guillaumin, Juergen Gall, Luc Van Gool |
CVPR | 2 |
| 2014 | Food-101 - Mining Discriminative Components with Random Forests
Lukas Bossard, Matthieu Guillaumin, Luc Van Gool |
ECCV (6) | 2 |
| 2014 | Closed-Form Approximate CRF Training for Scalable Image Segmentation
Alexander Kolesnikov 0003, Matthieu Guillaumin, Vittorio Ferrari, Christoph H. Lampert |
ECCV (3) | 2 |
| 2014 | Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations
Santiago Manen, Junseok Kwon, Matthieu Guillaumin, Luc Van Gool |
ECCV (5) | 3 |
| 2014 | Quantized Kernel Learning for Feature Matching
Danfeng Qin, Xuanli Chen, Matthieu Guillaumin, Luc Van Gool |
NIPS | 3 |
| 2014 | ImageNet Auto-Annotation with Segmentation Propagation
Matthieu Guillaumin, Daniel Küttel, Vittorio Ferrari |
Int. J. Comput. Vis. | 1 |
| 2013 | Fast Energy Minimization Using Learned State FiltersabstractPairwise discrete energies defined over graphs are ubiquitous in computer vision. Many algorithms have been proposed to minimize such energies, often concentrating on sparse graph topologies or specialized classes of pairwise potentials. However, when the graph is fully connected and the pairwise potentials are arbitrary, the complexity of even approximate minimization algorithms such as TRW-S grows quadratically both in the number of nodes and in the number of states a node can take. Moreover, recent applications are using more and more computationally expensive pairwise potentials. These factors make it very hard to employ fully connected models. In this paper we propose a novel, generic algorithm to approximately minimize any discrete pairwise energy function. Our method exploits tractable sub-energies to filter the domain of the function. The parameters of the filter are learnt from instances of the same class of energies with good candidate solutions. Compared to existing methods, it efficiently handles fully connected graphs, with many states per node, and arbitrary pairwise potentials, which might be expensive to compute. We demonstrate experimentally on two applications that our algorithm is much more efficient than other generic minimization algorithms such as TRW-S, while returning essentially identical solutions. Matthieu Guillaumin, Luc Van Gool, Vittorio Ferrari |
CVPR | 1 |
| 2013 | Event Recognition in Photo Collections with a Stopwatch HMMabstractThe task of recognizing events in photo collections is central for automatically organizing images. It is also very challenging, because of the ambiguity of photos across different event classes and because many photos do not convey enough relevant information. Unfortunately, the field still lacks standard evaluation data sets to allow comparison of different approaches. In this paper, we introduce and release a novel data set of personal photo collections containing more than 61,000 images in 807 collections, annotated with 14 diverse social event classes. Casting collections as sequential data, we build upon recent and state-of-the-art work in event recognition in videos to propose a latent sub-event approach for event recognition in photo collections. However, photos in collections are sparsely sampled over time and come in bursts from which transpires the importance of specific moments for the photographers. Thus, we adapt a discriminative hidden Markov model to allow the transitions between states to be a function of the time gap between consecutive images, which we coin as Stopwatch Hidden Markov model (SHMM). In our experiments, we show that our proposed model outperforms approaches based only on feature pooling or a classical hidden Markov model. With an average accuracy of 56%, we also highlight the difficulty of the data set and the need for future advances in event recognition in photo collections. Lukas Bossard, Matthieu Guillaumin, Luc Van Gool |
ICCV | 2 |
| 2013 | Prime Object Proposals with Randomized Prim's AlgorithmabstractGeneric object detection is the challenging task of proposing windows that localize all the objects in an image, regardless of their classes. Such detectors have recently been shown to benefit many applications such as speeding-up class-specific object detection, weakly supervised learning of object detectors and object discovery. In this paper, we introduce a novel and very efficient method for generic object detection based on a randomized version of Prim's algorithm. Using the connectivity graph of an image's super pixels, with weights modelling the probability that neighbouring super pixels belong to the same object, the algorithm generates random partial spanning trees with large expected sum of edge weights. Object localizations are proposed as bounding-boxes of those partial trees. Our method has several benefits compared to the state-of-the-art. Thanks to the efficiency of Prim's algorithm, it samples proposals very quickly: 1000 proposals are obtained in about 0.7s. With proposals bound to super pixel boundaries yet diversified by randomization, it yields very high detection rates and windows that tightly fit objects. In extensive experiments on the challenging PASCAL VOC 2007 and 2012 and SUN2012 benchmark datasets, we show that our method improves over state-of-the-art competitors for a wide range of evaluation scenarios. Santiago Manen, Matthieu Guillaumin, Luc Van Gool |
ICCV | 2 |
| 2012 | Large-scale knowledge transfer for object localization in ImageNetabstractImageNet is a large-scale database of object classes with millions of images. Unfortunately only a small fraction of them is manually annotated with bounding-boxes. This prevents useful developments, such as learning reliable object detectors for thousands of classes. In this paper we propose to automatically populate ImageNet with many more bounding-boxes, by leveraging existing manual annotations. The key idea is to localize objects of a target class for which annotations are not available, by transferring knowledge from related source classes with available annotations. We distinguish two kinds of source classes: ancestors and siblings. Each source provides knowledge about the plausible location, appearance and context of the target objects, which induces a probability distribution over windows in images of the target class. We learn to combine these distributions so as to maximize the location accuracy of the most probable window. Finally, we employ the combined distribution in a procedure to jointly localize objects in all images of the target class. Through experiments on 0.5 million images from 219 classes we show that our technique (i) annotates a wide range of classes with bounding-boxes; (ii) effectively exploits the hierarchical structure of ImageNet, since all sources and types of knowledge we propose contribute to the results; (iii) scales efficiently. Matthieu Guillaumin, Vittorio Ferrari |
CVPR | 1 |
| 2012 | Segmentation Propagation in ImageNet
Daniel Küttel, Matthieu Guillaumin, Vittorio Ferrari |
ECCV (7) | 2 |
| 2012 | Combining Image-Level and Segment-Level Models for Automatic Annotation
Daniel Küttel, Matthieu Guillaumin, Vittorio Ferrari |
MMM | 2 |
| 2012 | Face Recognition from Caption-Based Supervision
Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid |
Int. J. Comput. Vis. | 1 |
| 2010 | Multimodal semi-supervised learning for image classificationabstractIn image categorization the goal is to decide if an image belongs to a certain category or not. A binary classifier can be learned from manually labeled images; while using more labeled examples improves performance, obtaining the image labels is a time consuming process. We are interested in how other sources of information can aid the learning process given a fixed amount of labeled images. In particular, we consider a scenario where keywords are associated with the training images, e.g. as found on photo sharing websites. The goal is to learn a classifier for images alone, but we will use the keywords associated with labeled and unlabeled images to improve the classifier using semi-supervised learning. We first learn a strong Multiple Kernel Learning (MKL) classifier using both the image content and keywords, and use it to score unlabeled images. We then learn classifiers on visual features only, either support vector machines (SVM) or least-squares regression (LSR), from the MKL output values on both the labeled and unlabeled images. In our experiments on 20 classes from the PASCAL VOC'07 set and 38 from the MIR Flickr set, we demonstrate the benefit of our semi-supervised approach over only using the labeled images. We also present results for a scenario where we do not use any manual labeling but directly learn classifiers from the image tags. The semi-supervised approach also improves classification accuracy in this case. Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid |
CVPR | 1 |
| 2010 | Multiple Instance Metric Learning from Automatically Labeled Bags of Faces
Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid |
ECCV (1) | 1 |
| 2009 | TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotationabstractImage auto-annotation is an important open problem in computer vision. For this task we propose TagProp, a discriminatively trained nearest neighbor model. Tags of test images are predicted using a weighted nearest-neighbor model to exploit labeled training images. Neighbor weights are based on neighbor rank or distance. TagProp allows the integration of metric learning by directly maximizing the log-likelihood of the tag predictions in the training set. In this manner, we can optimally combine a collection of image similarity metrics that cover different aspects of image content, such as local shape descriptors, or global color histograms. We also introduce a word specific sigmoidal modulation of the weighted neighbor tag predictions to boost the recall of rare words. We investigate the performance of different variants of our model and compare to existing work. We present experimental results for three challenging data sets. On all three, TagProp makes a marked improvement as compared to the current state-of-the-art. Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid |
ICCV | 1 |
| 2009 | Is that you? Metric learning approaches for face identificationabstractFace identification is the problem of determining whether two face images depict the same person or not. This is difficult due to variations in scale, pose, lighting, background, expression, hairstyle, and glasses. In this paper we present two methods for learning robust distance measures: (a) a logistic discriminant approach which learns the metric from a set of labelled image pairs (LDML) and (b) a nearest neighbour approach which computes the probability for two images to belong to the same class (MkNN). We evaluate our approaches on the Labeled Faces in the Wild data set, a large and very challenging data set of faces from Yahoo! News. The evaluation protocol for this data set defines a restricted setting, where a fixed set of positive and negative image pairs is given, as well as an unrestricted one, where faces are labelled by their identity. We are the first to present results for the unrestricted setting, and show that our methods benefit from this richer training data, much more so than the current state-of-the-art method. Our results of 79.3% and 87.5% correct for the restricted and unrestricted setting respectively, significantly improve over the current state-of-the-art result of 78.5%. Confidence scores obtained for face identification can be used for many applications e.g. clustering or recognition from a single training example. We show that our learned metrics also improve performance for these tasks. Matthieu Guillaumin, Jakob Verbeek, Cordelia Schmid |
ICCV | 1 |
| 2008 | Automatic face naming with caption-based supervisionabstractWe consider two scenarios of naming people in databases of news photos with captions: (i) finding faces of a single person, and (ii) assigning names to all faces. We combine an initial text-based step, that restricts the name assigned to a face to the set of names appearing in the caption, with a second step that analyzes visual features of faces. By searching for groups of highly similar faces that can be associated with a name, the results of purely text-based search can be greatly ameliorated. We improve a recent graph-based approach, in which nodes correspond to faces and edges connect highly similar faces. We introduce constraints when optimizing the objective function, and propose improvements in the low-level methods used to construct the graphs. Furthermore, we generalize the graph-based approach to face naming in the full data set. In this multi-person naming case the optimization quickly becomes computationally demanding, and we present an important speed-up using graph-flows to compute the optimal name assignments in documents. Generative models have previously been proposed to solve the multi-person naming task. We compare the generative and graph-based methods in both scenarios, and find significantly better performance using the graph-based methods in both cases. Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid |
CVPR | 1 |