Maxime Oquab

dblp:151/8880 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Representation and self-supervised learning · 34% Deep learning architectures and training · 16% Image recognition and object detection · 16%

Topics — the 17 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
cross-modal alignment
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › joint embedding
joint embedding architecture
0.812024
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning
0.812024
Vision Transformers Need Registers · ICLR 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
Vision Transformers Need Registers · ICLR 2024
Computer vision › Image recognition and object detection
image classification
0.712023
Co-training 2L Submodels for Visual Recognition · CVPR 2023
Machine learning › Deep learning architectures and training › regularization
training regularization
0.712023
Co-training 2L Submodels for Visual Recognition · CVPR 2023
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization
0.522016
ContextLocNet: Context-Aware Deep Network Models for Weakly Supervised Localization · ECCV (5) 2016
Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.412019
Learning about an exponential amount of conditional distributions · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.412019
Learning about an exponential amount of conditional distributions · NeurIPS 2019
Machine learning › Generative modeling › diffusion model
conditional sampling
0.412019
Learning about an exponential amount of conditional distributions · NeurIPS 2019
Machine learning › Learning theory
hypothesis testing
0.312017
Revisiting Classifier Two-Sample Tests · ICLR (Poster) 2017
Machine learning › Learning theory › hypothesis testing
two-sample testing
0.312017
Revisiting Classifier Two-Sample Tests · ICLR (Poster) 2017
Machine learning › Efficient and distributed learning › efficient training
compute-constrained training
0.212024
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning · NeurIPS 2024
Computer vision › Image recognition and object detection
object localization
0.212015
Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212023
Co-training 2L Submodels for Visual Recognition · CVPR 2023
Computer vision › Image recognition and object detection › image classification
object classification
0.112015
Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.6self-supervised learning · 1.1text encoder alignment · 0.9lit training · 0.9stochastic depth · 0.7self-distillation · 0.7co-training · 0.7convolutional neural network · 0.4adversarial training · 0.4context modeling · 0.2
YearPublicationVenuePosition
2025 DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
abstract
Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP [67], self-supervised visual features are not readily aligned with language, hindering their adoption in open-vocabulary tasks. Our method, named dino.txt, unlocks this new ability for DINOv2 [63], a widely used self-supervised visual encoder. We build upon the LiT training strategy [97], which trains a text encoder to align with a frozen vision model but leads to unsatisfactory results on dense tasks. We propose several key ingredients to improve performance on both global and dense tasks, such as concatenating the [CLS] token with the patch average to train the alignment and curating data using both text and image modalities. With these, we successfully train a CLIP-like model with only a fraction of the computational cost compared to CLIP while achieving state-of-the-art results in zero-shot classification and open-vocabulary semantic segmentation.
Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre, Timothée Darcet, Hu Xu 0001, Daniel Li 0006, Marc Szafraniec, Michaël Ramamonjisoa, Maxime Oquab, Oriane Siméoni, Huy V. Vo, Patrick Labatut, Piotr Bojanowski
CVPR10
2024 Vision Transformers Need Registers
abstract
Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.
Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski
ICLR2
2024 You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
abstract
Self-Supervised learning (SSL) with Joint-Embedding Architectures (JEA) has led to outstanding performances. All instantiations of this paradigm were trained using strong and well-established hand-crafted data augmentations, leading to the general belief that they are required for the proper training and performance of such models. On the other hand, generative reconstruction-based models such as BEIT and MAE or Joint-Embedding Predictive Architectures such as I-JEPA have shown strong performance without using data augmentations except masking. In this work, we challenge the importance of invariance and data-augmentation in JEAs at scale. By running a case-study on a recent SSL foundation model -- DINOv2 -- we show that strong image representations can be obtained with JEAs and only cropping without resizing provided the training data is large enough, reaching state-of-the-art results and using the least amount of augmentation in the literature. Through this study, we also discuss the impact of compute constraints on the outcomes of experimental deep learning research, showing that they can lead to very different conclusions.
Théo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou, Piotr Bojanowski
NeurIPS2
2023 Co-training 2L Submodels for Visual Recognition
abstract
We introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, “submodels”, with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed “co-sub”, uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including RegNet, ViT, PiT, XCiT, Swin and ConvNext. Our training strategy improves their results in comparable settings. For instance, a ViT-B pretrained with cosub on ImageNet-21k obtains 87.4% top1 acc. @448 on ImageNet-val.
Hugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski, Jakob Verbeek, Hervé Jégou
CVPR3
2019 Consistent population control: generate plenty of points, but with a bit of resampling
abstract
Resampling methods, based on averaging the fitness of several clones, are the classical solution for dealing with noise. Population control has been proposed as a different tool for faster convergence of evolution strategies when the variance does not vanish around the optimum. However, we show that convergence may not hold even in the case of centered noise and construct a counterexample with variance dissymmetry, i.e. more variance on one side of the optimum than on the other. We propose a fix termed consistent population control and formally derive uniform constraints on deviations between averages and expectations within the proposed algorithm under either subgaussianity or finite variance assumptions on measurement noise. We prove convergence guarantees of consistent population control, verify it experimentally and show the effectiveness of population control in direct policy search, either with our fix or, in overparameterized cases, without our fix.
Vasil Khalidov, Maxime Oquab, Jérémy Rapin, Olivier Teytaud
FOGA2
2019 Learning about an exponential amount of conditional distributions
abstract
We introduce the Neural Conditioner (NC), a self-supervised machine able to learn about all the conditional distributions of a random vector X. The NC is a function NC(x⋅a,a,r) that leverages adversarial training to match each conditional distribution P(Xr|Xa=xa). After training, the NC generalizes to sample from conditional distributions never seen, including the joint distribution. The NC is also able to auto-encode examples, providing data representations useful for downstream classification tasks. In sum, the NC integrates different self-supervised tasks (each being the estimation of a conditional distribution) and levels of supervision (partially observed data) seamlessly into a single learning experience.
Mohamed Ishmael Belghazi, Maxime Oquab, David Lopez-Paz
NeurIPS2
2017 Revisiting Classifier Two-Sample Tests
David Lopez-Paz, Maxime Oquab
ICLR (Poster)2
2016 ContextLocNet: Context-Aware Deep Network Models for Weakly Supervised Localization
Vadim Kantorov, Maxime Oquab, Minsu Cho, Ivan Laptev
ECCV (5)2
2015 Is object localization for free? - Weakly-supervised learning with convolutional neural networks
abstract
Successful methods for visual object recognition typically rely on training datasets containing lots of richly annotated images. Detailed image annotation, e.g. by object bounding boxes, however, is both expensive and often subjective. We describe a weakly supervised convolutional neural network (CNN) for object classification that relies only on image-level labels, yet can learn from cluttered scenes containing multiple objects. We quantify its object classification and object location prediction performance on the Pascal VOC 2012 (20 object classes) and the much larger Microsoft COCO (80 object classes) datasets. We find that the network (i) outputs accurate image-level labels, (ii) predicts approximate locations (but not extents) of objects, and (iii) performs comparably to its fully-supervised counterparts using object bounding box annotation for training.
Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic
CVPR1
2014 Learning and Transferring Mid-level Image Representations Using Convolutional Neural Networks
abstract
Convolutional neural networks (CNN) have recently shown outstanding image classification performance in the large- scale visual recognition challenge (ILSVRC2012). The success of CNNs is attributed to their ability to learn rich mid-level image representations as opposed to hand-designed low-level features used in other image classification methods. Learning CNNs, however, amounts to estimating millions of parameters and requires a very large number of annotated image samples. This property currently prevents application of CNNs to problems with limited training data. In this work we show how image representations learned with CNNs on large-scale annotated datasets can be efficiently transferred to other visual recognition tasks with limited amount of training data. We design a method to reuse layers trained on the ImageNet dataset to compute mid-level image representation for images in the PASCAL VOC dataset. We show that despite differences in image statistics and tasks in the two datasets, the transferred representation leads to significantly improved results for object and action classification, outperforming the current state of the art on Pascal VOC 2007 and 2012 datasets. We also show promising results for object and action localization.
Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic
CVPR1