VLDB 2026 Research / reviewers in the wild / expert
Maxime Oquab
dblp:151/8880
· DBLP profile ↗
10ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Representation and self-supervised learning · 34% Deep learning architectures and training · 16% Image recognition and object detection · 16% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
cross-modal alignment |
0.9 | 1 | 2025 | DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model |
0.9 | 1 | 2025 | DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › joint embedding
joint embedding architecture |
0.8 | 1 | 2024 | You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning |
0.8 | 1 | 2024 | Vision Transformers Need Registers · ICLR 2024 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.8 | 1 | 2024 | Vision Transformers Need Registers · ICLR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 1 | 2023 | Co-training 2L Submodels for Visual Recognition · CVPR 2023 |
Machine learning › Deep learning architectures and training › regularization
training regularization |
0.7 | 1 | 2023 | Co-training 2L Submodels for Visual Recognition · CVPR 2023 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.5 | 2 | 2016 | ContextLocNet: Context-Aware Deep Network Models for Weakly Supervised Localization · ECCV (5) 2016 Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.4 | 1 | 2019 | Learning about an exponential amount of conditional distributions · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation |
0.4 | 1 | 2019 | Learning about an exponential amount of conditional distributions · NeurIPS 2019 |
Machine learning › Generative modeling › diffusion model
conditional sampling |
0.4 | 1 | 2019 | Learning about an exponential amount of conditional distributions · NeurIPS 2019 |
Machine learning › Learning theory
hypothesis testing |
0.3 | 1 | 2017 | Revisiting Classifier Two-Sample Tests · ICLR (Poster) 2017 |
Machine learning › Learning theory › hypothesis testing
two-sample testing |
0.3 | 1 | 2017 | Revisiting Classifier Two-Sample Tests · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning › efficient training
compute-constrained training |
0.2 | 1 | 2024 | You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning · NeurIPS 2024 |
Computer vision › Image recognition and object detection
object localization |
0.2 | 1 | 2015 | Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2023 | Co-training 2L Submodels for Visual Recognition · CVPR 2023 |
Computer vision › Image recognition and object detection › image classification
object classification |
0.1 | 1 | 2015 | Is object localization for free? - Weakly-supervised learning with convolutional neural networks · CVPR 2015 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.6self-supervised learning · 1.1text encoder alignment · 0.9lit training · 0.9stochastic depth · 0.7self-distillation · 0.7co-training · 0.7convolutional neural network · 0.4adversarial training · 0.4context modeling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language AlignmentabstractSelf-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP [67], self-supervised visual features are not readily aligned with language, hindering their adoption in open-vocabulary tasks. Our method, named dino.txt, unlocks this new ability for DINOv2 [63], a widely used self-supervised visual encoder. We build upon the LiT training strategy [97], which trains a text encoder to align with a frozen vision model but leads to unsatisfactory results on dense tasks. We propose several key ingredients to improve performance on both global and dense tasks, such as concatenating the [CLS] token with the patch average to train the alignment and curating data using both text and image modalities. With these, we successfully train a CLIP-like model with only a fraction of the computational cost compared to CLIP while achieving state-of-the-art results in zero-shot classification and open-vocabulary semantic segmentation. Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre, Timothée Darcet, Hu Xu 0001, Daniel Li 0006, Marc Szafraniec, Michaël Ramamonjisoa, Maxime Oquab, Oriane Siméoni, Huy V. Vo, Patrick Labatut, Piotr Bojanowski |
CVPR | 10 |
| 2024 | Vision Transformers Need RegistersabstractTransformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing. Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski |
ICLR | 2 |
| 2024 | You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised LearningabstractSelf-Supervised learning (SSL) with Joint-Embedding Architectures (JEA) has led to outstanding performances. All instantiations of this paradigm were trained using strong and well-established hand-crafted data augmentations, leading to the general belief that they are required for the proper training and performance of such models. On the other hand, generative reconstruction-based models such as BEIT and MAE or Joint-Embedding Predictive Architectures such as I-JEPA have shown strong performance without using data augmentations except masking. In this work, we challenge the importance of invariance and data-augmentation in JEAs at scale. By running a case-study on a recent SSL foundation model -- DINOv2 -- we show that strong image representations can be obtained with JEAs and only cropping without resizing provided the training data is large enough, reaching state-of-the-art results and using the least amount of augmentation in the literature. Through this study, we also discuss the impact of compute constraints on the outcomes of experimental deep learning research, showing that they can lead to very different conclusions. Théo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou, Piotr Bojanowski |
NeurIPS | 2 |
| 2023 | Co-training 2L Submodels for Visual RecognitionabstractWe introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, “submodels”, with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed “co-sub”, uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including RegNet, ViT, PiT, XCiT, Swin and ConvNext. Our training strategy improves their results in comparable settings. For instance, a ViT-B pretrained with cosub on ImageNet-21k obtains 87.4% top1 acc. @448 on ImageNet-val. Hugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski, Jakob Verbeek, Hervé Jégou |
CVPR | 3 |
| 2019 | Consistent population control: generate plenty of points, but with a bit of resamplingabstractResampling methods, based on averaging the fitness of several clones, are the classical solution for dealing with noise. Population control has been proposed as a different tool for faster convergence of evolution strategies when the variance does not vanish around the optimum. However, we show that convergence may not hold even in the case of centered noise and construct a counterexample with variance dissymmetry, i.e. more variance on one side of the optimum than on the other. We propose a fix termed consistent population control and formally derive uniform constraints on deviations between averages and expectations within the proposed algorithm under either subgaussianity or finite variance assumptions on measurement noise. We prove convergence guarantees of consistent population control, verify it experimentally and show the effectiveness of population control in direct policy search, either with our fix or, in overparameterized cases, without our fix. Vasil Khalidov, Maxime Oquab, Jérémy Rapin, Olivier Teytaud |
FOGA | 2 |
| 2019 | Learning about an exponential amount of conditional distributionsabstractWe introduce the Neural Conditioner (NC), a self-supervised machine able to learn about all the conditional distributions of a random vector X. The NC is a function NC(x⋅a,a,r) that leverages adversarial training to match each conditional distribution P(Xr|Xa=xa). After training, the NC generalizes to sample from conditional distributions never seen, including the joint distribution. The NC is also able to auto-encode examples, providing data representations useful for downstream classification tasks. In sum, the NC integrates different self-supervised tasks (each being the estimation of a conditional distribution) and levels of supervision (partially observed data) seamlessly into a single learning experience. Mohamed Ishmael Belghazi, Maxime Oquab, David Lopez-Paz |
NeurIPS | 2 |
| 2017 | Revisiting Classifier Two-Sample Tests
David Lopez-Paz, Maxime Oquab |
ICLR (Poster) | 2 |
| 2016 | ContextLocNet: Context-Aware Deep Network Models for Weakly Supervised Localization
Vadim Kantorov, Maxime Oquab, Minsu Cho, Ivan Laptev |
ECCV (5) | 2 |
| 2015 | Is object localization for free? - Weakly-supervised learning with convolutional neural networksabstractSuccessful methods for visual object recognition typically rely on training datasets containing lots of richly annotated images. Detailed image annotation, e.g. by object bounding boxes, however, is both expensive and often subjective. We describe a weakly supervised convolutional neural network (CNN) for object classification that relies only on image-level labels, yet can learn from cluttered scenes containing multiple objects. We quantify its object classification and object location prediction performance on the Pascal VOC 2012 (20 object classes) and the much larger Microsoft COCO (80 object classes) datasets. We find that the network (i) outputs accurate image-level labels, (ii) predicts approximate locations (but not extents) of objects, and (iii) performs comparably to its fully-supervised counterparts using object bounding box annotation for training. Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic |
CVPR | 1 |
| 2014 | Learning and Transferring Mid-level Image Representations Using Convolutional Neural NetworksabstractConvolutional neural networks (CNN) have recently shown outstanding image classification performance in the large- scale visual recognition challenge (ILSVRC2012). The success of CNNs is attributed to their ability to learn rich mid-level image representations as opposed to hand-designed low-level features used in other image classification methods. Learning CNNs, however, amounts to estimating millions of parameters and requires a very large number of annotated image samples. This property currently prevents application of CNNs to problems with limited training data. In this work we show how image representations learned with CNNs on large-scale annotated datasets can be efficiently transferred to other visual recognition tasks with limited amount of training data. We design a method to reuse layers trained on the ImageNet dataset to compute mid-level image representation for images in the PASCAL VOC dataset. We show that despite differences in image statistics and tasks in the two datasets, the transferred representation leads to significantly improved results for object and action classification, outperforming the current state of the art on Pascal VOC 2007 and 2012 datasets. We also show promising results for object and action localization. Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic |
CVPR | 1 |