EDBT 2026 Demo / reviewers in the wild / expert
Mert Bülent Sariyildiz
dblp:247/9362
· DBLP profile ↗
11ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Efficient and distributed learning · 18% Generative modeling · 16% Representation and self-supervised learning · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
multi-teacher distillation |
1.6 | 2 | 2025 | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers · CVPR 2025 UNIC: Universal Classification Models via Multi-teacher Distillation · ECCV (4) 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.1 | 2 | 2025 | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers · CVPR 2025 UNIC: Universal Classification Models via Multi-teacher Distillation · ECCV (4) 2024 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.9 | 2 | 2021 | Concept Generalization in Visual Representation Learning · ICCV 2021 Learning Visual Representations with Caption Annotations · ECCV (8) 2020 |
Robotics › Autonomous driving › perception
3d perception |
0.9 | 1 | 2025 | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning
deep encoder training |
0.9 | 1 | 2025 | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers · CVPR 2025 |
Robotics › Robot navigation and mapping
localization |
0.9 | 1 | 2025 | Kinaema: a recurrent sequence model for memory and pose in motion · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent memory |
0.9 | 1 | 2025 | Kinaema: a recurrent sequence model for memory and pose in motion · NeurIPS 2025 |
Computer vision › Image recognition and object detection
image classification |
0.8 | 2 | 2023 | Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones · CVPR 2023 Concept Generalization in Visual Representation Learning · ICCV 2021 |
Computer vision › 3D vision
visual localization |
0.8 | 1 | 2024 | Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones · CVPR 2023 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | No Reason for No Supervision: Improved Generalization in Supervised Models · ICLR 2023 |
Machine learning › Learning paradigms
supervised learning |
0.7 | 1 | 2023 | No Reason for No Supervision: Improved Generalization in Supervised Models · ICLR 2023 |
Machine learning › Generative modeling
synthetic training data |
0.7 | 1 | 2023 | Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones · CVPR 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.7 | 1 | 2023 | Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones · CVPR 2023 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.5 | 1 | 2021 | Concept Generalization in Visual Representation Learning · ICCV 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.4 | 1 | 2020 | Hard Negative Mixing for Contrastive Learning · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
data augmentation |
0.4 | 1 | 2020 | Hard Negative Mixing for Contrastive Learning · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › data augmentation
feature mixing |
0.4 | 1 | 2020 | Hard Negative Mixing for Contrastive Learning · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › contrastive learning › negative sampling
hard negative sampling |
0.4 | 1 | 2020 | Hard Negative Mixing for Contrastive Learning · NeurIPS 2020 |
Machine learning › Generative modeling › generative adversarial network
conditional GAN |
0.4 | 1 | 2019 | Gradient Matching Generative Networks for Zero-Shot Learning · CVPR 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Gradient Matching Generative Networks for Zero-Shot Learning · CVPR 2019 |
Machine learning › Efficient and distributed learning
gradient matching |
0.4 | 1 | 2019 | Gradient Matching Generative Networks for Zero-Shot Learning · CVPR 2019 |
Machine learning › Deep learning architectures and training
loss function design |
0.4 | 1 | 2019 | Gradient Matching Generative Networks for Zero-Shot Learning · CVPR 2019 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.4 | 1 | 2019 | Gradient Matching Generative Networks for Zero-Shot Learning · CVPR 2019 |
Information retrieval
image retrieval |
0.2 | 1 | 2024 | Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
text-to-image generation · 1.5geometric consistency · 1.5transformer · 0.9teacher-specific encoding · 0.9recurrent neural network · 0.9data-sharing strategies · 0.9co-distillation · 0.9multi-teacher distillation · 0.8synthetic data generation · 0.7prompt engineering · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersabstractRecent multi-teacher distillation methods have unified the encoders of multiple foundation models into a single encoder, achieving competitive performance on core vision tasks like classification, segmentation, and depth estimation. This led us to ask: Could similar success be achieved when the pool of teachers also includes vision models specialized in diverse tasks across both 2D and 3D perceptionƒ In this paper, we define and investigate the problem of heterogeneous teacher distillation, or co-distillation—a challenging multi-teacher distillation scenario where teacher models vary significantly in both (a) their design objectives and (b) the data they were trained on. We explore data-sharing strategies and teacher-specific encoding, and introduce DUNE, a single encoder excelling in 2D vision, 3D understanding, and 3D human perception. Our model achieves performance comparable to that of its larger teachers, sometimes even outperforming them, on their respective tasks. Notably, DUNE surpasses MASt3R in Map-free Visual Relocalization with a much smaller encoder. Mert Bülent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas 0002, Pau de Jorge, Diane Larlus, Yannis Kalantidis |
CVPR | 1 |
| 2025 | Kinaema: a recurrent sequence model for memory and pose in motionabstractOne key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves or previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where information observed before an actual episode start is exploited to optimize efficiency. We introduce a new model, "Kinaema" and agent, capable of integrating a stream of visual observations while moving in a potentially large scene, and upon request, processing a query image and predicting the relative position of the shown space with respect to its current position. Our model does not explicitly store an observation history, therefore does not have hard constraints on context length. It maintains an implicit latent memory, which is updated by a transformer in a recurrent way, compressing the history of sensor readings into a compact representation. We evaluate the impact of this model in a new downstream task we call "Mem-Nav", targeting continuous robotics operations. We show that our large-capacity recurrent model maintains a useful representation of the scene, navigates to goals observed before the actual episode start, and is computationally efficient, in particular compared to classical transformers with attention over an observation history. Mert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci, Christian Wolf 0001 |
NeurIPS | 1 |
| 2024 | UNIC: Universal Classification Models via Multi-teacher Distillation
Mert Bülent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas 0002, Diane Larlus, Yannis Kalantidis |
ECCV (4) | 1 |
| 2024 | Weatherproofing Retrieval for Localization with Generative AI and Geometric ConsistencyabstractState-of-the-art visual localization approaches generally rely on a first image retrieval step whose role is crucial. Yet, retrieval often struggles when facing varying conditions, due to e.g. weather or time of day, with dramatic consequences on the visual localization accuracy. In this paper, we improve this retrieval step and tailor it to the final localization task. Among the several changes we advocate for, we propose to synthesize variants of the training set images, obtained from generative text-to-image models, in order to automatically expand the training set towards a number of nameable variations that particularly hurt visual localization. After expanding the training set, we propose a training approach that leverages the specificities and the underlying geometry of this mix of real and synthetic images. We experimentally show that those changes translate into large improvements for the most challenging visual localization datasets. Yannis Kalantidis, Mert Bülent Sariyildiz, Rafael S. Rezende, Philippe Weinzaepfel, Diane Larlus, Gabriela Csurka |
ICLR | 2 |
| 2023 | Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesabstractRecent image generation models such as Stable Diffusion have exhibited an impressive ability to generate fairly realistic images starting from a simple text prompt. Could such models render real images obsolete for training image prediction models? In this paper, we answer part of this provocative question by investigating the need for real images when training models for ImageNet classification. Provided only with the class names that have been used to build the dataset, we explore the ability of Stable Diffusion to generate synthetic clones of ImageNet and measure how useful these are for training classification models from scratch. We show that with minimal and class-agnostic prompt engineering, ImageNet clones are able to close a large part of the gap between models produced by synthetic images and models trained with real images, for the several standard classification benchmarks that we consider in this study. More importantly, we show that models trained on synthetic images exhibit strong generalization properties and perform on par with models trained on real data for transfer. Project page: https://europe.naverlabs.com/imagenet-sd Mert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis Kalantidis |
CVPR | 1 |
| 2023 | No Reason for No Supervision: Improved Generalization in Supervised Models
Mert Bülent Sariyildiz, Yannis Kalantidis, Karteek Alahari, Diane Larlus |
ICLR | 1 |
| 2021 | Concept Generalization in Visual Representation LearningabstractMeasuring concept generalization, i.e., the extent to which models trained on a set of (seen) visual concepts can be leveraged to recognize a new set of (unseen) concepts, is a popular way of evaluating visual representations, especially in a self-supervised learning framework. Nonetheless, the choice of unseen concepts for such an evaluation is usually made arbitrarily, and independently from the seen concepts used to train representations, thus ignoring any semantic relationships between the two. In this paper, we argue that the semantic relationships between seen and unseen concepts affect generalization performance and propose ImageNet-CoG,1a novel benchmark on the ImageNet-21K (IN-21K) dataset that enables measuring concept generalization in a principled way. Our benchmark leverages expert knowledge that comes from WordNet in order to define a sequence of unseen IN-21K concept sets that are semantically more and more distant from the ImageNet-1K (IN-1K) subset, a ubiquitous training set. This allows us to benchmark visual representations learned on IN-1K out-of-the box. We conduct a large-scale study encompassing 31 convolution and transformer-based models and show how different architectures, levels of supervision, regularization techniques and use of web data impact the concept generalization performance. Mert Bülent Sariyildiz, Yannis Kalantidis, Diane Larlus, Karteek Alahari |
ICCV | 1 |
| 2020 | Learning Visual Representations with Caption Annotations
Mert Bülent Sariyildiz, Julien Perez, Diane Larlus |
ECCV (8) | 1 |
| 2020 | Hard Negative Mixing for Contrastive LearningabstractContrastive learning has become a key component of self-supervised learning approaches for computer vision. By learning to embed two augmented versions of the same image close to each other and to push the embeddings of different images apart, one can train highly transferable visual representations. As revealed by recent studies, heavy data augmentation and large sets of negatives are both crucial in learning such representations. At the same time, data mixing strategies, either at the image or the feature level, improve both supervised and semi-supervised learning by synthesizing novel examples, forcing networks to learn more robust features. In this paper, we argue that an important aspect of contrastive learning, i.e. the effect of hard negatives, has so far been neglected. To get more meaningful negative samples, current top contrastive self-supervised learning approaches either substantially increase the batch sizes, or keep very large memory banks; increasing memory requirements, however, leads to diminishing returns in terms of performance. We therefore start by delving deeper into a top-performing framework and show evidence that harder negatives are needed to facilitate better and faster learning. Based on these observations, and motivated by the success of data mixing, we propose hard negative mixing strategies at the feature level, that can be computed on-the-fly with a minimal computational overhead. We exhaustively ablate our approach on linear classification, object detection, and instance segmentation and show that employing our hard negative mixing procedure improves the quality of visual representations learned by a state-of-the-art self-supervised learning method. Yannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel, Diane Larlus |
NeurIPS | 2 |
| 2020 | Key protected classification for collaborative learning
Mert Bülent Sariyildiz, Ramazan Gokberk Cinbis, Erman Ayday |
Pattern Recognit. | 1 |
| 2019 | Gradient Matching Generative Networks for Zero-Shot LearningabstractZero-shot learning (ZSL) is one of the most promising problems where substantial progress can potentially be achieved through unsupervised learning, due to distributional differences between supervised and zero-shot classes. For this reason, several works investigate the incorporation of discriminative domain adaptation techniques into ZSL, which, however, lead to modest improvements in ZSL accuracy. In contrast, we propose a generative model that can naturally learn from unsupervised examples, and synthesize training examples for unseen classes purely based on their class embeddings, and therefore, reduce the zero-shot learning problem into a supervised classification task. The proposed approach consists of two important components: (i) a conditional Generative Adversarial Network that learns to produce samples that mimic the characteristics of unsupervised data examples, and (ii) the Gradient Matching (GM) loss that measures the quality of the gradient signal obtained from the synthesized examples. Using our GM loss formulation, we enforce the generator to produce examples from which accurate classifiers can be trained. Experimental results on several ZSL benchmark datasets show that our approach leads to significant improvements over the state of the art in generalized zero-shot classification. Mert Bülent Sariyildiz, Ramazan Gokberk Cinbis |
CVPR | 1 |