EDBT 2026 Demo / reviewers in the wild / expert
Tim Lebailly
dblp:276/0970
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-9814-7531ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Representation and self-supervised learning · 42% Segmentation and scene understanding · 14% Efficient and distributed learning · 11% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
1.5 | 2 | 2025 | Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025 Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
dense representation learning |
0.9 | 2 | 2024 | CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023 CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
feature decorrelation |
0.9 | 1 | 2025 | Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025 |
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation |
0.9 | 1 | 2025 | Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
non-contrastive self-supervised learning |
0.9 | 1 | 2025 | Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.9 | 1 | 2025 | Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
0.9 | 1 | 2025 | A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention |
0.9 | 1 | 2025 | Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025 |
Computer vision › Segmentation and scene understanding › open-world segmentation
zero-shot segmentation |
0.9 | 1 | 2025 | A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025 |
Natural language and speech › Information extraction and text analysis
bootstrapping |
0.8 | 1 | 2024 | CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024 |
Computer vision › Image recognition and object detection
nearest neighbor retrieval |
0.8 | 1 | 2024 | CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024 |
Computer vision › 3D vision › multi-view geometry
multi-view consistency |
0.7 | 1 | 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023 |
Machine learning › Time series and sequential data
online clustering |
0.7 | 1 | 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.7 | 1 | 2023 | Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning · ICCV 2023 |
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation |
0.2 | 1 | 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023 |
Computer vision › Video understanding and tracking
video object segmentation |
0.2 | 1 | 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
self-distillation · 1.5contrastive learning · 1.5text encoder alignment · 0.9slot attention · 0.9hyperdimensional computing · 0.9frozen vision encoder · 0.9exponential moving average · 0.9contrastive learning theory · 0.9nearest neighbor retrieval · 0.8cross-image bootstrapping · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Object-Centric Pretraining via Target Encoder BootstrappingabstractObject-centric representation learning has recently been successfully applied to real-world datasets. This success can be attributed to pretrained non-object-centric foundation models, whose features serve as reconstruction targets for slot attention. However, targets must remain frozen throughout the training, which sets an upper bound on the performance object-centric models can attain. Attempts to update the target encoder by bootstrapping result in large performance drops, which can be attributed to its lack of object-centric inductive biases, causing the object-centric model’s encoder to drift away from representations useful as reconstruction targets. To address these limitations, we propose **O**bject-**CE**ntric Pretraining by Target Encoder **BO**otstrapping, a self-distillation setup for training object-centric models from scratch, on real-world data, for the first time ever. In OCEBO, the target encoder is updated as an exponential moving average of the object-centric model, thus explicitly being enriched with object-centric inductive biases introduced by slot attention while removing the upper bound on performance present in other models. We mitigate the slot collapse caused by random initialization of the target encoder by introducing a novel cross-view patch filtering approach that limits the supervision to sufficiently informative patches. When pretrained on 241k images from COCO, OCEBO achieves unsupervised object discovery performance comparable to that of object-centric models with frozen non-object-centric target encoders pretrained on hundreds of millions of images. The code and pretrained models are publicly available at https://github.com/djukicn/ocebo. Nikola Dukic, Tim Lebailly, Tinne Tuytelaars |
ICLR | 2 |
| 2025 | A Simple Framework for Open-Vocabulary Zero-Shot SegmentationabstractZero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocabulary segmentation. This deficiency is often attributed to the absence of localization cues in captions and the intertwined nature of the learning process, which encompasses both image/text representation learning and cross-modality alignment. To tackle these issues, we propose SimZSS, a $\textbf{Sim}$ple framework for open-vocabulary $\textbf{Z}$ero-$\textbf{S}$hot $\textbf{S}$egmentation. The method is founded on two key principles: i) leveraging frozen vision-only models that exhibit spatial awareness while exclusively aligning the text encoder and ii) exploiting the discrete nature of text and linguistic knowledge to pinpoint local concepts within captions. By capitalizing on the quality of the visual representations, our method requires only image-caption pair datasets and adapts to both small curated and large-scale noisy datasets. When trained on COCO Captions across 8 GPUs, SimZSS achieves state-of-the-art results on 7 out of 8 benchmark datasets in less than 15 minutes. Our code and pretrained models are publicly available at https://github.com/tileb1/simzss. Thomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran |
ICLR | 2 |
| 2025 | Collapse-Proof Non-Contrastive Self-Supervised LearningabstractWe present a principled and simplified design of the projector and loss function for non-contrastive self-supervised learning based on hyperdimensional computing. We theoretically demonstrate that this design introduces an inductive bias that encourages representations to be simultaneously decorrelated and clustered, without explicitly enforcing these properties. This bias provably enhances generalization and suffices to avoid known training failure modes, such as representation, dimensional, cluster, and intracluster collapses. We validate our theoretical findings on image datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet-100. Our approach effectively combines the strengths of feature decorrelation and cluster-based self-supervised learning methods, overcoming training failure modes while achieving strong generalization in clustering and linear classification tasks. Emanuele Sansone, Tim Lebailly, Tinne Tuytelaars |
ICML | 2 |
| 2024 | CrIBo: Self-Supervised Learning via Cross-Image Object-Level BootstrappingabstractLeveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel $\textbf{Cr}$oss-$\textbf{I}$mage Object-Level $\textbf{Bo}$otstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo. Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars |
ICLR | 1 |
| 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation LearningabstractLearning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view consistency objective with an Online Clustering mechanism (CrOC) to discover and segment the semantics of the views. In the absence of hand-crafted priors, the resulting method is more generalizable and does not require a cumbersome pre-processing step. More importantly, the clustering algorithm conjointly operates on the features of both views, thereby elegantly bypassing the issue of content not represented in both views and the ambiguous matching of objects from one crop to the other. We demonstrate excellent performance on linear and unsupervised segmentation transfer tasks on various datasets and similarly for video object segmentation. Our code and pre-trained models are publicly available at https://github.com/stegmuel/CrOC. Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran |
CVPR | 2 |
| 2023 | Adaptive Similarity Bootstrapping for Self-Distillation based Representation LearningabstractMost self-supervised methods for representation learning leverage a cross-view consistency objective i.e. they maximize the representation similarity of a given image’s augmented views. Recent work NNCLR goes beyond the cross-view paradigm and uses positive pairs from different images obtained via nearest neighbor bootstrapping in a contrastive setting. We empirically show that as opposed to the contrastive learning setting which relies on negative samples, incorporating nearest neighbor bootstrapping in a self-distillation scheme can lead to a performance drop or even collapse. We scrutinize the reason for this unexpected behavior and provide a solution. We propose to adaptively bootstrap neighbors based on the estimated quality of the latent space. We report consistent improvements compared to the naive bootstrapping approach and the original baselines. Our approach leads to performance improvements for various self-distillation method/backbone combinations and standard downstream tasks. Our code is publicly available at https://github.com/tileb1/AdaSim. Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars |
ICCV | 1 |
| 2023 | Global-Local Self-Distillation for Visual Representation LearningabstractThe downstream accuracy of self-supervised methods is tightly linked to the proxy task solved during training and the quality of the gradients extracted from it. Richer and more meaningful gradients updates are key to allow self-supervised methods to learn better and in a more efficient manner. In a typical self-distillation framework, the representation of two augmented images are enforced to be coherent at the global level. Nonetheless, incorporating local cues in the proxy task can be beneficial and improve the model accuracy on downstream tasks. This leads to a dual objective in which, on the one hand, coherence between global-representations is enforced and on the other, coherence between local-representations is enforced. Unfortunately, an exact correspondence mapping between two sets of local-representations does not exist making the task of matching local-representations from one augmentation to another non-trivial. We propose to leverage the spatial information in the input images to obtain geometric matchings and compare this geometric approach against previous methods based on similarity matchings. Our study shows that not only 1) geometric matchings perform better than similarity based matchings in low-data regimes but also 2) that similarity based matchings are highly hurtful in low-data regimes compared to the vanilla baseline without local self-distillation. The code is available at https://github.com/tileb1/global-local-self-distillation. Tim Lebailly, Tinne Tuytelaars |
WACV | 1 |
| 2020 | Motion Prediction Using Temporal Inception Module
Tim Lebailly, Sena Kiciroglu, Mathieu Salzmann, Pascal Fua, Wei Wang 0108 |
ACCV (2) | 1 |