Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tim Lebailly

dblp:276/0970 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-9814-7531ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 42% Segmentation and scene understanding · 14% Efficient and distributed learning · 11%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
1.522025
Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025
Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
dense representation learning
0.922024
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023
CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning
feature decorrelation
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
non-contrastive self-supervised learning
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.912025
Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.912025
A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention
0.912025
Object-Centric Pretraining via Target Encoder Bootstrapping · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025
Computer vision › Segmentation and scene understanding › open-world segmentation
zero-shot segmentation
0.912025
A Simple Framework for Open-Vocabulary Zero-Shot Segmentation · ICLR 2025
Natural language and speech › Information extraction and text analysis
bootstrapping
0.812024
CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024
Computer vision › Image recognition and object detection
nearest neighbor retrieval
0.812024
CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping · ICLR 2024
Computer vision › 3D vision › multi-view geometry
multi-view consistency
0.712023
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023
Machine learning › Time series and sequential data
online clustering
0.712023
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning · ICCV 2023
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation
0.212023
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023
Computer vision › Video understanding and tracking
video object segmentation
0.212023
CrOC: Cross-View Online Clustering for Dense Visual Representation Learning · CVPR 2023

Methods — techniques the papers use, named apart from their topics

self-distillation · 1.5contrastive learning · 1.5text encoder alignment · 0.9slot attention · 0.9hyperdimensional computing · 0.9frozen vision encoder · 0.9exponential moving average · 0.9contrastive learning theory · 0.9nearest neighbor retrieval · 0.8cross-image bootstrapping · 0.8
YearPublicationVenuePosition
2025 Object-Centric Pretraining via Target Encoder Bootstrapping
abstract
Object-centric representation learning has recently been successfully applied to real-world datasets. This success can be attributed to pretrained non-object-centric foundation models, whose features serve as reconstruction targets for slot attention. However, targets must remain frozen throughout the training, which sets an upper bound on the performance object-centric models can attain. Attempts to update the target encoder by bootstrapping result in large performance drops, which can be attributed to its lack of object-centric inductive biases, causing the object-centric model’s encoder to drift away from representations useful as reconstruction targets. To address these limitations, we propose **O**bject-**CE**ntric Pretraining by Target Encoder **BO**otstrapping, a self-distillation setup for training object-centric models from scratch, on real-world data, for the first time ever. In OCEBO, the target encoder is updated as an exponential moving average of the object-centric model, thus explicitly being enriched with object-centric inductive biases introduced by slot attention while removing the upper bound on performance present in other models. We mitigate the slot collapse caused by random initialization of the target encoder by introducing a novel cross-view patch filtering approach that limits the supervision to sufficiently informative patches. When pretrained on 241k images from COCO, OCEBO achieves unsupervised object discovery performance comparable to that of object-centric models with frozen non-object-centric target encoders pretrained on hundreds of millions of images. The code and pretrained models are publicly available at https://github.com/djukicn/ocebo.
Nikola Dukic, Tim Lebailly, Tinne Tuytelaars
ICLR2
2025 A Simple Framework for Open-Vocabulary Zero-Shot Segmentation
abstract
Zero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocabulary segmentation. This deficiency is often attributed to the absence of localization cues in captions and the intertwined nature of the learning process, which encompasses both image/text representation learning and cross-modality alignment. To tackle these issues, we propose SimZSS, a $\textbf{Sim}$ple framework for open-vocabulary $\textbf{Z}$ero-$\textbf{S}$hot $\textbf{S}$egmentation. The method is founded on two key principles: i) leveraging frozen vision-only models that exhibit spatial awareness while exclusively aligning the text encoder and ii) exploiting the discrete nature of text and linguistic knowledge to pinpoint local concepts within captions. By capitalizing on the quality of the visual representations, our method requires only image-caption pair datasets and adapts to both small curated and large-scale noisy datasets. When trained on COCO Captions across 8 GPUs, SimZSS achieves state-of-the-art results on 7 out of 8 benchmark datasets in less than 15 minutes. Our code and pretrained models are publicly available at https://github.com/tileb1/simzss.
Thomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran
ICLR2
2025 Collapse-Proof Non-Contrastive Self-Supervised Learning
abstract
We present a principled and simplified design of the projector and loss function for non-contrastive self-supervised learning based on hyperdimensional computing. We theoretically demonstrate that this design introduces an inductive bias that encourages representations to be simultaneously decorrelated and clustered, without explicitly enforcing these properties. This bias provably enhances generalization and suffices to avoid known training failure modes, such as representation, dimensional, cluster, and intracluster collapses. We validate our theoretical findings on image datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet-100. Our approach effectively combines the strengths of feature decorrelation and cluster-based self-supervised learning methods, overcoming training failure modes while achieving strong generalization in clustering and linear classification tasks.
Emanuele Sansone, Tim Lebailly, Tinne Tuytelaars
ICML2
2024 CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping
abstract
Leveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel $\textbf{Cr}$oss-$\textbf{I}$mage Object-Level $\textbf{Bo}$otstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo.
Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars
ICLR1
2023 CrOC: Cross-View Online Clustering for Dense Visual Representation Learning
abstract
Learning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view consistency objective with an Online Clustering mechanism (CrOC) to discover and segment the semantics of the views. In the absence of hand-crafted priors, the resulting method is more generalizable and does not require a cumbersome pre-processing step. More importantly, the clustering algorithm conjointly operates on the features of both views, thereby elegantly bypassing the issue of content not represented in both views and the ambiguous matching of objects from one crop to the other. We demonstrate excellent performance on linear and unsupervised segmentation transfer tasks on various datasets and similarly for video object segmentation. Our code and pre-trained models are publicly available at https://github.com/stegmuel/CrOC.
Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran
CVPR2
2023 Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning
abstract
Most self-supervised methods for representation learning leverage a cross-view consistency objective i.e. they maximize the representation similarity of a given image’s augmented views. Recent work NNCLR goes beyond the cross-view paradigm and uses positive pairs from different images obtained via nearest neighbor bootstrapping in a contrastive setting. We empirically show that as opposed to the contrastive learning setting which relies on negative samples, incorporating nearest neighbor bootstrapping in a self-distillation scheme can lead to a performance drop or even collapse. We scrutinize the reason for this unexpected behavior and provide a solution. We propose to adaptively bootstrap neighbors based on the estimated quality of the latent space. We report consistent improvements compared to the naive bootstrapping approach and the original baselines. Our approach leads to performance improvements for various self-distillation method/backbone combinations and standard downstream tasks. Our code is publicly available at https://github.com/tileb1/AdaSim.
Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars
ICCV1
2023 Global-Local Self-Distillation for Visual Representation Learning
abstract
The downstream accuracy of self-supervised methods is tightly linked to the proxy task solved during training and the quality of the gradients extracted from it. Richer and more meaningful gradients updates are key to allow self-supervised methods to learn better and in a more efficient manner. In a typical self-distillation framework, the representation of two augmented images are enforced to be coherent at the global level. Nonetheless, incorporating local cues in the proxy task can be beneficial and improve the model accuracy on downstream tasks. This leads to a dual objective in which, on the one hand, coherence between global-representations is enforced and on the other, coherence between local-representations is enforced. Unfortunately, an exact correspondence mapping between two sets of local-representations does not exist making the task of matching local-representations from one augmentation to another non-trivial. We propose to leverage the spatial information in the input images to obtain geometric matchings and compare this geometric approach against previous methods based on similarity matchings. Our study shows that not only 1) geometric matchings perform better than similarity based matchings in low-data regimes but also 2) that similarity based matchings are highly hurtful in low-data regimes compared to the vanilla baseline without local self-distillation. The code is available at https://github.com/tileb1/global-local-self-distillation.
Tim Lebailly, Tinne Tuytelaars
WACV1
2020 Motion Prediction Using Temporal Inception Module
Tim Lebailly, Sena Kiciroglu, Mathieu Salzmann, Pascal Fua, Wei Wang 0108
ACCV (2)1