Timothée Darcet

dblp:344/5814 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Representation and self-supervised learning · 50% Vision and language · 27% Deep learning architectures and training · 23%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
cross-modal alignment
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning
0.812024
Vision Transformers Need Registers · ICLR 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
Vision Transformers Need Registers · ICLR 2024

Methods — techniques the papers use, named apart from their topics

text encoder alignment · 0.9lit training · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2025 DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
abstract
Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP [67], self-supervised visual features are not readily aligned with language, hindering their adoption in open-vocabulary tasks. Our method, named dino.txt, unlocks this new ability for DINOv2 [63], a widely used self-supervised visual encoder. We build upon the LiT training strategy [97], which trains a text encoder to align with a frozen vision model but leads to unsatisfactory results on dense tasks. We propose several key ingredients to improve performance on both global and dense tasks, such as concatenating the [CLS] token with the patch average to train the alignment and curating data using both text and image modalities. With these, we successfully train a CLIP-like model with only a fraction of the computational cost compared to CLIP while achieving state-of-the-art results in zero-shot classification and open-vocabulary semantic segmentation.
Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre, Timothée Darcet, Hu Xu 0001, Daniel Li 0006, Marc Szafraniec, Michaël Ramamonjisoa, Maxime Oquab, Oriane Siméoni, Huy V. Vo, Patrick Labatut, Piotr Bojanowski
CVPR5
2024 Vision Transformers Need Registers
abstract
Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.
Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski
ICLR1