EDBT 2026 Demo / reviewers in the wild / expert
Manuel Knott 0001
dblp:222/9651-1
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-5447-4549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Segmentation and scene understanding · 24% 3D vision · 24% Representation and self-supervised learning · 14% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.9 | 1 | 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025 |
Machine learning › Transfer learning and domain adaptation › transferability estimation
transfer evaluation |
0.9 | 1 | 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025 |
Computer vision › 3D vision
depth estimation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Segmentation and scene understanding › semantic segmentation
diffusion-based segmentation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Vision and language › cross-modal alignment
image-text alignment |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
linear probing · 0.9k-nearest neighbors · 0.9fine-tuning · 0.9model personalization · 0.8diffusion model · 0.8caption generation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Rapid Test for Accuracy and Bias of Face Recognition TechnologyabstractMeasuring the accuracy of face recognition (FR) systems is essential for improving performance and ensuring responsible use. Accuracy is typically estimated using large annotated datasets, which are costly and difficult to obtain. We propose a novel method for 1: 1 face verification that benchmarks FR systems quickly and without manual annotation, starting from approximate labels (e.g., from web search results). Unlike previous methods for training set label cleaning, ours leverages the embedding representation of the models being evaluated, achieving high accuracy in smaller-sized test datasets. Our approach reliably estimates FR accuracy and ranking, significantly reducing the time and cost of manual labeling. We also introduce the first public benchmark of five FR cloud services, revealing demographic biases, particularly lower accuracy for Asian women. Our rapid test method can democratize FR testing, promoting scrutiny and responsible use of the technology. Our method is provided as a publicly accessible tool at https://github.com/caltechvisionlab/frt-rapid-test. Manuel Knott 0001, Ignacio Serna, Ethan Mann, Pietro Perona |
WACV | 1 |
| 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image ClassificationabstractSelf-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w. Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona |
Int. J. Comput. Vis. | 2 |
| 2024 | Text-Image Alignment for Diffusion-Based PerceptionabstractDiffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusionbased semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our crossdomain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP Neehar Kondapaneni, Markus Marks, Manuel Knott 0001, Rogério Guimarães, Pietro Perona |
CVPR | 3 |