VLDB 2026 Research / reviewers in the wild / expert
Bertrand Delezoide
dblp:78/5050
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0002-1821-6697ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantifying User Coherence: A Unified Framework for Analyzing Recommender Systems Across DomainsabstractThe performance of Recommender Systems (RS) varies significantly across users, yet the underlying reasons for this variance remain poorly understood. This paper introduces a unified framework to analyze and explain this performance gap by quantifying user profile characteristics. We propose two novel, information-theoretic measures: Mean Surprise (𝑆(𝑢)), which captures a user's deviation from popular items and is closely related to popularity bias, and Mean Conditional Surprise (𝐂𝑆(𝑢)), which measures the internal coherence of a user's interactions in a domain-agnostic manner. Through extensive experiments on 7 algorithms and 9 datasets, we demonstrate that these measures are strong predictors of recommendation performance. Our analysis reveals that performance gains from complex models are concentrated on ''coherent'' users, while all algorithms perform poorly on ''incoherent'' users. We show how these measures provide practical utility for the Web community by: (1) enabling robust, stratified evaluation to identify model weaknesses; (2) facilitating a novel analysis of the behavioral alignment of recommendations; and (3) guiding targeted system design, which we validate by training a specialized model on a segment of ''coherent'' users that achieves superior performance for that group with significantly less data. This work provides a new lens for understanding user behavior and offers practical tools for building more robust and efficient large-scale recommender systems. Michaël Soumm, Alexandre Fournier-Montgieux, Adrian Popescu 0001, Bertrand Delezoide |
WWW | 4 |
| 2025 | Temporal Dynamics in Visual Data: Analyzing the Impact of Time on Classification AccuracyabstractVisual datasets are generally constructed from the samples available at the time of their collection and are not further updated. However, these static datasets do not reflect the distribution changes that occur in real data. We analyze how different collection times lead to a shift in class distribution by collecting a set of Flickr images published over 14 years. The proposed “Visual Classes through Time” (VCT-107) dataset contains images tagged by their publication date and includes 107 classes covering various topics (human-made objects, animals, plants, food, etc.). Images from each class are divided into five collection periods to study the impact of time on classification accuracy. When training different classification models using linear probing, we observe an accuracy loss when training on data from one period and testing on other periods. This happens even in the case of a strongly pre-trained model like DinoV2 ViT-B/14. Intuitively, the performance loss is generally more significant when the collection periods between the training and test data are further apart. Our analysis reveals that the temporal shift varies between classes, with the largest shifts observed for human-made objects and the smallest for natural concepts such as animal species. Our results stress the importance of regularly updating models to adapt to timeinduced changes in the distribution of visual classes, even when using a strongly pre-trained model. We release the VCT-107 dataset to facilitate research on temporal shifts. Tom Pégeot, Eva Feillet, Adrian Popescu 0001, Inna Kucher, Bertrand Delezoide |
WACV | 5 |
| 2024 | An Analysis of Initial Training Strategies for Exemplar-Free Class-Incremental LearningabstractClass-Incremental Learning (CIL) aims to build classification models from data streams. At each step of the CIL process, new classes must be integrated into the model. Due to catastrophic forgetting, CIL is particularly challenging when examples from past classes cannot be stored, the case on which we focus here. To date, most approaches are based exclusively on the target dataset of the CIL process. However, the use of models pre-trained in a self-supervised way on large amounts of data has recently gained momentum. The initial model of the CIL process may only use the first batch of the target dataset, or also use pre-trained weights obtained on an auxiliary dataset. The choice between these two initial learning strategies can significantly influence the performance of the incremental learning model, but has not yet been studied in depth. Performance is also influenced by the choice of the CIL algorithm, the neural architecture, the nature of the target task, the distribution of classes in the stream and the number of examples available for learning. We conduct a comprehensive experimental study to assess the roles of these factors. We present a statistical analysis framework that quantifies the relative contribution of each factor to incremental performance. Our main finding is that the initial training strategy is the dominant factor influencing the average incremental accuracy, but that the choice of CIL algorithm is more important in preventing forgetting. Based on this analysis, we propose practical recommendations for choosing the right initial training strategy for a given incremental learning use case. These recommendations are intended to facilitate the practical deployment of incremental learning. Grégoire Petit, Michaël Soumm, Eva Feillet, Adrian Popescu 0001, Bertrand Delezoide, David Picard, Céline Hudelot |
WACV | 5 |
| 2023 | FeTrIL: Feature Translation for Exemplar-Free Class-Incremental LearningabstractExemplar-free class-incremental learning is very challenging due to the negative effect of catastrophic forgetting. A balance between stability and plasticity of the incremental process is needed in order to obtain good accuracy for past as well as new classes. Existing exemplar-free class-incremental methods focus either on successive fine tuning of the model, thus favoring plasticity, or on using a feature extractor fixed after the initial incremental state, thus favoring stability. We introduce a method which combines a fixed feature extractor and a pseudo-features generator to improve the stability-plasticity balance. The generator uses a simple yet effective geometric translation of new class features to create representations of past classes, made of pseudo-features. The translation of features only requires the storage of the centroid representations of past classes to produce their pseudo-features. Actual features of new classes and pseudo-features of past classes are fed into a linear classifier which is trained incrementally to discriminate between all classes. The incremental process is much faster with the proposed method compared to mainstream ones which update the entire deep model. Experiments are performed with three challenging datasets, and different incremental settings. A comparison with ten existing methods shows that our method outperforms the others in most cases. FeTrIL code is available at https://github.com/GregoirePetit/FeTrIL. Grégoire Petit, Adrian Popescu 0001, Hugo Schindler, David Picard, Bertrand Delezoide |
WACV | 5 |
| 2023 | Vis2Rec: A Large-Scale Visual Dataset for Visit RecommendationabstractMost recommendation datasets for tourism are restricted to one world region and rely on explicit data such as checkins. However, in reality, tourists visit various places world-wide and document their trips primarily through photos. These images contain a wealth of raw information that can be used to capture users’ preferences and recommend personalized content. Visual content was already used in past works, but no large-scale publicly-available dataset that gives access to users’ personal images exists for recommender systems. As such a resource would open-up possibilities for new image-based recommendation algorithms, we introduce Vis2Rec, a new dataset based on visit data extracted from users’ Flickr photographic streams, which includes over 7 million photos, 36k recognizable points of interest, and 14k user profiles. Google Landmarks v2 is used as an auxiliary dataset to identify points of interest in users’ photos, using a state-of-the-art image-matching deep architecture. Image-based user profiles are then constituted by aggregating the points of interest detected for each user. In addition, ground truth visits were determined for the test subset in order to enable accurate evaluation. Finally, we benchmark Vis2Rec using various existing recommender systems, and discuss the possibilities opened up by the availability of user images, as well as the societal issues that come with them. Following good practice in dataset sharing, Vis2Rec is created using only freely distributable content, and additional anonymization is performed to ensure the privacy of users. The raw dataset and the preprocessed user profiles will be publicly available at https://github.com/MSoumm/Vis2Rec. Michaël Soumm, Adrian Popescu 0001, Bertrand Delezoide |
WACV | 3 |
| 2021 | Modelling relations with prototypes for visual relation detection
François Plesse, Alexandru-Lucian Gînsca, Bertrand Delezoide, Françoise J. Prêteux |
Multim. Tools Appl. | 3 |
| 2020 | Focusing Visual Relation Detection on Relevant Relations with Prior PotentialsabstractUnderstanding images relies on the understanding of how visible objects are linked to each other. Current approaches of Visual Relation Detection (VRD) are hindered by the high frequency of some relations: when an important focus is put on them, more meaningful ones are overlooked. We address this challenge by learning the relative relevance of relations, and integrating this term into a novel scene graph extraction scheme. We show that this allows our model to predict relations on fewer and more relevant object pairs. It outperforms MotifNet, a state of the art model, on the Visual Genome dataset. It increases the Class Macro recall, the metric we propose to use, from 38.1% to 44.4%. In addition, we propose a new split of Visual Genome, with a more balanced relation distribution, emphasizing on the detection of uncommon relations and validates the use of the previous metric. On this set, our model outperforms MotifNet on all metrics, e.g. from 39.6% to 44.0% at 10 predictions per image on the relation classification task. François Plesse, Alexandru-Lucian Gînsca, Bertrand Delezoide, Françoise J. Prêteux |
WACV | 3 |
| 2018 | Learning Prototypes for Visual Relationship DetectionabstractRelationships between objects drive our understanding of images, but modelling them poses several challenges due to the combinatorial nature of the problem and the complex structure of natural language. This paper tackles the task of predicting relationships in the form of (subject, predicate, object) triplets from still images. To address these issues, we propose a framework for learning predicate prototypes that aims to capture the multimodal nature of predicate distributions. Concurrently, a network is trained to define a space in which relationships with similar spatial layouts, interacting objects and predicates are clustered together. Finally, at test time, relationships are classified based on the similarity to their train nearest neighbor. We evaluate this approach on two well known datasets and observe comparable performance with state of the art models. Furthermore, we find that coupling prototype learning with a nearest neighbors approach increases the performance from 85.4 % to 87.6 % over a standard classification approach. François Plesse, Alexandru-Lucian Gînsca, Bertrand Delezoide, Françoise J. Prêteux |
CBMI | 3 |
| 2018 | Visual Relationship Detection Based on Guided Proposals and Semantic Knowledge DistillationabstractA thorough comprehension of image content demands a complex grasp of the interactions that may occur in the natural world. One of the key issues is to describe the visual relationships between objects. When dealing with real world data, capturing these very diverse interactions is a difficult problem. It can be alleviated by incorporating common sense in a network. For this, we propose a framework that makes use of semantic knowledge and estimates the relevance of object pairs during both training and test phases. Extracted from precomputed models and training annotations, this information is distilled into the neural network dedicated to this task. Using this approach, we observe a significant improvement on all classes of Visual Genome, a challenging visual relationship dataset. A 68.5 % relative gain on the recall at 100 is directly related to the relevance estimate and a 32.7% gain to the knowledge distillation. François Plesse, Alexandru-Lucian Gînsca, Bertrand Delezoide, Françoise J. Prêteux |
ICME | 3 |
| 2013 | Space-Time Robust Representation for Action RecognitionabstractWe address the problem of action recognition in unconstrained videos. We propose a novel content driven pooling that leverages space-time context while being robust toward global space-time transformations. Being robust to such transformations is of primary importance in unconstrained videos where the action localizations can drastically shift between frames. Our pooling identifies regions of interest using video structural cues estimated by differ ent saliency functions. To combine the different structural information, we introduce an iterative structure learning algorithm, WSVM (weighted SVM), that determines the optimal saliency layout of an action model through a sparse regularizer. A new optimization method is proposed to solve the WSVM' highly non-smooth objective function. We evaluate our approach on standard action datasets (KTH, UCF50 and HMDB). Most noticeably, the accuracy of our algorithm reaches 51.8% on the challenging HMDB dataset which outperforms the state-of-the-art of 7.3% relatively. Nicolas Ballas, Yi Yang 0001, Zhen-Zhong Lan, Bertrand Delezoide, Françoise J. Prêteux, Alex Hauptmann 0001 |
ICCV | 4 |
| 2012 | Trajectory signature for action recognition in videoabstractBag-of-Words representation based on trajectory local features and taking into account the spatio-temporal context through static segmentation grids is currently the leading paradigm to perform action annotation.While providing a coarse localization of low-level features, those approaches tend to be limited by the grid rigidity. In this work we propose two contributions on trajectory based signatures. First, we extend a local trajectory feature to characterize the acceleration in videos, leading to invariance to camera constant motion. We also introduce two new adaptive segmentation grids, namely Adaptive Grid (AG) and Deformable Adaptive Grid (DAG). AG is learnt from videos data, to fit a given dataset and overcome static grid rigidity. DAG is also learnt from video data. Moreover, it can be adapted to a specific video through a deformation operation. Our adaptive grids are then exploited by a Bag-of-Words model at the aggregation step for action recognition. Our proposal is evaluated on 4 publicly available datasets. Nicolas Ballas, Bertrand Delezoide, Françoise J. Prêteux |
ACM Multimedia | 2 |