EDBT 2026 Demo / reviewers in the wild / expert
Romaric Audigier
dblp:02/2645
· DBLP profile ↗
24ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-4757-2052ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CaMiT: A Time-Aware Car Model Dataset for Classification and GenerationabstractAI systems must adapt to the evolving visual landscape, especially in domains where object appearance shifts over time. While prior work on time-aware vision models has primarily addressed commonsense-level categories, we introduce Car Models in Time (CaMiT). This fine-grained dataset captures the temporal evolution of this representative subset of technological artifacts. CaMiT includes 787K labeled samples of 190 car models (2007–2023) and 5.1M unlabeled samples (2005–2023), supporting supervised and self-supervised learning. We show that static pretraining on in-domain data achieves competitive performance with large-scale generalist models, offering a more resource-efficient solution. However, accuracy degrades when testing a year's models backward and forward in time. To address this, we evaluate CaMiT in a time-incremental classification setting, a realistic continual learning scenario with emerging, evolving, and disappearing classes. We investigate two mitigation strategies: time-incremental pretraining, which updates the backbone model, and time-incremental classifier learning, which updates the final classification layer, with positive results in both cases. Finally, we introduce time-aware image generation by consistently using temporal metadata during training. Results indicate improved realism compared to standard generation. CaMiT provides a rich resource for exploring temporal adaptation in a fine-grained visual context for discriminative and generative AI systems. Frédéric Lin, Biruk Abere Ambaw, Adrian Popescu 0001, Hejer Ammar, Romaric Audigier, Hervé Le Borgne |
NeurIPS | 5 |
| 2024 | Early Feature Distributions Alignment in Visible-to-Thermal Unsupervised Domain Adaptation for Object Detection
Adrien Maglo, Romaric Audigier |
ICPR (17) | 2 |
| 2023 | Can Human Attribute Segmentation be More Robust to Operational Contexts Without New Labels?abstractHuman Attribute Segmentation (HAS) describes, pixel-wise, the different semantic parts of people in an image. This fine-grained description is useful for several applications (e.g. security, fashion). However, despite the good performance reached by supervised Semantic Segmentation (SS) approaches, they are usually biased by the source training dataset and suffer from a performance drop when applied on new domains. Pixelwise image annotation for each new encountered context is tedious and expensive. So how can HAS become more robust to new contexts without new annotations? In this first study of Unsupervised Domain Adaptation (UDA) for HAS, we present UDA-HPTR, a new method based on HPTR [1] (Human Parsing with TRansformers) combined with self-supervised and semi-supervised learning paradigms to deal with UDA. UDA-HPTR improves performance on both source (labeled) and target (unlabeled) datasets compared to the fully supervised version (HPTR). It also outperforms HRDA, a state-of-the-art UDA method in autonomous driving benchmarks, by +6.7 p.p. on the source and +8.8 p.p. on the target, when applied to HAS while using only half the number of parameters. Hejer Ammar, Angélique Loesch, Corentin Vannier, Romaric Audigier |
ICIP | 4 |
| 2023 | Generalized Pseudo-Labeling in Consistency Regularization for Semi-Supervised LearningabstractSemi-Supervised Learning (SSL) reduces annotation cost by exploiting large amounts of unlabeled data. A popular idea in SSL image classification is Pseudo-Labeling (PL), where the predictions of a network are used in order to assign a label to an unlabeled image. However, this practice exposes learning to confirmation bias. In this paper we propose Generalized Pseudo-Labeling (GPL), a simple and generic way to exploit negative pseudo-labels in consistency regularization, entailing minimal additional computational overhead and hyperpameter fine-tuning. GPL makes learning more robust by using the information that an image does not belong to a certain class, which is more abundant and reliable. We showcase GPL in the context of FixMatch. In the benchmark using only 40 labels of the CIFAR-10 dataset, adding GPL on top of FixMatch improves the error rate from 7.93% to 6.58%, and on CIFAR-100 with 2500 labels, from 28.02% to 26.85%. Nikolaos Karaliolios, Florian Chabot, Camille Dupont, Hervé Le Borgne, Quoc Cuong Pham, Romaric Audigier |
ICIP | 6 |
| 2023 | Proposal-Contrastive Pretraining for Object Detection from Fewer Data
Quentin Bouniot, Romaric Audigier, Angélique Loesch, Amaury Habrard |
ICLR | 2 |
| 2023 | Towards Few-Annotation Learning for Object Detection: Are Transformer-based Models More Efficient?abstractFor specialized and dense downstream tasks such as object detection, labeling data requires expertise and can be very expensive, making few-shot and semi-supervised models much more attractive alternatives. While in the few-shot setup we observe that transformer-based object detectors perform better than convolution-based two-stage models for a similar amount of parameters, they are not as effective when used with recent approaches in the semi-supervised setting. In this paper, we propose a semi-supervised method tailored for the current state-of-the-art object detector Deformable DETR in the few-annotation learning setup using a student-teacher architecture, which avoids relying on a sensitive post-processing of the pseudo-labels generated by the teacher model. We evaluate our method on the semi-supervised object detection benchmarks COCO and Pascal VOC, and it outperforms previous methods, especially when annotations are scarce. We believe that our contributions open new possibilities to adapt similar object detection methods in this setup as well. Quentin Bouniot, Angélique Loesch, Amaury Habrard, Romaric Audigier |
WACV | 4 |
| 2022 | Spatio-temporal predictive tasks for abnormal event detection in videosabstractAbnormal event detection in videos is a challenging problem, partly due to the multiplicity of abnormal patterns and the lack of their corresponding annotations. In this paper, we propose new constrained pretext tasks to learn object level normality patterns. Our approach consists in learning a mapping between down-scaled visual queries and their corresponding normal appearance and motion characteristics at the original resolution. The proposed tasks are more challenging than reconstruction and future frame prediction tasks which are widely used in the literature, since our model learns to jointly predict spatial and temporal features rather than reconstructing them. We believe that more constrained pretext tasks induce a better learning of normality patterns. Experiments on several benchmark datasets demonstrate the effectiveness of our approach to localize and track anomalies as it outperforms or reaches the current state-of-the-art on spatio-temporal evaluation metrics. Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier |
AVSS | 5 |
| 2022 | Improving Few-Shot Learning Through Multi-task Representation Learning Theory
Quentin Bouniot, Ievgen Redko, Romaric Audigier, Angélique Loesch, Amaury Habrard |
ECCV (20) | 3 |
| 2022 | Object-Centric and Memory-Guided Normality Reconstruction for Video Anomaly DetectionabstractThis paper addresses video anomaly detection problem for videosurveillance. Due to the inherent rarity and heterogeneity of abnormal events, the problem is viewed as a normality modeling strategy, in which our model learns object-centric normal patterns without seeing anomalous samples during training. The main contributions consist in coupling pre-trained object-level action features prototypes with a cosine distance-based anomaly estimation function, therefore extending previous methods by introducing additional constraints to the mainstream reconstruction-based strategy. Our framework leverages both appearance and motion information to learn object-level behavior and captures prototypical patterns within a memory module. Experiments on several well-known datasets demonstrate the effectiveness of our method as it outperforms current state-of-the-art on most relevant spatio-temporal evaluation metrics. Khalil Bergaoui, Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier |
ICIP | 6 |
| 2022 | A formal approach to good practices in Pseudo-Labeling for Unsupervised Domain Adaptive Re-Identification
Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu |
Comput. Vis. Image Underst. | 2 |
| 2021 | Detecting Human-to-Human-or-Object (H2O) Interactions with DIABOLOabstractDetecting human interactions is crucial for human behavior analysis. Many methods have been proposed to deal with Human-to-Object Interaction (HOI) detection, i.e., detecting in an image which person and object interact together and classifying the type of interaction. However, Human-to-Human Interactions, such as social and violent interactions, are generally not considered in available HOI training datasets. As we think these types of interactions cannot be ignored and decorrelated from HOI when analyzing human behavior, we propose a new interaction dataset to deal with both types of human interactions: Human-to-Human-or-Object (H2O). In addition, we introduce a novel taxonomy of verbs, intended to be closer to a description of human body attitude in relation to the surrounding targets of interaction, and more independent of the environment. Unlike some existing datasets, we strive to avoid defining synonymous verbs when their use highly depends on the target type or requires a high level of semantic interpretation. As H2O dataset includes V-COCO images annotated with this new taxonomy, images obviously contain more interactions. This can be an issue for HOI detection methods whose complexity depends on the number of people, targets or interactions. Thus, we propose DIABOLO (Detecting Inter Actions By Only Looking Once), an efficient subject-centric single-shot method to detect all interactions in one forward pass, with constant inference time independent of image content. In addition, this multi-task network simultaneously detects all people and objects. We show how sharing a network for these tasks does not only save computation resource but also improves performance collaboratively. Finally, DIABOLO is a strong baseline for the new proposed challenge of H2O- Interaction detection, as it outperforms all state-of-the-art methods when trained and evaluated on HOI dataset V-COCO. We hope that this new dataset and new baseline will foster future research. H2O is available on https:/lkalisteo.cea.fr/. Astrid Orcesi, Romaric Audigier, Fritz Poka Toukam, Bertrand Luvison |
FG | 2 |
| 2021 | Describe Me If You Can! Characterized Instance-Level Human ParsingabstractSeveral computer vision applications such as person search or online fashion rely on human description. The use of instance-level human parsing (HP) is therefore relevant since it localizes semantic attributes and body parts within a person. But how to characterize these attributes? To our knowledge, only some single-HP datasets describe attributes with some color, size and/or pattern characteristics. There is a lack of dataset for multi-HP in the wild with such characteristics. In this article, we propose the dataset CCIHP based on the multi-HP dataset CIHP, with 20 new labels covering these 3 kinds of characteristics.1In addition, we propose HPTR, a new bottom-up multi-task method based on transformers as a fast and scalable baseline. It is the fastest method of multi-HP state of the art while having precision comparable to the most precise bottom-up method. We hope this will encourage research for fast and accurate methods of precise human descriptions.1CCIHP is available on https://kalisteo.cea.fr/index.php/free-resources/ Angélique Loesch, Romaric Audigier |
ICIP | 2 |
| 2020 | Optimal Transport as a Defense Against Adversarial AttacksabstractDeep learning classifiers are now known to have flaws in the representations of their class. Adversarial attacks can find a human-imperceptible perturbation for a given image that will mislead a trained model. The most effective methods to defend against such attacks trains on generated adversarial examples to learn their distribution. Previous work aimed to align original and adversarial image representations in the same way as domain adaptation to improve robustness. Yet, they partially align the representations using approaches that do not reflect the geometry of space and distribution. In addition, it is difficult to accurately compare robustness between defended models. Until now, they have been evaluated using a fixed perturbation size. However, defended models may react differently to variations of this perturbation size. In this paper, the analogy of domain adaptation is taken a step further by exploiting optimal transport theory. We propose to use a loss between distributions that faithfully reflect the ground distance. This leads to SAT (Sinkhorn Adversarial Training), a more robust defense against adversarial attacks. Then, we propose to quantify more precisely the robustness of a model to adversarial attacks over a wide range of perturbation sizes using a different metric, the Area Under the Accuracy Curve (AUAC). We perform extensive experiments on both CIFAR-10 and CIFAR-100 datasets and show that our defense is globally more robust than the state-of-the-art. Quentin Bouniot, Romaric Audigier, Angélique Loesch |
ICPR | 2 |
| 2020 | Unsupervised Domain Adaptation for Person Re-Identification through Source-Guided Pseudo-LabelingabstractPerson Re-Identification (re-ID) aims at retrieving images of the same person taken by different cameras. A challenge for re-ID is the performance preservation when a model is used on data of interest (target data) which belong to a different domain from the training data domain (source data). Unsupervised Domain Adaptation (UDA) is an interesting research direction for this challenge as it avoids a costly annotation of the target data. Pseudo-labeling methods achieve the best results in UDA-based re-ID. They incrementally learn with identity pseudo-labels which are initialized by clustering features in the source reID encoder space. Surprisingly, labeled source data are discarded after this initialization step. However, we believe that pseudo-labeling could further leverage the labeled source data in order to improve the post-initialization training steps. In order to improve robustness against erroneous pseudo-labels, we advocate the exploitation of both labeled source data and pseudo-labeled target data during all training iterations. To support our guideline, we introduce a framework which relies on a two-branch architecture optimizing classification and triplet loss based metric learning in source and target domains, respectively, in order to allow adaptability to the target domain while ensuring robustness to noisy pseudo-labels. Indeed, shared low and mid-level parameters benefit from the source classification and triplet loss signal while high-level parameters of the target branch learn domain-specific features. Our method is simple enough to be easily combined with existing pseudo-labeling UDA approaches. We show experimentally that it is efficient and improves performance when the base method has no mechanism to deal with pseudo-label noise. Our approach reaches state-of-the-art performance when evaluated on commonly used datasets, Market-1501 and DukeMTMC-reID, and outperforms the state of the art when targeting the bigger and more challenging dataset MSMT. Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu |
ICPR | 2 |
| 2020 | Classifying All Interacting Pairs in a Single ShotabstractIn this paper, we introduce a novel human interaction detection approach, based on CALIPSO (Classifying ALl Interacting Pairs in a Single shOt), a classifier of human-object interactions. This new single-shot interaction classifier estimates interactions simultaneously for all human-object pairs, regardless of their number and class. State-of- the-art approaches adopt a multi-shot strategy based on a pairwise estimate of interactions for a set of human-object candidate pairs, which leads to a complexity depending, at least, on the number of interactions or, at most, on the number of candidate pairs. In contrast, the proposed method estimates the interactions on the whole image. Indeed, it simultaneously estimates all interactions between all human subjects and object targets by performing a single forward pass throughout the image. Consequently, it leads to a constant complexity and computation time independent of the number of subjects, objects or interactions in the image. In detail, interaction classification is achieved on a dense grid of anchors thanks to a joint multi-task network that learns three complementary tasks simultaneously: (i) prediction of the types of interaction, (ii) estimation of the presence of a target and (iii) learning of an embedding which maps interacting subject and target to a same representation, by using a metric learning strategy. In addition, we introduce an object-centric passive-voice verb estimation which significantly improves results. Evaluations on the two well-known Human-Object Interaction image datasets, V- COCO and HICO-DET, demonstrate the competitiveness of the proposed method (2nd place) compared to the state-of- the-art while having constant computation time regardless of the number of objects and interactions in the image. Sanaa Chafik, Astrid Orcesi, Romaric Audigier, Bertrand Luvison |
WACV | 3 |
| 2019 | End-To-End Person Search Sequentially Trained On Aggregated DatasetabstractIn video surveillance applications, person search is a challenging task consisting in detecting people and extracting features from their silhouette for re-identification (re-ID) purpose. We propose a new end-to-end model that jointly computes detection and feature extraction steps through a single deep Convolutional Neural Network architecture. Sharing feature maps between the two tasks for jointly describing people commonalities and specificities allows faster runtime, which is valuable in real-world applications. In addition to reaching state-of-the-art accuracy, this multi-task model can be sequentially trained task-by-task, which results in a broader acceptance of input dataset types. Indeed, we show that aggregating more pedestrian detection datasets without costly identity annotations makes the shared feature maps more generic, and improves re-ID precision. Moreover, these boosted shared feature maps result in re-ID features more robust to a cross-dataset scenario. Angélique Loesch, Jaonary Rabarisoa, Romaric Audigier |
ICIP | 3 |
| 2016 | Improving Multi-frame Data Association with Sparse Representations for Robust Near-online Multi-object Tracking
Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle |
ECCV (8) | 2 |
| 2016 | RIMOC, a feature to discriminate unstructured motions: Application to violence detection for video-surveillance
Pedro Ribeiro 0006, Romaric Audigier, Quoc Cuong Pham |
Comput. Vis. Image Underst. | 2 |
| 2015 | Collaboration and spatialization for an efficient multi-person tracking via sparse representationsabstractMulti-person tracking is a very difficult problem in Computer Vision as a tracking algorithm is facing several issues, such as appearance changes, targets' occlusions and similar appearances between people. In an online tracking-by-detection algorithm, robust and discriminative specific appearance models help handling these difficulties. As done in single object tracking, we use sparse representations to extract local features of the targets and study how these representations can be specifically employed for multi-person tracking. Experiments on several datasets show that considering spatial information is crucial in order to improve the tracking performances with local descriptions compared to holistic features. Using large collaborative representations also improve the tracking results by naturally discarding irrelevant local patches. Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle |
AVSS | 2 |
| 2015 | Online multi-person tracking based on global sparse collaborative representationsabstractMulti-person tracking is still a challenging problem due to recurrent occlusion, pose variation and similar appearances between people. Inspired by the success of sparse representations in single object tracking and face recognition, we propose in this paper an online tracking by detection framework based on collaborative sparse representations. We argue that collaborative representations can better differentiate people compared to target-specific models and therefore help to produce a more robust tracking system. We also show that despite the size of the dictionaries involved, these representations can be efficiently computed with large-scale optimization techniques to get a near real-time algorithm. Experiments show that the proposed approach compares well to other recent online tracking systems on various datasets. Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle |
ICIP | 2 |
| 2015 | Collaborative Tracking and Distributed Control for an IP-PTZ Camera Network
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham |
ICPRAM (2) | 2 |
| 2013 | IMM-Based Tracking and Latency Control with Off-the-Shelf IP PTZ Camera
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham |
ACIVS | 2 |
| 2010 | Relationships between some watershed definitions and their tie-zone transforms
Romaric Audigier, Roberto A. Lotufo |
Image Vis. Comput. | 1 |
| 2005 | The tie-zone watershed: definition, algorithm and applicationsabstractIn this work, a new type of watershed transform is introduced: the tie-zone watershed (TZWS). This region-based watershed transform does not depend on arbitrary implementation and provides a unique and optimal solution. Indeed, many solutions are sometimes possible when segmenting an image with a watershed algorithm. In this case, the TZWS assigns each pixel to a catchment basin (CB) if in all solutions it belongs to this CB. Otherwise, the pixel is said to belong to a tie-zone (TZ). We propose an efficient algorithm based on image foresting transform (IFT) which computes the TZWS transform as a shortest-path forest. Finally, two applications of this TZWS are presented: bounding intervals for segmented objects' extensions and a progressive segmentation procedure. Romaric Audigier, Roberto A. Lotufo, Michel Couprie |
ICIP (2) | 1 |