Romaric Audigier

dblp:02/2645 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-4757-2052ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
abstract
AI systems must adapt to the evolving visual landscape, especially in domains where object appearance shifts over time. While prior work on time-aware vision models has primarily addressed commonsense-level categories, we introduce Car Models in Time (CaMiT). This fine-grained dataset captures the temporal evolution of this representative subset of technological artifacts. CaMiT includes 787K labeled samples of 190 car models (2007–2023) and 5.1M unlabeled samples (2005–2023), supporting supervised and self-supervised learning. We show that static pretraining on in-domain data achieves competitive performance with large-scale generalist models, offering a more resource-efficient solution. However, accuracy degrades when testing a year's models backward and forward in time. To address this, we evaluate CaMiT in a time-incremental classification setting, a realistic continual learning scenario with emerging, evolving, and disappearing classes. We investigate two mitigation strategies: time-incremental pretraining, which updates the backbone model, and time-incremental classifier learning, which updates the final classification layer, with positive results in both cases. Finally, we introduce time-aware image generation by consistently using temporal metadata during training. Results indicate improved realism compared to standard generation. CaMiT provides a rich resource for exploring temporal adaptation in a fine-grained visual context for discriminative and generative AI systems.
Frédéric Lin, Biruk Abere Ambaw, Adrian Popescu 0001, Hejer Ammar, Romaric Audigier, Hervé Le Borgne
NeurIPS5
2024 Early Feature Distributions Alignment in Visible-to-Thermal Unsupervised Domain Adaptation for Object Detection
Adrien Maglo, Romaric Audigier
ICPR (17)2
2023 Can Human Attribute Segmentation be More Robust to Operational Contexts Without New Labels?
abstract
Human Attribute Segmentation (HAS) describes, pixel-wise, the different semantic parts of people in an image. This fine-grained description is useful for several applications (e.g. security, fashion). However, despite the good performance reached by supervised Semantic Segmentation (SS) approaches, they are usually biased by the source training dataset and suffer from a performance drop when applied on new domains. Pixelwise image annotation for each new encountered context is tedious and expensive. So how can HAS become more robust to new contexts without new annotations? In this first study of Unsupervised Domain Adaptation (UDA) for HAS, we present UDA-HPTR, a new method based on HPTR [1] (Human Parsing with TRansformers) combined with self-supervised and semi-supervised learning paradigms to deal with UDA. UDA-HPTR improves performance on both source (labeled) and target (unlabeled) datasets compared to the fully supervised version (HPTR). It also outperforms HRDA, a state-of-the-art UDA method in autonomous driving benchmarks, by +6.7 p.p. on the source and +8.8 p.p. on the target, when applied to HAS while using only half the number of parameters.
Hejer Ammar, Angélique Loesch, Corentin Vannier, Romaric Audigier
ICIP4
2023 Generalized Pseudo-Labeling in Consistency Regularization for Semi-Supervised Learning
abstract
Semi-Supervised Learning (SSL) reduces annotation cost by exploiting large amounts of unlabeled data. A popular idea in SSL image classification is Pseudo-Labeling (PL), where the predictions of a network are used in order to assign a label to an unlabeled image. However, this practice exposes learning to confirmation bias. In this paper we propose Generalized Pseudo-Labeling (GPL), a simple and generic way to exploit negative pseudo-labels in consistency regularization, entailing minimal additional computational overhead and hyperpameter fine-tuning. GPL makes learning more robust by using the information that an image does not belong to a certain class, which is more abundant and reliable. We showcase GPL in the context of FixMatch. In the benchmark using only 40 labels of the CIFAR-10 dataset, adding GPL on top of FixMatch improves the error rate from 7.93% to 6.58%, and on CIFAR-100 with 2500 labels, from 28.02% to 26.85%.
Nikolaos Karaliolios, Florian Chabot, Camille Dupont, Hervé Le Borgne, Quoc Cuong Pham, Romaric Audigier
ICIP6
2023 Proposal-Contrastive Pretraining for Object Detection from Fewer Data
Quentin Bouniot, Romaric Audigier, Angélique Loesch, Amaury Habrard
ICLR2
2023 Towards Few-Annotation Learning for Object Detection: Are Transformer-based Models More Efficient?
abstract
For specialized and dense downstream tasks such as object detection, labeling data requires expertise and can be very expensive, making few-shot and semi-supervised models much more attractive alternatives. While in the few-shot setup we observe that transformer-based object detectors perform better than convolution-based two-stage models for a similar amount of parameters, they are not as effective when used with recent approaches in the semi-supervised setting. In this paper, we propose a semi-supervised method tailored for the current state-of-the-art object detector Deformable DETR in the few-annotation learning setup using a student-teacher architecture, which avoids relying on a sensitive post-processing of the pseudo-labels generated by the teacher model. We evaluate our method on the semi-supervised object detection benchmarks COCO and Pascal VOC, and it outperforms previous methods, especially when annotations are scarce. We believe that our contributions open new possibilities to adapt similar object detection methods in this setup as well.
Quentin Bouniot, Angélique Loesch, Amaury Habrard, Romaric Audigier
WACV4
2022 Spatio-temporal predictive tasks for abnormal event detection in videos
abstract
Abnormal event detection in videos is a challenging problem, partly due to the multiplicity of abnormal patterns and the lack of their corresponding annotations. In this paper, we propose new constrained pretext tasks to learn object level normality patterns. Our approach consists in learning a mapping between down-scaled visual queries and their corresponding normal appearance and motion characteristics at the original resolution. The proposed tasks are more challenging than reconstruction and future frame prediction tasks which are widely used in the literature, since our model learns to jointly predict spatial and temporal features rather than reconstructing them. We believe that more constrained pretext tasks induce a better learning of normality patterns. Experiments on several benchmark datasets demonstrate the effectiveness of our approach to localize and track anomalies as it outperforms or reaches the current state-of-the-art on spatio-temporal evaluation metrics.
Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier
AVSS5
2022 Improving Few-Shot Learning Through Multi-task Representation Learning Theory
Quentin Bouniot, Ievgen Redko, Romaric Audigier, Angélique Loesch, Amaury Habrard
ECCV (20)3
2022 Object-Centric and Memory-Guided Normality Reconstruction for Video Anomaly Detection
abstract
This paper addresses video anomaly detection problem for videosurveillance. Due to the inherent rarity and heterogeneity of abnormal events, the problem is viewed as a normality modeling strategy, in which our model learns object-centric normal patterns without seeing anomalous samples during training. The main contributions consist in coupling pre-trained object-level action features prototypes with a cosine distance-based anomaly estimation function, therefore extending previous methods by introducing additional constraints to the mainstream reconstruction-based strategy. Our framework leverages both appearance and motion information to learn object-level behavior and captures prototypical patterns within a memory module. Experiments on several well-known datasets demonstrate the effectiveness of our method as it outperforms current state-of-the-art on most relevant spatio-temporal evaluation metrics.
Khalil Bergaoui, Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier
ICIP6
2022 A formal approach to good practices in Pseudo-Labeling for Unsupervised Domain Adaptive Re-Identification
Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu
Comput. Vis. Image Underst.2
2021 Detecting Human-to-Human-or-Object (H2O) Interactions with DIABOLO
abstract
Detecting human interactions is crucial for human behavior analysis. Many methods have been proposed to deal with Human-to-Object Interaction (HOI) detection, i.e., detecting in an image which person and object interact together and classifying the type of interaction. However, Human-to-Human Interactions, such as social and violent interactions, are generally not considered in available HOI training datasets. As we think these types of interactions cannot be ignored and decorrelated from HOI when analyzing human behavior, we propose a new interaction dataset to deal with both types of human interactions: Human-to-Human-or-Object (H2O). In addition, we introduce a novel taxonomy of verbs, intended to be closer to a description of human body attitude in relation to the surrounding targets of interaction, and more independent of the environment. Unlike some existing datasets, we strive to avoid defining synonymous verbs when their use highly depends on the target type or requires a high level of semantic interpretation. As H2O dataset includes V-COCO images annotated with this new taxonomy, images obviously contain more interactions. This can be an issue for HOI detection methods whose complexity depends on the number of people, targets or interactions. Thus, we propose DIABOLO (Detecting Inter Actions By Only Looking Once), an efficient subject-centric single-shot method to detect all interactions in one forward pass, with constant inference time independent of image content. In addition, this multi-task network simultaneously detects all people and objects. We show how sharing a network for these tasks does not only save computation resource but also improves performance collaboratively. Finally, DIABOLO is a strong baseline for the new proposed challenge of H2O- Interaction detection, as it outperforms all state-of-the-art methods when trained and evaluated on HOI dataset V-COCO. We hope that this new dataset and new baseline will foster future research. H2O is available on https:/lkalisteo.cea.fr/.
Astrid Orcesi, Romaric Audigier, Fritz Poka Toukam, Bertrand Luvison
FG2
2021 Describe Me If You Can! Characterized Instance-Level Human Parsing
abstract
Several computer vision applications such as person search or online fashion rely on human description. The use of instance-level human parsing (HP) is therefore relevant since it localizes semantic attributes and body parts within a person. But how to characterize these attributes? To our knowledge, only some single-HP datasets describe attributes with some color, size and/or pattern characteristics. There is a lack of dataset for multi-HP in the wild with such characteristics. In this article, we propose the dataset CCIHP based on the multi-HP dataset CIHP, with 20 new labels covering these 3 kinds of characteristics.1In addition, we propose HPTR, a new bottom-up multi-task method based on transformers as a fast and scalable baseline. It is the fastest method of multi-HP state of the art while having precision comparable to the most precise bottom-up method. We hope this will encourage research for fast and accurate methods of precise human descriptions.1CCIHP is available on https://kalisteo.cea.fr/index.php/free-resources/
Angélique Loesch, Romaric Audigier
ICIP2
2020 Optimal Transport as a Defense Against Adversarial Attacks
abstract
Deep learning classifiers are now known to have flaws in the representations of their class. Adversarial attacks can find a human-imperceptible perturbation for a given image that will mislead a trained model. The most effective methods to defend against such attacks trains on generated adversarial examples to learn their distribution. Previous work aimed to align original and adversarial image representations in the same way as domain adaptation to improve robustness. Yet, they partially align the representations using approaches that do not reflect the geometry of space and distribution. In addition, it is difficult to accurately compare robustness between defended models. Until now, they have been evaluated using a fixed perturbation size. However, defended models may react differently to variations of this perturbation size. In this paper, the analogy of domain adaptation is taken a step further by exploiting optimal transport theory. We propose to use a loss between distributions that faithfully reflect the ground distance. This leads to SAT (Sinkhorn Adversarial Training), a more robust defense against adversarial attacks. Then, we propose to quantify more precisely the robustness of a model to adversarial attacks over a wide range of perturbation sizes using a different metric, the Area Under the Accuracy Curve (AUAC). We perform extensive experiments on both CIFAR-10 and CIFAR-100 datasets and show that our defense is globally more robust than the state-of-the-art.
Quentin Bouniot, Romaric Audigier, Angélique Loesch
ICPR2
2020 Unsupervised Domain Adaptation for Person Re-Identification through Source-Guided Pseudo-Labeling
abstract
Person Re-Identification (re-ID) aims at retrieving images of the same person taken by different cameras. A challenge for re-ID is the performance preservation when a model is used on data of interest (target data) which belong to a different domain from the training data domain (source data). Unsupervised Domain Adaptation (UDA) is an interesting research direction for this challenge as it avoids a costly annotation of the target data. Pseudo-labeling methods achieve the best results in UDA-based re-ID. They incrementally learn with identity pseudo-labels which are initialized by clustering features in the source reID encoder space. Surprisingly, labeled source data are discarded after this initialization step. However, we believe that pseudo-labeling could further leverage the labeled source data in order to improve the post-initialization training steps. In order to improve robustness against erroneous pseudo-labels, we advocate the exploitation of both labeled source data and pseudo-labeled target data during all training iterations. To support our guideline, we introduce a framework which relies on a two-branch architecture optimizing classification and triplet loss based metric learning in source and target domains, respectively, in order to allow adaptability to the target domain while ensuring robustness to noisy pseudo-labels. Indeed, shared low and mid-level parameters benefit from the source classification and triplet loss signal while high-level parameters of the target branch learn domain-specific features. Our method is simple enough to be easily combined with existing pseudo-labeling UDA approaches. We show experimentally that it is efficient and improves performance when the base method has no mechanism to deal with pseudo-label noise. Our approach reaches state-of-the-art performance when evaluated on commonly used datasets, Market-1501 and DukeMTMC-reID, and outperforms the state of the art when targeting the bigger and more challenging dataset MSMT.
Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu
ICPR2
2020 Classifying All Interacting Pairs in a Single Shot
abstract
In this paper, we introduce a novel human interaction detection approach, based on CALIPSO (Classifying ALl Interacting Pairs in a Single shOt), a classifier of human-object interactions. This new single-shot interaction classifier estimates interactions simultaneously for all human-object pairs, regardless of their number and class. State-of- the-art approaches adopt a multi-shot strategy based on a pairwise estimate of interactions for a set of human-object candidate pairs, which leads to a complexity depending, at least, on the number of interactions or, at most, on the number of candidate pairs. In contrast, the proposed method estimates the interactions on the whole image. Indeed, it simultaneously estimates all interactions between all human subjects and object targets by performing a single forward pass throughout the image. Consequently, it leads to a constant complexity and computation time independent of the number of subjects, objects or interactions in the image. In detail, interaction classification is achieved on a dense grid of anchors thanks to a joint multi-task network that learns three complementary tasks simultaneously: (i) prediction of the types of interaction, (ii) estimation of the presence of a target and (iii) learning of an embedding which maps interacting subject and target to a same representation, by using a metric learning strategy. In addition, we introduce an object-centric passive-voice verb estimation which significantly improves results. Evaluations on the two well-known Human-Object Interaction image datasets, V- COCO and HICO-DET, demonstrate the competitiveness of the proposed method (2nd place) compared to the state-of- the-art while having constant computation time regardless of the number of objects and interactions in the image.
Sanaa Chafik, Astrid Orcesi, Romaric Audigier, Bertrand Luvison
WACV3
2019 End-To-End Person Search Sequentially Trained On Aggregated Dataset
abstract
In video surveillance applications, person search is a challenging task consisting in detecting people and extracting features from their silhouette for re-identification (re-ID) purpose. We propose a new end-to-end model that jointly computes detection and feature extraction steps through a single deep Convolutional Neural Network architecture. Sharing feature maps between the two tasks for jointly describing people commonalities and specificities allows faster runtime, which is valuable in real-world applications. In addition to reaching state-of-the-art accuracy, this multi-task model can be sequentially trained task-by-task, which results in a broader acceptance of input dataset types. Indeed, we show that aggregating more pedestrian detection datasets without costly identity annotations makes the shared feature maps more generic, and improves re-ID precision. Moreover, these boosted shared feature maps result in re-ID features more robust to a cross-dataset scenario.
Angélique Loesch, Jaonary Rabarisoa, Romaric Audigier
ICIP3
2016 Improving Multi-frame Data Association with Sparse Representations for Robust Near-online Multi-object Tracking
Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle
ECCV (8)2
2016 RIMOC, a feature to discriminate unstructured motions: Application to violence detection for video-surveillance
Pedro Ribeiro 0006, Romaric Audigier, Quoc Cuong Pham
Comput. Vis. Image Underst.2
2015 Collaboration and spatialization for an efficient multi-person tracking via sparse representations
abstract
Multi-person tracking is a very difficult problem in Computer Vision as a tracking algorithm is facing several issues, such as appearance changes, targets' occlusions and similar appearances between people. In an online tracking-by-detection algorithm, robust and discriminative specific appearance models help handling these difficulties. As done in single object tracking, we use sparse representations to extract local features of the targets and study how these representations can be specifically employed for multi-person tracking. Experiments on several datasets show that considering spatial information is crucial in order to improve the tracking performances with local descriptions compared to holistic features. Using large collaborative representations also improve the tracking results by naturally discarding irrelevant local patches.
Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle
AVSS2
2015 Online multi-person tracking based on global sparse collaborative representations
abstract
Multi-person tracking is still a challenging problem due to recurrent occlusion, pose variation and similar appearances between people. Inspired by the success of sparse representations in single object tracking and face recognition, we propose in this paper an online tracking by detection framework based on collaborative sparse representations. We argue that collaborative representations can better differentiate people compared to target-specific models and therefore help to produce a more robust tracking system. We also show that despite the size of the dictionaries involved, these representations can be efficiently computed with large-scale optimization techniques to get a near real-time algorithm. Experiments show that the proposed approach compares well to other recent online tracking systems on various datasets.
Loïc Fagot-Bouquet, Romaric Audigier, Yoann Dhome, Frédéric Lerasle
ICIP2
2015 Collaborative Tracking and Distributed Control for an IP-PTZ Camera Network
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham
ICPRAM (2)2
2013 IMM-Based Tracking and Latency Control with Off-the-Shelf IP PTZ Camera
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham
ACIVS2
2010 Relationships between some watershed definitions and their tie-zone transforms
Romaric Audigier, Roberto A. Lotufo
Image Vis. Comput.1
2005 The tie-zone watershed: definition, algorithm and applications
abstract
In this work, a new type of watershed transform is introduced: the tie-zone watershed (TZWS). This region-based watershed transform does not depend on arbitrary implementation and provides a unique and optimal solution. Indeed, many solutions are sometimes possible when segmenting an image with a watershed algorithm. In this case, the TZWS assigns each pixel to a catchment basin (CB) if in all solutions it belongs to this CB. Otherwise, the pixel is said to belong to a tie-zone (TZ). We propose an efficient algorithm based on image foresting transform (IFT) which computes the TZWS transform as a shortest-path forest. Finally, two applications of this TZWS are presented: bounding intervals for segmented objects' extensions and a progressive segmentation procedure.
Romaric Audigier, Roberto A. Lotufo, Michel Couprie
ICIP (2)1