Angélique Loesch

dblp:173/5969 · also Angelique Loesch · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0001-5427-3010ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2024 MonoProb: Self-Supervised Monocular Depth Estimation with Interpretable Uncertainty
abstract
Self-supervised monocular depth estimation methods aim to be used in critical applications such as autonomous vehicles for environment analysis. To circumvent the potential imperfections of these approaches, a quantification of the prediction confidence is crucial to guide decision-making systems that rely on depth estimation. In this paper, we propose MonoProb, a new unsupervised monocular depth estimation method that returns an interpretable uncertainty, which means that the uncertainty reflects the expected error of the network in its depth predictions. We rethink the stereo or the structure-from-motion paradigms used to train unsupervised monocular depth models as a probabilistic problem. Within a single forward pass inference, this model provides a depth prediction and a measure of its confidence, without increasing the inference time. We then improve the performance on depth and uncertainty with a novel self-distillation loss for which a student is supervised by a pseudo ground truth that is a probability distribution on depth output by a teacher. To quantify the performance of our models we design new metrics that, unlike traditional ones, measure the absolute performance of uncertainty predictions. Our experiments highlight enhancements achieved by our method on standard depth and uncertainty metrics as well as on our tailored metrics. https://github.com/CEA-LIST/MonoProb
Rémi Marsal, Florian Chabot, Angélique Loesch, William Grolleau, Hichem Sahbi
WACV3
2023 Can Human Attribute Segmentation be More Robust to Operational Contexts Without New Labels?
abstract
Human Attribute Segmentation (HAS) describes, pixel-wise, the different semantic parts of people in an image. This fine-grained description is useful for several applications (e.g. security, fashion). However, despite the good performance reached by supervised Semantic Segmentation (SS) approaches, they are usually biased by the source training dataset and suffer from a performance drop when applied on new domains. Pixelwise image annotation for each new encountered context is tedious and expensive. So how can HAS become more robust to new contexts without new annotations? In this first study of Unsupervised Domain Adaptation (UDA) for HAS, we present UDA-HPTR, a new method based on HPTR [1] (Human Parsing with TRansformers) combined with self-supervised and semi-supervised learning paradigms to deal with UDA. UDA-HPTR improves performance on both source (labeled) and target (unlabeled) datasets compared to the fully supervised version (HPTR). It also outperforms HRDA, a state-of-the-art UDA method in autonomous driving benchmarks, by +6.7 p.p. on the source and +8.8 p.p. on the target, when applied to HAS while using only half the number of parameters.
Hejer Ammar, Angélique Loesch, Corentin Vannier, Romaric Audigier
ICIP2
2023 Proposal-Contrastive Pretraining for Object Detection from Fewer Data
Quentin Bouniot, Romaric Audigier, Angélique Loesch, Amaury Habrard
ICLR3
2023 Towards Few-Annotation Learning for Object Detection: Are Transformer-based Models More Efficient?
abstract
For specialized and dense downstream tasks such as object detection, labeling data requires expertise and can be very expensive, making few-shot and semi-supervised models much more attractive alternatives. While in the few-shot setup we observe that transformer-based object detectors perform better than convolution-based two-stage models for a similar amount of parameters, they are not as effective when used with recent approaches in the semi-supervised setting. In this paper, we propose a semi-supervised method tailored for the current state-of-the-art object detector Deformable DETR in the few-annotation learning setup using a student-teacher architecture, which avoids relying on a sensitive post-processing of the pseudo-labels generated by the teacher model. We evaluate our method on the semi-supervised object detection benchmarks COCO and Pascal VOC, and it outperforms previous methods, especially when annotations are scarce. We believe that our contributions open new possibilities to adapt similar object detection methods in this setup as well.
Quentin Bouniot, Angélique Loesch, Amaury Habrard, Romaric Audigier
WACV2
2023 BrightFlow: Brightness-Change-Aware Unsupervised Learning of Optical Flow
abstract
Unsupervised optical flow estimation relies on the assumption that pixels characterizing the same observed object should exhibit a stable appearance across video frames. With this assumption, the long-standing principle behind flow estimation consists in optimizing a photometric loss that maximizes the similarity between paired pixels in successive frames. However, these frames could be subject to strong brightness changes due to the radiometric properties of scenes as well as their viewing conditions.In this paper, we present BrightFlow, a new method to train any optical flow estimation network in an unsupervised manner. It consists in training two networks that jointly estimate optical flow and brightness changes. These changes are then compensated in the photometric loss so that reconstruction errors due to shadows or reflections will not affect negatively the training. As this compensation mechanism is only used at training stage, our method does not impact the number of parameters or the complexity at inference. Extensive experiments conducted on standard datasets and optical flow architectures show a consistent gain of our method. Source code is available at https://github.com/CEA-LIST/BrightFlow.
Rémi Marsal, Florian Chabot, Angélique Loesch, Hichem Sahbi
WACV3
2022 Spatio-temporal predictive tasks for abnormal event detection in videos
abstract
Abnormal event detection in videos is a challenging problem, partly due to the multiplicity of abnormal patterns and the lack of their corresponding annotations. In this paper, we propose new constrained pretext tasks to learn object level normality patterns. Our approach consists in learning a mapping between down-scaled visual queries and their corresponding normal appearance and motion characteristics at the original resolution. The proposed tasks are more challenging than reconstruction and future frame prediction tasks which are widely used in the literature, since our model learns to jointly predict spatial and temporal features rather than reconstructing them. We believe that more constrained pretext tasks induce a better learning of normality patterns. Experiments on several benchmark datasets demonstrate the effectiveness of our approach to localize and track anomalies as it outperforms or reaches the current state-of-the-art on spatio-temporal evaluation metrics.
Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier
AVSS3
2022 Improving Few-Shot Learning Through Multi-task Representation Learning Theory
Quentin Bouniot, Ievgen Redko, Romaric Audigier, Angélique Loesch, Amaury Habrard
ECCV (20)4
2022 Object-Centric and Memory-Guided Normality Reconstruction for Video Anomaly Detection
abstract
This paper addresses video anomaly detection problem for videosurveillance. Due to the inherent rarity and heterogeneity of abnormal events, the problem is viewed as a normality modeling strategy, in which our model learns object-centric normal patterns without seeing anomalous samples during training. The main contributions consist in coupling pre-trained object-level action features prototypes with a cosine distance-based anomaly estimation function, therefore extending previous methods by introducing additional constraints to the mainstream reconstruction-based strategy. Our framework leverages both appearance and motion information to learn object-level behavior and captures prototypical patterns within a memory module. Experiments on several well-known datasets demonstrate the effectiveness of our method as it outperforms current state-of-the-art on most relevant spatio-temporal evaluation metrics.
Khalil Bergaoui, Yassine Naji, Aleksandr Setkov, Angélique Loesch, Michèle Gouiffès, Romaric Audigier
ICIP4
2022 A formal approach to good practices in Pseudo-Labeling for Unsupervised Domain Adaptive Re-Identification
Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu
Comput. Vis. Image Underst.3
2021 Describe Me If You Can! Characterized Instance-Level Human Parsing
abstract
Several computer vision applications such as person search or online fashion rely on human description. The use of instance-level human parsing (HP) is therefore relevant since it localizes semantic attributes and body parts within a person. But how to characterize these attributes? To our knowledge, only some single-HP datasets describe attributes with some color, size and/or pattern characteristics. There is a lack of dataset for multi-HP in the wild with such characteristics. In this article, we propose the dataset CCIHP based on the multi-HP dataset CIHP, with 20 new labels covering these 3 kinds of characteristics.1In addition, we propose HPTR, a new bottom-up multi-task method based on transformers as a fast and scalable baseline. It is the fastest method of multi-HP state of the art while having precision comparable to the most precise bottom-up method. We hope this will encourage research for fast and accurate methods of precise human descriptions.1CCIHP is available on https://kalisteo.cea.fr/index.php/free-resources/
Angélique Loesch, Romaric Audigier
ICIP1
2020 Optimal Transport as a Defense Against Adversarial Attacks
abstract
Deep learning classifiers are now known to have flaws in the representations of their class. Adversarial attacks can find a human-imperceptible perturbation for a given image that will mislead a trained model. The most effective methods to defend against such attacks trains on generated adversarial examples to learn their distribution. Previous work aimed to align original and adversarial image representations in the same way as domain adaptation to improve robustness. Yet, they partially align the representations using approaches that do not reflect the geometry of space and distribution. In addition, it is difficult to accurately compare robustness between defended models. Until now, they have been evaluated using a fixed perturbation size. However, defended models may react differently to variations of this perturbation size. In this paper, the analogy of domain adaptation is taken a step further by exploiting optimal transport theory. We propose to use a loss between distributions that faithfully reflect the ground distance. This leads to SAT (Sinkhorn Adversarial Training), a more robust defense against adversarial attacks. Then, we propose to quantify more precisely the robustness of a model to adversarial attacks over a wide range of perturbation sizes using a different metric, the Area Under the Accuracy Curve (AUAC). We perform extensive experiments on both CIFAR-10 and CIFAR-100 datasets and show that our defense is globally more robust than the state-of-the-art.
Quentin Bouniot, Romaric Audigier, Angélique Loesch
ICPR3
2020 Unsupervised Domain Adaptation for Person Re-Identification through Source-Guided Pseudo-Labeling
abstract
Person Re-Identification (re-ID) aims at retrieving images of the same person taken by different cameras. A challenge for re-ID is the performance preservation when a model is used on data of interest (target data) which belong to a different domain from the training data domain (source data). Unsupervised Domain Adaptation (UDA) is an interesting research direction for this challenge as it avoids a costly annotation of the target data. Pseudo-labeling methods achieve the best results in UDA-based re-ID. They incrementally learn with identity pseudo-labels which are initialized by clustering features in the source reID encoder space. Surprisingly, labeled source data are discarded after this initialization step. However, we believe that pseudo-labeling could further leverage the labeled source data in order to improve the post-initialization training steps. In order to improve robustness against erroneous pseudo-labels, we advocate the exploitation of both labeled source data and pseudo-labeled target data during all training iterations. To support our guideline, we introduce a framework which relies on a two-branch architecture optimizing classification and triplet loss based metric learning in source and target domains, respectively, in order to allow adaptability to the target domain while ensuring robustness to noisy pseudo-labels. Indeed, shared low and mid-level parameters benefit from the source classification and triplet loss signal while high-level parameters of the target branch learn domain-specific features. Our method is simple enough to be easily combined with existing pseudo-labeling UDA approaches. We show experimentally that it is efficient and improves performance when the base method has no mechanism to deal with pseudo-label noise. Our approach reaches state-of-the-art performance when evaluated on commonly used datasets, Market-1501 and DukeMTMC-reID, and outperforms the state of the art when targeting the bigger and more challenging dataset MSMT.
Fabian Dubourvieux, Romaric Audigier, Angélique Loesch, Samia Ainouz 0001, Stéphane Canu
ICPR3
2019 End-To-End Person Search Sequentially Trained On Aggregated Dataset
abstract
In video surveillance applications, person search is a challenging task consisting in detecting people and extracting features from their silhouette for re-identification (re-ID) purpose. We propose a new end-to-end model that jointly computes detection and feature extraction steps through a single deep Convolutional Neural Network architecture. Sharing feature maps between the two tasks for jointly describing people commonalities and specificities allows faster runtime, which is valuable in real-world applications. In addition to reaching state-of-the-art accuracy, this multi-task model can be sequentially trained task-by-task, which results in a broader acceptance of input dataset types. Indeed, we show that aggregating more pedestrian detection datasets without costly identity annotations makes the shared feature maps more generic, and improves re-ID precision. Moreover, these boosted shared feature maps result in re-ID features more robust to a cross-dataset scenario.
Angélique Loesch, Jaonary Rabarisoa, Romaric Audigier
ICIP1
2018 Localization of 3D objects using model-constrained SLAM
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Olivier Gomez, Michel Dhome
Mach. Vis. Appl.1
2016 A Hybrid Structure/Trajectory Constraint for Visual SLAM
abstract
This paper presents a hybrid structure/trajectory constraint, that uses output camera poses of a model-based tracker, for object localization with SLAM algorithm. This constraint takes into account the structure information given by a CAD model while relying on the formalism of trajectory constraints. It has the advantages to be compact in memory and to accelerate the SLAM optimization process. The accuracy and robustness of the resulting localization as well as the memory and time gains are evaluated on synthetic and real data. Videos are available as supplementary material.
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome
3DV1
2015 Generic edgelet-based tracking of 3D objects in real-time
abstract
This paper addresses the challenging issue of real-time camera localization relative to any object that have texture or not, sharp edges or occluding contours. 3D contour points, dynamically extracted from a CAD model by Analysis-by-Synthesis on the graphics hardware, are combined with a keyframe-based SLAM algorithm to estimate camera poses. Our tracking solution is accurate, robust to sudden motions and to occlusions, as demonstrated on synthetic and real data. This solution is also easy to deploy since it only uses an RGB camera and a CAD model of the object of interest, requires no manual intervention on this model and runs on a consumer tablet at a frequency of 40Hz on a HD video-stream. Videos are available as supplemental material.
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome
IROS1