EDBT 2026 Demo / reviewers in the wild / expert
Arnaud Dapogny
dblp:165/8156
· DBLP profile ↗
39ranked-venue papers
12as first author
22since 2021 · last 2025
0000-0002-0074-8719ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 12 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 first-author · 10 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analyzing Fine-Tuning Representation Shift for Multimodal LLMs Steering
Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord |
ICCV | 4 |
| 2025 | Learning to Steer: Input-dependent Steering for Multimodal LLMsabstractSteering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior.
However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as \textit{mean} steering, rely on a single steering vector, applied independently of the input query. This paradigm faces limitations when the desired behavior is dependent on the example at hand. For example, a safe answer may consist in abstaining from answering when asked for an illegal activity, or may point to external resources or consultation with an expert when asked about medical advice. In this paper, we investigate a fine-grained steering that uses an input-specific linear shift. This shift is computed using contrastive input-specific prompting. However, the input-specific prompts required for this approach are not known at test time. Therefore, we propose to train a small auxiliary module to predict the input-specific steering vector. Our approach, dubbed as L2S (Learn-to-Steer), demonstrates that it reduces hallucinations and enforces safety in MLLMs, outperforming other static baselines. We will open-source our code. Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny, Alasdair Newson, Matthieu Cord |
NeurIPS | 4 |
| 2025 | NUPES: Non-Uniform Post-Training Quantization via Power Exponent SearchabstractDeep neural network (DNN) deployment has been confined to larger hardware devices due to their expensive computational requirements. This challenge has recently reached another scale with the emergence of large language models (LLMs). In order to reduce both their memory footprint and latency, a promising technique is quantization. It consists in converting floating point representations to low bit-width fixed point representations, usually by assuming a uniform mapping onto a regular grid. This process, referred to in the literature as uniform quantization, may however be ill-suited as most DNN weights and activations follow a bell-shaped distribution. This is even worse on LLMs whose weight distributions are known to exhibit large, high impact, outlier values. In this work, we propose an improvement over the most commonly adopted way to tackle this limitation in deep learning models quantization, namely, non-uniform quantization. NUPES leverages automorphisms to preserve the scalar multiplications. Such transformations are derived from power functions. However, the optimization of the exponent parameter and weight values remains a challenging and novel problem which could not be solved with previous post training optimization techniques which only learn to round up or down weight values in order to preserve the predictive function. We circumvent this limitation with a new paradigm: learning new quantized weights over the entire quantized space. Similarly, we enable the optimization of the power exponent, i.e. the optimization of the quantization operator itself during training by alleviating all the numerical instabilities. The resulting predictive function is compatible with integer-only low-bit inference. We show the ability of the method to achieve state-of-the-art compression rates in both, data-free and data-driven configurations. Our empirical benchmarks highlight the ability of NUPES to circumvent the limitations of previous post-training quantization techniques on transformers and large language models in particular. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Network Memory Footprint Compression Through Jointly Learnable Codebooks and MappingsabstractThe massive interest in deep neural networks (DNNs) for both computer vision and natural language processing has been sparked by the growth in computational power. However, this led to an increase in the memory footprint, to a point where it can be challenging to simply load a model on commodity devices such as mobile phones. To address this limitation, quantization is a favored solution as it maps high precision tensors to a low precision, memory efficient format. In terms of memory footprint reduction, its most effective variants are based on codebooks. These methods, however, suffer from two limitations. First, they either define a single codebook for each tensor, or use a memory-expensive mapping to multiple codebooks. Second, gradient descent optimization of the mapping favors jumps toward extreme values, hence not defining a proximal search. In this work, we propose to address these two limitations. First, we initially group similarly distributed neurons and leverage the re-ordered structure to either apply different scale factors to the different groups, or map weights that fall in these groups to several codebooks, without any mapping overhead. Second, stemming from this initialization, we propose a joint learning of the codebook and weight mappings that bears similarities with recent gradient-based post-training quantization techniques. Third, drawing estimation from straight-through estimation techniques, we introduce a novel gradient update definition to enable a proximal search of the codebooks and their mappings. The proposed jointly learnable codebooks and mappings (JLCM) method allows a very efficient approximation of any DNN: as such, a Llama 7B can be compressed down to 2Go and loaded on 5-year-old smartphones. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICLR | 2 |
| 2024 | Ontology-Guided Deep Metric Learning and Applications to Obstetrics
Jules Bonnard, Arnaud Dapogny, Ferdinand Dhombres, Kevin Bailly |
ICPR (9) | 2 |
| 2024 | PIPE: Parallelized inference through ensembling of residual quantization expansions
Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
Pattern Recognit. | 2 |
| 2024 | Prior-Guided Attribution of Deep Neural Networks for Obstetrics and GynecologyabstractObstetrics and gynecology (OB/GYN) are areas of medicine that specialize in the care of women during pregnancy and childbirth and in the diagnosis of diseases of the female reproductive system. Ultrasound scanning has become ubiquitous in these branches of medicine, as breast or fetal ultrasound images can lead the sonographer and guide him through his diagnosis. However, ultrasound scan images require a lot of resources to annotate and are often unavailable for training purposes because of confidentiality reasons, which explains why deep learning methods are still not as commonly used to solve OB/GYN tasks as in other computer vision tasks. In order to tackle this lack of data for training deep neural networks in this context, we propose Prior-Guided Attribution (PGA), a novel method that takes advantage of prior spatial information during training by guiding part of its attribution towards these salient areas. Furthermore, we introduce a novel prior allocation strategy method to take into account several spatial priors at the same time while providing the model enough degrees of liberty to learn relevant features by itself. The proposed method only uses the additional information during training, without needing it during inference. After validating the different elements of the method as well as its genericity on a facial analysis problem, we demonstrate that the proposed PGA method constantly outperforms existing baselines on two ultrasound imaging OB/GYN tasks: breast cancer detection and scan plane detection with segmentation prior maps. Jules Bonnard, Arnaud Dapogny, Richard Zsamboki, Lucrezia De Braud, Davor Jurkovic, Kevin Bailly, Ferdinand Dhombres |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | RULe: Relocalization-Uniformization-Landmark Estimation Network for Real-Time Face Alignment in Degraded ConditionsabstractFace alignment refers to the process of estimating the position of a number of salient landmarks on face images or videos, such as mouth and eye corners, nose tip, etc. With the availability of large annotated databases and the rise of deep learning-based methods, face alignment as a domain has matured to a point where it can be applied in more or less unconstrained conditions, e.g. non-frontal head poses, presence of heavy make-up or partial occlusions. However, when considering real-case alignment on videos with possibly low frame rates, we need to make sure that the algorithms are robust to jittering of the face bounding box localization, low-resolution of the face crops, possible bad environmental lighting, brightness, and presence of noise. To tackle these issues, we propose RULe, a three-staged Relocalization-Uniformization-Landmark Estimation network. In the first stage, an initial loosely localized bounding box gets refined to output a well centered face crop, thus reducing the variability of the images prior to passing them to the subsequent stage. Then, in the second stage, the face style is uniformized (using adversarial learning as well as perceptual losses) to correct low resolution or variations of brightness/contrast. Finally, the third stage outputs a precise landmark estimation given such enhanced face crop using a cascaded compact model trained using hint-based knowledge distillation. We show through a variety of experiments that RULe achieves real-time face alignment with state-of-the-art precision in heavily degraded conditions. Arnaud Dapogny, Gauthier Tallec, Jules Bonnard, Edouard Yvinec, Kevin Bailly |
FG | 1 |
| 2023 | Adversarial Deep Multi-Task Learning Using Semantically Orthogonal Spaces and Application to Facial Attributes PredictionabstractDeep learning-based multi-task approaches usually rely on factorizing representation layers up to a certain point, where the network splits into several heads, each one addressing a specific task. Depending on the inter-task correlation, such naive model may or may not allow the tasks to benefit from each others. In this paper, we propose a novel Semantic Orthogonality Spaces (SOS) method for multi-task problems, where each task is predicted using the information from a common subspace that factorizes information among all tasks, as well as a task-specific subspace. We enforce orthogonality between these tasks by applying soft orthogonality constraints, as well as adversarially-learned semantic orthogonality objectives that ensures that predicting one task requires the specific information related to that task. We demonstrate the effectiveness of SOS on synthetic data, as well as for large-scale facial attributes prediction. In particular, we use SOS to craft a lightweight architecture that provides high-end accuracies on CelebA database. Arnaud Dapogny, Gauthier Tallec, Jules Bonnard, Edouard Yvinec, Kevin Bailly |
FG | 1 |
| 2023 | Fighting Over-Fitting with Quantization for Learning Deep Neural Networks on Noisy LabelsabstractThe rising performance of deep neural networks is often empirically attributed to an increase in the available computational power, which allows complex models to be trained upon large amounts of annotated data. However, increased model complexity leads to costly deployment of modern neural networks, while gathering such amounts of data requires huge costs to avoid label noise. In this work, we study the ability of compression methods to tackle both of these problems at once. We hypothesize that quantization-aware training, by restricting the expressivity of neural networks, behaves as a regularization. Thus, it may help fighting overfitting on noisy data while also allowing for the compression of the model at inference. We first validate this claim on a controlled test with manually introduced label noise. Furthermore, we also test the proposed method on Facial Action Unit detection, where labels are typically noisy due to the subtlety of the task. In all cases, our results suggests that quantization significantly improve the results compared with existing baselines, regularization as well as other compression methods. Gauthier Tallec, Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICIP | 3 |
| 2023 | Designing Strong Baselines for Ternary Neural Network Quantization through Support and Mass EqualizationabstractDeep neural networks (DNNs) offer the highest performance in a wide range of applications in computer vision. These results rely on over-parameterized backbones, which are expensive to run. This computational burden can be dramatically reduced by quantizing (in either data-free (DFQ), post-training (PTQ) or quantization-aware training (QAT) scenarios) floating point values to ternary values (2 bits, with each weight taking value in {−1,0,1}). In this context, we observe that rounding to nearest minimizes the expected error given a uniform distribution and thus does not account for the skewness and kurtosis of the weight distribution, which strongly affects ternary quantization performance. This raises the following question: shall one minimize the highest or average quantization error? To answer this, we design two operators: TQuant and MQuant that correspond to these respective minimization tasks. We show experimentally that our approach allows to significantly improve the performance of ternary quantization through a variety of scenarios in DFQ, PTQ and QAT and give strong insights to pave the way for future research in deep neural network quantization. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICIP | 2 |
| 2023 | PowerQuant: Automorphism Search for Non-Uniform Quantization
Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICLR | 2 |
| 2023 | REx: Data-Free Residual Quantization Error ExpansionabstractDeep neural networks (DNNs) are ubiquitous in computer vision and natural language processing, but suffer from high inference cost. This problem can be addressed by quantization, which consists in converting floating point operations into a lower bit-width format. With the growing concerns on privacy rights, we focus our efforts on data-free methods. However, such techniques suffer from their lack of adaptability to the target devices, as a hardware typically only supports specific bit widths. Thus, to adapt to a variety of devices, a quantization method shall be flexible enough to find good accuracy v.s. speed trade-offs for every bit width and target device. To achieve this, we propose REx, a quantization method that leverages residual error expansion, along with group sparsity.
We show experimentally that REx enables better trade-offs (in terms of accuracy given any target bit-width) on both convnets and transformers for computer vision, as well as NLP models. In particular, when applied to large language models, we show that REx elegantly solves the outlier problem that hinders state-of-the-art quantization methods.
In addition, REx is backed off by strong theoretical guarantees on the preservation of the predictive function of the original model. Lastly, we show that REx is agnostic to the quantization operator and can be used in combination with previous quantization work. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 2 |
| 2023 | SPIQ: Data-Free Per-Channel Static Input QuantizationabstractComputationally expensive neural networks are ubiquitous in computer vision and solutions for efficient inference have drawn a growing attention in the machine learning community. Examples of such solutions comprise quantization, i.e. converting the processing values (weights and inputs) from floating point into integers e.g. int8 or int4. Concurrently, the rise of privacy concerns motivated the study of less invasive acceleration methods, such as data-free quantization of pre-trained models weights and activations. Previous approaches either exploit statistical information to deduce scalar ranges and scaling factors for the activations in a static manner, or dynamically adapt this range on-the-fly for each input of each layer (also referred to as activations): the latter generally being more accurate at the expense of significantly slower inference. In this work, we argue that static input quantization can reach the accuracy levels of dynamic methods by means of a per-channel input quantization scheme that allows one to more finely preserve cross-channel dynamics. We show through a thorough empirical evaluation on multiple computer vision problems (e.g. ImageNet classification, Pascal VOC object detection as well as CityScapes semantic segmentation) that the proposed method, dubbed SPIQ, achieves accuracies rivalling dynamic approaches with static-level inference speed, significantly outperforming state-of-the-art quantization methods on every benchmark. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
WACV | 2 |
| 2023 | RED++ : Data-Free Pruning of Deep Neural Networks via Input Splitting and Output MergingabstractPruning Deep Neural Networks (DNNs) is a prominent field of study in the goal of inference runtime acceleration. In this paper, we introduce a novel data-free pruning protocol RED++. Only requiring a trained neural network, and not specific to any particular DNN, we exploit an adaptive data-free scalar hashing which exhibits redundancies among neuron weight values. We study the theoretical and empirical guarantees on the preservation of the accuracy from the hashing as well as the expected pruning ratio resulting from the exploitation of said redundancies. We propose a novel data-free pruning technique of DNN layers which removes the input-wise redundant operations. This algorithm is straightforward, parallelizable and offers novel perspective on DNN pruning by shifting the burden of large computation to efficient memory access and allocation. We provide theoretical guarantees on RED++ performance and empirically demonstrate its superiority over other data-free pruning methods and its competitiveness with data-driven ones on ResNets, MobileNets, and EfficientNets. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | THIN: THrowable Information Networks and Application for Facial Expression Recognition in the WildabstractFor a number of machine learning problems, an exogenous variable can be identified such that it heavily influences the appearance of the different classes, and an ideal classifier should be invariant to this variable. An example of such exogenous variable is identity if facial expression recognition (FER) is considered. In this paper, we propose a dual exogenous/endogenous representation. The former captures the exogenous variable whereas the second one models the task at hand (e.g. facial expression). We design a prediction layer that uses a tree-gated deep ensemble conditioned by the exogenous representation. We also propose an exogenous dispelling loss to remove the exogenous information from the endogenous representation. Thus, the exogenous information is used two times in a throwable fashion, first as a conditioning variable for the target task, and second to create invariance within the endogenous representation. We call this method THIN, standing for THrowable Information Networks. We experimentally validate THIN in several contexts where an exogenous information can be identified, such as digit recognition under large rotations and shape recognition at multiple scales. We also apply it to FER with identity as the exogenous variable. We demonstrate that THIN significantly outperforms state-of-the-art approaches on several challenging datasets. Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Multi-Order Networks for Action Unit DetectionabstractAction Units (AU) are muscular activations used to describe facial expressions. Therefore accurate AU recognition unlocks unbiaised face representation which can improve face-based affective computing applications. From a learning standpoint AU detection is a multi-task problem with strong inter-task dependencies. To solve such problem, most approaches either rely on weight sharing, or add explicit dependency modelling by decomposing the joint task distribution using Bayes chain rule. If the latter strategy yields comprehensive inter-task relationships modelling, it requires imposing an arbitrary order into an unordered task set. Crucially, this ordering choice has been identified as a source of performance variations. In this paper, we present Multi-Order Network (MONET), a multi-task method with joint task order optimization. MONET uses a differentiable order selection to jointly learn task-wise modules with their optimal chaining order. Furthermore, we introduce warmup and order dropout to enhance order selection by encouraging order exploration. Experimentally, we first demonstrate MONET capacity to retrieve the optimal order in a toy environment. Second, we validate MONET architecture by showing that MONET outperforms existing multi-task baselines on multiple attribute detection problems chosen for their wide range of dependency settings. More importantly, we demonstrate that MONET significantly extends state-of-the-art performance in AU detection. Gauthier Tallec, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Privileged Attribution Constrained Deep Networks for Facial Expression RecognitionabstractFacial Expression Recognition (FER) is crucial in many research domains because it enables machines to better understand human behaviours. FER methods face the problems of relatively small datasets and noisy data that don’t allow classical networks to generalize well. To alleviate these issues, we guide the model to concentrate on specific facial areas like the eyes, the mouth or the eyebrows, which we argue are decisive to recognise facial expressions. We propose the Privileged Attribution Loss (PAL), a method that directs the focus of the model towards the most salient facial regions by encouraging its attribution maps to correspond to a heatmap formed by facial landmarks. Furthermore, we introduce several channel strategies that allow the model to have more degrees of freedom. The proposed method is independent of the backbone architecture and doesn’t need additional semantic information at test time. Finally, experimental results show that the proposed PAL method outperforms current state-of-the-art methods on both RAF-DB and AffectNet. Jules Bonnard, Arnaud Dapogny, Ferdinand Dhombres, Kevin Bailly |
ICPR | 2 |
| 2022 | To Fold or Not to Fold: a Necessary and Sufficient Condition on Batch-Normalization Layers FoldingabstractBatch-Normalization (BN) layers have become fundamental components in the evermore complex deep neural network architectures. Such models require acceleration processes for deployment on edge devices. However, BN layers add computation bottlenecks due to the sequential operation processing: thus, a key, yet often overlooked component of the acceleration process is BN layers folding. In this paper, we demonstrate that the current BN folding approaches are suboptimal in terms of how many layers can be removed. We therefore provide a necessary and sufficient condition for BN folding and a corresponding optimal algorithm. The proposed approach systematically outperforms existing baselines and allows to dramatically reduce the inference time of deep neural networks. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
IJCAI | 2 |
| 2022 | SInGE: Sparsity via Integrated Gradients Estimation of Neuron RelevanceabstractThe leap in performance in state-of-the-art computer vision methods is attributed to the development of deep neural networks. However it often comes at a computational price which may hinder their deployment. To alleviate this limitation, structured pruning is a well known technique which consists in removing channels, neurons or filters, and is commonly applied in order to produce more compact models. In most cases, the computations to remove are selected based on a relative importance criterion. At the same time, the need for explainable predictive models has risen tremendously and motivated the development of robust attribution methods that highlight the relative importance of pixels of an input image or feature map. In this work, we discuss the limitations of existing pruning heuristics, among which magnitude and gradient-based methods. We draw inspiration from attribution methods to design a novel integrated gradient pruning criterion, in which the relevance of each neuron is defined as the integral of the gradient variation on a path towards this neuron removal. Furthermore, We propose an entwined DNN pruning and fine-tuning flowchart to better preserve DNN accuracy while removing parameters. We show through extensive validation on several datasets, architectures as well as pruning scenarios that the proposed method, dubbed SInGE, significantly outperforms existing state-of-the-art DNN pruning methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 2 |
| 2021 | PLOP: Learning Without Forgetting for Continual Semantic SegmentationabstractDeep learning approaches are nowadays ubiquitously used to tackle computer vision tasks such as semantic segmentation, requiring large datasets and substantial computational power. Continual learning for semantic segmentation (CSS) is an emerging trend that consists in updating an old model by sequentially adding new classes. However, continual learning methods are usually prone to catastrophic forgetting. This issue is further aggravated in CSS where, at each step, old classes from previous iterations are collapsed into the background. In this paper, we propose Local POD, a multi-scale pooling distillation scheme that preserves long- and short-range spatial relationships at feature level. Furthermore, we design an entropy-based pseudo-labelling of the background w.r.t. classes predicted by the old model to deal with background shift and avoid catastrophic forgetting of the old classes. Our approach, called PLOP, significantly outperforms state-of-the-art methods in existing CSS scenarios, as well as in newly proposed challenging benchmarks1. Arthur Douillard, Arnaud Dapogny, Matthieu Cord |
CVPR | 3 |
| 2021 | RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksabstractDeep Neural Networks (DNNs) are ubiquitous in today's computer vision landscape, despite involving considerable computational costs. The mainstream approaches for runtime acceleration consist in pruning connections (unstructured pruning) or, better, filters (structured pruning), both often requiring data to retrain the model. In this paper, we present RED, a data-free, unified approach to tackle structured pruning. First, we propose a novel adaptive hashing of the scalar DNN weight distribution densities to increase the number of identical neurons represented by their weight vectors. Second, we prune the network by merging redundant neurons based on their relative similarities, as defined by their distance. Third, we propose a novel uneven depthwise separation technique to further prune convolutional layers. We demonstrate through a large variety of benchmarks that RED largely outperforms other data-free pruning methods, often reaching performance similar to unconstrained, data-driven methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 2 |
| 2020 | The Missing Data Encoder: Cross-Channel Image Completion with Hide-and-Seek Adversarial NetworkabstractImage completion is the problem of generating whole images from fragments only. It encompasses inpainting (generating a patch given its surrounding), reverse inpainting/extrapolation (generating the periphery given the central patch) as well as colorization (generating one or several channels given other ones). In this paper, we employ a deep network to perform image completion, with adversarial training as well as perceptual and completion losses, and call it the “missing data encoder” (MDE). We consider several configurations based on how the seed fragments are chosen. We show that training MDE for “random extrapolation and colorization” (MDE-REC), i.e. using random channel-independent fragments, allows a better capture of the image semantics and geometry. MDE training makes use of a novel “hide-and-seek” adversarial loss, where the discriminator seeks the original non-masked regions, while the generator tries to hide them. We validate our models qualitatively and quantitatively on several datasets, showing their interest for image completion, representation learning as well as face occlusion handling. Arnaud Dapogny, Matthieu Cord, Patrick Pérez |
AAAI | 1 |
| 2020 | MoDuL: Deep Modal and Dual Landmark-wise Gated Network for Facial Expression RecognitionabstractAutomatic facial expression recognition (FER) is a challenging computer vision problem that finds a number of applications in human-computer interaction. Most recent FER approaches are deep-learning based and involve the extraction of two types of features from face images: geometric features (e.g. distances between aligned facial landmarks) and appearance features extracted using convolutional neural networks applied on patches extracted around each landmark. In this paper, we explore the use of gating networks to learn an optimal combination of these two modalities (modal gate). Furthermore, we also design landmark-wise gates to adaptively weight each landmark as well as the corresponding patch contribution. The proposed MoDuL architecture achieves state-of-the-art results on several FER databases with negligible computational overhead. Sacha Bernheim, Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
FG | 3 |
| 2020 | Deep Entwined Learning Head Pose and Face Alignment Inside an Attentional Cascade with Doubly-Conditional fusionabstractHead pose estimation and face alignment constitute a backbone preprocessing for many applications relying on face analysis. While both are closely related tasks, they are generally addressed separately, e.g. by deducing the head pose from the landmark locations. In this paper, we propose to entwine face alignment and head pose tasks inside an attentional cascade. This cascade uses a geometry transfer network for integrating heterogeneous annotations to enhance landmark localization accuracy. Furthermore, we propose a doubly-conditional fusion scheme to select relevant feature maps, and regions thereof, based on a current head pose and landmark localization estimate. We empirically show the benefit of entwining head pose and landmark localization objectives inside our architecture, and that the proposed AC-DC model enhances the state-of-the-art accuracy on multiple databases for both face alignment and head pose estimation tasks. Arnaud Dapogny, Kevin Bailly, Matthieu Cord |
FG | 1 |
| 2020 | DeeSCo: Deep heterogeneous ensemble with Stochastic Combinatory loss for gaze estimationabstractFrom medical research to gaming applications, gaze estimation is becoming a valuable tool. While there exists a number of hardware-based solutions, recent deep learning-based approaches, coupled with the availability of large-scale databases, have allowed to provide a precise gaze estimate using only consumer sensors. However, there remains a number of questions, regarding the problem formulation, architectural choices and learning paradigms for designing gaze estimation systems in order to bridge the gap between geometry-based systems involving specific hardware and approaches using consumer sensors only. In this paper, we introduce a deep, end-to-end trainable ensemble of heatmap-based weak predictors for 2D/3D gaze estimation. We show that, through heterogeneous architectural design of these weak predictors, we can improve the decorrelation between the latter predictors to design more robust deep ensemble models. Furthermore, we propose a stochastic combinatory loss that consists in randomly sampling combinations of weak predictors at train time. This allows to train better individual weak predictors, with lower correlation between them. This, in turns, allows to significantly enhance the performance of the deep ensemble. We show that our Deep heterogeneous ensemble with Stochastic Combinatory loss (DeeSCo) outperforms state-of-the-art approaches for 2D/3D gaze estimation on multiple datasets. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
FG | 2 |
| 2020 | SEMEDA: Enhancing segmentation precision with semantic edge aware loss
Arnaud Dapogny, Matthieu Cord |
Pattern Recognit. | 2 |
| 2019 | Tree-gated Deep Regressor Ensemble For Face Alignment In The WildabstractFace alignment consists in aligning a shape model on a face in an image. It is an active domain in computer vision as it is a preprocessing for applications like facial expression recognition, face recognition and tracking, face animation, etc. Current state-of-the-art methods already perform well on "easy" datasets, i.e. those that present moderate variations in head pose, expression, illumination or partial occlusions, but may not be robust to "in-the-wild" data. In this paper, we address this problem by using an ensemble of deep regressors instead of a single large regressor. Furthermore, instead of averaging the ouputs of each regressor, we propose an adaptative weighting scheme that uses a tree-structured gate. Experiments on several challenging face datasets demonstrate that our approach outperforms the state-of-the-art methods. Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
FG | 2 |
| 2019 | DeCaFA: Deep Convolutional Cascade for Face Alignment in the WildabstractFace Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce an end-to-end deep convolutional cascade (DeCaFA) architecture for face alignment. Face Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce DeCaFA, an end-to-end deep convolutional cascade architecture for face alignment. DeCaFA uses fully-convolutional stages to keep full spatial resolution throughout the cascade. Between each cascade stage, DeCaFA uses multiple chained transfer layers with spatial softmax to produce landmark-wise attention maps for each of several landmark alignment tasks. Weighted intermediate supervision, as well as efficient feature fusion between the stages allow to learn to progressively refine the attention maps in an end-to-end manner. We show experimentally that DeCaFA significantly outperforms existing approaches on 300W, CelebA and WFLW databases. In addition, we show that DeCaFA can learn fine alignment with reasonable accuracy from very few images using coarsely annotated data. Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICCV | 1 |
| 2019 | Delving Deep into Interpreting Neural Nets with Piece-Wise Affine RepresentationabstractDeep convolutional neural networks (CNNs) are now ubiquitous in computer vision problems. However, these models usually describe very complicated functions of the input images. For a number of application, it is of utmost importance to be able to explain the decisions of a network, e.g. by highlighting the most relevant pixels in an image or a feature map w.r.t. a particular class. In this paper, we show that CNNs locally describe piece-wise affine functions of each pixel, whose coefficient and bias can be retrieved analytically. We apply our methodology on several popular CNNs and draw interesting conclusions on the relative contributions of pixels and biases for these networks. Antoine Saporta, Arnaud Dapogny, Matthieu Cord |
ICIP | 3 |
| 2019 | Dynamic Pose-Robust Facial Expression Recognition by Multi-View Pairwise Conditional Random ForestsabstractAutomatic facial expression classification (FER) from videos is a critical problem for the development of intelligent human-computer interaction systems. Still, it is a challenging problem that involves capturing high-dimensional spatio-temporal patterns describing the variation of one's appearance over time. Such representation undergoes great variability of the facial morphology and environmental factors as well as head pose variations. In this paper, we use Conditional Random Forests to capture low-level expression transition patterns. More specifically, heterogeneous derivative features (e.g., feature point movements or texture variations) are evaluated upon pairs of images. When testing on a video frame, pairs are created between this current frame and previous ones and predictions for each previous frame are used to draw trees from Pairwise Conditional Random Forests (PCRF) whose pairwise outputs are averaged over time to produce robust estimates. Moreover, PCRF collections can also be conditioned on head pose estimation for multi-view dynamic FER. As such, our approach appears as a natural extension of Random Forests for learning spatio-temporal patterns, potentially from multiple viewpoints. Experiments on popular datasets show that our method leads to significant improvements over standard Random Forests as well as state-of-the-art approaches on several scenarios, including a novel multi-view video corpus generated from a publicly available database. Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Investigating Deep Neural Forests for Facial Expression RecognitionabstractFacial Expression Recognition (FER) usually involves learning intermediate representations from high-dimensional, potentially noisy images, that can be adapted to unknown face morphologies. Moreover, it needs to fulfil the real-time constraint to be useful, e.g. for consumer robotics or healthcare systems. To tackle this issue, Random Forests (RFs) are convenient predictors, as the hard decisions at each node allow non-linear subdivisions of the space with a very fast evaluation runtime. However RFs are non-differenciable predictors, making it impossible to back-propagate the error down to upstream feature extraction layers (e.g. CNN). In this paper, we investigate the adaptation of deep Neural Forests (NFs) to FER. The latter allows to combine the best of the two worlds, i.e. a differentiable classification model that can be trained with backpropagation and stochastic gradient descent but has the runtime of a RF. We show that NFs provides competitive results for FER as compared to state-of-the-art approaches when trained upon various combinations of geometric/CNN-based features. Arnaud Dapogny, Kevin Bailly |
FG | 1 |
| 2018 | JEMImE: A Serious Game to Teach Children with ASD How to Adequately Produce Facial ExpressionsabstractBeing able to produce facial expressions (FEs) that are adequate given a social context is key to harmonious social development, particularly in the case of children plagued with autism spectrum disorder (ASD). In this paper, we introduce JEMImE, a serious game solution that aims at teaching children how to produce FEs. JEMImE is based on a FE recognition module that is learned on a large video corpus of children performing FEs. This module is validated and incorporated through multiple scenarios of gradual difficulty, ranging from a training phase where children have to perform the FEs on request, with or without an avatar model, to an in-context phase that involves many emotion-eliciting social situations with virtual characters. Arnaud Dapogny, Charline Grossard, Stéphanie Hun, Sylvie Serret, Jeremy Bourgeois, Hedy Jean-Marie, Pierre Foulon, Huaxiong Ding, Liming Chen 0002, Séverine Dubuisson, Ouriel Grynszpan, Kevin Bailly |
FG | 1 |
| 2018 | Confidence-Weighted Local Expression Predictions for Occlusion Handling in Expression Recognition and Action Unit Detection
Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
Int. J. Comput. Vis. | 1 |
| 2018 | Face alignment with cascaded semi-parametric deep greedy neural forests
Arnaud Dapogny, Kevin Bailly |
Pattern Recognit. Lett. | 1 |
| 2017 | Multi-Output Random Forests for Facial Action Unit DetectionabstractInternational audience Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
FG | 1 |
| 2017 | Sequential recognition of in-hand object shape using a collection of neural forestsabstractTactile object shape identification is important for robotic hands to perform dexterous manipulation. Most of the proposed approaches concentrate on specific object recognition. This limits their application into more realistic environments where a larger amount of objects are present. We present a method that performs object shape identification independently on the size and location of the object within the hand. This method allows sequential learning of new shapes. The method combines proprioceptive signatures and contact normal information to build a descriptor that reduces the impact of the size and pose of the object on recognition rate. Sequential training is performed with a collection of Neural Forests (NF). This allows adding new objects to the training set so that training the model from scratch is avoided. Extensive experiments reveal that the combination of multiple modalities (e.g. contact normals, proprioceptive information) is beneficial to the system accuracy, and that results for sequential learning are on par with its batch counterpart. This makes sequential training advantageous because it is less time consuming. Results showed that both techniques depict similar results and perform with at least 83% in a 7-shapes case scenario in a simulated environment. Experiments are made with a real shadow hand using a NF trained on simulated data. Alex Vásquez, Arnaud Dapogny, Kevin Bailly, Véronique Perdereau |
IROS | 2 |
| 2016 | On leveraging crowdsourced data for automatic perceived stress detectionabstractResorting to crowdsourcing platforms is a popular way to obtain annotations. Multiple potentially noisy answers can thus be aggregated to retrieve an underlying ground truth. However, it may be irrelevant to look for a unique ground truth when we ask crowd workers for opinions, notably when dealing with subjective phenomena such as stress. In this paper, we discuss how we can better use crowdsourced annotations with an application to automatic detection of perceived stress. Towards this aim, we first acquired video data from 44 subjects in a stressful situation and gathered answers to a binary question using a crowdsourcing platform. Then, we propose to integrate two measures derived from the set of gathered answers into the machine learning framework. First, we highlight that using the consensus level among crowd worker answers substantially increases classification accuracies. Then, we show that it is suitable to directly predict for each video the proportion of positive answers to the question from the different crowd workers. Hence, we propose a thorough study on how crowdsourced annotations can be used to enhance performance of classification and regression methods. Jonathan Aigrain, Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson, Marcin Detyniecki, Mohamed Chetouani |
ICMI | 2 |
| 2015 | Pairwise Conditional Random Forests for Facial Expression RecognitionabstractFacial expression can be seen as the dynamic variation of one's appearance over time. Successful recognition thus involves finding representations of high-dimensional spatiotemporal patterns that can be generalized to unseen facial morphologies and variations of the expression dynamics. In this paper, we propose to learn Random Forests from heterogeneous derivative features (e.g. facial fiducial point movements or texture variations) upon pairs of images. Those forests are conditioned on the expression label of the first frame to reduce the variability of the ongoing expression transitions. When testing on a specific frame of a video, pairs are created between this frame and the previous ones. Predictions for each previous frame are used to draw trees from Pairwise Conditional Random Forests (PCRF) whose pairwise outputs are averaged over time to produce robust estimates. As such, PCRF appears as a natural extension of Random Forests to learn spatio-temporal patterns, that leads to significant improvements over standard Random Forests as well as state-of-the-art approaches on several facial expression benchmarks. Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
ICCV | 1 |