EDBT 2026 Demo / reviewers in the wild / expert
Kevin Bailly
dblp:41/3712 · also Kévin Bailly
· DBLP profile ↗
55ranked-venue papers
7as first author
21since 2021 · last 2025
0000-0001-7802-3673ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 5Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NUPES: Non-Uniform Post-Training Quantization via Power Exponent SearchabstractDeep neural network (DNN) deployment has been confined to larger hardware devices due to their expensive computational requirements. This challenge has recently reached another scale with the emergence of large language models (LLMs). In order to reduce both their memory footprint and latency, a promising technique is quantization. It consists in converting floating point representations to low bit-width fixed point representations, usually by assuming a uniform mapping onto a regular grid. This process, referred to in the literature as uniform quantization, may however be ill-suited as most DNN weights and activations follow a bell-shaped distribution. This is even worse on LLMs whose weight distributions are known to exhibit large, high impact, outlier values. In this work, we propose an improvement over the most commonly adopted way to tackle this limitation in deep learning models quantization, namely, non-uniform quantization. NUPES leverages automorphisms to preserve the scalar multiplications. Such transformations are derived from power functions. However, the optimization of the exponent parameter and weight values remains a challenging and novel problem which could not be solved with previous post training optimization techniques which only learn to round up or down weight values in order to preserve the predictive function. We circumvent this limitation with a new paradigm: learning new quantized weights over the entire quantized space. Similarly, we enable the optimization of the power exponent, i.e. the optimization of the quantization operator itself during training by alleviating all the numerical instabilities. The resulting predictive function is compatible with integer-only low-bit inference. We show the ability of the method to achieve state-of-the-art compression rates in both, data-free and data-driven configurations. Our empirical benchmarks highlight the ability of NUPES to circumvent the limitations of previous post-training quantization techniques on transformers and large language models in particular. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Network Memory Footprint Compression Through Jointly Learnable Codebooks and MappingsabstractThe massive interest in deep neural networks (DNNs) for both computer vision and natural language processing has been sparked by the growth in computational power. However, this led to an increase in the memory footprint, to a point where it can be challenging to simply load a model on commodity devices such as mobile phones. To address this limitation, quantization is a favored solution as it maps high precision tensors to a low precision, memory efficient format. In terms of memory footprint reduction, its most effective variants are based on codebooks. These methods, however, suffer from two limitations. First, they either define a single codebook for each tensor, or use a memory-expensive mapping to multiple codebooks. Second, gradient descent optimization of the mapping favors jumps toward extreme values, hence not defining a proximal search. In this work, we propose to address these two limitations. First, we initially group similarly distributed neurons and leverage the re-ordered structure to either apply different scale factors to the different groups, or map weights that fall in these groups to several codebooks, without any mapping overhead. Second, stemming from this initialization, we propose a joint learning of the codebook and weight mappings that bears similarities with recent gradient-based post-training quantization techniques. Third, drawing estimation from straight-through estimation techniques, we introduce a novel gradient update definition to enable a proximal search of the codebooks and their mappings. The proposed jointly learnable codebooks and mappings (JLCM) method allows a very efficient approximation of any DNN: as such, a Llama 7B can be compressed down to 2Go and loaded on 5-year-old smartphones. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICLR | 3 |
| 2024 | Ontology-Guided Deep Metric Learning and Applications to Obstetrics
Jules Bonnard, Arnaud Dapogny, Ferdinand Dhombres, Kevin Bailly |
ICPR (9) | 4 |
| 2024 | PIPE: Parallelized inference through ensembling of residual quantization expansions
Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
Pattern Recognit. | 3 |
| 2024 | Prior-Guided Attribution of Deep Neural Networks for Obstetrics and GynecologyabstractObstetrics and gynecology (OB/GYN) are areas of medicine that specialize in the care of women during pregnancy and childbirth and in the diagnosis of diseases of the female reproductive system. Ultrasound scanning has become ubiquitous in these branches of medicine, as breast or fetal ultrasound images can lead the sonographer and guide him through his diagnosis. However, ultrasound scan images require a lot of resources to annotate and are often unavailable for training purposes because of confidentiality reasons, which explains why deep learning methods are still not as commonly used to solve OB/GYN tasks as in other computer vision tasks. In order to tackle this lack of data for training deep neural networks in this context, we propose Prior-Guided Attribution (PGA), a novel method that takes advantage of prior spatial information during training by guiding part of its attribution towards these salient areas. Furthermore, we introduce a novel prior allocation strategy method to take into account several spatial priors at the same time while providing the model enough degrees of liberty to learn relevant features by itself. The proposed method only uses the additional information during training, without needing it during inference. After validating the different elements of the method as well as its genericity on a facial analysis problem, we demonstrate that the proposed PGA method constantly outperforms existing baselines on two ultrasound imaging OB/GYN tasks: breast cancer detection and scan plane detection with segmentation prior maps. Jules Bonnard, Arnaud Dapogny, Richard Zsamboki, Lucrezia De Braud, Davor Jurkovic, Kevin Bailly, Ferdinand Dhombres |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | RULe: Relocalization-Uniformization-Landmark Estimation Network for Real-Time Face Alignment in Degraded ConditionsabstractFace alignment refers to the process of estimating the position of a number of salient landmarks on face images or videos, such as mouth and eye corners, nose tip, etc. With the availability of large annotated databases and the rise of deep learning-based methods, face alignment as a domain has matured to a point where it can be applied in more or less unconstrained conditions, e.g. non-frontal head poses, presence of heavy make-up or partial occlusions. However, when considering real-case alignment on videos with possibly low frame rates, we need to make sure that the algorithms are robust to jittering of the face bounding box localization, low-resolution of the face crops, possible bad environmental lighting, brightness, and presence of noise. To tackle these issues, we propose RULe, a three-staged Relocalization-Uniformization-Landmark Estimation network. In the first stage, an initial loosely localized bounding box gets refined to output a well centered face crop, thus reducing the variability of the images prior to passing them to the subsequent stage. Then, in the second stage, the face style is uniformized (using adversarial learning as well as perceptual losses) to correct low resolution or variations of brightness/contrast. Finally, the third stage outputs a precise landmark estimation given such enhanced face crop using a cascaded compact model trained using hint-based knowledge distillation. We show through a variety of experiments that RULe achieves real-time face alignment with state-of-the-art precision in heavily degraded conditions. Arnaud Dapogny, Gauthier Tallec, Jules Bonnard, Edouard Yvinec, Kevin Bailly |
FG | 5 |
| 2023 | Adversarial Deep Multi-Task Learning Using Semantically Orthogonal Spaces and Application to Facial Attributes PredictionabstractDeep learning-based multi-task approaches usually rely on factorizing representation layers up to a certain point, where the network splits into several heads, each one addressing a specific task. Depending on the inter-task correlation, such naive model may or may not allow the tasks to benefit from each others. In this paper, we propose a novel Semantic Orthogonality Spaces (SOS) method for multi-task problems, where each task is predicted using the information from a common subspace that factorizes information among all tasks, as well as a task-specific subspace. We enforce orthogonality between these tasks by applying soft orthogonality constraints, as well as adversarially-learned semantic orthogonality objectives that ensures that predicting one task requires the specific information related to that task. We demonstrate the effectiveness of SOS on synthetic data, as well as for large-scale facial attributes prediction. In particular, we use SOS to craft a lightweight architecture that provides high-end accuracies on CelebA database. Arnaud Dapogny, Gauthier Tallec, Jules Bonnard, Edouard Yvinec, Kevin Bailly |
FG | 5 |
| 2023 | Fighting Over-Fitting with Quantization for Learning Deep Neural Networks on Noisy LabelsabstractThe rising performance of deep neural networks is often empirically attributed to an increase in the available computational power, which allows complex models to be trained upon large amounts of annotated data. However, increased model complexity leads to costly deployment of modern neural networks, while gathering such amounts of data requires huge costs to avoid label noise. In this work, we study the ability of compression methods to tackle both of these problems at once. We hypothesize that quantization-aware training, by restricting the expressivity of neural networks, behaves as a regularization. Thus, it may help fighting overfitting on noisy data while also allowing for the compression of the model at inference. We first validate this claim on a controlled test with manually introduced label noise. Furthermore, we also test the proposed method on Facial Action Unit detection, where labels are typically noisy due to the subtlety of the task. In all cases, our results suggests that quantization significantly improve the results compared with existing baselines, regularization as well as other compression methods. Gauthier Tallec, Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICIP | 4 |
| 2023 | Designing Strong Baselines for Ternary Neural Network Quantization through Support and Mass EqualizationabstractDeep neural networks (DNNs) offer the highest performance in a wide range of applications in computer vision. These results rely on over-parameterized backbones, which are expensive to run. This computational burden can be dramatically reduced by quantizing (in either data-free (DFQ), post-training (PTQ) or quantization-aware training (QAT) scenarios) floating point values to ternary values (2 bits, with each weight taking value in {−1,0,1}). In this context, we observe that rounding to nearest minimizes the expected error given a uniform distribution and thus does not account for the skewness and kurtosis of the weight distribution, which strongly affects ternary quantization performance. This raises the following question: shall one minimize the highest or average quantization error? To answer this, we design two operators: TQuant and MQuant that correspond to these respective minimization tasks. We show experimentally that our approach allows to significantly improve the performance of ternary quantization through a variety of scenarios in DFQ, PTQ and QAT and give strong insights to pave the way for future research in deep neural network quantization. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
ICIP | 3 |
| 2023 | PowerQuant: Automorphism Search for Non-Uniform Quantization
Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICLR | 4 |
| 2023 | REx: Data-Free Residual Quantization Error ExpansionabstractDeep neural networks (DNNs) are ubiquitous in computer vision and natural language processing, but suffer from high inference cost. This problem can be addressed by quantization, which consists in converting floating point operations into a lower bit-width format. With the growing concerns on privacy rights, we focus our efforts on data-free methods. However, such techniques suffer from their lack of adaptability to the target devices, as a hardware typically only supports specific bit widths. Thus, to adapt to a variety of devices, a quantization method shall be flexible enough to find good accuracy v.s. speed trade-offs for every bit width and target device. To achieve this, we propose REx, a quantization method that leverages residual error expansion, along with group sparsity.
We show experimentally that REx enables better trade-offs (in terms of accuracy given any target bit-width) on both convnets and transformers for computer vision, as well as NLP models. In particular, when applied to large language models, we show that REx elegantly solves the outlier problem that hinders state-of-the-art quantization methods.
In addition, REx is backed off by strong theoretical guarantees on the preservation of the predictive function of the original model. Lastly, we show that REx is agnostic to the quantization operator and can be used in combination with previous quantization work. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 4 |
| 2023 | SPIQ: Data-Free Per-Channel Static Input QuantizationabstractComputationally expensive neural networks are ubiquitous in computer vision and solutions for efficient inference have drawn a growing attention in the machine learning community. Examples of such solutions comprise quantization, i.e. converting the processing values (weights and inputs) from floating point into integers e.g. int8 or int4. Concurrently, the rise of privacy concerns motivated the study of less invasive acceleration methods, such as data-free quantization of pre-trained models weights and activations. Previous approaches either exploit statistical information to deduce scalar ranges and scaling factors for the activations in a static manner, or dynamically adapt this range on-the-fly for each input of each layer (also referred to as activations): the latter generally being more accurate at the expense of significantly slower inference. In this work, we argue that static input quantization can reach the accuracy levels of dynamic methods by means of a per-channel input quantization scheme that allows one to more finely preserve cross-channel dynamics. We show through a thorough empirical evaluation on multiple computer vision problems (e.g. ImageNet classification, Pascal VOC object detection as well as CityScapes semantic segmentation) that the proposed method, dubbed SPIQ, achieves accuracies rivalling dynamic approaches with static-level inference speed, significantly outperforming state-of-the-art quantization methods on every benchmark. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
WACV | 4 |
| 2023 | RED++ : Data-Free Pruning of Deep Neural Networks via Input Splitting and Output MergingabstractPruning Deep Neural Networks (DNNs) is a prominent field of study in the goal of inference runtime acceleration. In this paper, we introduce a novel data-free pruning protocol RED++. Only requiring a trained neural network, and not specific to any particular DNN, we exploit an adaptive data-free scalar hashing which exhibits redundancies among neuron weight values. We study the theoretical and empirical guarantees on the preservation of the accuracy from the hashing as well as the expected pruning ratio resulting from the exploitation of said redundancies. We propose a novel data-free pruning technique of DNN layers which removes the input-wise redundant operations. This algorithm is straightforward, parallelizable and offers novel perspective on DNN pruning by shifting the burden of large computation to efficient memory access and allocation. We provide theoretical guarantees on RED++ performance and empirically demonstrate its superiority over other data-free pruning methods and its competitiveness with data-driven ones on ResNets, MobileNets, and EfficientNets. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | THIN: THrowable Information Networks and Application for Facial Expression Recognition in the WildabstractFor a number of machine learning problems, an exogenous variable can be identified such that it heavily influences the appearance of the different classes, and an ideal classifier should be invariant to this variable. An example of such exogenous variable is identity if facial expression recognition (FER) is considered. In this paper, we propose a dual exogenous/endogenous representation. The former captures the exogenous variable whereas the second one models the task at hand (e.g. facial expression). We design a prediction layer that uses a tree-gated deep ensemble conditioned by the exogenous representation. We also propose an exogenous dispelling loss to remove the exogenous information from the endogenous representation. Thus, the exogenous information is used two times in a throwable fashion, first as a conditioning variable for the target task, and second to create invariance within the endogenous representation. We call this method THIN, standing for THrowable Information Networks. We experimentally validate THIN in several contexts where an exogenous information can be identified, such as digit recognition under large rotations and shape recognition at multiple scales. We also apply it to FER with identity as the exogenous variable. We demonstrate that THIN significantly outperforms state-of-the-art approaches on several challenging datasets. Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Multi-Order Networks for Action Unit DetectionabstractAction Units (AU) are muscular activations used to describe facial expressions. Therefore accurate AU recognition unlocks unbiaised face representation which can improve face-based affective computing applications. From a learning standpoint AU detection is a multi-task problem with strong inter-task dependencies. To solve such problem, most approaches either rely on weight sharing, or add explicit dependency modelling by decomposing the joint task distribution using Bayes chain rule. If the latter strategy yields comprehensive inter-task relationships modelling, it requires imposing an arbitrary order into an unordered task set. Crucially, this ordering choice has been identified as a source of performance variations. In this paper, we present Multi-Order Network (MONET), a multi-task method with joint task order optimization. MONET uses a differentiable order selection to jointly learn task-wise modules with their optimal chaining order. Furthermore, we introduce warmup and order dropout to enhance order selection by encouraging order exploration. Experimentally, we first demonstrate MONET capacity to retrieve the optimal order in a toy environment. Second, we validate MONET architecture by showing that MONET outperforms existing multi-task baselines on multiple attribute detection problems chosen for their wide range of dependency settings. More importantly, we demonstrate that MONET significantly extends state-of-the-art performance in AU detection. Gauthier Tallec, Arnaud Dapogny, Kevin Bailly |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Neural Architecture Search for Fracture ClassificationabstractThe adoption by radiologists of deep-learning based solutions to the bone fracture problem has helped improved diagnostic performances and patient care. The base models behind these tools were initially designed to solve problems on natural images, favoring transfer learning between standard image datasets and sets of radiographs. Those architectures could yet be made more specific to radiographs using neural architecture search (NAS). Unfortunately, current NAS approaches do not benefit from transfer learning. In this paper, we introduce an efficient scheme to exploit transfer learning when performing NAS. Using our approach, we validate the architecture tailoring paradigm to radiographs. On a custom fracture classification task, we find a new model with improved performances and reduced computational overhead over its counterparts pre-trained on ImageNet. Aloïs Pourchot, Kevin Bailly, Alexis Ducarouge, Olivier Sigaud |
ICIP | 2 |
| 2022 | Privileged Attribution Constrained Deep Networks for Facial Expression RecognitionabstractFacial Expression Recognition (FER) is crucial in many research domains because it enables machines to better understand human behaviours. FER methods face the problems of relatively small datasets and noisy data that don’t allow classical networks to generalize well. To alleviate these issues, we guide the model to concentrate on specific facial areas like the eyes, the mouth or the eyebrows, which we argue are decisive to recognise facial expressions. We propose the Privileged Attribution Loss (PAL), a method that directs the focus of the model towards the most salient facial regions by encouraging its attribution maps to correspond to a heatmap formed by facial landmarks. Furthermore, we introduce several channel strategies that allow the model to have more degrees of freedom. The proposed method is independent of the backbone architecture and doesn’t need additional semantic information at test time. Finally, experimental results show that the proposed PAL method outperforms current state-of-the-art methods on both RAF-DB and AffectNet. Jules Bonnard, Arnaud Dapogny, Ferdinand Dhombres, Kevin Bailly |
ICPR | 4 |
| 2022 | To Fold or Not to Fold: a Necessary and Sufficient Condition on Batch-Normalization Layers FoldingabstractBatch-Normalization (BN) layers have become fundamental components in the evermore complex deep neural network architectures. Such models require acceleration processes for deployment on edge devices. However, BN layers add computation bottlenecks due to the sequential operation processing: thus, a key, yet often overlooked component of the acceleration process is BN layers folding. In this paper, we demonstrate that the current BN folding approaches are suboptimal in terms of how many layers can be removed. We therefore provide a necessary and sufficient condition for BN folding and a corresponding optimal algorithm. The proposed approach systematically outperforms existing baselines and allows to dramatically reduce the inference time of deep neural networks. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
IJCAI | 3 |
| 2022 | SInGE: Sparsity via Integrated Gradients Estimation of Neuron RelevanceabstractThe leap in performance in state-of-the-art computer vision methods is attributed to the development of deep neural networks. However it often comes at a computational price which may hinder their deployment. To alleviate this limitation, structured pruning is a well known technique which consists in removing channels, neurons or filters, and is commonly applied in order to produce more compact models. In most cases, the computations to remove are selected based on a relative importance criterion. At the same time, the need for explainable predictive models has risen tremendously and motivated the development of robust attribution methods that highlight the relative importance of pixels of an input image or feature map. In this work, we discuss the limitations of existing pruning heuristics, among which magnitude and gradient-based methods. We draw inspiration from attribution methods to design a novel integrated gradient pruning criterion, in which the relevance of each neuron is defined as the integral of the gradient variation on a path towards this neuron removal. Furthermore, We propose an entwined DNN pruning and fine-tuning flowchart to better preserve DNN accuracy while removing parameters. We show through extensive validation on several datasets, architectures as well as pruning scenarios that the proposed method, dubbed SInGE, significantly outperforms existing state-of-the-art DNN pruning methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 4 |
| 2022 | An extensive appraisal of weight-sharing on the NAS-Bench-101 benchmark
Aloïs Pourchot, Kevin Bailly, Alexis Ducarouge, Olivier Sigaud |
Neurocomputing | 2 |
| 2021 | RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksabstractDeep Neural Networks (DNNs) are ubiquitous in today's computer vision landscape, despite involving considerable computational costs. The mainstream approaches for runtime acceleration consist in pruning connections (unstructured pruning) or, better, filters (structured pruning), both often requiring data to retrain the model. In this paper, we present RED, a data-free, unified approach to tackle structured pruning. First, we propose a novel adaptive hashing of the scalar DNN weight distribution densities to increase the number of identical neurons represented by their weight vectors. Second, we prune the network by merging redundant neurons based on their relative similarities, as defined by their distance. Third, we propose a novel uneven depthwise separation technique to further prune convolutional layers. We demonstrate through a large variety of benchmarks that RED largely outperforms other data-free pruning methods, often reaching performance similar to unconstrained, data-driven methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 4 |
| 2020 | MoDuL: Deep Modal and Dual Landmark-wise Gated Network for Facial Expression RecognitionabstractAutomatic facial expression recognition (FER) is a challenging computer vision problem that finds a number of applications in human-computer interaction. Most recent FER approaches are deep-learning based and involve the extraction of two types of features from face images: geometric features (e.g. distances between aligned facial landmarks) and appearance features extracted using convolutional neural networks applied on patches extracted around each landmark. In this paper, we explore the use of gating networks to learn an optimal combination of these two modalities (modal gate). Furthermore, we also design landmark-wise gates to adaptively weight each landmark as well as the corresponding patch contribution. The proposed MoDuL architecture achieves state-of-the-art results on several FER databases with negligible computational overhead. Sacha Bernheim, Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
FG | 4 |
| 2020 | Deep Entwined Learning Head Pose and Face Alignment Inside an Attentional Cascade with Doubly-Conditional fusionabstractHead pose estimation and face alignment constitute a backbone preprocessing for many applications relying on face analysis. While both are closely related tasks, they are generally addressed separately, e.g. by deducing the head pose from the landmark locations. In this paper, we propose to entwine face alignment and head pose tasks inside an attentional cascade. This cascade uses a geometry transfer network for integrating heterogeneous annotations to enhance landmark localization accuracy. Furthermore, we propose a doubly-conditional fusion scheme to select relevant feature maps, and regions thereof, based on a current head pose and landmark localization estimate. We empirically show the benefit of entwining head pose and landmark localization objectives inside our architecture, and that the proposed AC-DC model enhances the state-of-the-art accuracy on multiple databases for both face alignment and head pose estimation tasks. Arnaud Dapogny, Kevin Bailly, Matthieu Cord |
FG | 2 |
| 2020 | DeeSCo: Deep heterogeneous ensemble with Stochastic Combinatory loss for gaze estimationabstractFrom medical research to gaming applications, gaze estimation is becoming a valuable tool. While there exists a number of hardware-based solutions, recent deep learning-based approaches, coupled with the availability of large-scale databases, have allowed to provide a precise gaze estimate using only consumer sensors. However, there remains a number of questions, regarding the problem formulation, architectural choices and learning paradigms for designing gaze estimation systems in order to bridge the gap between geometry-based systems involving specific hardware and approaches using consumer sensors only. In this paper, we introduce a deep, end-to-end trainable ensemble of heatmap-based weak predictors for 2D/3D gaze estimation. We show that, through heterogeneous architectural design of these weak predictors, we can improve the decorrelation between the latter predictors to design more robust deep ensemble models. Furthermore, we propose a stochastic combinatory loss that consists in randomly sampling combinations of weak predictors at train time. This allows to train better individual weak predictors, with lower correlation between them. This, in turns, allows to significantly enhance the performance of the deep ensemble. We show that our Deep heterogeneous ensemble with Stochastic Combinatory loss (DeeSCo) outperforms state-of-the-art approaches for 2D/3D gaze estimation on multiple datasets. Edouard Yvinec, Arnaud Dapogny, Kevin Bailly |
FG | 3 |
| 2020 | Face and Gesture Analysis for Health InformaticsabstractThe goal of Face and Gesture Analysis for Health Informatics's workshop is to share and discuss the achievements as well as the challenges in using computer vision and machine learning for automatic human behavior analysis and modeling for clinical research and healthcare applications. The workshop aims to promote current research and support growth of multidisciplinary collaborations to advance this groundbreaking research. The meeting gathers scientists working in related areas of computer vision and machine learning, multi-modal signal processing and fusion, human centered computing, behavioral sensing, assistive technologies, and medical tutoring systems for healthcare applications and medicine. Zakia Hammal, Di Huang 0001, Kevin Bailly, Liming Chen 0002, Mohamed Daoudi |
ICMI | 3 |
| 2020 | The POTUS Corpus, a Database of Weekly Addresses for the Study of Stance in Politics and Virtual AgentsabstractOne of the main challenges in the field of Embodied Conversational Agent (ECA) is to generate socially believable agents. The common strategy for agent behaviour synthesis is to rely on dedicated corpus analysis. Such a corpus is composed of multimedia files of socio-emotional behaviors which have been annotated by external observers. The underlying idea is to identify interaction information for the agent’s socio-emotional behavior by checking whether the intended socio-emotional behavior is actually perceived by humans. Then, the annotations can be used as learning classes for machine learning algorithms applied to the social signals. This paper introduces the POTUS Corpus composed of high-quality audio-video files of political addresses to the American people. Two protagonists are present in this database. First, it includes speeches of former president Barack Obama to the American people. Secondly, it provides videos of these same speeches given by a virtual agent named Rodrigue. The ECA reproduces the original address as closely as possible using social signals automatically extracted from the original one. Both are annotated for social attitudes, providing information about the stance observed in each file. It also provides the social signals automatically extracted from Obama’s addresses used to generate Rodrigue’s ones. Thomas Janssoone, Kevin Bailly, Gaël Richard, Chloé Clavel |
LREC | 2 |
| 2019 | Tree-gated Deep Regressor Ensemble For Face Alignment In The WildabstractFace alignment consists in aligning a shape model on a face in an image. It is an active domain in computer vision as it is a preprocessing for applications like facial expression recognition, face recognition and tracking, face animation, etc. Current state-of-the-art methods already perform well on "easy" datasets, i.e. those that present moderate variations in head pose, expression, illumination or partial occlusions, but may not be robust to "in-the-wild" data. In this paper, we address this problem by using an ensemble of deep regressors instead of a single large regressor. Furthermore, instead of averaging the ouputs of each regressor, we propose an adaptative weighting scheme that uses a tree-structured gate. Experiments on several challenging face datasets demonstrate that our approach outperforms the state-of-the-art methods. Estephe Arnaud, Arnaud Dapogny, Kevin Bailly |
FG | 3 |
| 2019 | DeCaFA: Deep Convolutional Cascade for Face Alignment in the WildabstractFace Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce an end-to-end deep convolutional cascade (DeCaFA) architecture for face alignment. Face Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce DeCaFA, an end-to-end deep convolutional cascade architecture for face alignment. DeCaFA uses fully-convolutional stages to keep full spatial resolution throughout the cascade. Between each cascade stage, DeCaFA uses multiple chained transfer layers with spatial softmax to produce landmark-wise attention maps for each of several landmark alignment tasks. Weighted intermediate supervision, as well as efficient feature fusion between the stages allow to learn to progressively refine the attention maps in an end-to-end manner. We show experimentally that DeCaFA significantly outperforms existing approaches on 300W, CelebA and WFLW databases. In addition, we show that DeCaFA can learn fine alignment with reasonable accuracy from very few images using coarsely annotated data. Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICCV | 3 |
| 2019 | Dynamic Pose-Robust Facial Expression Recognition by Multi-View Pairwise Conditional Random ForestsabstractAutomatic facial expression classification (FER) from videos is a critical problem for the development of intelligent human-computer interaction systems. Still, it is a challenging problem that involves capturing high-dimensional spatio-temporal patterns describing the variation of one's appearance over time. Such representation undergoes great variability of the facial morphology and environmental factors as well as head pose variations. In this paper, we use Conditional Random Forests to capture low-level expression transition patterns. More specifically, heterogeneous derivative features (e.g., feature point movements or texture variations) are evaluated upon pairs of images. When testing on a video frame, pairs are created between this current frame and previous ones and predictions for each previous frame are used to draw trees from Pairwise Conditional Random Forests (PCRF) whose pairwise outputs are averaged over time to produce robust estimates. Moreover, PCRF collections can also be conditioned on head pose estimation for multi-view dynamic FER. As such, our approach appears as a natural extension of Random Forests for learning spatio-temporal patterns, potentially from multiple viewpoints. Experiments on popular datasets show that our method leads to significant improvements over standard Random Forests as well as state-of-the-art approaches on several scenarios, including a novel multi-view video corpus generated from a publicly available database. Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | Investigating Deep Neural Forests for Facial Expression RecognitionabstractFacial Expression Recognition (FER) usually involves learning intermediate representations from high-dimensional, potentially noisy images, that can be adapted to unknown face morphologies. Moreover, it needs to fulfil the real-time constraint to be useful, e.g. for consumer robotics or healthcare systems. To tackle this issue, Random Forests (RFs) are convenient predictors, as the hard decisions at each node allow non-linear subdivisions of the space with a very fast evaluation runtime. However RFs are non-differenciable predictors, making it impossible to back-propagate the error down to upstream feature extraction layers (e.g. CNN). In this paper, we investigate the adaptation of deep Neural Forests (NFs) to FER. The latter allows to combine the best of the two worlds, i.e. a differentiable classification model that can be trained with backpropagation and stochastic gradient descent but has the runtime of a RF. We show that NFs provides competitive results for FER as compared to state-of-the-art approaches when trained upon various combinations of geometric/CNN-based features. Arnaud Dapogny, Kevin Bailly |
FG | 2 |
| 2018 | JEMImE: A Serious Game to Teach Children with ASD How to Adequately Produce Facial ExpressionsabstractBeing able to produce facial expressions (FEs) that are adequate given a social context is key to harmonious social development, particularly in the case of children plagued with autism spectrum disorder (ASD). In this paper, we introduce JEMImE, a serious game solution that aims at teaching children how to produce FEs. JEMImE is based on a FE recognition module that is learned on a large video corpus of children performing FEs. This module is validated and incorporated through multiple scenarios of gradual difficulty, ranging from a training phase where children have to perform the FEs on request, with or without an avatar model, to an in-context phase that involves many emotion-eliciting social situations with virtual characters. Arnaud Dapogny, Charline Grossard, Stéphanie Hun, Sylvie Serret, Jeremy Bourgeois, Hedy Jean-Marie, Pierre Foulon, Huaxiong Ding, Liming Chen 0002, Séverine Dubuisson, Ouriel Grynszpan, Kevin Bailly |
FG | 13 |
| 2018 | Confidence-Weighted Local Expression Predictions for Occlusion Handling in Expression Recognition and Action Unit Detection
Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
Int. J. Comput. Vis. | 2 |
| 2018 | Face alignment with cascaded semi-parametric deep greedy neural forests
Arnaud Dapogny, Kevin Bailly |
Pattern Recognit. Lett. | 2 |
| 2017 | Multi-Output Random Forests for Facial Action Unit DetectionabstractInternational audience Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
FG | 2 |
| 2017 | Sequential recognition of in-hand object shape using a collection of neural forestsabstractTactile object shape identification is important for robotic hands to perform dexterous manipulation. Most of the proposed approaches concentrate on specific object recognition. This limits their application into more realistic environments where a larger amount of objects are present. We present a method that performs object shape identification independently on the size and location of the object within the hand. This method allows sequential learning of new shapes. The method combines proprioceptive signatures and contact normal information to build a descriptor that reduces the impact of the size and pose of the object on recognition rate. Sequential training is performed with a collection of Neural Forests (NF). This allows adding new objects to the training set so that training the model from scratch is avoided. Extensive experiments reveal that the combination of multiple modalities (e.g. contact normals, proprioceptive information) is beneficial to the system accuracy, and that results for sequential learning are on par with its batch counterpart. This makes sequential training advantageous because it is less time consuming. Results showed that both techniques depict similar results and perform with at least 83% in a 7-shapes case scenario in a simulated environment. Experiments are made with a real shadow hand using a NF trained on simulated data. Alex Vásquez, Arnaud Dapogny, Kevin Bailly, Véronique Perdereau |
IROS | 3 |
| 2016 | On leveraging crowdsourced data for automatic perceived stress detectionabstractResorting to crowdsourcing platforms is a popular way to obtain annotations. Multiple potentially noisy answers can thus be aggregated to retrieve an underlying ground truth. However, it may be irrelevant to look for a unique ground truth when we ask crowd workers for opinions, notably when dealing with subjective phenomena such as stress. In this paper, we discuss how we can better use crowdsourced annotations with an application to automatic detection of perceived stress. Towards this aim, we first acquired video data from 44 subjects in a stressful situation and gathered answers to a binary question using a crowdsourcing platform. Then, we propose to integrate two measures derived from the set of gathered answers into the machine learning framework. First, we highlight that using the consensus level among crowd worker answers substantially increases classification accuracies. Then, we show that it is suitable to directly predict for each video the proportion of positive answers to the question from the different crowd workers. Hence, we propose a thorough study on how crowdsourced annotations can be used to enhance performance of classification and regression methods. Jonathan Aigrain, Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson, Marcin Detyniecki, Mohamed Chetouani |
ICMI | 3 |
| 2016 | Using Temporal Association Rules for the Synthesis of Embodied Conversational Agents with a Specific Stance
Thomas Janssoone, Chloé Clavel, Kevin Bailly, Gaël Richard |
IVA | 3 |
| 2016 | Real-time facial action unit intensity prediction with regularized metric learning
Jérémie Nicolle, Kevin Bailly, Mohamed Chetouani |
Image Vis. Comput. | 2 |
| 2015 | Pairwise Conditional Random Forests for Facial Expression RecognitionabstractFacial expression can be seen as the dynamic variation of one's appearance over time. Successful recognition thus involves finding representations of high-dimensional spatiotemporal patterns that can be generalized to unseen facial morphologies and variations of the expression dynamics. In this paper, we propose to learn Random Forests from heterogeneous derivative features (e.g. facial fiducial point movements or texture variations) upon pairs of images. Those forests are conditioned on the expression label of the first frame to reduce the variability of the ongoing expression transitions. When testing on a specific frame of a video, pairs are created between this frame and the previous ones. Predictions for each previous frame are used to draw trees from Pairwise Conditional Random Forests (PCRF) whose pairwise outputs are averaged over time to produce robust estimates. As such, PCRF appears as a natural extension of Random Forests to learn spatio-temporal patterns, that leads to significant improvements over standard Random Forests as well as state-of-the-art approaches on several facial expression benchmarks. Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson |
ICCV | 2 |
| 2014 | A Pose-Adaptive Constrained Local Model for Accurate Head Pose TrackingabstractRobust and precise face tracking under unconstrained imaging conditions is still a challenging task. Recently, the Constrained Local Model (CLM) framework has proven to be very powerful to track frontal and near frontal facial movements. In this paper, we introduce a Pose-Adaptive CLM which is able to accurately track large 3D head rotations. This model relies on two main parts: (1) an adaptive 3D Point Distribution Model that ensures consistency between a tracked point in the image and the corresponding point in the shape model and (2) an adaptive appearance model that deals with appearance variation of a point under different viewing angle. We present comparative experimental results highlighting the improvement in both robustness and accuracy of our method. We also introduce a new challenging dataset with accurate head pose annotation. Lucas Zamuner, Kevin Bailly, Erwan Bigorgne |
ICPR | 2 |
| 2014 | Impact of action unit detection in automatic emotion recognition
Thibaud Senechal, Kevin Bailly, Lionel Prevost |
Pattern Anal. Appl. | 2 |
| 2013 | Locating facial landmarks with binary map cross-correlationsabstractPrecise facial landmark localization in still images is a key step for many face analysis applications, such as biometrics or automatic emotion recognition. In this paper, we propose a framework for facial point detection in frontal and near-frontal images. We introduce a new appearance model based on binary map cross-correlations that efficiently uses LBP and LPQ in a localization context. Inclusion of shape-related constraints is performed by a nonparametric voting method using relational properties within triplets of points, designed to correct outliers without losing precision for accurately detected points. We tested our system's performance on the widely used as benchmark BioID database obtaining state-of-the-art results. We also discuss evaluation metrics used to compare facial landmarking systems and which have been mixed up in recent literature. Jérémie Nicolle, Kevin Bailly, Vincent Rapp, Mohamed Chetouani |
ICIP | 2 |
| 2013 | Multi-Kernel Appearance Model
Vincent Rapp, Kevin Bailly, Thibaud Senechal, Lionel Prevost |
Image Vis. Comput. | 2 |
| 2012 | Robust continuous prediction of human emotions using multiscale dynamic cuesabstractDesigning systems able to interact with humans in a natural manner is a complex and far from solved problem. A key aspect of natural interaction is the ability to understand and appropriately respond to human emotions. This paper details our response to the Audio/Visual Emotion Challenge (AVEC'12) whose goal is to continuously predict four affective signals describing human emotions (namely valence, arousal, expectancy and power). The proposed method uses log-magnitude Fourier spectra to extract multiscale dynamic descriptions of signals characterizing global and local face appearance as well as head movements and voice. We perform a kernel regression with very few representative samples selected via a supervised weighted-distance-based clustering, that leads to a high generalization power. For selecting features, we introduce a new correlation-based measure that takes into account a possible delay between the labels and the data and significantly increases robustness. We also propose a particularly fast regressor-level fusion framework to merge systems based on different modalities. Experiments have proven the efficiency of each key point of the proposed method and we obtain very promising results. Jérémie Nicolle, Vincent Rapp, Kevin Bailly, Lionel Prevost, Mohamed Chetouani |
ICMI | 3 |
| 2012 | Learning global cost function for face alignment
Kevin Bailly, Maurice Milgram, Philippe Phothisane, Erwan Bigorgne |
ICPR | 1 |
| 2012 | Facial Action Recognition Combining Heterogeneous Features via Multikernel LearningabstractThis paper presents our response to the first international challenge on facial emotion recognition and analysis. We propose to combine different types of features to automatically detect action units (AUs) in facial images. We use one multikernel support vector machine (SVM) for each AU we want to detect. The first kernel matrix is computed using local Gabor binary pattern histograms and a histogram intersection kernel. The second kernel matrix is computed from active appearance model coefficients and a radial basis function kernel. During the training step, we combine these two types of features using the recently proposed SimpleMKL algorithm. SVM outputs are then averaged to exploit temporal information in the sequence. To evaluate our system, we perform deep experimentation on several key issues: influence of features and kernel function in histogram-based SVM approaches, influence of spatially independent information versus geometric local appearance information and benefits of combining both, sensitivity to training data, and interest of temporal context adaptation. We also compare our results with those of the other participants and try to explain why our method had the best performance during the facial expression recognition and analysis challenge. Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2011 | Multiple kernel learning SVM and statistical validation for facial landmark detectionabstractIn this paper we present a robust and accurate method to detect 17 facial landmarks in expressive face images. We introduce a new multi-resolution framework based on the recent multiple kernel algorithm. Low resolution patches carry the global information of the face and give a coarse but robust detection of the desired landmark. High resolution patches, using local details, refine this location. This process is combined with a bootstrap process and a statistical validation, both improving the system robustness. Combining independent point detection and prior knowledge on the point distribution, the proposed detector is robust to variable lighting conditions and facial expressions. This detector is tested on several databases and the results reported can be compared favorably with the current state of the art point detectors. Vincent Rapp, Thibaud Senechal, Kevin Bailly, Lionel Prevost |
FG | 3 |
| 2011 | Combining AAM coefficients with LGBP histograms in the multi-kernel SVM framework to detect facial action unitsabstractThis study presents a combination of geometric and appearance features used to automatically detect Action Units in face images. We use one multi-kernel SVM for each Action Unit we want to detect. The first kernel matrix is computed using Local Gabor Binary Pattern (LGBP) histograms and a histogram intersection kernel. The second kernel matrix is computed from AAM coefficients and a RBF kernel. During the training step, we combine these two type s of features using the recent SimpleMKL algorithm. SVM outputs are then filtered to exploit dynamic relationships between Action Units. Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost |
FG | 5 |
| 2010 | Automatic Facial Action Detection Using Histogram Variation Between Emotional StatesabstractThis article presents an appearance based method to detect automatically facial actions. Our approach focuses on reducing features sensitivity to identity of the subject. We compute from an expressive image a Local Gabor Binary Pattern (LGBP) histogram and synthesize a LGBP histogram approaching the one we would compute on a neutral face. Difference between these two histograms are used as inputs of Support Vector Machine (SVM) binary detectors associated with a new kernel: the Histogram Difference Intersection (HDI) kernel. Experimental results carried out for 16 Action Units (AUs) on the benchmark Cohn-Kanade database can be compared favorably with two state-of-the-art methods. Thibaud Senechal, Kevin Bailly, Lionel Prevost |
ICPR | 2 |
| 2009 | Head Pose Estimation by a Stepwise Nonlinear Regression
Kevin Bailly, Maurice Milgram, Philippe Phothisane |
CAIP | 1 |
| 2009 | Head pan angle estimation by a nonlinear regression on selected featuresabstractHead pose is a crucial step for numerous face applications such as gaze tracking and face recognition. In this paper, we introduce a new method to learn the mapping between a set of features and the corresponding head pose. It combines a filter based feature selection and a Generalized Regression Neural Network where inputs are sequentially selected through a boosting process. We propose the Fuzzy Functional Criterion, a new filter used to select relevant features. At each step, features are evaluated using weights on examples computed using the error produced by the neural network at the previous step. This boosting strategy helps to focus on hard examples and selects a set of complementary features. Results are compared with two state-of-the-art methods on the Pointing 04 database. Kevin Bailly, Maurice Milgram |
ICIP | 1 |
| 2009 | BISAR: Boosted input selection algorithm for regressionabstractWe present in this paper a new regression method adapted to problems dealing with a huge set of potential features like in pattern recognition. This method combines a boosted forward feature selection algorithm and a generalized regression neural network. The feature selection uses a new criterion, the fuzzy functional criterion, to evaluate the relevance of each feature. It is well suited to measure to what extent a random variable y can be viewed as a function of another random variable x. We explain how this measure is more appropriate than the classical mutual information. At each step, features are evaluated using weights on examples computed from the error produced by the neural network at the previous step. This boosting strategy helps our system to focus on hard examples during the feature selection process. The application is head pose estimation, a challenging problem in pattern recognition. Test are carried out on the commonly used Pointing 04 database and compared with state-of-the-art results. Kevin Bailly, Maurice Milgram |
IJCNN | 1 |
| 2009 | Boosting feature selection for Neural Network based regression
Kevin Bailly, Maurice Milgram |
Neural Networks | 1 |
| 2008 | Head Pose Determination Using Synthetic Images
Kevin Bailly, Maurice Milgram |
ACIVS | 1 |
| 2008 | Recursive Shape and Pose Determination Using Deformable Model
Kevin Bailly, Maurice Milgram |
CIARP | 1 |