Renaud Séguier

dblp:27/2387 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-7199-7563ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 1
YearPublicationVenuePosition
2026 Dyna-Westdrive - VR-Based Multimodal Dataset for Emotion Recognition
abstract
National audience
Amdjed Belaref, Nathanel Vertchik, Quentin Guay, Zineb Noumir, Renaud Séguier
FG5
2025 JanusGAN: GANs Disentangled Editing with Two Discriminators
abstract
Generative Adversarial Networks (GANs) have found applications in image editing. However, GANs tend to synthesize and manipulate global features, such as age, rather than focusing on local features such as facial wrinkles. Consequently, when a specific wrinkle is edited, all age-related features change as well. This paper proposes a new method that allows a GAN to learn a specific global or local disentangled edit. The method involves fine-tuning a pre-trained GAN using two discriminators, each trained on a specific dataset representing distinct states of a single disentangled feature. This approach facilitates the generator’s ability to learn features along a defined editing direction within the latent space. Importantly, to avoid interfering with prior GAN knowledge, the editing direction is defined in the meaningless dimension of the GAN latent space. Although our primary focus is on local editing, our method can be extended to global features, such as age editing. Quantitative and qualitative results show that our method provides a better balance between feature accuracy and disentanglement than other state-of-the-art methods for both local and global features. The code is available on GitHub: https://github.com/Neilstid/JanusGAN.
Neil Farmer, Catherine Soladié, Gabriel Cazorla, Renaud Séguier
FG4
2025 A vector quantized masked autoencoder for audiovisual speech emotion recognition
abstract
An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised multimodal model that leverages masked autoencoders to learn representations of audiovisual speech without labels. The model includes vector quantized variational autoencoders that compress raw audio and visual speech data into discrete tokens. The audiovisual speech tokens are used to train a multimodal masked autoencoder that consists of an encoder–decoder architecture with attention mechanisms. The model is designed to extract both local (i.e., at the frame level) and global (i.e., at the sequence level) representations of audiovisual speech. During self-supervised pre-training, the VQ-MAE-AV model is trained on a large-scale unlabeled dataset of audiovisual speech, for the task of reconstructing randomly masked audiovisual speech tokens and with a contrastive learning strategy. During this pre-training, the encoder learns to extract a representation of audiovisual speech that can be subsequently leveraged for emotion recognition. During the supervised fine-tuning stage, a small classification model is trained on top of the VQ-MAE-AV encoder for an emotion recognition task. The proposed approach achieves state-of-the-art emotion recognition results across several datasets in both controlled and in-the-wild conditions. • We present a self-supervised model for audiovisual speech emotion recognition. • The model operates on discrete audiovisual speech tokens. • A multimodal masked autoencoder with attention fuses the audio and visual modalities. • The model achieves state-of-the-art audiovisual speech emotion recognition results. • Ablation studies reveal the importance of each model component.
Samir Sadok, Simon Leglaive, Renaud Séguier
Comput. Vis. Image Underst.3
2024 Pivotal Tuning Editing: Towards Disentangled Wrinkle Editing with GANs
abstract
Generative Adversarial Networks (GANs) enable image editing by manipulating image features. However, these manipulations still lack disentanglement. For example, when a specific wrinkle is edited, other age-related features or facial expressions are often changed as well. This paper proposes a new method for disentangled editing. The presented approach is based on two pivot images that allow learning an editing direction for an input image. These pivots are based on a real image (the input) and a synthetic modification of the real image along the desired editing direction. Although our primary focus is on wrinkle editing applications, our method can be extended to other editing tasks, such as hair color or lipstick editing. Qualitative and quantitative results show that our Pivotal Tuning Editing (PTE) provides a higher level of disentanglement and a more realistic editing than state-of-the-art methods. The code is available on GitHub1.
Neil Farmer, Catherine Soladié, Gabriel Cazorla, Renaud Séguier
FG4
2024 A multimodal dynamical variational autoencoder for audiovisual speech representation learning
abstract
High-dimensional data such as natural images or speech signals exhibit some form of regularity, preventing their dimensions from varying independently. This suggests that there exists a lower dimensional latent representation from which the high-dimensional observed data were generated. Uncovering the hidden explanatory features of complex data is the goal of representation learning, and deep latent variable generative models have emerged as promising unsupervised approaches. In particular, the variational autoencoder (VAE) which is equipped with both a generative and an inference model allows for the analysis, transformation, and generation of various types of data. Over the past few years, the VAE has been extended to deal with data that are either multimodal or dynamical (i.e., sequential). In this paper, we present a multimodal and dynamical VAE (MDVAE) applied to unsupervised audiovisual speech representation learning. The latent space is structured to dissociate the latent dynamical factors that are shared between the modalities from those that are specific to each modality. A static latent variable is also introduced to encode the information that is constant over time within an audiovisual speech sequence. The model is trained in an unsupervised manner on an audiovisual emotional speech dataset, in two stages. In the first stage, a vector quantized VAE (VQ-VAE) is learned independently for each modality, without temporal modeling. The second stage consists in learning the MDVAE model on the intermediate representation of the VQ-VAEs before quantization. The disentanglement between static versus dynamical and modality-specific versus modality-common information occurs during this second training stage. Extensive experiments are conducted to investigate how audiovisual speech latent factors are encoded in the latent space of MDVAE. These experiments include manipulating audiovisual speech, audiovisual facial image denoising, and audiovisual speech emotion recognition. The results show that MDVAE effectively combines the audio and visual information in its latent space. They also show that the learned static representation of audiovisual speech can be used for emotion recognition with few labeled data, and with better accuracy compared with unimodal baselines and a state-of-the-art supervised model based on an audiovisual transformer architecture.
Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier
Neural Networks5
2023 Exploring Mental Prototypes by an Efficient Interdisciplinary Approach: Interactive Microbial Genetic Algorithm
abstract
Facial expression-based technologies have flooded our daily lives. However, most technologies are limited to Ekman's basic facial expressions and rarely deal with more than ten emotional states. This is not only due to the lack of prototypes for complex emotions but also the time-consuming and laborious task of building an extensive labeled database. To remove these obstacles, we were inspired by a psychophysical approach for affective computing, so-called the reverse correlation process (RevCor), to extract mental prototypes of what a given emotion should look like for an observer. We proposed a novel, efficient, and interdisciplinary approach called Interactive Microbial Genetic Algorithm (IMGA) by integrating the concepts of RevCor into an interactive genetic algorithm (IGA). Our approach achieves four challenges: online feedback loop, expertise-free, velocity, and diverse results. Experimental results show that for each observer, with limited trials, our approach can provide diverse mental prototypes for both basic emotions and emotions that are not available in existing deep-learning databases. Our work is available at https://yansen0508.github.io/Interactive-Microbial-Genetic-Algorithm/.
Sen Yan 0004, Catherine Soladié, Renaud Séguier
FG3
2023 Motion-DVAE: Unsupervised learning for fast human motion denoising
abstract
Pose and motion priors are crucial for recovering realistic and accurate human motion from noisy observations. Substantial progress has been made on pose and shape estimation from images, and recent works showed impressive results using priors to refine frame-wise predictions. However, a lot of motion priors only model transitions between consecutive poses and are used in time-consuming optimization procedures, which is problematic for many applications requiring real-time motion capture. We introduce Motion-DVAE, a motion prior to capture the short-term dependencies of human motion. As part of the dynamical variational autoencoder (DVAE) models family, Motion-DVAE combines the generative capability of VAE models and the temporal modeling of recurrent architectures. Together with Motion-DVAE, we introduce an unsupervised learned denoising method unifying regression- and optimization-based approaches in a single framework for real-time 3D human pose estimation. Experiments show that the proposed approach reaches competitive performance with state-of-the-art methods while being much faster.
Guénolé Fiche, Simon Leglaive, Xavier Alameda-Pineda, Renaud Séguier
MIG4
2023 SwimXYZ: A large-scale dataset of synthetic swimming motions and videos
abstract
Technologies play an increasingly important role in sports and become a real competitive advantage for the athletes who benefit from it. Among them, the use of motion capture is developing in various sports to optimize sporting gestures. Unfortunately, traditional motion capture systems are expensive and constraining. Recently developed computer vision-based approaches also struggle in certain sports, like swimming, due to the aquatic environment. One of the reasons for the gap in performance is the lack of labeled datasets with swimming videos. In an attempt to address this issue, we introduce SwimXYZ, a synthetic dataset of swimming motions and videos. SwimXYZ contains 3.4 million frames annotated with ground truth 2D and 3D joints, as well as 240 sequences of swimming motions in the SMPL parameters format. In addition to making this dataset publicly available, we present use cases for SwimXYZ in swimming stroke clustering and 2D pose estimation.
Guénolé Fiche, Vincent Sevestre, Camila Gonzalez-Barral, Simon Leglaive, Renaud Séguier
MIG5
2023 MES-Loss: Mutually equidistant separation metric learning loss function
Yasser Boutaleb, Catherine Soladié, Nam-Duong Duong, Amine Kacete, Jérôme Royan, Renaud Séguier
Pattern Recognit. Lett.6
2023 Learning and controlling the source-filter representation of speech with a variational autoencoder
abstract
National audience
Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier
Speech Commun.5
2023 Local Temporal Pattern and Data Augmentation for Spotting Micro-Expressions
abstract
Micro-expressions (MEs) are very important nonverbal communication clues. However, due to their local and short nature, spotting them is challenging. In this article, we address this problem by using a dedicated local and temporal pattern (LTP) of facial movement. This pattern has a specific shape (an S-pattern) when MEs are displayed. Thus, by using a classic classification algorithm (SVM), MEs can be distinguished from other facial movements. We also propose a global final fusion analysis covering the whole face to improve the distinction between ME (local) and head (global) movements. However, the learning of S-patterns is limited by the small number of ME databases and the low volume of ME samples. Hammerstein models (HMs) are known to effectively approximate muscle movements. By approximating each S-pattern with an HM, we can both filter out outliers and generate new similar S-patterns. In this way, we augment the dataset for S-pattern training and improve the ability to differentiate MEs from other movements. The spotting results, performed in the CASMEI and CASMEII databases, show that our proposed LTP outperforms the most popular spotting method in terms of the F1-score. Adding a fusion process and data augmentation improves the spotting performance even further.
Jingting Li 0001, Catherine Soladié, Renaud Séguier
IEEE Trans. Affect. Comput.3
2021 Micro-expression recognition from local facial regions
Mouath Aouayeb, Wassim Hamidouche, Catherine Soladié, Kidiyo Kpalma, Renaud Séguier
Signal Process. Image Commun.5
2020 Realistic Transformation of Facial and Vocal Smiles in Real-Time Audiovisual Streams
abstract
Research in affective computing and cognitive science has shown the importance of emotional facial and vocal expressions during human-computer and human-human interactions. But, while models exist to control the display and interactive dynamics of emotional expressions, such as smiles, in embodied agents, these techniques can not be applied to video interactions between humans. In this work, we propose an audiovisual smile transformation algorithm able to manipulate an incoming video stream in real-time to parametrically control the amount of smile seen on the user's face and heard in their voice, while preserving other characteristics such as the user's identity or the timing and content of the interaction. The transformation is composed of separate audio and visual pipelines, both based on a warping technique informed by real-time detection of audio and visual landmarks. Taken together, these two parts constitute a unique audiovisual algorithm which, in addition to providing simultaneous real-time transformations of a real person's face and voice, allows to investigate the integration of both modalities of smiles in real-world social interactions.
Pablo Arias 0003, Catherine Soladié, Oussema Bouafif, Axel Röbel, Renaud Séguier, Jean-Julien Aucouturier
IEEE Trans. Affect. Comput.5
2020 Unsupervised Adaptation of a Person-Specific Manifold of Facial Expressions
abstract
In order to analyze expressions that are different from the prototypic expressions defined by Ekman, manifold learning has been proposed to build person-specific continuous representations of facial expressions. Yet, it is still a challenging problem to build such a manifold with no prior knowledge on the morphology of the subject. Here, we propose a method to build a person-specific manifold of facial expressions able to adapt to the morphology of the subject in an unsupervised manner. The manifold is initialized with the facial landmarks of the neutral face and 5 synthesized basic expressions. Our first contribution is to detect automatically the neutral face of the subject so that we can build the manifold in an unsupervised manner. Our second and main contribution is to adapt in an unsupervised manner the initialized manifold to the morphology of the subject by detecting the real basic expressions of the subject while maintaining constraints in the manifold. Our third contribution is to perform the adaptation on spontaneous expressions with typical head pose variation for human-computer interaction. The experiments show that the adaptation works well on posed expressions and that the constraints for the adaptation on spontaneous expressions is efficient when head pose variation is considered.
Raphaël Weber, Vincent Barrielle, Catherine Soladié, Renaud Séguier
IEEE Trans. Affect. Comput.4
2019 Spotting Micro-Expressions on Long Videos Sequences
abstract
This paper presents two methods for the first Micro-Expression Spotting Challenge 2019 by evaluating local temporal pattern (LTP) and local binary pattern (LBP) on two most recent databases, i.e. SAMM and CAS(ME)2. First we propose LTP-ML method as the baseline results for the challenge and then we compare the results with the LBP-χ2-distance method. The LTP patterns are extracted by applying PCA in a temporal window on several facial local regions. The micro-expression sequences are then spotted by a local classification of LTP and a global fusion. The LBP-χ2-distance method is to compare the feature difference by calculating χ2distance of LBP in a time window, the facial movements are then detected with a threshold. The performance is evaluated by Leave-One-Subject-Out cross validation. The overlap frames are used to determine the True Positives and the metric F1-score is used to compare the spotting performance of the databases. The F1-score of LTP-ML result for SAMM and CAS(ME)2are 0.0316 and 0.0179, respectively. The results show our proposed LTP-ML method outperformed LBP-χ2-distance method in terms of F1-score on both databases.
Jingting Li 0001, Catherine Soladié, Renaud Séguier, Moi Hoon Yap
FG3
2019 Face aging simulation with a new wrinkle oriented active appearance model
abstract
The use of computer simulation to understand how human faces age has been a growing area of research since decades. It has been applied to the search for missing children as well as to the fields of entertainment, cosmetics and dermatology research. Our objective is to elaborate a model for the age-related changes of visual cues which affect the perception of age, so that we may better predict them. Traditional approaches based on the Active Appearance Model (AAM) tend to blurry appearance and wipe out texture details such as wrinkles. We introduce Wrinkle Oriented Active Appearance Model (WOAAM) where a new channel is added to the AAM dedicated to analyze wrinkles. Firstly, we propose to represent both the shape and texture of each wrinkle on a face by a compact and interpretable vector. Afterwards, to model the distribution of wrinkles on a face, we introduce a new way to approximate an empiric joint probability density by creating an ensemble of joint probability densities estimated by Kernel Density Estimation. Finally, we show how to create new samples from such an ensemble of densities, and thus synthesize new plausible wrinkles. In comparison to other methods which add wrinkles at post-processing level, our method fully integrates them in AAM. Thereby, the wrinkles generated are statistically representative of a specific age in terms of number, length, shape and intensity. With an age estimation Convolutional Neural Network, we found that age-progressed faces produced by the WOAAM better reduces the gap between the expected age and the estimated age than those produced by a classic AAM.
Victor Martin, Renaud Séguier, Aurélie Porcheron, Frédérique Morizot
Multim. Tools Appl.2
2018 LTP-ML: Micro-Expression Detection by Recognition of Local Temporal Pattern of Facial Movements
abstract
The Micro-expressions (MEs) carry specific nonverbal information, for example the facial movement caused by pain. However, as a consequence of their local and short nature, it is difficult to detect MEs. This paper presents a novel detection method by recognizing a local and temporal pattern (LTP) of facial movement. In our system, with the purpose of improving the detection accuracy, temporal local features are generated from the video in a sliding window of 300ms (mean duration of a ME). These features are extracted from a projection in PCA space and form a specific pattern during ME which is the same for all MEs. Using a classical classification algorithm (SVM), MEs are then distinguished from other facial movements. Finally, a global fusion analysis is applied on the whole face to eliminate false positives. Experiments are performed on two databases: CASME I and CASME II. The detection results show that the proposed method outperforms the most popular detection method in terms of F1-score according to the analysis of multiple metrics.
Jingting Li 0001, Catherine Soladié, Renaud Séguier
FG3
2018 A survey on face modeling: building a bridge between face analysis and synthesis
Hanan Salam, Renaud Séguier
Vis. Comput.2
2016 Unconstrained Gaze Estimation Using Random Forest Regression Voting
Amine Kacete, Renaud Séguier, Michel Collobert, Jérôme Royan
ACCV (3)2
2016 Real-time eye pupil localization using Hough regression forest
abstract
Eyes are one of the most salient features of the human face, and the location of the pupil allows access to important information which can be used in several computer vision applications. Several commercial eye-trackers can estimate with good accuracy the pupil location, but need complex hardware specifications and a controlled user environment (high eye image resolution, good illumination, small head pose variations) making these solutions difficult to use in an arbitrary environment. In this paper, we present an approach based on Hough randomized regression trees. We demonstrate, by several evaluations on challenging public datasets that our approach is very robust to illumination, scale, eye movements and high head pose variations and yields a significant improvement compared to a wide range of state-of-the-art methods.
Amine Kacete, Jérôme Royan, Renaud Séguier, Michel Collobert, Catherine Soladié
WACV3
2015 3D facial clone based on depth patches
abstract
3D face clones can be used in many areas such as Human-Computer Interaction and as preprocessing in applications, such as emotion analysis. However, such clones should be structured and the model facial shape accurately while keeping the attributes of individuals. A structured mesh is a mesh with a known semantic and topological structure. We use a face model designed from a database of 3D face examples. These global models can produce structured clones but they do not often retain the specifics of the analyzed person. Indeed, methods using models are very dependent on their databases. In our technique, we use an RGB-D sensor to get the attributes of individuals and a 3D Morphable Face Model to mark facial shape. We reverse the process classically used: we first perform fitting and then data fusion. For each depth frame, we retain the suitable data parts called Patches. This selection is performed using a distance error and the direction of the normal vectors. Depending on the location, we merge either sensor data or 3D Morphable Face Model data. We compare our method with state of the art fitting processes. The qualitative and quantitative tests show that our results are more accurate than an current fitting method and our clone has both the attributes of the person and the shape of the face well modeled.
Jérôme Manceau, Catherine Soladié, Renaud Séguier
VCIP3
2013 Bilinear decomposition for blended expressions representation
abstract
This paper proposes a new method for the analysis of blended expressions with varying intensity. The method is based on an asymmetric bilinear model learned on a small amount of expressions. In the resulting expression space, a blended unknown expression has a signature, that can be interpreted as a mixture of the basic expressions used in the creation of the space. Three methods are compared: a traditional method based on active appearance vectors, the asymmetric bilinear model on person-independent appearance vectors and the asymmetric bilinear model on person-specific appearance vectors. Experimental results on the recognition of 14 blended unknown expressions show the relevance of the bilinear models compared to appearance-based methods and the robustness of the person-specific models according to the types of parameters (shape and/or texture).
Catherine Soladié, Renaud Séguier, Nicolas Stoiber
VCIP2
2013 Invariant representation of facial expressions for blended expression recognition on unknown subjects
Catherine Soladié, Nicolas Stoiber, Renaud Séguier
Comput. Vis. Image Underst.3
2012 A multi-texture approach for estimating iris positions in the eye using 2.5D Active Appearance Models
abstract
This paper describes a new approach for the detection of the iris center. Starting from a learning base that only contains people in frontal view and looking in front of them, our model (based on 2.5D Active Appearance Models (AAM)) is capable of capturing the iris movements for both people in frontal view and with different head poses. We merge an iris model and a local eye model where holes are put in the place of the white-iris region. The iris texture slides under the eye hole permitting to synthesize and thus analyze any gaze direction. We propose a multi-objective optimization technique to deal with large head poses. We compared our method to a 2.5D AAM trained on faces with different gaze directions and showed that our proposition outperforms it in robustness and accuracy of detection specifically when head pose varies and with subjects wearing eyeglasses.
Hanan Salam, Nicolas Stoiber, Renaud Séguier
ICIP3
2012 A new invariant representation of facial expressions: Definition and application to blended expression recognition
abstract
This paper proposes a novel method to perform accurate facial expression recognition by transforming the appearance space into an expression space. The expression space is computed from the person-independent organization of the facial expressions found out from data. The dimension of the expression space is reduced by the projection on a manifold compliant with the organization of the expressions. Experimental results on 14 different blended expressions show that the proposed organization based method improve the facial expression recognition performance compared to appearance based methods by 13%.
Catherine Soladié, Nicolas Stoiber, Renaud Séguier
ICIP3
2012 A multimodal fuzzy inference system using a continuous facial expression representation for emotion detection
abstract
This paper presents a multimodal fuzzy inference system for emotion detection. The system extracts and merges visual, acoustic and context relevant features. The experiments have been performed as part of the AVEC 2012 challenge. Facial expressions play an important role in emotion detection. However, having an automatic system to detect facial emotional expressions on unknown subjects is still a challenging problem. Here, we propose a method that adapts to the morphology of the subject and that is based on an invariant representation of facial expressions. Our method relies on 8 key expressions of emotions of the subject. In our system, each image of a video sequence is defined by its relative position to these 8 expressions. These 8 expressions are synthesized for each subject from plausible distortions learnt on other subjects and transferred on the neutral face of the subject. Expression recognition in a video sequence is performed in this space with a basic intensity-area detector. The emotion is described in the 4 dimensions: valence, arousal, power and expectancy. The results show that the duration of high intensity smile is an expression that is meaningful for continuous valence detection and can also be used to improve arousal detection. The main variations in power and expectancy are given by context data.
Catherine Soladié, Hanan Salam, Catherine Pelachaud, Nicolas Stoiber, Renaud Séguier
ICMI5
2012 Facial Action Recognition Combining Heterogeneous Features via Multikernel Learning
abstract
This paper presents our response to the first international challenge on facial emotion recognition and analysis. We propose to combine different types of features to automatically detect action units (AUs) in facial images. We use one multikernel support vector machine (SVM) for each AU we want to detect. The first kernel matrix is computed using local Gabor binary pattern histograms and a histogram intersection kernel. The second kernel matrix is computed from active appearance model coefficients and a radial basis function kernel. During the training step, we combine these two types of features using the recently proposed SimpleMKL algorithm. SVM outputs are then averaged to exploit temporal information in the sequence. To evaluate our system, we perform deep experimentation on several key issues: influence of features and kernel function in histogram-based SVM approaches, influence of spatially independent information versus geometric local appearance information and benefits of combining both, sensitivity to training data, and interest of temporal context adaptation. We also compare our results with those of the other participants and try to explain why our method had the best performance during the facial expression recognition and analysis challenge.
Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost
IEEE Trans. Syst. Man Cybern. Part B4
2011 Combining AAM coefficients with LGBP histograms in the multi-kernel SVM framework to detect facial action units
abstract
This study presents a combination of geometric and appearance features used to automatically detect Action Units in face images. We use one multi-kernel SVM for each Action Unit we want to detect. The first kernel matrix is computed using Local Gabor Binary Pattern (LGBP) histograms and a histogram intersection kernel. The second kernel matrix is computed from AAM coefficients and a RBF kernel. During the training step, we combine these two type s of features using the recent SimpleMKL algorithm. SVM outputs are then filtered to exploit dynamic relationships between Action Units.
Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost
FG4
2010 Facial animation retargeting and control based on a human appearance space
abstract
Abstract Expressive facial animations are essential to enhance the realism and the credibility of virtual characters. Parameter‐based animation methods offer a precise control over facial configurations while performance‐based animation benefits from the naturalness of captured human motion. In this paper, we propose an animation system that gathers the advantages of both approaches. By analyzing a database of facial motion, we create the human appearance space. The appearance space provides a coherent and continuous parameterization of human facial movements, while encapsulating the coherence of real facial deformations. We present a method to optimally construct an analogous appearance face for a synthetic character. The link between both appearance spaces makes it possible to retarget facial animation on a synthetic face from a video source. Moreover, the topological characteristics of the appearance space allow us to detect the principal variation patterns of a face and automatically reorganize them on a low‐dimensional control space. The control space acts as an interactive user‐interface to manipulate the facial expressions of any synthetic face. This interface makes it simple and intuitive to generate still facial configurations for keyframe animation, as well as complete temporal sequences of facial movements. The resulting animations combine the flexibility of a parameter‐based system and the realism of real human motion. Copyright © 2010 John Wiley & Sons, Ltd.
Nicolas Stoiber, Renaud Séguier, Gaspard Breton
Comput. Animat. Virtual Worlds2
2009 Automatic design of a control interface for a synthetic face
abstract
Getting synthetic faces to display natural facial expressions is essential to enhance the interaction between human users and virtual characters. Yet traditional facial control techniques provide precise but complex sets of control parameters, which are not adapted for non-expert users. In this article, we present a system that generates a simple, 2-Dimensional interface that offers an efficient control over the facial expressions of any synthetic character. The interface generation process relies on the analysis of the deformation of a real human face. The principal geometrical and textural variation patterns of the real face are detected and automatically reorganized onto a low-dimensional space. This control space can then be easily adapted to pilot the deformations of synthetic faces. The resulting virtual character control interface makes it easy to produce varied emotional facial expressions, both extreme and subtle. In addition, the continuous nature of the interface allows the production of coherent temporal sequences of facial animation.
Nicolas Stoiber, Renaud Séguier, Gaspard Breton
IUI2
2009 Fast Simplex Optimization for Active Appearance Model
Yasser Aidarous, Renaud Séguier
PSIVT2
2008 MVAAM (multi-view active appearance model) optimized by multi-objective genetic algorithm
abstract
This paper presents an efficient algorithm of face alignment of multiple images by 2.5D Active Appearance Model (AAM). Currently with wide availability of inexpensive webcams a multi-view system is as practical as mono-view. To manage these multiple information obtained from multiview system we propose a new optimization technique of AAM. Our technique is based on Pareto multi-objective genetic optimization of NSGA-II (Non-dominated Sorting Genetic Algorithm). Our approach of multi-view AAM outperforms conventional single-view AAM (SVAAM) and is more accurate, robust and capable of performing face alignment with large lateral movements of a face. Sometimes face is oriented such that one of the camera of multi-view system do not hold a valid face information thus system has to discard the information from this camera and focus on to other camera. Our algorithm has this capability which makes it time efficient with respect to applying multiple instances of conventional SVAAM on multiple images. Algorithm is applied on number of multi-view images and compared with a system based on mono-view images. Results obtained validate our proposition.
Abdul Sattar 0003, Renaud Séguier
FG2
2008 GAGM-AAM: A genetic optimization with Gaussian mixtures for Active Appearance Models
abstract
This paper proposes an optimization technique of genetic algorithm (GA) combined with Gaussian mixtures (GAGM) to make a robust, efficient and real time face alignment application for embedded systems. It uses 2.5D Active Appearance Model (AAM) for the face search, the model is generated by taking 3D landmarks and 2D texture of the face image. 3D face alignment requires to optimize 6 DOF (Degrees of Freedom) pose and appearance parameters of AAM. These parameters span in a huge face search space. In order to optimize them GA (due to its exploration property) is taken as an optimization technique, but unfortunately it suffers from massive computations. Thanks to the clustering of appearance parameters by Gaussian Mixture, GA optimization becomes time efficient and accurate. We compare it with other technique of simplex, which is found to be more efficient than classical AAM.
Abdul Sattar 0003, Yasser Aidarous, Renaud Séguier
ICIP3
2007 Avatar Puppetry Using Real-Time Audio and Video Analysis
Sylvain Le Gallou, Gaspard Breton, Renaud Séguier, Christophe Garcia
IVA3
2006 Time series modeling for IDS alert management
abstract
Intrusion detection systems create large amounts of alerts. Significant part of these alerts can be seen as background noise of an operational information system, and its quantity typically overwhelms the user. In this paper we have three points to make. First, we present our findings regarding the causes of this noise. Second, we provide some reasoning why one would like to keep an eye on the noise despite the large number of alerts. Finally, one approach for monitoring the noise with reasonable user load is proposed. The approach is based on modeling regularities in alert flows with classical time series methods. We present experimentations and results obtained using real world data.
Jouni Viinikka, Hervé Debar, Ludovic Mé, Renaud Séguier
AsiaCCS4
2006 Scale Normalization for the Distance Maps AAM
abstract
The Active Apearence Models (AAM) are often used in Man-Machine Interaction for their ability to align the faces. We propose a new normalization method for AAM based on distance map in order to strengthen their robustness to differences in illumination. Our normalization do not use the photometric normalization protocol classically used in AAM and is much more simpler to implement. Compared to Distance Map AAM performances of [1] and other AAM implementation which use CLAHE [2] normalization or gradient information, our proposition is at the same time much robust to illumination and AAM initialization. The tests have been drive in the context of generalization: 10 persons with frontal illumination from M2VTS database [3] were considered to build the AAM, and 17 persons under 21 different illuminations from CMU database [4] were used for the testing base.
Denis Giri, Maxime Rosenwald, Benjamin Villeneuve, Sylvain Le Gallou, Renaud Séguier
ICARCV5
2002 Audio-Visual Speech Recognition One Pass Learning with Spiking Neurons
Renaud Séguier, David Mercier
ICANN1
2000 A Set of Neural Tools for Human-Computer Interactions: Application to the Handwritten Character Recognition, and Visual Speech Recognition Problems
Gilles Vaucher, Abdul Rauf Baig, Renaud Séguier
Neural Comput. Appl.3