VLDB 2026 Research / reviewers in the wild / expert
Catherine Soladié
dblp:120/4009
· DBLP profile ↗
22ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-1694-4706ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | JanusGAN: GANs Disentangled Editing with Two DiscriminatorsabstractGenerative Adversarial Networks (GANs) have found applications in image editing. However, GANs tend to synthesize and manipulate global features, such as age, rather than focusing on local features such as facial wrinkles. Consequently, when a specific wrinkle is edited, all age-related features change as well. This paper proposes a new method that allows a GAN to learn a specific global or local disentangled edit. The method involves fine-tuning a pre-trained GAN using two discriminators, each trained on a specific dataset representing distinct states of a single disentangled feature. This approach facilitates the generator’s ability to learn features along a defined editing direction within the latent space. Importantly, to avoid interfering with prior GAN knowledge, the editing direction is defined in the meaningless dimension of the GAN latent space. Although our primary focus is on local editing, our method can be extended to global features, such as age editing. Quantitative and qualitative results show that our method provides a better balance between feature accuracy and disentanglement than other state-of-the-art methods for both local and global features. The code is available on GitHub: https://github.com/Neilstid/JanusGAN. Neil Farmer, Catherine Soladié, Gabriel Cazorla, Renaud Séguier |
FG | 2 |
| 2024 | Dense Trajectory Fields: Consistent and Efficient Spatio-Temporal Pixel Tracking
Marc Tournadre, Catherine Soladié, Nicolas Stoiber, Pierre-Yves Richard |
ACCV (2) | 2 |
| 2024 | Pivotal Tuning Editing: Towards Disentangled Wrinkle Editing with GANsabstractGenerative Adversarial Networks (GANs) enable image editing by manipulating image features. However, these manipulations still lack disentanglement. For example, when a specific wrinkle is edited, other age-related features or facial expressions are often changed as well. This paper proposes a new method for disentangled editing. The presented approach is based on two pivot images that allow learning an editing direction for an input image. These pivots are based on a real image (the input) and a synthetic modification of the real image along the desired editing direction. Although our primary focus is on wrinkle editing applications, our method can be extended to other editing tasks, such as hair color or lipstick editing. Qualitative and quantitative results show that our Pivotal Tuning Editing (PTE) provides a higher level of disentanglement and a more realistic editing than state-of-the-art methods. The code is available on GitHub1. Neil Farmer, Catherine Soladié, Gabriel Cazorla, Renaud Séguier |
FG | 2 |
| 2023 | Exploring Mental Prototypes by an Efficient Interdisciplinary Approach: Interactive Microbial Genetic AlgorithmabstractFacial expression-based technologies have flooded our daily lives. However, most technologies are limited to Ekman's basic facial expressions and rarely deal with more than ten emotional states. This is not only due to the lack of prototypes for complex emotions but also the time-consuming and laborious task of building an extensive labeled database. To remove these obstacles, we were inspired by a psychophysical approach for affective computing, so-called the reverse correlation process (RevCor), to extract mental prototypes of what a given emotion should look like for an observer. We proposed a novel, efficient, and interdisciplinary approach called Interactive Microbial Genetic Algorithm (IMGA) by integrating the concepts of RevCor into an interactive genetic algorithm (IGA). Our approach achieves four challenges: online feedback loop, expertise-free, velocity, and diverse results. Experimental results show that for each observer, with limited trials, our approach can provide diverse mental prototypes for both basic emotions and emotions that are not available in existing deep-learning databases. Our work is available at https://yansen0508.github.io/Interactive-Microbial-Genetic-Algorithm/. Sen Yan 0004, Catherine Soladié, Renaud Séguier |
FG | 2 |
| 2023 | MES-Loss: Mutually equidistant separation metric learning loss function
Yasser Boutaleb, Catherine Soladié, Nam-Duong Duong, Amine Kacete, Jérôme Royan, Renaud Séguier |
Pattern Recognit. Lett. | 2 |
| 2023 | Local Temporal Pattern and Data Augmentation for Spotting Micro-ExpressionsabstractMicro-expressions (MEs) are very important nonverbal communication clues. However, due to their local and short nature, spotting them is challenging. In this article, we address this problem by using a dedicated local and temporal pattern (LTP) of facial movement. This pattern has a specific shape (an S-pattern) when MEs are displayed. Thus, by using a classic classification algorithm (SVM), MEs can be distinguished from other facial movements. We also propose a global final fusion analysis covering the whole face to improve the distinction between ME (local) and head (global) movements. However, the learning of S-patterns is limited by the small number of ME databases and the low volume of ME samples. Hammerstein models (HMs) are known to effectively approximate muscle movements. By approximating each S-pattern with an HM, we can both filter out outliers and generate new similar S-patterns. In this way, we augment the dataset for S-pattern training and improve the ability to differentiate MEs from other movements. The spotting results, performed in the CASMEI and CASMEII databases, show that our proposed LTP outperforms the most popular spotting method in terms of the F1-score. Adding a fusion process and data augmentation improves the spotting performance even further. Jingting Li 0001, Catherine Soladié, Renaud Séguier |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Micro-expression recognition from local facial regions
Mouath Aouayeb, Wassim Hamidouche, Catherine Soladié, Kidiyo Kpalma, Renaud Séguier |
Signal Process. Image Commun. | 3 |
| 2020 | Intuitive Facial Animation Editing Based On A Generative RNN FrameworkabstractAbstract For the last decades, the concern of producing convincing facial animation has garnered great interest, that has only been accelerating with the recent explosion of 3D content in both entertainment and professional activities. The use of motion capture and retargeting has arguably become the dominant solution to address this demand. Yet, despite high level of quality and automation performance‐based animation pipelines still require manual cleaning and editing to refine raw results, which is a time‐ and skill‐demanding process. In this paper, we look to leverage machine learning to make facial animation editing faster and more accessible to non‐experts. Inspired by recent image inpainting methods, we design a generative recurrent neural network that generates realistic motion into designated segments of an existing facial animation, optionally following user‐provided guiding constraints. Our system handles different supervised or unsupervised editing scenarios such as motion filling during occlusions, expression corrections, semantic content modifications, and noise filtering. We demonstrate the usability of our system on several animation editing use cases. Eloïse Berson, Catherine Soladié, Nicolas Stoiber |
Comput. Graph. Forum | 2 |
| 2020 | Efficient multi-output scene coordinate prediction for fast and accurate camera relocalization from a single RGB image
Nam-Duong Duong, Catherine Soladié, Amine Kacete, Pierre-Yves Richard, Jérôme Royan |
Comput. Vis. Image Underst. | 2 |
| 2020 | Realistic Transformation of Facial and Vocal Smiles in Real-Time Audiovisual StreamsabstractResearch in affective computing and cognitive science has shown the importance of emotional facial and vocal expressions during human-computer and human-human interactions. But, while models exist to control the display and interactive dynamics of emotional expressions, such as smiles, in embodied agents, these techniques can not be applied to video interactions between humans. In this work, we propose an audiovisual smile transformation algorithm able to manipulate an incoming video stream in real-time to parametrically control the amount of smile seen on the user's face and heard in their voice, while preserving other characteristics such as the user's identity or the timing and content of the interaction. The transformation is composed of separate audio and visual pipelines, both based on a warping technique informed by real-time detection of audio and visual landmarks. Taken together, these two parts constitute a unique audiovisual algorithm which, in addition to providing simultaneous real-time transformations of a real person's face and voice, allows to investigate the integration of both modalities of smiles in real-world social interactions. Pablo Arias 0003, Catherine Soladié, Oussema Bouafif, Axel Röbel, Renaud Séguier, Jean-Julien Aucouturier |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Unsupervised Adaptation of a Person-Specific Manifold of Facial ExpressionsabstractIn order to analyze expressions that are different from the prototypic expressions defined by Ekman, manifold learning has been proposed to build person-specific continuous representations of facial expressions. Yet, it is still a challenging problem to build such a manifold with no prior knowledge on the morphology of the subject. Here, we propose a method to build a person-specific manifold of facial expressions able to adapt to the morphology of the subject in an unsupervised manner. The manifold is initialized with the facial landmarks of the neutral face and 5 synthesized basic expressions. Our first contribution is to detect automatically the neutral face of the subject so that we can build the manifold in an unsupervised manner. Our second and main contribution is to adapt in an unsupervised manner the initialized manifold to the morphology of the subject by detecting the real basic expressions of the subject while maintaining constraints in the manifold. Our third contribution is to perform the adaptation on spontaneous expressions with typical head pose variation for human-computer interaction. The experiments show that the adaptation works well on posed expressions and that the constraints for the adaptation on spontaneous expressions is efficient when head pose variation is considered. Raphaël Weber, Vincent Barrielle, Catherine Soladié, Renaud Séguier |
IEEE Trans. Affect. Comput. | 3 |
| 2019 | Spotting Micro-Expressions on Long Videos SequencesabstractThis paper presents two methods for the first Micro-Expression Spotting Challenge 2019 by evaluating local temporal pattern (LTP) and local binary pattern (LBP) on two most recent databases, i.e. SAMM and CAS(ME)2. First we propose LTP-ML method as the baseline results for the challenge and then we compare the results with the LBP-χ2-distance method. The LTP patterns are extracted by applying PCA in a temporal window on several facial local regions. The micro-expression sequences are then spotted by a local classification of LTP and a global fusion. The LBP-χ2-distance method is to compare the feature difference by calculating χ2distance of LBP in a time window, the facial movements are then detected with a threshold. The performance is evaluated by Leave-One-Subject-Out cross validation. The overlap frames are used to determine the True Positives and the metric F1-score is used to compare the spotting performance of the databases. The F1-score of LTP-ML result for SAMM and CAS(ME)2are 0.0316 and 0.0179, respectively. The results show our proposed LTP-ML method outperformed LBP-χ2-distance method in terms of F1-score on both databases. Jingting Li 0001, Catherine Soladié, Renaud Séguier, Moi Hoon Yap |
FG | 2 |
| 2019 | Person-Specific Joy Expression Synthesis with Geometric MethodabstractSmiling has a psychiatric effect in emotional state and may hold tremendous potential for clinical remediation in psychiatric disorders. A few researchers in image synthesis work on acting on the emotional state of subjects by automatically deforming their faces to synthesize joyful expression. However, to generate these expressions they apply the same deformation for the subjects while each person smiles differently. In this paper, we head towards a personalized synthesis of the joy expression. We have studied the trajectories of the face landmarks during a smile on the CK, Oulu-CASIA and MMI databases. The obtained results show that the smile is personal, straight and it occurs in a different way for each person. That is why we propose a system that can photo-realistically transform in real-time a detected face into a personalized joyful one using a geometric method. Both visual fidelity and a statistical study demonstrate that our person-specific method can generate personalized joy expressions closer to the ground truth than two non-personal state-of-the-art approaches. Sarra Zaied, Catherine Soladié, Pierre-Yves Richard |
ICIP | 2 |
| 2019 | A Robust Interactive Facial Animation Editing SystemabstractOver the past few years, the automatic generation of facial animation for virtual characters has garnered interest among the animation research and industry communities. Recent research contributions leverage machine-learning approaches to enable impressive capabilities at generating plausible facial animation from audio and/or video signals. However, these approaches do not address the problem of animation edition, meaning the need for correcting an unsatisfactory baseline animation or modifying the animation content itself. In facial animation pipelines, the process of editing an existing animation is just as important and time-consuming as producing a baseline. In this work, we propose a new learning-based approach to easily edit a facial animation from a set of intuitive control parameters. To cope with high-frequency components in facial movements and preserve a temporal coherency in the animation, we use a resolution-preserving fully convolutional neural network that maps control parameters to blendshapes coefficients sequences. We stack an additional resolution-preserving animation autoencoder after the regressor to ensure that the system outputs natural-looking animation. The proposed system is robust and can handle coarse, exaggerated edits from non-specialist users. It also retains the high-frequency motion of the facial animation. The training and the tests are performed on an extension of the B3D(AC)23032 database [Fanelli et al. 2010], that we make available with this paper at http://www.rennes.centralesupelec.fr/biwi3D. Eloïse Berson, Catherine Soladié, Vincent Barrielle, Nicolas Stoiber |
MIG | 2 |
| 2018 | Accurate Sparse Feature Regression Forest Learning for Real-Time Camera RelocalizationabstractCamera relocalization is needed in several applications such as augmented reality or robot navigation. However, it is still challenging to have a both real-time and accurate method. In this paper, we present our hybrid method combing machine learning approach and geometric approach for real-time camera relocalization from a single RGB image. We introduce our sparse feature regression forest to improve the machine learning part. In our regression forest, we propose a novel split function, that uses a whole feature vector instead of classical binary test function to improve the accuracy of 2D-3D point correspondences. Moreover, we use sparse feature extraction (SURF features) to reduce time processing. The results indicate that our method is the only real-time hybrid method (50ms per frame). We also achieve results as accurate as the best state-of-the-art methods (hybrid methods) and outperform machine learning based and sparse feature based methods. Nam-Duong Duong, Amine Kacete, Catherine Soladié, Pierre-Yves Richard, Jérôme Royan |
3DV | 3 |
| 2018 | LTP-ML: Micro-Expression Detection by Recognition of Local Temporal Pattern of Facial MovementsabstractThe Micro-expressions (MEs) carry specific nonverbal information, for example the facial movement caused by pain. However, as a consequence of their local and short nature, it is difficult to detect MEs. This paper presents a novel detection method by recognizing a local and temporal pattern (LTP) of facial movement. In our system, with the purpose of improving the detection accuracy, temporal local features are generated from the video in a sliding window of 300ms (mean duration of a ME). These features are extracted from a projection in PCA space and form a specific pattern during ME which is the same for all MEs. Using a classical classification algorithm (SVM), MEs are then distinguished from other facial movements. Finally, a global fusion analysis is applied on the whole face to eliminate false positives. Experiments are performed on two databases: CASME I and CASME II. The detection results show that the proposed method outperforms the most popular detection method in terms of F1-score according to the analysis of multiple metrics. Jingting Li 0001, Catherine Soladié, Renaud Séguier |
FG | 2 |
| 2016 | Real-time eye pupil localization using Hough regression forestabstractEyes are one of the most salient features of the human face, and the location of the pupil allows access to important information which can be used in several computer vision applications. Several commercial eye-trackers can estimate with good accuracy the pupil location, but need complex hardware specifications and a controlled user environment (high eye image resolution, good illumination, small head pose variations) making these solutions difficult to use in an arbitrary environment. In this paper, we present an approach based on Hough randomized regression trees. We demonstrate, by several evaluations on challenging public datasets that our approach is very robust to illumination, scale, eye movements and high head pose variations and yields a significant improvement compared to a wide range of state-of-the-art methods. Amine Kacete, Jérôme Royan, Renaud Séguier, Michel Collobert, Catherine Soladié |
WACV | 5 |
| 2015 | 3D facial clone based on depth patchesabstract3D face clones can be used in many areas such as Human-Computer Interaction and as preprocessing in applications, such as emotion analysis. However, such clones should be structured and the model facial shape accurately while keeping the attributes of individuals. A structured mesh is a mesh with a known semantic and topological structure. We use a face model designed from a database of 3D face examples. These global models can produce structured clones but they do not often retain the specifics of the analyzed person. Indeed, methods using models are very dependent on their databases. In our technique, we use an RGB-D sensor to get the attributes of individuals and a 3D Morphable Face Model to mark facial shape. We reverse the process classically used: we first perform fitting and then data fusion. For each depth frame, we retain the suitable data parts called Patches. This selection is performed using a distance error and the direction of the normal vectors. Depending on the location, we merge either sensor data or 3D Morphable Face Model data. We compare our method with state of the art fitting processes. The qualitative and quantitative tests show that our results are more accurate than an current fitting method and our clone has both the attributes of the person and the shape of the face well modeled. Jérôme Manceau, Catherine Soladié, Renaud Séguier |
VCIP | 2 |
| 2013 | Bilinear decomposition for blended expressions representationabstractThis paper proposes a new method for the analysis of blended expressions with varying intensity. The method is based on an asymmetric bilinear model learned on a small amount of expressions. In the resulting expression space, a blended unknown expression has a signature, that can be interpreted as a mixture of the basic expressions used in the creation of the space. Three methods are compared: a traditional method based on active appearance vectors, the asymmetric bilinear model on person-independent appearance vectors and the asymmetric bilinear model on person-specific appearance vectors. Experimental results on the recognition of 14 blended unknown expressions show the relevance of the bilinear models compared to appearance-based methods and the robustness of the person-specific models according to the types of parameters (shape and/or texture). Catherine Soladié, Renaud Séguier, Nicolas Stoiber |
VCIP | 1 |
| 2013 | Invariant representation of facial expressions for blended expression recognition on unknown subjects
Catherine Soladié, Nicolas Stoiber, Renaud Séguier |
Comput. Vis. Image Underst. | 1 |
| 2012 | A new invariant representation of facial expressions: Definition and application to blended expression recognitionabstractThis paper proposes a novel method to perform accurate facial expression recognition by transforming the appearance space into an expression space. The expression space is computed from the person-independent organization of the facial expressions found out from data. The dimension of the expression space is reduced by the projection on a manifold compliant with the organization of the expressions. Experimental results on 14 different blended expressions show that the proposed organization based method improve the facial expression recognition performance compared to appearance based methods by 13%. Catherine Soladié, Nicolas Stoiber, Renaud Séguier |
ICIP | 1 |
| 2012 | A multimodal fuzzy inference system using a continuous facial expression representation for emotion detectionabstractThis paper presents a multimodal fuzzy inference system for emotion detection. The system extracts and merges visual, acoustic and context relevant features. The experiments have been performed as part of the AVEC 2012 challenge. Facial expressions play an important role in emotion detection. However, having an automatic system to detect facial emotional expressions on unknown subjects is still a challenging problem. Here, we propose a method that adapts to the morphology of the subject and that is based on an invariant representation of facial expressions. Our method relies on 8 key expressions of emotions of the subject. In our system, each image of a video sequence is defined by its relative position to these 8 expressions. These 8 expressions are synthesized for each subject from plausible distortions learnt on other subjects and transferred on the neutral face of the subject. Expression recognition in a video sequence is performed in this space with a basic intensity-area detector. The emotion is described in the 4 dimensions: valence, arousal, power and expectancy. The results show that the duration of high intensity smile is an expression that is meaningful for continuous valence detection and can also be used to improve arousal detection. The main variations in power and expectancy are given by context data. Catherine Soladié, Hanan Salam, Catherine Pelachaud, Nicolas Stoiber, Renaud Séguier |
ICMI | 1 |