Michel F. Valstar

dblp:00/2794 · also Michel François Valstar · DBLP profile ↗
← Back
95ranked-venue papers
13as first author
21since 2021 · last 2025
0000-0003-2414-161XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 8 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 24 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia9
2025 Two-Stage Temporal Modelling Framework for Video-Based Depression Recognition Using Graph Representation
abstract
Video-based automatic depression analysis provides a fast, objective and repeatable self-assessment solution, which has been widely developed in recent years. While depression cues may be reflected by human facial behaviours of various temporal scales, most existing approaches either focused on modelling depression from short-term or video-level facial behaviours. In this sense, we propose a two-stage framework that models depression severity from multi-scale short-term and video-level facial behaviours. The short-term depressive behaviour modelling stage first deep learns depression-related facial behavioural features from multiple short temporal scales, where a Depression Feature Enhancement (DFE) module is proposed to enhance the depression-related cues for all temporal scales and remove non-depression related noise. Two novel graph encoding strategies are proposed in the video-level depressive behavior modeling stage, i.e., Sequential Graph Representation (SEG) and Spectral Graph Representation (SPG), to re-encode all short-term features of the target video into a video-level graph representation, summarizing depression-related multi-scale video-level temporal information. As a result, the produced graph representations predict depression severity using both short-term and long-term facial behaviour patterns. The experimental results on AVEC 2013, AVEC 2014 and AVEC 2019 datasets show that the proposed DFE module constantly enhanced the depression severity estimation performance for various CNN models while the SPG is superior than other video-level modelling methods. More importantly, the result achieved for the proposed two-stage framework shows its promising and solid performance compared to widely-used one-stage modelling approaches.
Hatice Gunes, Keerthy Kusumam, Michel F. Valstar, Siyang Song
IEEE Trans. Affect. Comput.4
2024 REACT 2024: the Second Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024.
Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
FG8
2024 Loss Relaxation Strategy for Noisy Facial Video-based Automatic Depression Recognition
abstract
Automatic depression analysis has been widely investigated on face videos that have been carefully collected and annotated in lab conditions. However, videos collected under real-world conditions may suffer from various types of noise due to challenging data acquisition conditions and lack of annotators. Although deep learning (DL) models frequently show excellent depression analysis performances on datasets collected in controlled lab conditions, such noise may degrade their generalization abilities for real-world depression analysis tasks. In this article, we uncovered that noisy facial data and annotations consistently change the distribution of training losses for facial depression DL models; i.e., noisy data–label pairs cause larger loss values compared to clean data–label pairs. Since different loss functions could be applied depending on the employed model and task, we propose a generic loss function relaxation strategy that can jointly reduce the negative impact of various noisy data and annotation problems occurring in both classification and regression loss functions for face video-based depression analysis, where the parameters of the proposed strategy can be automatically adapted during depression model training. The experimental results on 25 different artificially created noisy depression conditions (i.e., five noise types with five different noise levels) show that our loss relaxation strategy can clearly enhance both classification and regression loss functions, enabling the generation of superior face video-based depression analysis models under almost all noisy conditions. Our approach is robust to its main variable settings and can adaptively and automatically obtain its parameters during training.
Siyang Song, Tugba Tümer, Changzeng Fu, Michel F. Valstar, Hatice Gunes
ACM Trans. Comput. Heal.5
2024 COLD Fusion: Calibrated and Ordinal Latent Distribution Fusion for Uncertainty-Aware Multimodal Emotion Recognition
abstract
Automatically recognising apparent emotions from face and voice is hard, in part because of various sources of uncertainty, including in the input data and the labels used in a machine learning framework. This paper introduces an uncertainty-aware multimodal fusion approach that quantifies modality-wise aleatoric or data uncertainty towards emotion prediction. We propose a novel fusion framework, in which latent distributions over unimodal temporal context are learned by constraining their variance. These variance constraints, Calibration and Ordinal Ranking, are designed such that the variance estimated for a modality can represent how informative the temporal context of that modality is w.r.t. emotion recognition. When well-calibrated, modality-wise uncertainty scores indicate how much their corresponding predictions are likely to differ from the ground truth labels. Well-ranked uncertainty scores allow the ordinal ranking of different frames across different modalities. To jointly impose both these constraints, we propose a softmax distributional matching loss. Our evaluation on AVEC 2019 CES, CMU-MOSEI, and IEMOCAP datasets shows that the proposed multimodal fusion method not only improves the generalisation performance of emotion recognition models and their predictive uncertainty estimates, but also makes the models robust to novel noise patterns encountered at test time.
Mani Kumar Tellamekala, Shahin Amiriparian, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Guest Editorial: Ethics in Affective Computing
abstract
Stunning advances in machine learning are heralding a new era in sensing, interpreting, simulating and stimulating human emotion. In the human sciences, research is increasingly highlighting the explanatory power of emotions, feelings, and other affective processes to predict how we think and behave. This is beginning to translate into an explosion of applications that can improve human wellbeing including methods to reduce stress and improve emotion regulation skills, techniques to support healthier social media use, pain monitoring in neonates, and decision-support tools that recognize emotional bias.
Jonathan Gratch, Gretchen Greene, Rosalind W. Picard, Lachlan Urquhart, Michel F. Valstar
IEEE Trans. Affect. Comput.5
2024 Are 3D Face Shapes Expressive Enough for Recognising Continuous Emotions and Action Unit Intensities?
abstract
Recognising continuous emotions and action unit (AU) intensities from face videos, requires a spatial and temporal understanding of expression dynamics. Existing works primarily rely on 2D face appearance features to extract such dynamics. This work focuses on a promising alternative based on parametric 3D face alignment models, which disentangle different factors of variation, including expression-induced shape variations. We aim to understand how expressive 3D face shapes are in estimating valence-arousal and AU intensities compared to the state-of-the-art 2D appearance-based models. We benchmark five recent 3D face models: ExpNet, 3DDFA-V2, RingNet, DECA, and EMOCA. In valence-arousal estimation, expression features of 3D face models consistently surpassed previous works and yielded an average concordance correlation of. 745 and. 574 on SEWA and AVEC 2019 CES corpora, respectively. We also study how 3D face shapes performed on AU intensity estimation on BP4D and DISFA datasets, and report that 3D face features were on par with 2D appearance features in recognising AUs 4, 6, 10, 12, and 25, but not the entire set of AUs. To understand this discrepancy, we conduct a correspondence analysis between valence-arousal and AUs, which points out that accurate prediction of valence-arousal may require the knowledge of only a few AUs.
Mani Kumar Tellamekala, Ömer Sümer, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Affect. Comput.6
2023 REACT2023: The First Multiple Appropriate Facial Reaction Generation Challenge
abstract
The Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023.
Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
ACM Multimedia7
2023 A Transfer Learning Approach to Heatmap Regression for Action Unit Intensity Estimation
abstract
Action Units (AUs) are geometrically-based atomic facial muscle movements known to produce appearance changes at specific facial locations. Motivated by this observation we propose a novel AU modelling problem that consists of jointly estimating their localisation and intensity. To this end, we propose a simple yet efficient approach based on Heatmap Regression that merges both problems into a single task. A Heatmap models whether an AU occurs or not at a given spatial location. To accommodate the joint modelling of AUs intensity, we propose variable size heatmaps, with their amplitude and size varying according to the labelled intensity. Using Heatmap Regression, we can inherit from the progress recently witnessed in facial landmark localisation. Building upon the similarities between both problems, we devise a transfer learning approach where we exploit the knowledge of a network trained on large-scale facial landmark datasets. In particular, we explore different alternatives for transfer learning through a) fine-tuning, b) adaptation layers, c) attention maps, and d) reparametrisation. Our approach effectively inherits the rich facial features produced by a strong face alignment network, with minimal extra computational cost. We empirically validate that our system sets a new state-of-the-art on three popular datasets, namely BP4D, DISFA, and FERA2017.
Ioanna Ntinou, Enrique Sánchez-Lozano, Adrian Bulat, Michel F. Valstar, Georgios Tzimiropoulos
IEEE Trans. Affect. Comput.4
2023 Self-Supervised Learning of Person-Specific Facial Dynamics for Automatic Personality Recognition
abstract
This article aims to solve two important issues that frequently occur in existing automatic personality analysis systems: 1. Attempting to use very short video segments or even single frames, rather than long-term behaviour, to infer personality traits; 2. Lack of methods to encode person-specific facial dynamics for personality recognition. To deal with these issues, this paper first proposes a novel Rank Loss which utilizes the natural temporal evolution of facial actions, rather than personality labels, for self-supervised learning of facial dynamics. Our approach first trains a generic U-net style model that can infer general facial dynamics learned from a set of unlabelled face videos. Then, the generic model is frozen, and a set of intermediate filters are incorporated into this architecture. The self-supervised learning is then resumed with only person-specific videos. This way, the learned filters’ weights are person-specific, making them a valuable source for modeling person-specific facial dynamics. We then propose to concatenate the weights of the learned filters as a person-specific representation, which can be directly used to predict the personality traits without needing other parts of the network. We evaluate the proposed approach on both self-reported personality and apparent personality datasets. In addition to achieving promising results in the estimation of personality trait scores from videos, we show that the tasks conducted by the subject in the video matters, that fusion of a combination of tasks reaches highest accuracy, and that multi-scale dynamics are more informative than single-scale dynamics.
Siyang Song, Shashank Jaiswal, Enrique Sánchez-Lozano, Georgios Tzimiropoulos, LinLin Shen, Michel F. Valstar
IEEE Trans. Affect. Comput.6
2023 Learning Person-Specific Cognition From Facial Reactions for Automatic Personality Recognition
abstract
This article proposes to recognise the true (self-reported) personality traits from the target subject's cognition simulated from facial reactions. This approach builds on the following two findings in cognitive science: (i) human cognition partially determines expressed behaviour and is directly linked to true personality traits; and (ii) in dyadic interactions, individuals’ nonverbal behaviours are influenced by their conversational partner's behaviours. In this context, we hypothesise that during a dyadic interaction, a target subject's facial reactions are driven by two main factors: their internal (person-specific) cognitive process, and the externalised nonverbal behaviours of their conversational partner. Consequently, we propose to represent the target subject's (defined as the listener) person-specific cognition in the form of a person-specific CNN architecture that has unique architectural parameters and depth, which takes audio-visual non-verbal cues displayed by the conversational partner (defined as the speaker) as input, and is able to reproduce the target subject's facial reactions. Each person-specific CNN is explored by the Neural Architecture Search (NAS) and a novel adaptive loss function, which is then represented as a graph representation for recognising the target subject's true personality. Experimental results not only show that the produced graph representations are well associated with target subjects’ personality traits in both human-human and human-machine interaction scenarios, and outperform the existing approaches with significant advantages, but also demonstrate that the proposed novel strategies help in learning more reliable personality representations.
Siyang Song, Zilong Shao, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes
IEEE Trans. Affect. Comput.5
2023 Modelling Stochastic Context of Audio-Visual Expressive Behaviour With Affective Processes
abstract
Recognising apparent emotion from audio-visual signals in naturalistic conditions remains an open problem. Existing methods that build on recurrent models, or in the modelling of contextual dependencies at the feature level using self-attention fail to model the long-term dependencies that subtly occur at different levels of abstraction. Affective Processes have emerged as a novel paradigm to the modelling of temporal dynamics through a probabilistic global latent variable that captures context and induces dependencies in the outputs, showing superior performance with little complexity. Despite its impressive results on visual data, Affective Processes remain unexplored in the domain of audio data, known to crucially influence the perception of emotions. In this paper, we first revisit and extend Affective Processes to the speech domain, identifying the key components and learning procedures for their efficient training. We then extend Affective Processes to audio-visual affect recognition, using modality-specific context encoders. Finally, we propose a novel application of Affective Processes in the domain of Cooperative Machine Learning for propagating affect labels in videos using sparse human supervision. We conduct extensive ablation studies, identifying the main components behind the success of Affective Processes, as well as comparisons against existing works in a variety of datasets.
Mani Kumar Tellamekala, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Affect. Comput.3
2022 Spectral Representation of Behaviour Primitives for Depression Analysis
abstract
Depression is a serious mental disorder affecting millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and require extensive participation of clinicians. Recent advances in automatic depression analysis systems promise a future where these shortcomings are addressed by objective, repeatable, and readily available diagnostic tools to aid health professionals in their work. Yet there remain a number of barriers to the development of such tools. One barrier is that existing automatic depression analysis algorithms base their predictions on very brief sequential segments, sometimes as little as one frame. Another barrier is that existing methods do not take into account what the context of the measured behaviour is. In this article, we extract multi-scale video-level features for video-based automatic depression analysis. We propose to use automatically detected human behaviour primitives as the low-dimensional descriptor for each frame. We also propose two novel spectral representations, i.e., spectral heatmaps and spectral vectors, to represent video-level multi-scale temporal dynamics of expressive behaviour. Constructed spectral representations are fed to Convolution Neural Networks (CNNs) and Artificial Neural Networks (ANNs) for depression analysis. We conducted experiments on the AVEC 2013 and AVEC 2014 benchmark datasets to investigate the influence of interview tasks on depression analysis. In addition to achieving state of the art accuracy in severity of depression estimation, we show that the task conducted by the user matters, that fusion of a combination of tasks reaches highest accuracy, and that longer tasks are more informative than shorter tasks, up to a point.
Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar
IEEE Trans. Affect. Comput.4
2022 Dimensional Affect Uncertainty Modelling for Apparent Personality Recognition
abstract
Despite achieving impressive performance, dimensional affect or emotion recognition from faces is largely based on uncertainty-unaware models that predict only point estimates. Modelling uncertainty is important to learn reliable facial emotion recognition models with the abilities to (a). holistically quantify predictive uncertainty estimates and (b). propagate those estimates to the benefit of downstream behavioural analysis tasks. In this work, we first quantify uncertainties in dimensional emotion recognition by adopting the framework of epistemic (model) and aleatoric (data) uncertainty categorisation. Then for evaluating the practical utility of uncertainty-aware emotion predictions, we introduce them in learning an important downstream task, apparent personality recognition. To this end, we ask two questions: how to effectively (a). use already known behavioural attributes (emotions) in a downstream task (personality recognition) and (b). summarise global temporal context from uncertainty-aware emotion predictions fused with image embeddings. Answering these questions, we learn a conditional latent variable model building on recently proposed neural latent variable models. Our experiments on two in-the-wild datasets, SEWA for emotion recognition and ChaLearn for personality recognition, demonstrate that fusion of epistemic and aleatoric emotion uncertainties significantly improves personality recognition performance, with$\sim$42% relative improvement in Pearson correlation coefficient, leading to a new state-of-the-art.
Mani Kumar Tellamekala, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Affect. Comput.3
2021 Affective Processes: Stochastic Modelling of Temporal Context for Emotion and Facial Expression Recognition
abstract
Temporal context is key to the recognition of expressions of emotion. Existing methods, that rely on recurrent or self-attention models to enforce temporal consistency, work on the feature level, ignoring the task-specific temporal dependencies, and fail to model context uncertainty. To alleviate these issues, we build upon the framework of Neural Processes to propose a method for apparent emotion recognition with three key novel components: (a) probabilistic contextual representation with a global latent variable model; (b) temporal context modelling using task-specific predictions in addition to features; and (c) smart temporal context selection. We validate our approach on four databases, two for Valence and Arousal estimation (SEWA and AffWild2), and two for Action Unit intensity estimation (DISFA and BP4D). Results show a consistent improvement over a series of strong baselines as well as over state-of-the-art methods.
Enrique Sanchez, Mani Kumar Tellamekala, Michel F. Valstar, Georgios Tzimiropoulos
CVPR3
2021 Apparent Personality Recognition from Uncertainty-Aware Facial Emotion Predictions using Conditional Latent Variable Models
abstract
We propose two key ideas to improve the performance of apparent personality traits estimation from face videos: 1. using dimensional emotion predictions fused with face image embeddings as input features and 2. effectively aggregating global temporal context related to personality traits from the input feature sequence. In the former, we propose to use uncertainty-aware predictions of valence and arousal as additional input features along with the face image embeddings. To this end, we first build uncertainty-aware facial emotion recognition models by adopting epistemic (model) and aleatoric (data) uncertainty categorisation framework. In terms of improvement in the personality recognition performance, we show that uncertainty-aware emotion predictions outperform the point estimates of emotions by significant margins. On the other hand, for effectively aggregating the temporal context from the input feature sequence, we propose to use a conditional latent variable model that builds on some recently proposed neural latent variable methods for global context aggregation. By combining these two ideas, our proposed personality recognition method achieves state-of-the-art results on a large-scale in-the-wild dataset, ChaLearn, with ~42 % relative performance improvement over the best of existing benchmarks.
Mani Kumar Tellamekala, Timo Giesbrecht, Michel F. Valstar
FG3
2021 Few-Shot Learning for Postnatal Gestational Age Estimation
abstract
A baby's gestational age determines whether or not they are premature, which helps clinicians decide treatment. The most accurate dating methods use Ultrasound Scans, but these are expensive, require trained personnel and cannot always be deployed to remote areas. In the absence of such methods, the Ballard Score, a postnatal clinical examination, can be used. However, this method is highly subjective and results vary widely depending on the experience of the examiner. In the last decade, there have been efforts to exploit machine learning methods to create reliable postnatal methods deployable anywhere in the world. However, current state of the art methods remain influenced by the quality of their input data, which is a major issue in areas where data collection is difficult or impossible. Gathering images of newborns, especially those who are premature, is a very challenging task, due to the intrusiveness of taking photographs inside an ICU. This paper explores few-shot learning for gestational age estimation and investigates whether it can be used as an alternative way to explicitly deal with a lack of data. We compare two popular methods for few-shot learning, Model Agnostic Meta-Learning and Prototypical Networks, with a small dataset collected as part of the GesATional Project. Experimental results show that few-shot learning can be used to estimate gestational age postnatally. As a novel contribution to deal with the lack of data, models are trained with CelebFaces Attributes, to aid the procedure. We demonstrate that the novel integration of additional meta-sets improves performance by an average of 2.3% for both models comparing to just using a single dataset.
Stepan Romanov, Heda Song, Michel F. Valstar, Don Sharkey, Caz Henry, Isaac Triguero, Mercedes Torres Torres
IJCNN3
2021 Stochastic Process Regression for Cross-Cultural Speech Emotion Recognition
Mani Kumar Tellamekala, Enrique Sanchez, Georgios Tzimiropoulos, Timo Giesbrecht, Michel F. Valstar
Interspeech5
2021 Design and Evaluation of Virtual Human Mediated Tasks for Assessment of Depression and Anxiety
abstract
Virtual human technologies are now being widely explored as therapy tools for mental health disorders including depression and anxiety. These technologies leverage the ability of the virtual agents to engage in naturalistic social interactions with a user to elicit behavioural expressions which are indicative of depression and anxiety. Research efforts have focused on optimising the human-like expressive capabilities of the virtual human, but less attention has been given to investigating the effect of virtual human mediation on the expressivity of the user. In addition, it is still not clear what an optimal task is or what task characteristics are likely to sustain long term user engagement. To this end, this paper describes the design and evaluation of virtual human-mediated tasks in a user study of 56 participants. Half the participants complete tasks guided by a virtual human, while the other half are guided by text on screen. Self-reported PHQ9 scores, biosignals and participants' ratings of tasks are collected. Findings show that virtual-human mediation influences behavioural expressiveness and this observation differs for different depression severity levels. It further shows that virtual human mediation improves users' disposition towards tasks.
Joy Egede, Dominic Price, Deepa B. Krishnan, Shashank Jaiswal, Natasha Elliot, Richard F. Morriss, Maria Jose Galvez Trigo, Neil Nixon, Peter Liddle, Christopher Greenhalgh, Michel F. Valstar
IVA11
2021 Designing an Adaptive Embodied Conversational Agent for Health Literacy: a User Study
abstract
Access to healthcare advice is crucial to promote healthy societies. Many factors shape how access might be constrained, such as economic status, education or, as the COVID-19 pandemic has shown, remote consultations with health practitioners. Our work focuses on providing pre/post-natal advice to maternal women. A salient factor of our work concerns the design and deployment of embodied conversation agents (ECAs) which can sense the (health) literacy of users and adapt to scaffold user engagement in this setting. We present an account of a Wizard of Oz user study of 'ALTCAI', an ECA with three modes of interaction (i.e., adaptive speech and text, adaptive ECA, and non-adaptive ECA). We compare reported engagement with these modes from 44 maternal women who have differing levels of literacy. The study shows that a combination of embodiment and adaptivity scaffolds reported engagement, but matters of health-literacy and language introduce nuanced considerations for the design of ECAs.
Joy Egede, Maria Jose Galvez Trigo, Adrian Hazzard, Martin Porcheron, Edgar Bodiaj, Joel E. Fischer, Christopher Greenhalgh, Michel F. Valstar
IVA8
2021 Personality Recognition by Modelling Person-specific Cognitive Processes using Graph Representation
abstract
Recent research shows that in dyadic and group interactions individuals' nonverbal behaviours are influenced by the behaviours of their conversational partner(s). Therefore, in this work we hypothesise that during a dyadic interaction, the target subject's facial reactions are driven by two main factors: (i) their internal (person-specific) cognition, and (ii) the externalised nonverbal behaviours of their conversational partner. Subsequently, our novel proposition is to simulate and represent the target subject's (i.e., the listener) cognitive process in the form of a person-specific CNN architecture whose input is the audio-visual non-verbal cues displayed by the conversational partner (i.e., the speaker), and the output is the target subject's (i.e., the listener) facial reactions. We then undertake a search for the optimal CNN architecture whose results are used to create a person-specific graph representation for recognising the target subject's personality. The graph representation, fortified with a novel end-to-end edge feature learning strategy, helps with retaining both the unique parameters of the person-specific CNN and the geometrical relationship between its layers. Consequently, the proposed approach is the first work that aims to recognize the true (self-reported) personality of a target subject (i.e., the listener) from the learned simulation of their cognitive process (i.e., parameters of the person-specific CNN). The experimental results show that the CNN architectures are well associated with target subjects' personality traits and the proposed approach clearly outperforms multiple existing approaches that predict personality directly from non-verbal behaviours. In light of these findings, this work opens up a new avenue of research for predicting and recognizing socio-emotional phenomena (personality, affect, engagement etc.) from simulations of person-specific cognitive processes.
Zilong Shao, Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes
ACM Multimedia5
2020 EMOPAIN Challenge 2020: Multimodal Pain Evaluation from Facial and Bodily Expressions
abstract
The EmoPain 2020 Challenge is the first international competition aimed at creating a uniform platform for the comparison of multi-modal machine learning and multimedia processing methods of chronic pain assessment from human expressive behaviour, and also the identification of pain-related behaviours. The objective of the challenge is to promote research in the development of assistive technologies that help improve the quality of life for people with chronic pain via real-time monitoring and feedback to help manage their condition and remain physically active. The challenge also aims to encourage the use of the relatively underutilised, albeit vital bodily expression signals for automatic pain and pain-related emotion recognition. This paper presents a description of the challenge, competition guidelines, bench-marking dataset, and the baseline systems' architecture and performance on the Challenge's three sub-tasks: pain estimation from facial expressions, pain recognition from multimodal movement, and protective movement behaviour detection.
Joy Egede, Siyang Song, Temitayo A. Olugbade, Amanda C. de C. Williams, Hongying Meng, M. S. Hane Aung, Nicholas D. Lane, Michel F. Valstar, Nadia Bianchi-Berthouze
FG9
2020 A recurrent cycle consistency loss for progressive face-to-face synthesis
abstract
This paper addresses a major flaw of the cycle consistency loss when used to preserve the input appearance in the face-to-face synthesis domain. In particular, we show that the images generated by a network trained using this loss conceal a noise that hinders their use for further tasks. To overcome this limitation, we propose a “recurrent cycle consistency loss” which for different sequences of target attributes minimises the distance between the output images, independent of any intermediate step. We empirically validate not only that our loss enables the re-use of generated images, but that it also improves their quality. In addition, we propose the very first network that covers the task of unconstrained landmark-guided face-to-face synthesis. Contrary to previous works, our proposed approach enables the transfer of a particular set of input features to a large span of poses and expressions, whereby the target landmarks become the ground-truth points. We then evaluate the consistency of our proposed approach to synthesise faces at the target landmarks. To the best of our knowledge, we are the first to propose a loss to overcome the limitation of the cycle consistency loss, and the first to propose an “in-the-wild” landmark guided synthesis approach.
Enrique Sanchez, Michel F. Valstar
FG2
2020 Self-supervised learning of Dynamic Representations for Static Images
abstract
Facial actions are spatio-temporal signals by nature, and therefore their modeling is crucially dependent on the availability of temporal information. In this paper, we focus on inferring such temporal dynamics of facial actions when no explicit temporal information is available, i.e. from still images. We present a novel self-supervised learning approach to capture multiple scales of temporal dynamics, with an application to facial Action Unit (AU) intensity estimation and dimensional affect estimation. In particular: 1. We propose a framework that infers a dynamic representation (DR) from a still image, capturing the bi-directional flow of time within a short time-window centered at the input image; 2. We show that the proposed rank loss can apply facial temporal evolution to self-supervise the training process without using target representations, allowing the network to represent dynamics more broadly; 3. We propose a multiple temporal scale approach that infers DRs for different window lengths (MDR) from a still image. We empirically validate the value of our approach on the task of frame ranking, and show how our proposed MDR attains state of the art results on BP4D for AU intensity estimation and on SEMAINE for dimensional affect estimation, using only still images at test time.
Siyang Song, Enrique Sanchez, LinLin Shen, Michel F. Valstar
ICPR4
2020 Audio-Visual Predictive Coding for Self-Supervised Visual Representation Learning
abstract
Self-supervised learning has emerged as a candidate approach to learn semantic visual features from unlabeled video data. In self-supervised learning, intrinsic correspondences between data points are used to define a proxy task that forces the model to learn semantic representations. Most existing proxy tasks applied to video data exploit only either intra-modal (e.g. temporal) or cross-modal (e.g. audio-visual) correspondences separately. In theory, jointly learning both these correspondences may result in richer visual features; but, as we show in this work, doing so is non-trivial in practice. To address this problem, we introduce `Audio-Visual Permutative Predictive Coding' (AV-PPC), a multi-task learning framework designed to fully leverage the temporal and cross-modal correspondences as natural supervision signals. In AV-PPC, the model is trained to simultaneously learn multiple intra- and cross-modal predictive coding sub-tasks. By using visual speech recognition (lip-reading) as the downstream evaluation task, we show that our proposed proxy task can learn higher quality visual features than existing proxy tasks. We also show that AV-PPC visual features are highly data-efficient. Without further finetuning, AV-PPC visual encoder achieves 80.30% spoken word classification rate on the LRW dataset, performing on par with directly supervised visual encoders that are learned from large amounts of labeled data.
Mani Kumar Tellamekala, Michel F. Valstar, Michael P. Pound, Timo Giesbrecht
ICPR2
2019 Automatic Neonatal Pain Estimation: An Acute Pain in Neonates Database
abstract
Pain assessment is a vital part of newborn treatment in Intensive Care Units. However, clinical pain assessment is highly subjective and does not support continual pain monitoring. Automated tools have been introduced to address this problem, but their performance is limited by inadequate training data and unsuitable pain annotations. In addition, current automated tools focus on pain detection rather than severity estimation, which is the unmet need in medical treatment. In this paper, we present: a) the Acute Pain in Neonates (APN-db) database, a public dataset to support research in this field and allow for benchmarking of new automated tools; b) a novel L1-point Neonatal Face and Limb Acute Pain Scale (NFLAPS), a visual behaviour-centric pain measurement tool, which is an adaptation of the Neonatal Infant Pain Scale (NIPS) and the Neonatal Facial Coding System (NFCS); and c) a system for neonatal pain assessment which encodes pain indicative-features using handcrafted algorithms and deep-learned features. Experiments show that our system performs well with an RMSE of 1.9 compared to human error of 1.65 on the same dataset, demonstrating its potential application to newborn health care.
Joy Egede, Michel F. Valstar, Mercedes Torres Torres, Don Sharkey
ACII2
2019 Automatic prediction of Depression and Anxiety from behaviour and personality attributes
abstract
Current vision based approaches for automatic prediction of mental health conditions like depression and anxiety, rely on models that use behavioural features (usually extracted from faces) only and do not take personality into account. However, there is a considerable amount of evidence that people with certain personality traits are more prone to depression and anxiety disorders. In order to exploit the underlying relationship between personality and these mental health conditions, we propose to use a combination of features consisting of observed facial behaviour and self-reported personality scores. This combination of features is employed for training deep neural networks for predicting depression and anxiety scores. The proposed method was evaluated on a new dataset consisting of personality scores and interview videos from a total of 55 people. The results show that the proposed combination of features significantly improves the prediction performance compared to using the behavioural features alone. This improvement in performance is shown for 2 different kinds of video feature extraction methods and 3 different kinds of interview questionnaires.
Shashank Jaiswal, Siyang Song, Michel F. Valstar
ACII3
2019 Temporally Coherent Visual Representations for Dimensional Affect Recognition
abstract
The success of end-to-end supervised representation learning and regression in recent years has shifted the main focus of continuous emotion recognition approaches to learning visual representations directly from large scale labeled datasets. Supervised representation learning is highly effective for learning tasks with unambiguous labels. But annotating dimensional affect data is inherently subjective which more often than not leads to ambiguous labels. Relying only on such ambiguous labels for representation learning does not result in robust features with good generalization capacity, as we will show in this work. To address this fundamental problem, we propose to apply a constrained representation learning method that encourages the latent features to be less sensitive to ambiguous emotion annotations by using a generic representation learning prior called `temporal coherency' or `temporal smoothness'. This approach forces the latent features to be temporally coherent, by adding a first-order temporal coherency regularization constraint to the supervised learning loss function. To assess the utility of temporally coherent representations, we trained both unconstrained and temporal coherency constrained models on the Aff-wild dataset. Temporal coherency constrained models outperformed the unconstrained models by a significant margin. This performance improvement was also observed in the case of cross-dataset evaluation results on AFEW -VA database. Most notably, temporally coherent visual features produced state-of-the-art performance on the Aff-wild benchmark without using additional inputs such as facial landmarks.
Mani Kumar Tellamekala, Michel F. Valstar
ACII2
2019 Virtual Human Questionnaire for Analysis of Depression, Anxiety and Personality
abstract
Self-report questionnaires like PHQ-9 and GAD-7 are often used in the field of psychology to detect mental health problems and measure their severity. Their validity has been well established by a number of previous studies. However, most previous studies have used self-administration method to study their validity. In the context of automating the administering of questionnaires as a part of an interaction scenario led by a virtual human, we investigate if the virtual human administration of these questionnaires can be considered equivalent to the self-administration of these questionnaires (through an electronic form). Additionally, we also examine the equivalence when the virtual human is replaced by a real human. The interaction with real human is studied in two different modes: face-to-face and through a video-conferencing link. Statistical analysis of the scores obtained from our user study (consisting of 55 participants) revealed that human/virtual-human administered questionnaires can be considered practically equivalent to self administered ones. In most cases, the differences in the scores from self-administered questionnaires and human/virtual-human administered questionnaires, were found to be smaller than any effect size which could be considered practically significant.
Shashank Jaiswal, Michel F. Valstar, Keerthy Kusumam, Christopher Greenhalgh
IVA2
2019 AVEC'19: Audio/Visual Emotion Challenge and Workshop
abstract
The ninth Audio-Visual Emotion Challenge and workshop AVEC 2019 was held in conjunction with ACM Multimedia'19. This year, the AVEC series addressed major novelties with three distinct tasks: State-of-Mind Sub-challenge (SoMS), Detecting Depression with Artificial Intelligence Sub-challenge (DDS), and Cross-cultural Emotion Sub-challenge (CES). The SoMS was based on a novel dataset (USoM corpus) that includes self-reported mood (10-point Likert scale) after the narrative of personal stories (two positive and two negative). The DDS was based on a large extension of the DAIC-WOZ corpus (c.f. AVEC 2016) that includes new recordings of patients suffering from depression with the virtual agent conducting the interview being, this time, wholly driven by AI, i.e., without any human intervention. The CES was based on the SEWA dataset (c.f. AVEC 2018) that has been extended with the inclusion of new participants in order to investigate how emotion knowledge of Western European cultures (German, Hungarian) can be transferred to the Chinese culture. In this summary, we mainly describe participation and conditions of the AVEC Challenge.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Nicholas Cummins, Roddy Cowie, Maja Pantic
ACM Multimedia3
2019 Postnatal gestational age estimation of newborns using Small Sample Deep Learning
abstract
A baby's gestational age determines whether or not they are premature, which helps clinicians decide on suitable post-natal treatment. The most accurate dating methods use Ultrasound Scan (USS) machines, but these are expensive, require trained personnel and cannot always be deployed to remote areas. In the absence of USS, the Ballard Score, a postnatal clinical examination, can be used. However, this method is highly subjective and results vary widely depending on the experience of the examiner. Our main contribution is a novel system for automatic postnatal gestational age estimation using small sets of images of a newborn's face, foot and ear. Our two-stage architecture makes the most out of Convolutional Neural Networks trained on small sets of images to predict broad classes of gestational age, and then fuses the outputs of these discrete classes with a baby's weight to make fine-grained predictions of gestational age using Support Vector Regression. On a purpose-collected dataset of 130 babies, experiments show that our approach surpasses current automatic state-of-the-art postnatal methods and attains an expected error of 6 days. It is three times more accurate than the Ballard method. Making use of images improves predictions by 33% compared to using weight only. This indicates that even with a very small set of data, our method is a viable candidate for postnatal gestational age estimation in areas were USS is not available.
Mercedes Torres Torres, Michel F. Valstar, Caroline Henry, Carole Ward, Don Sharkey
Image Vis. Comput.2
2019 Automatic Analysis of Facial Actions: A Survey
abstract
As one of the most comprehensive and objective ways to describe facial expressions, the Facial Action Coding System (FACS) has recently received significant attention. Over the past 30 years, extensive research has been conducted by psychologists and neuroscientists on various aspects of facial expression analysis using FACS. Automating FACS coding would make this research faster and more widely applicable, opening up new avenues to understanding how we communicate through facial expressions. Such an automated process can also potentially increase the reliability, precision and temporal resolution of coding. This paper provides a comprehensive survey of research into machine analysis of facial actions. We systematically review all components of such systems: pre-processing, feature extraction and machine coding of facial actions. In addition, the existing FACS-coded facial expression databases are summarised. Finally, challenges that have to be addressed to make automatic facial action analysis applicable in real-life situations are extensively discussed. There are two underlying motivations for us to write this survey paper: the first is to provide an up-to-date review of the existing literature, and the second is to offer some insights into the future of machine recognition of facial actions: what are the challenges and opportunities that researchers in the field face.
Brais Martínez, Michel F. Valstar, Bihan Jiang, Maja Pantic
IEEE Trans. Affect. Comput.2
2018 Joint Action Unit localisation and intensity estimation through heatmap regression
Enrique Sánchez-Lozano, Georgios Tzimiropoulos, Michel F. Valstar
BMVC3
2018 Deep Learned Cumulative Attribute Regression
abstract
Learning regression-based machine learning models for computer vision problems is a challenging task due to noisy features, variation in pose and illumination, occlusion, etc. Typically the problem is compounded by the non-uniform distribution of labels in the training data, resulting in parts of the label space that suffer from data sparsity and a problem of label imbalance in general. Deep Convolutional Neural Networks (CNN) have shown remarkable success on a number of computer vision tasks such as object classification and face recognition. However, they too suffer from sparse and imbalanced training datasets for regression problems, even when those datasets are very large. Cumulative Attributes have previously been proposed to address the issue of label imbalance, but to date this concept has not been integrated with Deep Learning. In this work, we propose a CNN-based framework for learning regression models by using Cumulative Attributes as intermediate features. We evaluate our method on a number of tasks which includes pain intensity estimation, Facial Action Unit intensity estimation and age estimation. Our results show that the proposed method is robust to imbalance and sparsity present in the training datasets, and performs significantly better than the current methods where CNNs are learnt directly for regression.
Shashank Jaiswal, Joy Egede, Michel F. Valstar
FG3
2018 Human Behaviour-Based Automatic Depression Analysis Using Hand-Crafted Statistics and Deep Learned Spectral Features
abstract
Depression is a serious mental disorder that affects millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and need extensive participation of experts. Audio-visual automatic depression analysis systems predominantly base their predictions on very brief sequential segments, sometimes as little as one frame. Such data contains much redundant information, causes a high computational load, and negatively affects the detection accuracy. Final decision making at the sequence level is then based on the fusion of frame or segment level predictions. However, this approach loses longer term behavioural correlations, as the behaviours themselves are abstracted away by the frame-level predictions. We propose to on the one hand use automatically detected human behaviour primitives such as Gaze directions, Facial action units (AU), etc. as low-dimensional multi-channel time series data, which can then be used to create two sequence descriptors. The first calculates the sequence-level statistics of the behaviour primitives and the second casts the problem as a Convolutional Neural Network problem operating on a spectral representation of the multichannel behaviour signals. The results of depression detection (binary classification) and severity estimation (regression) experiments conducted on the AVEC 2016 DAIC-WOZ database show that both methods achieved significant improvement compared to the previous state of the art in terms of the depression severity estimation.
Siyang Song, LinLin Shen, Michel F. Valstar
FG3
2018 Predicting Folds in Poker Using Action Unit Detectors and Decision Trees
abstract
Predicting how a person will respond can be very useful, for instance when designing a strategy for negotiations. We investigate whether it is possible for machine learning and computer vision techniques to recognize a person's intentions and predict their actions based on their visually expressive behaviour, where in this paper we focus on the face. We have chosen as our setting pairs of humans playing a simplified version of poker, where the players are behaving naturally and spontaneously, albeit mediated through a computer connection. In particular, we ask if we can automatically predict whether a player is going to fold or not. We also try to answer the question of at what time point the signal for predicting if a player will fold is strongest. We use state-of-the-art FACS Action Unit detectors to automatically annotate the players facial expressions, which have been recorded on video. In addition, we use timestamps of when the player received their card and when they placed their bets, as well as the amounts they bet. Thus, the system is fully automated. We are able to predict whether a person will fold or not significantly better than chance based solely on their expressive behaviour starting three seconds before they fold.
Doratha E. Drake Vinkemeier, Michel F. Valstar, Jonathan Gratch
FG2
2018 Noise Invariant Frame Selection: A Simple Method to Address the Background Noise Problem for Text-independent Speaker Verification
abstract
The performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a simple pre-processing method called Noise Invariant Frame Selection (NIFS). Based on several noisy constraints, it selects noise invariant frames from utterances to represent speakers. Experiments conducted on the TIMIT database showed that the NIFS can significantly improve the performance of Vector Quantization (VQ), Gaussian Mixture Model-Universal Background Model (GMM-UBM) and i-vector-based speaker verification systems in different unknown noisy environments with different SNRs, in comparison to their baselines. Meanwhile, the proposed NIFS-based speaker verification systems achieves similar performance when we change the constraints (hyper-parameters) or features, which indicates that it is robust and easy to reproduce. Since NIFS is designed as a general algorithm, it could be further applied to other similar tasks.
Siyang Song, Shuimei Zhang, Björn W. Schuller, LinLin Shen, Michel F. Valstar
IJCNN5
2018 Summary for AVEC 2018: Bipolar Disorder and Cross-Cultural Affect Recognition
abstract
The eighth Audio-Visual Emotion Challenge and workshop AVEC 2018 was held in conjunction with ACM Multimedia'18. This year, the AVEC series addressed major novelties with three distinct sub-challenges: bipolar disorder classification, cross-cultural dimensional emotion recognition, and emotional label generation from individual ratings. The Bipolar Disorder Sub-challenge was based on a novel dataset of structured interviews of patients suffering from bipolar disorder (BD corpus), the Cross-cultural Emotion Sub-challenge relied on an extension of the SEWA dataset, which includes human-human interactions recorded 'in-the-wild' for the German and the Hungarian cultures, and the Gold-standard Emotion Sub-challenge was based on the RECOLA dataset, which was previously used in the AVEC series for emotion recognition. In this summary, we mainly describe participation and conditions of the AVEC Challenge.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ACM Multimedia3
2018 Guest Editorial: The Computational Face
abstract
The papers in this special section examine the concept of automated face analysis (AFA). AFA has received special attention from the computer vision and pattern recognition communities. Research progress often gives the impression that problems such as face recognition and face detection are solved, at least for some scenarios. Several aspects of face analysis remain open problems, including the implementation of large scale face recognition/detection methods for in the wild images, emotion recognition, micro-expression analysis, and others. The community keeps making rapid progress on these topics, with continual improvement of current methods and creation of new ones that push the state-of-the-art. Applications are countless, including security and video surveillance, human computer/robot interaction, communication, entertainment, and commerce, while having an important social impact in assistive technologies for education and health. The importance of face analysis, together with the vast amount of work on the subject and the latest developments in the field, motivated us to organize a special section on this theme. The scope of the compilation comprises all aspects of face analysis from a computer vision perspective. Including, but not limited to: recognition, detection, alignment, reconstruction of faces, pose estimation of faces, gaze analysis, age, emotion, gender, and facial attributes estimation, and applications among others.
Sergio Escalera, Xavier Baró, Isabelle Guyon, Hugo Jair Escalante, Georgios Tzimiropoulos, Michel F. Valstar, Maja Pantic, Jeffrey F. Cohn, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 A Functional Regression Approach to Facial Landmark Tracking
abstract
Linear regression is a fundamental building block in many face detection and tracking algorithms, typically used to predict shape displacements from image features through a linear mapping. This paper presents a Functional Regression solution to the least squares problem, which we coin Continuous Regression, resulting in the first real-time incremental face tracker. Contrary to prior work in Functional Regression, in which B-splines or Fourier series were used, we propose to approximate the input space by its first-order Taylor expansion, yielding a closed-form solution for the continuous domain of displacements. We then extend the continuous least squares problem to correlated variables, and demonstrate the generalisation of our approach. We incorporate Continuous Regression into the cascaded regression framework, and show its computational benefits for both training and testing. We then present a fast approach for incremental learning within Cascaded Continuous Regression, coined iCCR, and show that its complexity allows real-time face tracking, being 20 times faster than the state of the art. To the best of our knowledge, this is the first incremental face tracker that is shown to operate in real-time. We show that iCCR achieves state-of-the-art performance on the 300-VW dataset, the most recent, large-scale benchmark for face tracking.
Enrique Sánchez-Lozano, Georgios Tzimiropoulos, Brais Martínez, Fernando De la Torre, Michel F. Valstar
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Introduction to the Special Section on Multimedia Computing and Applications of Socio-Affective Behaviors in the Wild
abstract
No abstract available.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Jonathan Gratch, Roddy Cowie, Maja Pantic
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Fusing Deep Learned and Hand-Crafted Features of Appearance, Shape, and Dynamics for Automatic Pain Estimation
abstract
Automatic continuous time, continuous value assessment of a patient's pain from face video is highly sought after by the medical profession. Despite the recent advances in deep learning that attain impressive results in many domains, pain estimation risks not being able to benefit from this due to the difficulty in obtaining data sets of considerable size. In this work we propose a combination of hand-crafted and deep-learned features that makes the most of deep learning techniques in small sample settings. Encoding shape, appearance, and dynamics, our method significantly outperforms the current state of the art, attaining a RMSE error of less than 1 point on a 16-level pain scale, whilst simultaneously scoring a 67.3% Pearson correlation coefficient between our predicted pain level time series and the ground truth.
Joy Egede, Michel F. Valstar, Brais Martínez
FG2
2017 Automatic Detection of ADHD and ASD from Expressive Behaviour in RGBD Data
abstract
Attention Deficit Hyperactivity Disorder (ADHD) and Autism Spectrum Disorder (ASD) are neurodevelopmental conditions which impact on a significant number of children and adults. Currently, the diagnosis of such disorders is done by experts who employ standard questionnaires and look for certain behavioural markers through manual observation. Such methods for their diagnosis are not only subjective, difficult to repeat, and costly but also extremely time consuming. In this work, we present a novel methodology to aid diagnostic predictions about the presence/absence of ADHD and ASD by automatic visual analysis of a persons behaviour. To do so, we conduct the questionnaires in a computer-mediated way while recording participants with modern RGBD (Colour+Depth) sensors. In contrast to previous automatic approaches which have focussed only on detecting certain behavioural markers, our approach provides a fully automatic end-to-end system to directly predict ADHD and ASD in adults. Using state of the art facial expression analysis based on Dynamic Deep Learning and 3D analysis of behaviour, we attain classification rates of 96% for Controls vs Condition (ADHD/ASD) groups and 94% for Comorbid (ADHD+ASD) vs ASD only group. We show that our system is a potentially useful time saving contribution to the clinical diagnosis of ADHD and ASD.
Shashank Jaiswal, Michel F. Valstar, Alinda Gillott, David Daley
FG2
2017 Small Sample Deep Learning for Newborn Gestational Age Estimation
abstract
A baby's gestational age determines whether or not they are preterm, which helps clinicians decide on suitable post-natal treatment. The most accurate dating methods use Ultrasound Scan (USS) machines, but these machines are expensive, require trained personnel and cannot always be deployed to remote areas. In the absence of USS, the Ballard Score can be used, which is a manual postnatal dating method. However, this method is highly subjective and results can vary widely depending on the experience of the rater. In this paper, we present an automatic system for postnatal gestational age estimation aimed to be deployed on mobile phones, using small sets of images of a newborn's face, foot and ear. We present a novel two-stage approach that makes the most out of Convolutional Neural Networks trained on small sets of images to predict broad classes of gestational age, and then fuse the outputs of these discrete classes with a baby's weight to make fine-grained predictions of gestational age. On a purpose=collected dataset of 88 babies, experiments show that our approach attains an expected error of 6 days and is three times more accurate than the manual postnatal method (Ballard). Making use of images improves predictions by 30% compared to using weight only. This indicates that even with a very small set of data, our method is a viable candidate for postnatal gestational age estimation in areas were USS is not available.
Mercedes Torres, Michel F. Valstar, Caroline Henry, Carole Ward, Don Sharkey
FG2
2017 FERA 2017 - Addressing Head Pose in the Third Facial Expression Recognition and Analysis Challenge
abstract
The field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal recordings, or both. This makes it hard to tell how existing expression recognition approaches perform under conditions where faces appear in a wide range of poses (or camera views), displaying ecologically valid expressions. The main obstacle for assessing this is the availability of suitable data, and the challenge proposed here addresses this limitation. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) extends FERA 2015 to the estimation of Action Units occurrence and intensity under different camera views. In this paper we present the third challenge in automatic recognition of facial expressions, to be held in conjunction with the 12th IEEE conference on Face and Gesture Recognition, May 2017, in Washington, United States. Two sub-challenges are defined: the detection of AU occurrence, and the estimation of AU intensity. In this work we outline the evaluation protocol, the data used, and the results of a baseline method for both sub-challenges.
Michel F. Valstar, Enrique Sánchez-Lozano, Jeffrey F. Cohn, László A. Jeni, Jeffrey M. Girard, Zheng Zhang 0023, Lijun Yin 0001, Maja Pantic
FG1
2017 The NoXi database: multimodal recordings of mediated novice-expert interactions
abstract
We present a novel multi-lingual database of natural dyadic novice-expert interactions, named NoXi, featuring screen-mediated dyadic human interactions in the context of information exchange and retrieval. NoXi is designed to provide spontaneous interactions with emphasis on adaptive behaviors and unexpected situations (e.g. conversational interruptions). A rich set of audio-visual data, as well as continuous and discrete annotations are publicly available through a web interface. Descriptors include low level social signals (e.g. gestures, smiles), functional descriptors (e.g. turn-taking, dialogue acts) and interaction descriptors (e.g. engagement, interest, and fluidity).
Angelo Cafaro, Johannes Wagner 0001, Tobias Baur 0001, Soumia Dermouche, Mercedes Torres, Catherine Pelachaud, Elisabeth André, Michel F. Valstar
ICMI8
2017 Cumulative attributes for pain intensity estimation
abstract
Pain estimation from face video is a hard problem in automatic behaviour understanding. One major obstacle is the difficulty of collecting sufficient amounts of data, with balanced amounts of data for all pain intensity levels. To overcome this, we propose to adopt Cumulative Attributes, which assume that attributes for high pain levels with few examples are a superset of all attributes of lower pain levels. Experimental results show a consistent relative performance increase in the order of 20% regardless of features used. Our final system significantly outperforms the state of the art on the UNBC McMaster Shoulder Pain database by using cumulative attributes with Relevance Vector Regression on a combination of features, including appearance, geometric, and deep learned features.
Joy Egede, Michel F. Valstar
ICMI2
2017 Summary for AVEC 2017: Real-life Depression and Affect Challenge and Workshop
abstract
The seventh Audio-Visual Emotion Challenge and workshop AVEC 2017 was held in conjunction with ACM Multimedia'17. This year, the AVEC series addresses two distinct sub-challenges: emotion recognition and depression detection. The Affect Sub-Challenge is based on a novel dataset of human-human interactions recorded 'in-the-wild', whereas the Depression Sub-Challenge is based on the same dataset as the one used in AVEC 2016, with human-agent interactions. In this summary, we mainly describe participation and its conditions.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Jonathan Gratch, Roddy Cowie, Maja Pantic
ACM Multimedia3
2016 Cascaded Continuous Regression for Real-Time Incremental Face Tracking
Enrique Sánchez-Lozano, Brais Martínez, Georgios Tzimiropoulos, Michel F. Valstar
ECCV (8)4
2016 Ask Alice: an artificial retrieval of information agent
abstract
We present a demonstration of the ARIA framework, a modular approach for rapid development of virtual humans for information retrieval that have linguistic, emotional, and social skills and a strong personality. We demonstrate the framework's capabilities in a scenario where `Alice in Wonderland', a popular English literature book, is embodied by a virtual human representing Alice. The user can engage in an information exchange dialogue, where Alice acts as the expert on the book, and the user as an interested novice. Besides speech recognition, sophisticated audio-visual behaviour analysis is used to inform the core agent dialogue module about the user's state and intentions, so that it can go beyond simple chat-bot dialogue. The behaviour generation module features a unique new capability of being able to deal gracefully with interruptions of the agent.
Michel F. Valstar, Tobias Baur 0001, Angelo Cafaro, Alexandru Ghitulescu, Blaise Potard, Johannes Wagner 0001, Elisabeth André, Laurent Durieu, Matthew P. Aylett, Soumia Dermouche, Catherine Pelachaud, Eduardo Coutinho, Björn W. Schuller, Yue Zhang 0014, Dirk Heylen, Mariët Theune, Jelte van Waterschoot
ICMI1
2016 Playing with Social and Emotional Game Companions
abstract
This paper presents the findings of an empirical study that explores player game experience by implementing the ERiSA Framework in games. A study with Action Role-Playing Game (RPG) was designed to evaluate player interactions with game companions, who were imbued with social and emotional skill by the ERiSA Framework. Players had to complete a quest in the Skyrim game, in which players had to use social and emotional skills to obtain a sword. The results clearly show that game companions who are capable of perceiving and exhibit emotions, are perceived to have personality and can forge relationships with the players, enhancing the player experience during the game.
Andry Chowanda, Martin Flintham, Peter Blanchfield, Michel F. Valstar
IVA4
2016 Topic Switch Models for Dialogue Management in Virtual Humans
Wenjue Zhu, Andry Chowanda, Michel F. Valstar
IVA3
2016 Summary for AVEC 2016: Depression, Mood, and Emotion Recognition Workshop and Challenge
abstract
The sixth Audio-Visual Emotion Challenge and workshop AVEC 2016 was held in conjunction ACM Multimedia'16. This year the AVEC series addresses two distinct sub-challenges, multi-modal emotion recognition and audio-visual depression detection. Both sub-challenges are in a way a return to AVEC's past editions: the emotion sub-challenge is based on the same dataset as the one used in AVEC 2015, and depression analysis was previously addressed in AVEC 2013/2014. In this summary, we mainly describe participation and its conditions.
Michel F. Valstar, Jonathan Gratch, Björn W. Schuller, Fabien Ringeval, Roddy Cowie, Maja Pantic
ACM Multimedia1
2016 Deep learning the dynamic appearance and shape of facial action units
abstract
Spontaneous facial expression recognition under uncontrolled conditions is a hard task. It depends on multiple factors including shape, appearance and dynamics of the facial features, all of which are adversely affected by environmental noise and low intensity signals typical of such conditions. In this work, we present a novel approach to Facial Action Unit detection using a combination of Convolutional and Bi-directional Long Short-Term Memory Neural Networks (CNN-BLSTM), which jointly learns shape, appearance and dynamics in a deep learning manner. In addition, we introduce a novel way to encode shape features using binary image masks computed from the locations of facial landmarks. We show that the combination of dynamic CNN features and Bi-directional Long Short-Term Memory excels at modelling the temporal information. We thoroughly evaluate the contributions of each component in our system and show that it achieves state-of-the-art performance on the FERA-2015 Challenge dataset.
Shashank Jaiswal, Michel F. Valstar
WACV2
2016 L2, 1-based regression and prediction accumulation across views for robust facial landmark detection
Brais Martínez, Michel F. Valstar
Image Vis. Comput.2
2016 Cascaded regression with sparsified feature covariance matrix for facial landmark detection
abstract
This paper explores the use of context on regression-based methods for facial landmarking. Regression based methods have revolutionised facial landmarking solutions. In particular those that implicitly infer the whole shape of a structured object have quickly become the state-of-the-art. The most notable exemplar is the Supervised Descent Method (SDM). Its main characteristics are the use of the cascaded regression approach, the use of the full appearance as the inference input, and the aforementioned aim to directly predict the full shape. In this article we argue that the key aspects responsible for the success of SDM are the use of cascaded regression and the avoidance of the constrained optimisation problem that characterised most of the previous approaches. We show that, surprisingly, it is possible to achieve comparable or superior performance using only landmark-specific predictors, which are linearly combined. We reason that augmenting the input with too much context (of which using the full appearance is the extreme case) can be harmful. In fact, we experimentally found that there is a relation between the data variance and the benefits of adding context to the input. We finally devise a simple greedy procedure that makes use of this fact to obtain superior performance to the SDM, while maintaining the simplicity of the algorithm. We show extensive results both for intermediate stages devised to prove the main aspects of the argumentative line, and to validate the overall performance of two models constructed based on these considerations.
Enrique Sánchez-Lozano, Brais Martínez, Michel F. Valstar
Pattern Recognit. Lett.3
2016 The Automatic Detection of Chronic Pain-Related Expression: Requirements, Challenges and the Multimodal EmoPain Dataset
abstract
Pain-related emotions are a major barrier to effective self rehabilitation in chronic pain. Automated coaching systems capable of detecting these emotions are a potential solution. This paper lays the foundation for the development of such systems by making three contributions. First, through literature reviews, an overview of how pain is expressed in chronic pain and the motivation for detecting it in physical rehabilitation is provided. Second, a fully labelled multimodal dataset (named `EmoPain') containing high resolution multiple-view face videos, head mounted and room audio signals, full body 3D motion capture and electromyographic signals from back muscles is supplied. Natural unconstrained pain related facial expressions and body movement behaviours were elicited from people with chronic pain carrying out physical exercises. Both instructed and non-instructed exercises were considered to reflect traditional scenarios of physiotherapist directed therapy and home-based self-directed therapy. Two sets of labels were assigned: level of pain from facial expressions annotated by eight raters and the occurrence of six pain-related body behaviours segmented by four experts. Third, through exploratory experiments grounded in the data, the factors and challenges in the automated recognition of such expressions and behaviour are described, the paper concludes by discussing potential avenues in the context of these findings also highlighting differences for the two exercise scenarios addressed.
M. S. Hane Aung, Sebastian Kaltwang, Bernardino Romera-Paredes, Brais Martínez, Aneesha Singh, Matteo Cella, Michel F. Valstar, Hongying Meng, Andrew Kemp, Moshen Shafizadeh, Aaron C. Elkins, Natalie Kanakam, Amschel de Rothschild, Nick Tyler, Paul J. Watson, Amanda C. de C. Williams, Maja Pantic, Nadia Bianchi-Berthouze
IEEE Trans. Affect. Comput.7
2015 Building autonomous sensitive artificial listeners (Extended abstract)
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and non-verbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and non-verbal behaviours required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and non-verbal behaviour, since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on non-verbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling etc. We first summarise three prototype versions of the SAL scenario, in which the behaviour of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analysing and synthesising the respective behaviours. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behaviour, dialogue management, and synthesis of speaker and listener behaviour of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse, and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
ACII14
2015 Learning to Transfer: Transferring Latent Task Structures and Its Application to Person-Specific Facial Action Unit Detection
abstract
In this article we explore the problem of constructing person-specific models for the detection of facial Action Units (AUs), addressing the problem from the point of view of Transfer Learning and Multi-Task Learning. Our starting point is the fact that some expressions, such as smiles, are very easily elicited, annotated, and automatically detected, while others are much harder to elicit and to annotate. We thus consider a novel problem: all AU models for the target subject are to be learnt using person-specific annotated data for a reference AU (AU12 in our case), and no data or little data regarding the target AU. In order to design such a model, we propose a novel Multi-Task Learning and the associated Transfer Learning framework, in which we consider both relations across subjects and AUs. That is to say, we consider a tensor structure among the tasks. Our approach hinges on learning the latent relations among tasks using one single reference AU, and then transferring these latent relations to other AUs. We show that we are able to effectively make use of the annotated data for AU12 when learning other person-specific AU models, even in the absence of data for the target task. Finally, we show the excellent performance of our method when small amounts of annotated data for the target tasks are made available.
Timur R. Almaev, Brais Martínez, Michel F. Valstar
ICCV3
2015 TRIC-track: Tracking by Regression with Incrementally Learned Cascades
abstract
This paper proposes a novel approach to part-based tracking by replacing local matching of an appearance model by direct prediction of the displacement between local image patches and part locations. We propose to use cascaded regression with incremental learning to track generic objects without any prior knowledge of an object's structure or appearance. We exploit the spatial constraints between parts by implicitly learning the shape and deformation parameters of the object in an online fashion. We integrate a multiple temporal scale motion model to initialise our cascaded regression search close to the target and to allow it to cope with occlusions. Experimental results show that our tracker ranks first on the CVPR 2013 Benchmark.
Michel F. Valstar, Brais Martínez, Muhammad Haris Khan, Tony P. Pridmore
ICCV2
2015 AVEC 2015: The 5th International Audio/Visual Emotion Challenge and Workshop
abstract
The fifth Audio-Visual Emotion Challenge and workshop AVEC 2015 was held in conjunction ACM Multimedia'15. Like the previous editions of AVEC, the workshop/challenge addresses the detection of affective signals represented in audio-visual data in terms of high-level continuous dimensions. A major novelty was further introduced this year by the inclusion of the physiological modality - along with the audio and the video modalities - in the dataset. In this summary, we mainly describe participation and its conditions.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ACM Multimedia3
2014 MTS: A Multiple Temporal Scale Tracker Handling Occlusion and Abrupt Motion Variation
Muhammad Haris Khan, Michel F. Valstar, Tony P. Pridmore
ACCV (5)2
2014 Decision Level Fusion of Domain Specific Regions for Facial Action Recognition
abstract
In this paper we propose a new method for the detection of action units that relies on a novel region-based face representation and a mid-level decision layer that combines region-specific information. Different from other approaches, we do not represent the face as a regular grid based on the face location alone (holistic representation), nor by using small patches centred at iducial facial point locations (local representation). Instead, we propose to use domain knowledge regarding AU-specific facial muscle contractions to define a set of face regions covering the whole face. Therefore, as opposed to local appearance models, our face representation makes use of the full facial appearance, while the use of facial point locations to define the regions means that we obtain better-registered descriptors compared to holistic representations. Finally, we propose an AU-specific weighted sum model is used as a decision-level fusion layer in charge of combining region-specific probabilistic information. This configuration allows each classier to learning the typical appearance changes for a specific face part and reduces the dimensionality of the problem thus proving to be more robust. Our approach is evaluated on the DISFA and GEMEP-FERA datasets using two histogram-based appearance features, Local Binary Pattern and Local Phase Quantisation. We show superior performance for both the domain-specific region definition and the decision-level fusion respect to the standard approaches when it comes to automatic facial action unit detection.
Bihan Jiang, Brais Martínez, Michel F. Valstar, Maja Pantic
ICPR3
2014 A Generalized Search Method for Multiple Competing Hypotheses in Visual Tracking
abstract
Visual tracking frameworks have traditionally relied upon a single motion model such as Random Walk, and a fixed, embedded search method like Particle Filter. As a single motion model can't reliably handle various target motion types, the interest toward multiple motion models has grown over the years. The existence of multiple competing hypotheses or predictions by the multiple motion models opens up the possibility of a wider range of search methods. To search for the target in a fixed grid of equal sized cells, an integration of the Wang-Landau method and the Markov Chain Monte Carlo (MCMC) method has recently been introduced. In this paper, we generalize this search method to cells of variable size and location, where the cells are formed around the predictions generated by multiple motion models. The effectiveness of the proposed method is tested by adopting a multiple motion model tracker. Experiments show that the modified tracker has improved accuracy and better consistency over different runs compared to its original, and superior performance over state-of-the-art trackers in challenging video sequences.
Muhammad Haris Khan, Michel F. Valstar, Tony P. Pridmore
ICPR2
2014 ERiSA: Building Emotionally Realistic Social Game-Agents Companions
abstract
We propose an integrated framework for social and emotional game-agents to enhance their believability and quality of interaction, in particular by allowing an agent to forge social relations and make appropriate use of social signals. The framework is modular including sensing, interpretation, behaviour generation, and game components. We propose a generic formulation of action selection rules based on observed social and emotional signals, the agent’s personality, and the social relation between agent and player. The rules are formulated such that its variables can easily be obtained from real data. We illustrate and evaluate our framework using a simple social game called The Smile Game.
Andry Chowanda, Peter Blanchfield, Martin Flintham, Michel F. Valstar
IVA4
2014 AVEC 2014: the 4th international audio/visual emotion challenge and workshop
abstract
The fourth Audio-Visual Emotion Challenge and workshop AVEC 2014 was held in conjunction ACM Multimedia'14. Like the 2013 edition of AVEC, the workshop/challenge addresses the interpretation of social signals represented in both audio and video in terms of high-level continuous dimensions from a large number of clinically depressed patients and controls, with a sub-challenge in self-reported severity of depression estimation. In this summary, we mainly describe participation and its conditions.
Michel F. Valstar, Björn W. Schuller, Jarek Krajewski, Roddy Cowie, Maja Pantic
ACM Multimedia1
2014 A Dynamic Appearance Descriptor Approach to Facial Actions Temporal Modeling
abstract
Both the configuration and the dynamics of facial expressions are crucial for the interpretation of human facial behavior. Yet to date, the vast majority of reported efforts in the field either do not take the dynamics of facial expressions into account, or focus only on prototypic facial expressions of six basic emotions. Facial dynamics can be explicitly analyzed by detecting the constituent temporal segments in Facial Action Coding System (FACS) Action Units (AUs)-onset, apex, and offset. In this paper, we present a novel approach to explicit analysis of temporal dynamics of facial actions using the dynamic appearance descriptor Local Phase Quantization from Three Orthogonal Planes (LPQ-TOP). Temporal segments are detected by combining a discriminative classifier for detecting the temporal segments on a frame-by-frame basis with Markov Models that enforce temporal consistency over the whole episode. The system is evaluated in detail over the MMI facial expression database, the UNBC-McMaster pain database, the SAL database, the GEMEP-FERA dataset in database-dependent experiments, in cross-database experiments using the Cohn-Kanade, and the SEMAINE databases. The comparison with other state-of-the-art methods shows that the proposed LPQ-TOP method outperforms the other approaches for the problem of AU temporal segment detection, and that overall AU activation detection benefits from dynamic appearance information.
Bihan Jiang, Michel F. Valstar, Brais Martínez, Maja Pantic
IEEE Trans. Cybern.2
2013 Local Gabor Binary Patterns from Three Orthogonal Planes for Automatic Facial Expression Recognition
abstract
Facial actions cause local appearance changes over time, and thus dynamic texture descriptors should inherently be more suitable for facial action detection than their static variants. In this paper we propose the novel dynamic appearance descriptor Local Gabor Binary Patterns from Three Orthogonal Planes (LGBP-TOP), combining the previous success of LGBP-based expression recognition with TOP extensions of other descriptors. LGBP-TOP combines spatial and dynamic texture analysis with Gabor filtering to achieve unprecedented levels of recognition accuracy in real-time. While TOP features risk being sensitive to misalignment of consecutive face images, a rigorous analysis of the descriptor shows the relative robustness of LGBP-TOP to face registration errors caused by errors in rotational alignment. Experiments on the MMI Facial Expression and Cohn-Kanade databases show that for the problem of FACS Action Unit detection, LGBP-TOP outperforms both its static variant LGBP and the related dynamic appearance descriptor LBP-TOP.
Timur R. Almaev, Michel F. Valstar
ACII2
2013 Distribution-based iterative pairwise classification of emotions in the wild using LGBP-TOP
abstract
Automatic facial expression analysis promises to be a game-changer in many application areas. But before this promise can be fulfilled, it has to move from the laboratory into the wild. The Emotion Recognition in the Wild challenge provides an opportunity to develop approaches in this direction. We propose a novel Distribution-based Pairwise Iterative Classification scheme, which outperforms standard multi-class classification on this challenge data. We also verify that the recently proposed dynamic appearance descriptor, Local Gabor Patterns on Three Orthogonal Planes, performs well on this real-world data, indicating that it is robust to the type of facial misalignments that can be expected in such scenarios. Finally, we provide details of ACTC, our affective computing tools on the cloud, which is a new resource for researchers in the field of affective computing.
Timur R. Almaev, Anil Yüce, Alexandru Ghitulescu, Michel F. Valstar
ICMI4
2013 Workshop summary for the 3rd international audio/visual emotion challenge and workshop (AVEC'13)
abstract
The third Audio-Visual Emotion Challenge and workshop AVEC 2013 will be held in conjunction ACM Multimedia'13. Like the 2012 edition of AVEC, the workshop/challenge addresses the interpretation of social signals represented in both audio and video in terms of the high-level continuous dimensions arousal and valence, but importantly this year the data is that of a large number of clinically depressed patients and controls, with a sub-challenge in self-reported severity of depression estimation. Like both previous AVECs, the aim is to bring together the audio and video analysis communities.
Michel F. Valstar, Björn W. Schuller, Jarek Krajewski, Roddy Cowie, Maja Pantic
ACM Multimedia1
2013 Local Evidence Aggregation for Regression-Based Facial Point Detection
abstract
We propose a new algorithm to detect facial points in frontal and near-frontal face images. It combines a regression-based approach with a probabilistic graphical model-based face shape model that restricts the search to anthropomorphically consistent regions. While most regression-based approaches perform a sequential approximation of the target location, our algorithm detects the target location by aggregating the estimates obtained from stochastically selected local appearance information into a single robust prediction. The underlying assumption is that by aggregating the different estimates, their errors will cancel out as long as the regressor inputs are uncorrelated. Once this new perspective is adopted, the problem is reformulated as how to optimally select the test locations over which the regressors are evaluated. We propose to extend the regression-based model to provide a quality measure of each prediction, and use the shape model to restrict and correct the sampling region. Our approach combines the low computational cost typical of regression-based approaches with the robustness of exhaustive-search approaches. The proposed algorithm was tested on over 7,500 images from five databases. Results showed significant improvement over the current state of the art.
Brais Martínez, Michel F. Valstar, Xavier Binefa, Maja Pantic
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 AVEC 2012: the continuous audio/visual emotion challenge - an introduction
abstract
The second international Audio/Visual Emotion Challenge and Workshop 2012 (AVEC 2012) is introduced shortly. 34 teams from 12 countries signed up for the Challenge. The SEMAINE database serves for prediction of four-dimensional continuous affect in audio and video. For the eligible participants, final scores for the Fully-Continuous Sub-Challenge ranged between a correlation coefficient between gold standard and prediction of 0.174 and 0.456, and for Word-Level Sub-Challenge between 0.113 and 0.280.
Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ICMI2
2012 AVEC 2012: the continuous audio/visual emotion challenge
abstract
We present the second Audio-Visual Emotion recognition Challenge and workshop (AVEC 2012), which aims to bring together researchers from the audio and video analysis communities around the topic of emotion recognition. The goal of the challenge is to recognise four continuously valued affective dimensions: arousal, expectancy, power, and valence. There are two sub-challenges: in the Fully Continuous Sub-Challenge participants have to predict the values of the four dimensions at every moment during the recordings, while for the Word-Level Sub-Challenge a single prediction has to be given per word uttered by the user. This paper presents the challenge guidelines, the common data used, and the performance of the baseline system on the two tasks.
Björn W. Schuller, Michel F. Valstar, Florian Eyben, Roddy Cowie, Maja Pantic
ICMI2
2012 The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent
abstract
SEMAINE has created a large audiovisual database as a part of an iterative approach to building Sensitive Artificial Listener (SAL) agents that can engage a person in a sustained, emotionally colored conversation. Data used to build the agents came from interactions between users and an "operator” simulating a SAL agent, in different configurations: Solid SAL (designed so that operators displayed an appropriate nonverbal behavior) and Semi-automatic SAL (designed so that users' experience approximated interacting with a machine). We then recorded user interactions with the developed system, Automatic SAL, comparing the most communicatively competent version to versions with reduced nonverbal skills. High quality recording was provided by five high-resolution, high-framerate cameras, and four microphones, recorded synchronously. Recordings total 150 participants, for a total of 959 conversations with individual SAL characters, lasting approximately 5 minutes each. Solid SAL recordings are transcribed and extensively annotated: 6-8 raters per clip traced five affective dimensions and 27 associated categories. Other scenarios are labeled on the same pattern, but less fully. Additional information includes FACS annotation on selected extracts, identification of laughs, nods, and shakes, and measures of user engagement with the automatic system. The material is available through a web-accessible database.
Gary McKeown, Michel F. Valstar, Roddy Cowie, Maja Pantic, Marc Schröder 0001
IEEE Trans. Affect. Comput.2
2012 Building Autonomous Sensitive Artificial Listeners
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and nonverbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and nonverbal behaviors required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and nonverbal behavior since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on nonverbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling, etc. We first report on three prototype versions of the SAL scenario in which the behavior of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analyzing and synthesizing the respective behaviors. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behavior, dialogue management, and synthesis of speaker and listener behavior of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
IEEE Trans. Affect. Comput.14
2012 Meta-Analysis of the First Facial Expression Recognition Challenge
abstract
Automatic facial expression recognition has been an active topic in computer science for over two decades, in particular facial action coding system action unit (AU) detection and classification of a number of discrete emotion states from facial expressive imagery. Standardization and comparability have received some attention; for instance, there exist a number of commonly used facial expression databases. However, lack of a commonly accepted evaluation protocol and, typically, lack of sufficient details needed to reproduce the reported individual results make it difficult to compare systems. This, in turn, hinders the progress of the field. A periodical challenge in facial expression recognition would allow such a comparison on a level playing field. It would provide an insight on how far the field has come and would allow researchers to identify new goals, challenges, and targets. This paper presents a meta-analysis of the first such challenge in automatic recognition of facial expressions, held during the IEEE conference on Face and Gesture Recognition 2011. It details the challenge data, evaluation protocol, and the results attained in two subchallenges: AU detection and classification of facial expression imagery in terms of a number of discrete emotion categories. We also summarize the lessons learned and reflect on the future of the field of facial expression recognition in general and on possible future challenges in particular.
Michel F. Valstar, Marc Mehu, Bihan Jiang, Maja Pantic, Klaus R. Scherer
IEEE Trans. Syst. Man Cybern. Part B1
2012 Fully Automatic Recognition of the Temporal Phases of Facial Actions
abstract
Past work on automatic analysis of facial expressions has focused mostly on detecting prototypic expressions of basic emotions like happiness and anger. The method proposed here enables the detection of a much larger range of facial behavior by recognizing facial muscle actions [action units (AUs)] that compound expressions. AUs are agnostic, leaving the inference about conveyed intent to higher order decision making (e.g., emotion recognition). The proposed fully automatic method not only allows the recognition of 22 AUs but also explicitly models their temporal characteristics (i.e., sequences of temporal segments: neutral, onset, apex, and offset). To do so, it uses a facial point detector based on Gabor-feature-based boosted classifiers to automatically localize 20 facial fiducial points. These points are tracked through a sequence of images using a method called particle filtering with factorized likelihoods. To encode AUs and their temporal activation models based on the tracking data, it applies a combination of GentleBoost, support vector machines, and hidden Markov models. We attain an average AU recognition rate of 95.3% when tested on a benchmark set of deliberately displayed facial expressions and 72% when tested on spontaneous expressions.
Michel F. Valstar, Maja Pantic
IEEE Trans. Syst. Man Cybern. Part B1
2011 The First Audio/Visual Emotion Challenge and Workshop - An Introduction
Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ACII (2)2
2011 AVEC 2011-The First International Audio/Visual Emotion Challenge
Björn W. Schuller, Michel F. Valstar, Florian Eyben, Gary McKeown, Roddy Cowie, Maja Pantic
ACII (2)2
2011 A Multimodal Database for Mimicry Analysis
Xiaofan Sun, Jeroen Lichtenauer, Michel F. Valstar, Anton Nijholt, Maja Pantic
ACII (1)3
2011 String-based audiovisual fusion of behavioural events for the assessment of dimensional affect
abstract
The automatic assessment of affect is mostly based on feature-level approaches, such as distances between facial points or prosodic and spectral information when it comes to audiovisual analysis. However, it is known and intuitive that behavioural events such as smiles, head shakes or laughter and sighs also bear highly relevant information regarding a subject's affective display. Accordingly, we propose a novel string-based prediction approach to fuse such events and to predict human affect in a continuous dimensional space. Extensive analysis and evaluation has been conducted using the newly released SEMAINE database of human-to-agent communication. For a thorough understanding of the obtained results, we provide additional benchmarks by more conventional feature-level modelling, and compare these and the string-based approach to fusion of signal-based features and string-based events. Our experimental results show that the proposed string-based approach is the best performing approach for automatic prediction of Valence and Expectation dimensions, and improves prediction performance for the other dimensions when combined with at least acoustic signal-based features.
Florian Eyben, Martin Wöllmer, Michel F. Valstar, Hatice Gunes, Björn W. Schuller, Maja Pantic
FG3
2011 Action unit detection using sparse appearance descriptors in space-time video volumes
abstract
Recently developed appearance descriptors offer the opportunity for efficient and robust facial expression recognition. In this paper we investigate the merits of the family of local binary pattern descriptors for FACS Action-Unit (AU) detection. We compare Local Binary Patterns (LBP) and Local Phase Quantisation (LPQ) for static AU analysis. To encode facial expression dynamics, we extend the purely spatial representation LPQ to a dynamic texture descriptor which we call Local Phase Quantisation from Three Orthogonal Planes (LPQ-TOP), and compare this with the Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). The efficiency of these descriptors is evaluated by a fully automatic AU detection system and tested on posed and spontaneous expression data collected from the MMI and SEMAINE databases. Results show that the systems based on LPQ achieve higher accuracy rate than those using LBP, and that the systems that utilise dynamic appearance descriptors outperform those that use static appearance descriptors. Overall, our proposed LPQ-TOP method outperformed all other tested methods.
Bihan Jiang, Michel F. Valstar, Maja Pantic
FG2
2011 Come and have an emotional workout with sensitive artificial listeners!
abstract
This demonstration aims to showcase the recently completed SEMAINE system. The SEMAINE system is a publicly available, fully autonomous Sensitive Artificial Listeners (SAL) system that consists of virtual dialog partners based on audiovisual analysis and synthesis (see http://semaine.opendfki.de/wiki). The system runs in real-time, and combines incremental analysis of user behavior, dialog management, and synthesis of speaker and listener behavior of a SAL character, displayed as a virtual agent. The SAL characters intend to engage the user in a conversation by paying attention to the user's emotions and nonverbal expressions. The characters have their own emotionally defined personality. During an interaction, the characters attempt to create an emotional workout for the user by drawing her/him towards their dominant emotion, through a combination of verbal and nonverbal expressions.
Marc Schröder 0001, Sathish Pammi, Hatice Gunes, Maja Pantic, Michel F. Valstar, Roddy Cowie, Gary McKeown, Dirk Heylen, Mark ter Maat, Florian Eyben, Björn W. Schuller, Martin Wöllmer, Elisabetta Bevacqua, Catherine Pelachaud, Etienne de Sevin
FG5
2011 The first facial expression recognition and analysis challenge
Michel F. Valstar, Bihan Jiang, Marc Mehu, Maja Pantic, Klaus R. Scherer
FG1
2011 Cost-effective solution to synchronised audio-visual data capture using multiple sensors
Jeroen Lichtenauer, Jie Shen 0008, Michel F. Valstar, Maja Pantic
Image Vis. Comput.3
2010 Facial point detection using boosted regression and graph models
abstract
Finding fiducial facial points in any frame of a video showing rich naturalistic facial behaviour is an unsolved problem. Yet this is a crucial step for geometric-feature-based facial expression analysis, and methods that use appearance-based features extracted at fiducial facial point locations. In this paper we present a method based on a combination of Support Vector Regression and Markov Random Fields to drastically reduce the time needed to search for a point's location and increase the accuracy and robustness of the algorithm. Using Markov Random Fields allows us to constrain the search space by exploiting the constellations that facial points can form. The regressors on the other hand learn a mapping between the appearance of the area surrounding a point and the positions of these points, which makes detection of the points very fast and can make the algorithm robust to variations of appearance due to facial expression and moderate changes in head pose. The proposed point detection algorithm was tested on 1855 images, the results of which showed we outperform current state of the art point detectors.
Michel F. Valstar, Brais Martínez, Xavier Binefa, Maja Pantic
CVPR1
2010 The SEMAINE corpus of emotionally coloured character interactions
abstract
We have recorded a new corpus of emotionally coloured conversations. Users were recorded while holding conversations with an operator who adopts in sequence four roles designed to evoke emotional reactions. The operator and the user are seated in separate rooms; they see each other through teleprompter screens, and hear each other through speakers. To allow high quality recording, they are recorded by five high-resolution, high framerate cameras, and by four microphones. All sensor information is recorded synchronously, with an accuracy of 25 μs. In total, we have recorded 20 participants, for a total of 100 character conversational and 50 non-conversational recordings of approximately 5 minutes each. All recorded conversations have been fully transcribed and annotated for five affective dimensions and partially annotated for 27 other dimensions. The corpus has been made available to the scientific community through a web-accessible database.
Gary McKeown, Michel F. Valstar, Roddy Cowie, Maja Pantic
ICME2
2010 The Detection of Concept Frames Using Clustering Multi-instance Learning
abstract
The classification of sequences requires the combination of information from different time points. In this paper the detection of facial expressions is considered. Experiments on the detection of certain facial muscle activations in videos show that it is not always required to model the sequences fully, but that the presence of specific frames (the concept frame) can be sufficient for a reliable detection of certain facial expression classes. For the detection of these concept frames a standard classifier is often sufficient, although a more advanced clustering approach performs better in some cases.
David M. J. Tax, E. Hendriks, Michel F. Valstar, Maja Pantic
ICPR3
2009 Cost-Effective Solution to Synchronized Audio-Visual Capture Using Multiple Sensors
abstract
Applications such as surveillance and human motion capture require high-bandwidth recording from multiple cameras. Furthermore, the recent increase in research on sensor fusion has raised the demand on synchronization accuracy between video, audio and other sensor modalities. Previously, capturing synchronized, high resolution video from multiple cameras required complex, inflexible and expensive solutions. Our experiments show that a single PC, built from contemporary low-cost computer hardware, could currently handle up to 470MB/s of input data. This allows capturing from 18 cameras of 780x580pixels at 60fps each, or 36 cameras at 30fps. Furthermore, we achieve accurate synchronization between audio, video and additional sensors, by recording audio together with sensor trigger- or timestamp signals, using a multi-channel audio input. In this way, each sensor modality can be captured with separate software and hardware, allowing maximal flexibility with minimal cost.
Jeroen Lichtenauer, Michel F. Valstar, Jie Shen 0008, Maja Pantic
AVSS2
2008 Emotionally aware automated portrait painting demonstration
abstract
We propose to demonstrate the emotionally aware painting fool, a novel system that combines a machine vision system able to recognise emotions with a non-photorealistic rendering (NPR) system to automatically produce portraits of the sitter in an emotionally enhanced style. During the demonstration, the vision system records a short video clip of a person showing a basic emotion. The system then analyses this video clip, locating facial features and tracking their motion. Using this tracking data, the system analyses which emotion was expressed and what the temporal dynamics of the expression were. This information is then passed to the NPR software. The detected emotion is used to choose appropriate (simulated) art materials, colour palettes, abstraction methods and painting styles, so that the rendered image may heighten the emotion being expressed. The live demonstration shows how each element of the emotionally aware painting fool functions and produces a portrait in approximately 7 minutes.
Michel F. Valstar, Simon Colton, Maja Pantic
FG1
2008 Multivariate Statistical Analysis of Whole Brain Structural Networks Obtained Using Probabilistic Tractography
Emma C. Robinson, Michel F. Valstar, Alexander Hammers, Anders Ericsson, A. David Edwards, Daniel Rueckert
MICCAI (1)2
2007 How to distinguish posed from spontaneous smiles using geometric features
abstract
Automatic distinction between posed and spontaneous expressions is an unsolved problem. Previously cognitive sciences' studies indicated that the automatic separation of posed from spontaneous expressions is possible using the face modality alone. However, little is known about the information contained in head and shoulder motion. In this work, we propose to (i) distinguish between posed and spontaneous smiles by fusing the head, face, and shoulder modalities, (ii) investigate which modalities carry important information and how the information of the modalities relate to each other, and (iii) to which extent the temporal dynamics of these signals attribute to solving the problem. We use a cylindrical head tracker to track the head movements and two particle filtering techniques to track the facial and shoulder movements. Classification is performed by kernel methods combined with ensemble learning techniques. We investigated two aspects of multimodal fusion: the level of abstraction (i.e., early, mid-level, and late fusion) and the fusion rule used (i.e., sum, product and weight criteria). Experimental results from 100 videos displaying posed smiles and 102 videos displaying spontaneous smiles are presented. Best results were obtained with late fusion of all modalities when 94.0% of the videos were classified correctly.
Michel F. Valstar, Hatice Gunes, Maja Pantic
ICMI1
2006 Biologically vs. Logic Inspired Encoding of Facial Actions and Emotions in Video
abstract
Automatic facial expression analysis is an important aspect of human machine interaction as the face is an important communicative medium. We use our face to signal interest, disagreement, intentions or mood through subtle facial motions and expressions. Work on automatic facial expression analysis can roughly be divided into the recognition of prototypic facial expressions such as the six basic emotional states and the recognition of atomic facial muscle actions (action units, AUs). Detection of AUs rather than emotions makes facial expression detection independent of culture-dependent interpretation, reduces the dimensionality of the problem and reduces the amount of training data required. Classic psychological studies suggest that humans consciously map AUs onto the basic emotion categories using a finite number of rules. On the other hand, recent studies suggest that humans recognize emotions unconsciously with a process that is perhaps best modeled by artificial neural networks (ANNs). This paper investigates these two claims. A comparison is made between detection of emotions directly from features vs. a two-step approach where we first detect AUs and use the AUs as input to either a rulebase or an ANN to recognize emotions. The results suggest that the two-step approach is possible with a small loss of accuracy and that biologically inspired classification techniques outperform those that approach the classification problem from a logical perspective, suggesting that biologically inspired classifiers are more suitable for computer-based analysis of facial behavior than logic inspired methods
Michel F. Valstar, Maja Pantic
ICME1
2006 Spontaneous vs. posed facial behavior: automatic analysis of brow actions
abstract
Past research on automatic facial expression analysis has focused mostly on the recognition of prototypic expressions of discrete emotions rather than on the analysis of dynamic changes over time, although the importance of temporal dynamics of facial expressions for interpretation of the observed facial behavior has been acknowledged for over 20 years. For instance, it has been shown that the temporal dynamics of spontaneous and volitional smiles are fundamentally different from each other. In this work, we argue that the same holds for the temporal dynamics of brow actions and show that velocity, duration, and order of occurrence of brow actions are highly relevant parameters for distinguishing posed from spontaneous brow actions. The proposed system for discrimination between volitional and spontaneous brow actions is based on automatic detection of Action Units (AUs) and their temporal segments (onset, apex, offset) produced by movements of the eyebrows. For each temporal segment of an activated AU, we compute a number of mid-level feature parameters including the maximal intensity, duration, and order of occurrence. We use Gentle Boost to select the most important of these parameters. The selected parameters are used further to train Relevance Vector Machines to determine per temporal segment of an activated AU whether the action was displayed spontaneously or volitionally. Finally, a probabilistic decision function determines the class (spontaneous or posed) for the entire brow action. When tested on 189 samples taken from three different sets of spontaneous and volitional facial data, we attain a 90.7% correct recognition rate.
Michel F. Valstar, Maja Pantic, Zara Ambadar, Jeffrey F. Cohn
ICMI1
2005 Web-based database for facial expression analysis
abstract
In the last decade, the research topic of automatic analysis of facial expressions has become a central topic in machine vision research. Nonetheless, there is a glaring lack of a comprehensive, readily accessible reference set of face images that could be used as a basis for benchmarks for efforts in the field. This lack of easily accessible, suitable, common testing resource forms the major impediment to comparing and extending the issues concerned with automatic facial expression analysis. In this paper, we discuss a number of issues that make the problem of creating a benchmark facial expression database difficult. We then present the MMI facial expression database, which includes more than 1500 samples of both static images and image sequences of faces in frontal and in profile view displaying various expressions of emotion, single and multiple facial muscle activation. It has been built as a Web-based direct-manipulation application, allowing easy access and easy search of the available images. This database represents the most comprehensive reference set of images for studies on facial expression analysis to date.
Maja Pantic, Michel F. Valstar, Ron Rademaker, Ludo Maat
ICME2