Fabien Ringeval

dblp:39/4933 · DBLP profile ↗
← Back
50ranked-venue papers
14as first author
16since 2021 · last 2025
0000-0002-9213-4529ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 26 · 6 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals
abstract
This study investigates the ability of GPT models (ChatGPT, GPT-4 and GPT-4o) to generate dialogue summaries that adhere to human guidelines. Our evaluation involved experimenting with various prompts to guide the models in complying with guidelines on two datasets: DialogSum (English social conversations) and DECODA (French call center interactions). Human evaluation, based on summarization guidelines, served as the primary assessment method, complemented by extensive quantitative and qualitative analyses. Our findings reveal a preference for GPT-generated summaries over those from task-specific pre-trained models and reference summaries, highlighting GPT models’ ability to follow human guidelines despite occasionally producing longer outputs and exhibiting divergent lexical and structural alignment with references. The discrepancy between ROUGE, BERTScore, and human evaluation underscores the need for more reliable automatic evaluation metrics.
Yongxin Zhou 0004, Fabien Ringeval, François Portet
INLG2
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia12
2025 THERADIA WoZ: An Ecological Corpus for Appraisal-Based Affect Research in Healthcare
abstract
We present THERADIA WoZ, an ecological corpus designed for audiovisual research on affect in healthcare. Two groups of senior individuals, consisting of 52 healthy participants and 9 individuals with Mild Cognitive Impairment (MCI), performed Computerised Cognitive Training (CCT) exercises while receiving support from a virtual assistant, tele-operated by a human in the role of a Wizard-of-Oz (WoZ). The audiovisual expressions produced by the participants were fully transcribed, and partially annotated based on dimensions derived from recent appraisal theory models, including novelty, intrinsic pleasantness, goal conduciveness, and coping. Additionally, the annotations included 23 affective labels from the literature of achievement affects. We present the data collection, transcription, and annotation protocols, alongside a detailed analysis of the annotated dimensions and labels. Baseline methods and results for their automatic prediction are also presented. Results reveal that the dimensions of appraisal theory can be predicted, with the performance varying across different modalities. The corpus aims to serve as a valuable resource for researchers in affective computing, and is made available to both industry and academia.
Hippolyte Fournier, Sina Alisamir, Safaa Azzakhnini, Isabella Zsoldos, Eléonore Trân, Gérard Bailly, Frédéric Elisei, Béatrice Bouchot, Brice Varini, Patrick Constant, Joan Fruitet, Franck Tarpin-Bernard, Solange Rossato, François Portet, Olivier Koenig, Hanna Chainay, Fabien Ringeval
IEEE Trans. Affect. Comput.17
2024 PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
abstract
Automatic dialogue summarization is a well-established task with the goal of distilling the most crucial information from human conversations into concise textual summaries. However, most existing research has predominantly focused on summarizing factual information, neglecting the affective content, which can hold valuable insights for analyzing, monitoring, or facilitating human interactions. In this paper, we introduce and assess a set of measures PSentScore, aimed at quantifying the preservation of affective content in dialogue summaries. Our findings indicate that state-of-the-art summarization models do not preserve well the affective content within their summaries. Moreover, we demonstrate that a careful selection of the training set for dialogue samples can lead to improved preservation of affective content in the generated summaries, albeit with a minor reduction in content-related metrics.
Yongxin Zhou 0004, Fabien Ringeval, François Portet
LREC/COLING2
2024 Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains
abstract
Pretrained Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important domains: transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations: continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets.
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Audibert, Cécile Macaire, Adrien Pupier, Yongxin Zhou 0004, Mathilde Aguiar, Felix Herron, Magali Norré, Massih-Reza Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
LREC/COLING22
2024 REACT 2024: the Second Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024.
Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
FG10
2024 EVAC 2024 - Empathic Virtual Agent Challenge: Appraisal-based Recognition of Affective States
abstract
As autonomous interactive agents become increasingly prevalent, it is crucial for these virtual agents to understand and respond to both our verbal content and emotions, enabling deeper interactions. Despite significant advancements in the automatic recognition and understanding of human speech, challenges remain in accurately identifying and addressing the nuances of human emotions, hindering the development of more empathic virtual agents. We believe that empathic virtual agents should excel in three key tasks: (i) recognising spontaneous emotional expressions alongside understanding verbal content, (ii) generating timely and appropriate responses, and (iii) providing insightful feedback while comprehending user responses. To advance the development of empathic agents, we introduce the first Empathic Virtual Agent Challenge (EVAC). The inaugural edition focuses on robustly recognising spontaneous human expressions during interactions with a virtual agent, using the newly introduced THERADIA WoZ dataset. This paper provides an overview of the baseline systems operated on the pseudonymised version of the corpus on the two following modeling tasks: core affect presence and intensity, and appraisal based dimensions.
Fabien Ringeval, Björn W. Schuller, Gérard Bailly, Safaa Azzakhnini, Hippolyte Fournier
ICMI1
2024 LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Comput. Speech Lang.20
2023 Speaker Identification Enhancement Using Emotional Features
Jihed Jabnoun, Ahmed Zrigui, Anwer Slimi, Fabien Ringeval, Didier Schwab, Mounir Zrigui
ICCCI4
2023 Verbal and nonverbal feedback signals in response to increasing levels of miscommunication
abstract
International audience
Maëva Garnier, Éric Le Ferrand, Fabien Ringeval
INTERSPEECH3
2023 REACT2023: The First Multiple Appropriate Facial Reaction Generation Challenge
abstract
The Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023.
Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
ACM Multimedia9
2022 Multi-Corpus Affect Recognition with Emotion Embeddings and Self-Supervised Representations of Speech
abstract
Speech emotion recognition systems use data-driven machine learning techniques that rely on annotated corpora. To achieve a usable performance in real-life, we need to exploit multiple different datasets since each one can shed the light on some specific expression of affect. However, different corpora use subjectively defined annotation schemes, which poses a challenge to train a model that can sense similar emotions across different corpora. Here, we propose a method that can relate similar emotions across corpora without being explicitly trained for it. Our method relies on self-supervised representations, which can provide us with highly contextualised speech representations, and multi-task learning paradigms. This allows to train on different corpora without changing their labelling schemes. The results show that by fine-tuning self-supervised representations on each corpus separately, we can significantly improve the state of the art within-corpus performance. We further demonstrate that by using multiple corpora during the training of the same model, we can improve the cross-corpus performance, and show that our emotion embeddings can effectively recognise the same emotions across different corpora.
Sina Alisamir, Fabien Ringeval, François Portet
ACII2
2022 Effectiveness of French Language Models on Abstractive Dialogue Summarization Task
abstract
Pre-trained language models have established the state-of-the-art on various natural language processing tasks, including dialogue summarization, which allows the reader to quickly access key information from long conversations in meetings, interviews or phone calls. However, such dialogues are still difficult to handle with current models because the spontaneity of the language involves expressions that are rarely present in the corpora used for pre-training the language models. Moreover, the vast majority of the work accomplished in this field has been focused on English. In this work, we present a study on the summarization of spontaneous oral dialogues in French using several language specific pre-trained models: BARThez, and BelGPT-2, as well as multilingual pre-trained models: mBART, mBARThez, and mT5. Experiments were performed on the DECODA (Call Center) dialogue corpus whose task is to generate abstractive synopses from call center conversations between a caller and one or several agents depending on the situation. Results show that the BARThez models offer the best performance far above the previous state-of-the-art on DECODA. We further discuss the limits of such pre-trained models and the challenges that must be addressed for summarizing spontaneous dialogues.
Yongxin Zhou 0004, François Portet, Fabien Ringeval
LREC3
2021 LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech
abstract
Self-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech.
Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Interspeech16
2021 Social Signals and Multimedia: Past, Present, Future
abstract
The rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events.
Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency
ACM Multimedia5
2021 SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild
abstract
Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life. Accurately annotated real-world data are the crux in devising such systems. However, existing databases usually consider controlled settings, low demographic variability, and a single task. In this paper, we introduce the SEWA database of more than 2,000 minutes of audio-visual data of 398 people coming from six cultures, 50 percent female, and uniformly spanning the age range of 18 to 65 years old. Subjects were recorded in two different contexts: while watching adverts and while discussing adverts in a video chat. The database includes rich annotations of the recordings in terms of facial landmarks, facial action units (FAU), various vocalisations, mirroring, and continuously valued valence, arousal, liking, agreement, and prototypic examples of (dis)liking. This database aims to be an extremely valuable resource for researchers in affective computing and automatic human sensing and is expected to push forward the research in human behaviour analysis, including cultural studies. Along with the database, we provide extensive baseline experiments for automatic FAU detection and automatic valence, arousal, and (dis)liking intensity estimation.
Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen 0008, Maximilian Schmitt, Fabien Ringeval, Jing Han 0010, Vedhas Pandit, Antoine Toisoul, Björn W. Schuller, Kam Star, Elnar Hajiyev, Maja Pantic
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 AVEC'19: Audio/Visual Emotion Challenge and Workshop
abstract
The ninth Audio-Visual Emotion Challenge and workshop AVEC 2019 was held in conjunction with ACM Multimedia'19. This year, the AVEC series addressed major novelties with three distinct tasks: State-of-Mind Sub-challenge (SoMS), Detecting Depression with Artificial Intelligence Sub-challenge (DDS), and Cross-cultural Emotion Sub-challenge (CES). The SoMS was based on a novel dataset (USoM corpus) that includes self-reported mood (10-point Likert scale) after the narrative of personal stories (two positive and two negative). The DDS was based on a large extension of the DAIC-WOZ corpus (c.f. AVEC 2016) that includes new recordings of patients suffering from depression with the virtual agent conducting the interview being, this time, wholly driven by AI, i.e., without any human intervention. The CES was based on the SEWA dataset (c.f. AVEC 2018) that has been extended with the inclusion of new participants in order to investigate how emotion knowledge of Western European cultures (German, Hungarian) can be transferred to the Chinese culture. In this summary, we mainly describe participation and conditions of the AVEC Challenge.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Nicholas Cummins, Roddy Cowie, Maja Pantic
ACM Multimedia1
2019 From Speech to Facial Activity: Towards Cross-modal Sequence-to-Sequence Attention Networks
abstract
Multimodal data sources offer the possibility to capture and model interactions between modalities, leading to an improved understanding of underlying relationships. In this regard, the work presented in this paper explores the relationship between facial muscle movements and speech signals. Specifically, we explore the efficacy of different sequence-to-sequence neural network architectures for the task of predicting Facial Action Coding System Action Units (AUs) from one of two acoustic feature representations extracted from speech signals, namely the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPs) or the Interspeech Computational Paralinguistics Challenge features set (ComParE). Furthermore, these architectures were enhanced by two different attention mechanisms (intra- and inter-attention) and various state-of-the-art network settings to improve prediction performance. Results indicate that a sequence-to-sequence model with inter-attention can achieve on average an Unweighted Average Recall (UAR) of 65.9 % for AU onset, 67.8 % for AU apex (both eGeMAPs), 79.7 % for AU offset and 65.3 % for AU occurrence (both ComParE) detection over all AUs.
Lukas Stappen, Vincent Karas, Nicholas Cummins, Fabien Ringeval, Klaus R. Scherer, Björn W. Schuller
MMSP4
2019 Affective and behavioural computing: Lessons learnt from the First Computational Paralinguistics Challenge
Björn W. Schuller, Felix Weninger, Yue Zhang 0014, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli, Klaus R. Scherer, Mohamed Chetouani, Marcello Mortillaro
Comput. Speech Lang.4
2018 Towards Conditional Adversarial Training for Predicting Emotions from Speech
abstract
Motivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consists of two networks, trained in an adversarial manner: The first network tries to predict emotion from acoustic features, while the second network aims at distinguishing between the predictions provided by the first network and the emotion labels from the database using the acoustic features as conditional information. We evaluate the performance of the proposed conditional adversarial training framework on the widely used emotion database RECOLA. Experimental results show that the proposed training strategy outperforms the conventional training method, and is comparable with, or even superior to other recently reported approaches, including deep and end-to-end learning.
Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Fabien Ringeval, Björn W. Schuller
ICASSP4
2018 Automatic Recognition of Affective Laughter in Spontaneous Dyadic Interactions from Audiovisual Signals
abstract
Laughter is a highly spontaneous behavior that frequently occurs during social interactions. It serves as an expressive-communicative social signal which conveys a large spectrum of affect display. Even though many studies have been performed on the automatic recognition of laughter -- or emotion -- from audiovisual signals, very little is known about the automatic recognition of emotion conveyed by laughter. In this contribution, we provide insights on emotional laughter by extensive evaluations carried out on a corpus of dyadic spontaneous interactions, annotated with dimensional labels of emotion (arousal and valence). We evaluate, by automatic recognition experiments and correlation based analysis, how different categories of laughter, such as unvoiced laughter, voiced laughter, speech laughter, and speech (non-laughter) can be differentiated from audiovisual features, and to which extent they might convey different emotions. Results show that voiced laughter performed best in the automatic recognition of arousal and valence for both audio and visual features. The context of production is further analysed and results show that, acted and spontaneous expressions of laughter produced by a same person can be differentiated from audiovisual signals, and multilingual induced expressions can be differentiated from those produced during interactions.
Reshmashree B. Kantharaju, Fabien Ringeval, Laurent Besacier
ICMI2
2018 Bags in Bag: Generating Context-Aware Bags for Tracking Emotions from Speech
abstract
International audience
Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Zhao Ren, Fabien Ringeval, Björn W. Schuller
INTERSPEECH5
2018 Summary for AVEC 2018: Bipolar Disorder and Cross-Cultural Affect Recognition
abstract
The eighth Audio-Visual Emotion Challenge and workshop AVEC 2018 was held in conjunction with ACM Multimedia'18. This year, the AVEC series addressed major novelties with three distinct sub-challenges: bipolar disorder classification, cross-cultural dimensional emotion recognition, and emotional label generation from individual ratings. The Bipolar Disorder Sub-challenge was based on a novel dataset of structured interviews of patients suffering from bipolar disorder (BD corpus), the Cross-cultural Emotion Sub-challenge relied on an extension of the SEWA dataset, which includes human-human interactions recorded 'in-the-wild' for the German and the Hungarian cultures, and the Gold-standard Emotion Sub-challenge was based on the RECOLA dataset, which was previously used in the AVEC series for emotion recognition. In this summary, we mainly describe participation and conditions of the AVEC Challenge.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ACM Multimedia1
2018 A closed-form solution to the graph total variation problem for continuous emotion profiling in noisy environment
Shaoling Jing, Xia Mao, Lijiang Chen, Maria Colomba Comes, Arianna Mencattini, Grazia Raguso, Fabien Ringeval, Björn W. Schuller, Corrado Di Natale, Eugenio Martinelli
Speech Commun.7
2018 Introduction to the Special Section on Multimedia Computing and Applications of Socio-Affective Behaviors in the Wild
abstract
No abstract available.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Jonathan Gratch, Roddy Cowie, Maja Pantic
ACM Trans. Multim. Comput. Commun. Appl.1
2017 Reconstruction-error-based learning for continuous emotion recognition in speech
abstract
To advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoencoder for reconstructing the original features, and the second is employed to perform emotion prediction. The RE of the original features is used as a complementary descriptor, which is merged with the original features and fed to the second model. The assumption of this framework is that the system has the ability to learn its `drawback' which is expressed by the RE. Experimental results on the RECOLA database show that the proposed framework significantly outperforms the baseline systems without any RE information in terms of Concordance Correlation Coefficient (.729 vs .710 for arousal, .360 vs .237 for valence), and also significantly overcomes other state-of-the-art methods.
Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller
ICASSP3
2017 Prediction-based learning for continuous emotion recognition in speech
abstract
In this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different regression models cooperatively. To this end, we take two widely used regression models for example, i. e., support vector regression and bidirectional long short-term memory recurrent neural network. We concatenate the two models in a tandem structure by different ways, forming a united cascaded framework. The outputs predicted by the former model are combined together with the original features as the input of the following model for final predictions. The experimental results on a time- and value-continuous spontaneous emotion database (RECOLA) show that, the prediction-based learning framework significantly outperforms the individual models for both arousal and valence dimensions, and provides significantly better results in comparison to other state-of-the-art methodologies on this corpus.
Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller
ICASSP3
2017 End-to-end learning for dimensional emotion recognition from physiological signals
abstract
Dimensional emotion recognition from physiological signals is a highly challenging task. Common methods rely on hand-crafted features that do not yet provide the performance necessary for real-life application. In this work, we exploit a series of convolutional and recurrent neural networks to predict affect from physiological signals, such as electrocardiogram and electrodermal activity, directly from the raw time representation. The motivation behind this so-called end-to-end approach is that, ultimately, the network learns an intermediate representation of the physiological signals that better suits the task at hand. Experimental evaluations show that, this very first study on end-to-end learning of emotion based on physiology, yields significantly better performance in comparison to existing work on the challenging RECOLA database, which includes fully spontaneous affective behaviors displayed during naturalistic interactions. Furthermore, we gain better understanding of the models' inner representations, by demonstrating that some cells' activations in the convolutional network are correlated to a large extent with hand-crafted features.
Gil Keren, Tobias Kirschstein, Erik Marchi, Fabien Ringeval, Björn W. Schuller
ICME4
2017 Summary for AVEC 2017: Real-life Depression and Affect Challenge and Workshop
abstract
The seventh Audio-Visual Emotion Challenge and workshop AVEC 2017 was held in conjunction with ACM Multimedia'17. This year, the AVEC series addresses two distinct sub-challenges: emotion recognition and depression detection. The Affect Sub-Challenge is based on a novel dataset of human-human interactions recorded 'in-the-wild', whereas the Depression Sub-Challenge is based on the same dataset as the one used in AVEC 2016, with human-agent interactions. In this summary, we mainly describe participation and its conditions.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Jonathan Gratch, Roddy Cowie, Maja Pantic
ACM Multimedia1
2017 Strength modelling for real-worldautomatic continuous affect recognition from audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Nicholas Cummins, Fabien Ringeval, Björn W. Schuller
Image Vis. Comput.4
2017 Continuous Estimation of Emotions in Speech by Dynamic Cooperative Speaker Models
abstract
Research on automatic emotion recognition from speech has recently focused on the prediction of time-continuous dimensions (e.g., arousal and valence) of spontaneous and realistic expressions of emotion, as found in real-life interactions. However, the automatic prediction of such emotions poses several challenges, such as the subjectivity found in the definition of a gold-standard from a pool of raters and the issue of data scarcity in training models. In this work, we introduce a novel emotion recognition system, based on ensembles of single-speaker-regression-models. The estimation of emotion is provided by combining a subset of the initial pool of single-speaker-regression-models selecting those that are most concordant among them. The proposed approach allows the addition or removal of speakers from the ensemble without the necessity to re-build the entire recognition system. The simplicity of this aggregation strategy, coupled with the flexibility assured by the modular architecture, and the promising results observed on the RECOLA database highlight the potential implications of the proposed method in a real-life scenario and in particular in web-based applications.
Arianna Mencattini, Eugenio Martinelli, Fabien Ringeval, Björn W. Schuller, Corrado Di Natale
IEEE Trans. Affect. Comput.3
2016 Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network
abstract
The automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers while being able to model the context. Whereas the latter has been tackled by the use of Long Short-Term Memory (LSTM) networks, the former is still under very active investigations, even though more than a decade of research has provided a large set of acoustic descriptors. In this paper, we propose a solution to the problem of ‘context-aware’ emotional relevant feature extraction, by combining Convolutional Neural Networks (CNNs) with LSTM networks, in order to automatically learn the best representation of the speech signal directly from the raw time representation. In this novel work on the so-called end-to-end speech emotion recognition, we show that the use of the proposed topology significantly outperforms the traditional approaches based on signal processing techniques for the prediction of spontaneous and natural emotions on the RECOLA database.
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou
ICASSP2
2016 Enhanced semi-supervised learning for multimodal emotion recognition
abstract
Semi-Supervised Learning (SSL) techniques have found many applications where labeled data is scarce and/or expensive to obtain. However, SSL suffers from various inherent limitations that limit its performance in practical applications. A central problem is that the low performance that a classifier can deliver on challenging recognition tasks reduces the trustability of the automatically labeled data. Another related issue is the noise accumulation problem - instances that are misclassified by the system are still used to train it in future iterations. In this paper, we propose to address both issues in the context of emotion recognition. Initially, we exploit the complementarity between audio-visual features to improve the performance of the classifier during the supervised phase. Then, we iteratively re-evaluate the automatically labeled instances to correct possibly mislabeled data and this enhances the overall confidence of the system's predictions. Experimental results performed on the RECOLA database demonstrate that our methodology delivers a strong performance in the classification of high/low emotional arousal (UAR = 76.5%), and significantly outperforms traditional SSL methods by at least 5.0% (absolute gain).
Zixing Zhang 0001, Fabien Ringeval, Eduardo Coutinho, Erik Marchi, Björn W. Schuller
ICASSP2
2016 Discriminatively Trained Recurrent Neural Networks for Continuous Dimensional Emotion Recognition from Audio
Felix Weninger, Fabien Ringeval, Erik Marchi, Björn W. Schuller
IJCAI2
2016 Automatic Analysis of Typical and Atypical Encoding of Spontaneous Emotion in the Voice of Children
abstract
International audience
Fabien Ringeval, Erik Marchi, Charline Grossard, Jean Xavier, Mohamed Chetouani, Björn W. Schuller
INTERSPEECH1
2016 At the Border of Acoustics and Linguistics: Bag-of-Audio-Words for the Recognition of Emotions in Speech
abstract
Recognition of natural emotion in speech is a challenging task.Different methods have been proposed to tackle this complex task, such as acoustic feature brute-forcing or even endto-end learning.Recently, bag-of-audio-words (BoAW) representations of acoustic low-level descriptors (LLDs) have been employed successfully in the domain of acoustic event classification and other audio recognition tasks.In this approach, feature vectors of acoustic LLDs are quantised according to a learnt codebook of audio words.Then, a histogram of the occurring 'words' is built.Despite their massive potential, BoAW have not been thoroughly studied in emotion recognition.Here, we propose a method using BoAW created only of mel-frequency cepstral coefficients (MFCCs).Support vector regression is then used to predict emotion continuously in time and value, such as in the dimensions arousal and valence.We compare this approach with the computation of functionals based on the MFCCs and perform extensive evaluations on the RECOLA database, which features spontaneous and natural emotions.Results show that, BoAW representation of MFCCs does not only perform significantly better than functionals, but also outperforms by far most of recently published deep learning approaches, including convolutional and recurrent networks.
Maximilian Schmitt, Fabien Ringeval, Björn W. Schuller
INTERSPEECH2
2016 Facing Realism in Spontaneous Emotion Recognition from Speech: Feature Enhancement by Autoencoder with LSTM Neural Networks
abstract
International audience
Zixing Zhang 0001, Fabien Ringeval, Jing Han 0010, Erik Marchi, Björn W. Schuller
INTERSPEECH2
2016 Spectral and Cepstral Audio Noise Reduction Techniques in Speech Emotion Recognition
abstract
Signal noise reduction can improve the performance of machine learning systems dealing with time signals such as audio. Real-life applicability of these recognition technologies requires the system to uphold its performance level in variable, challenging conditions such as noisy environments. In this contribution, we investigate audio signal denoising methods in cepstral and log-spectral domains and compare them with common implementations of standard techniques. The different approaches are first compared generally using averaged acoustic distance metrics. They are then applied to automatic recognition of spontaneous and natural emotions under simulated smartphone-recorded noisy conditions. Emotion recognition is implemented as support vector regression for continuous-valued prediction of arousal and valence on a realistic multimodal database. In the experiments, the proposed methods are found to generally outperform standard noise reduction algorithms.
Jouni Pohjalainen, Fabien Ringeval, Zixing Zhang 0001, Björn W. Schuller
ACM Multimedia2
2016 Summary for AVEC 2016: Depression, Mood, and Emotion Recognition Workshop and Challenge
abstract
The sixth Audio-Visual Emotion Challenge and workshop AVEC 2016 was held in conjunction ACM Multimedia'16. This year the AVEC series addresses two distinct sub-challenges, multi-modal emotion recognition and audio-visual depression detection. Both sub-challenges are in a way a return to AVEC's past editions: the emotion sub-challenge is based on the same dataset as the one used in AVEC 2015, and depression analysis was previously addressed in AVEC 2013/2014. In this summary, we mainly describe participation and its conditions.
Michel F. Valstar, Jonathan Gratch, Björn W. Schuller, Fabien Ringeval, Roddy Cowie, Maja Pantic
ACM Multimedia4
2015 Face reading from speech - predicting facial action units from audio cues
abstract
The automatic recognition of facial behaviours is usually achieved through the detection of particular FACS Action Unit (AU), which then makes it possible to analyse the affective behaviours expressed in the face.Despite the fact that advanced techniques have been proposed to extract relevant facial descriptors, the processing of real-life data, i. e., recorded in unconstrained environments, makes the automatic detection of FACS AU much more challenging compared to constrained recordings, such as posed faces, and even impossible when the corresponding parts of the face are masked or subject to low or no illumination.We present in this paper the very first attempt in using acoustic cues for the automatic detection of FACS AU, as an alternative way to obtain information from the face when such data are not available.Results show that features extracted from the voice can be effectively used to predict different types of FACS AU, and that the best performance are obtained for the prediction of the apex, in comparison to the prediction of onset, offset and occurrence.
Fabien Ringeval, Erik Marchi, Marc Mehu, Klaus R. Scherer, Björn W. Schuller
INTERSPEECH1
2015 AVEC 2015: The 5th International Audio/Visual Emotion Challenge and Workshop
abstract
The fifth Audio-Visual Emotion Challenge and workshop AVEC 2015 was held in conjunction ACM Multimedia'15. Like the previous editions of AVEC, the workshop/challenge addresses the detection of affective signals represented in audio-visual data in terms of high-level continuous dimensions. A major novelty was further introduced this year by the inclusion of the physiological modality - along with the audio and the video modalities - in the dataset. In this summary, we mainly describe participation and its conditions.
Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Roddy Cowie, Maja Pantic
ACM Multimedia1
2015 Prediction of asynchronous dimensional emotion ratings from audiovisual and physiological data
Fabien Ringeval, Florian Eyben, Eleni Kroupi, Anil Yüce, Jean-Philippe Thiran, Touradj Ebrahimi, Denis Lalanne, Björn W. Schuller
Pattern Recognit. Lett.1
2014 Emotion Recognition in the Wild: Incorporating Voice and Lip Activity in Multimodal Decision-Level Fusion
abstract
In this paper, we investigate the relevance of using voice and lip activity to improve performance of audiovisual emotion recognition in unconstrained settings, as part of the 2014 Emotion Recognition in the Wild Challenge (EmotiW14). Indeed, the dataset provided by the organisers contains movie excerpts with highly challenging variability in terms of audiovisual content; e.g., speech and/or face of the subject expressing the emotion can be absent in the data. We therefore propose to tackle this issue by incorporating both voice and lip activity as additional features in a decision-level fusion. Results obtained on the blind test set show that the decision-level fusion can improve the best mono-modal approach, and that the addition of both voice and lip activity in the feature set leads to the best performance (UAR=35.27%), with an absolute improvement of 5.36% over the baseline.
Fabien Ringeval, Shahin Amiriparian, Florian Eyben, Klaus R. Scherer, Björn W. Schuller
ICMI1
2014 The INTERSPEECH 2014 computational paralinguistics challenge: cognitive & physical load
abstract
The INTERSPEECH 2014 Computational Paralinguistics Challenge provides for the first time a unified test-bed for the automatic recognition of speakers’ cognitive and physical load in speech. In this paper, we describe these two Sub-Challenges, their conditions, baseline results and experimental procedures, as well as the COMPARE baseline features generated with the openSMILE toolkit and provided to the participants in the Challenge.
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julien Epps, Florian Eyben, Fabien Ringeval, Erik Marchi, Yue Zhang 0014
INTERSPEECH6
2013 On the Influence of Emotional Feedback on Emotion Awareness and Gaze Behavior
abstract
This paper examines how emotion feedback influences emotion awareness and gaze behavior. Simulating a videoconference setup, 36 participants watched 12 emotional video sequences that were selected from the SEMAINE database. All participants wore an eye-tracker to measure gaze behavior and were asked to rate the perceived emotion for each video sequence. 3 conditions were tested: (c1) no feedback, i.e., the original video-sequences, (c2) correct feedback, i.e., an emoticon is integrated in the video to show the emotion depicted by the person in the video and (c3) random feedback, i.e., the emoticon displays at random an emotional state that may or may not correspond to the one of the person. The results showed that emotion feedback had a significant influence on gaze behavior, e.g., over time random feedback led to a decrease in the frequency of episodes of gaze. No effect of emotion display was observed for emotion recognition. However, experiments on the automatic emotion recognition using gaze behavior provided good performance, with better score on arousal than valence, and a very good performance was obtained in the automatic recognition of the correctness of the emotion feedback.
Fabien Ringeval, Andreas Sonderegger, Basilio Noris, Aude Billard, Jürgen S. Sauer, Denis Lalanne
ACII1
2013 Computer-Supported Work in Partially Distributed and Co-located Teams: The Influence of Mood Feedback
Andreas Sonderegger, Denis Lalanne, Luisa Bergholz, Fabien Ringeval, Jürgen S. Sauer
INTERACT (2)4
2013 The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism
abstract
International audience
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim
INTERSPEECH6
2012 Novel Metrics of Speech Rhythm for the Assessment of Emotion
abstract
International audience
Fabien Ringeval, Mohamed Chetouani, Björn W. Schuller
INTERSPEECH1
2011 Automatic Intonation Recognition for the Prosodic Assessment of Language-Impaired Children
abstract
This study presents a preliminary investigation into the automatic assessment of language-impaired children's (LIC) prosodic skills in one grammatical aspect: sentence modalities. Three types of language impairments were studied: autism disorder (AD), pervasive developmental disorder-not otherwise specified (PDD-NOS), and specific language impairment (SLI). A control group of typically developing (TD) children that was both age and gender matched with LIC was used for the analysis. All of the children were asked to imitate sentences that provided different types of intonation (e.g., descending and rising contours). An automatic system was then used to assess LIC's prosodic skills by comparing the intonation recognition scores with those obtained by the control group. The results showed that all LIC have difficulties in reproducing intonation contours because they achieved significantly lower recognition scores than TD children on almost all studied intonations (p <; 0.05). Regarding the “Rising” intonation, only SLI children had high recognition scores similar to TD children, which suggests a more pronounced pragmatic impairment in AD and PDD-NOS children. The automatic approach used in this study to assess LIC's prosodic skills confirms the clinical descriptions of the subjects' communication impairments.
Fabien Ringeval, Jean Demouy, György Szaszák, Mohamed Chetouani, L. Robel, Jean Xavier, Monique Plaza
IEEE Trans. Speech Audio Process.1
2008 A vowel based approach for acted emotion recognition
abstract
This paper is devoted to the description of a new approach for emotion recognition. Our contribution is based on both the extraction and the characterization of phonemic units such as vowels and consonants, which are provided by a pseudophonetic speech segmentation phase combined with a vowel detector. Concerning the emotion recognition task, we explore acoustic and prosodic features from these pseudo-phonetic segments (vowels and consonants), and we compare this approach with traditional voiced and unvoiced segments. The classification is realized by the well-known k-nn classifier (k nearest neighbors) from two different emotional speech databases: Berlin (German) and Aholab (Basque).
Fabien Ringeval, Mohamed Chetouani
INTERSPEECH1