EDBT 2026 Demo / reviewers in the wild / expert
Alice Baird
dblp:194/1299 · also Alice E. Baird
· DBLP profile ↗
40ranked-venue papers
10as first author
11since 2021 · last 2024
0000-0002-7003-5650ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 27 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species RecognitionabstractIn ornithology, bird species are known to have variedit's widely acknowledged that bird species display diverse dialects in their calls across different regions.Consequently, computational methods to identify bird species onsolely through their calls face critsignificalnt challenges.There is growing interest in understanding the impact of species-specific dialects on the effectiveness of bird species recognition methods.Despite potential mitigation through the expansion of dialect datasets, the absence of publicly available testing data currently impedes robust benchmarking efforts.This paper presents the Dialect Dominated Dataset of Bird Vocalisation (D3BV), the first crosscorpus dataset that focuses on dialects in bird vocalisations.The D3BV comprises more than 25 hours of audio recordings from 10 bird species distributed across three distinct regions in the contiguous United States (CONUS).In addition to presenting the dataset, we conduct analyses and establish baseline models for cross-corpus bird recognition.The data and code are publicly available online: https://zenodo.org/ Xin Jing 0001, Jiangjian Xie, Alexander Gebhard 0001, Alice Baird, Björn W. Schuller |
INTERSPEECH | 5 |
| 2023 | Large-Scale Nonverbal Vocalization Detection Using TransformersabstractDetecting emotionally expressive nonverbal vocalizations is essential to developing technologies that can converse fluently with humans. The affective computing community has largely focused on understanding the intonation of emotional speech and language. However, advances in the study of vocal emotional behavior suggest that emotions may be more readily conveyed not by speech but by nonverbal vocalizations such as laughs, sighs, shrieks, and grunts – vocalizations that often occur in lieu of speech. The task of detecting such emotional vocalizations has been largely overlooked by researchers, likely due to the limited availability of data capturing a sufficiently wide variety of vocalizations. Most studies in the literature focus on detecting laughter or cries. In this paper, we present the first, to the best of our knowledge, nonverbal vocalization detection model trained to detect as many as 67 types of emotional vocalizations. For our purposes, we use the large-scale and in-the-wild HUME-VB dataset that provides more than 156 h of data. We thoroughly investigate the use of pre-trained audio transformer models, such as Wav2Vec2 and Whisper, and provide useful insights for the task at hand using different types of noise signals. Panagiotis Tzirakis, Alice Baird, Jeffrey A. Brooks, Chris Gagne 0001, Lauren Kim, Michael Opara, Christopher B. Gregory, Jacob Metrick, Garrett Boseck, Vineet Tiruvadi, Björn W. Schuller, Dacher Keltner, Alan Cowen |
ICASSP | 2 |
| 2023 | The ACM Multimedia 2023 Computational Paralinguistics Challenge: Emotion Share & RequestsabstractThe ACM Multimedia 2023 Computational Paralinguistics Challenge addresses two different problems for the first time in a research competition under well-defined conditions: In the Emotion Share Sub-Challenge, a regression on speech has to be made; and in the Requests Sub-Challenges, requests and complaints need to be detected. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' ComPaRE features, the auDeep toolkit, and deep feature extraction from pre-trained CNNs using the DeepSpectRum toolkit; in addition, wav2vec2 models are used. Björn W. Schuller, Anton Batliner, Shahin Amiriparian, Alexander Barnhill, Maurice Gerczuk, Andreas Triantafyllopoulos, Alice Baird, Panagiotis Tzirakis, Chris Gagne 0001, Alan Cowen, Nikola Lackovic, Marie-José Caraty, Claude Montacié |
ACM Multimedia | 7 |
| 2023 | Ethical Awareness in Paralinguistics: A Taxonomy of ApplicationsabstractSince the end of the last century, the automatic processing of paralinguistics has been investigated widely and put into practice in many applications, on wearables, smartphones, and computers. In this contribution, we address ethical awareness for paralinguistic applications, by establishing taxonomies for data representations, system designs for and a typology of applications, and users/test sets and subject areas. These are related to an “ethical grid” consisting of the most relevant ethical cornerstones, based on principalism. The characteristics of and the interdependencies between these taxonomies are described and exemplified. This makes it possible to assess more or less critical “ethical constellations.” To the best of our knowledge, this is the first attempt of its kind. Anton Batliner, Michael Neumann 0001, Felix Burkhardt, Alice Baird, Sarina Meyer, Ngoc Thang Vu, Björn W. Schuller |
Int. J. Hum. Comput. Interact. | 4 |
| 2023 | The Multimodal Sentiment Analysis in Car Reviews (MuSe-CaR) Dataset: Collection, Insights and ImprovementsabstractTruly real-life data presents a strong, but exciting challenge for sentiment and emotion research. The high variety of possible ‘in-the-wild’ properties makes large datasets such as these indispensable with respect to building robust machine learning models. A sufficient quantity of data covering a deep variety in the challenges of each modality to force the exploratory analysis of the interplay of all modalities has not yet been made available in this context. In this contribution, we present MuSe-CaR, a first of its kind multimodal dataset. The data is publicly available as it recently served as the testing bed for the 1st Multimodal Sentiment Analysis Challenge, and focused on the tasks of emotion, emotion-target engagement, and trustworthiness recognition by means of comprehensively integrating the audio-visual and language modalities. Furthermore, we give a thorough overview of the dataset in terms of collection and annotation, including annotation tiers not used in this year's MuSe 2020. In addition, for one of the sub-challenges – predicting the level of trustworthiness – no participant outperformed the baseline model, and so we propose a simple, but highly efficient Multi-Head-Attention network that exceeds using multimodal fusion the baseline by around 0.2 CCC (almost 50 percent improvement). Lukas Stappen, Alice Baird, Lea Schumann, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | State & Trait Measurement from Nonverbal Vocalizations: A Multi-Task Joint Learning Approach
Alice Baird, Panagiotis Tzirakis, Jeffrey A. Brooks, Lauren Kim, Michael Opara, Christopher B. Gregory, Jacob Metrick, Garrett Boseck, Dacher Keltner, Alan Cowen |
INTERSPEECH | 1 |
| 2022 | Face mask recognition from audio: The MASC database and an overview on the mask challenge
Mostafa M. Mohamed, Mina A. Nessiem, Anton Batliner, Christian Bergler, Simone Hantke, Maximilian Schmitt, Alice Baird, Adria Mallol-Ragolta, Vincent Karas, Shahin Amiriparian, Björn W. Schuller |
Pattern Recognit. | 7 |
| 2021 | A Prototypical Network Approach for Evaluating Generated Emotional SpeechabstractThe collection of emotional speech data is a time-consuming and costly endeavour.Generative networks can be applied to augment the limited audio data artificially.However, it is challenging to evaluate generated audio for its similarity to source data, as current quantitative metrics are not necessarily suited to the audio domain.We explore the use of a prototypical network to evaluate four classes of generated emotional audio with this in mind.We first extract spectrogram images from WAVEGAN generated audio and other audio augmentation approaches, comparing similarity to the class prototype and diversity within the embedding space.Furthermore, we augment the source training set with each augmentation type and perform a classification to explore the generated audio plausibility.Results suggest that quality and diversity can be quantitatively observed with this approach.In the chosen context, we see that WAVEGAN generated data is recognisable as a source data class (F1-score 43.6 %), and the samples add similar diversity as unseen source data.This result leads to more plausible data for augmentation of the source training set -achieving up to 63.9 % F1 which is a 3.5 % improvement over the source data baseline. Alice Baird, Silvan Mertes, Manuel Milling, Lukas Stappen, Thomas Wiest, Elisabeth André, Björn W. Schuller |
Interspeech | 1 |
| 2021 | The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & PrimatesabstractThe INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the COVID-19 Cough and COVID-19 Speech Sub-Challenges, a binary classification on COVID-19 infection has to be made based on coughing sounds and speech; in the Escalation SubChallenge, a three-way assessment of the level of escalation in a dialogue is featured; and in the Primates Sub-Challenge, four species vs background need to be classified. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AuDeep toolkit, and deep feature extraction from pre-trained CNNs using the Deep Spectrum toolkit; in addition, we add deep end-to-end sequential modelling, and partially linguistic analysis. Björn W. Schuller, Anton Batliner, Christian Bergler, Cecilia Mascolo, Jing Han 0010, Iulia Lefter, Heysem Kaya, Shahin Amiriparian, Alice Baird, Lukas Stappen, Sandra Ottl, Maurice Gerczuk, Panagiotis Tzirakis, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Léon J. M. Rothkrantz, Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp |
Interspeech | 9 |
| 2021 | Evaluating Deep Music Generation Methods Using Data AugmentationabstractDespite advances in deep algorithmic music generation, evaluation of generated samples often relies on human evaluation, which is subjective and costly. We focus on designing a homogeneous, objective framework for evaluating samples of algorithmically generated music. Any engineered measures to evaluate generated music typically attempt to define the samples’ musicality, but do not capture qualities of music such as theme or mood. We do not seek to assess the musical merit of generated music, but instead explore whether generated samples contain meaningful information pertaining to emotion or mood/theme. We achieve this by measuring the change in predictive performance of a music mood/theme classifier after augmenting its training data with generated samples. We analyse music samples generated by three models – SampleRNN, Jukebox, and DDSP – and employ a homogeneous framework across all methods to allow for objective comparison. This is the first attempt at augmenting a music genre classification dataset with conditionally generated music. We investigate the classification performance improvement using deep music generation and the ability of the generators to make emotional music by using an additional, emotion annotation of the dataset. Finally, we use a classifier trained on real data to evaluate the label validity of class-conditionally generated samples. Toby Godwin, Georgios Rizos, Alice Baird, Najla Al Futaisi, Vincent Brisse, Björn W. Schuller |
MMSP | 3 |
| 2021 | Emotion Recognition in Public Speaking Scenarios Utilising An LSTM-RNN Approach with AttentionabstractSpeaking in public can be a cause of fear for many people. Research suggests that there are physical markers such as an increased heart rate and vocal tremolo that indicate an individual's state of wellbeing during a public speech. In this study, we explore the advantages of speech-based features for continuous recognition of the emotional dimensions of arousal and valence during a public speaking scenario. Furthermore, we explore biological signal fusion, and perform cross-language (German and English) analysis by training language-independent models and testing them on speech from various native and non-native speaker groupings. For the emotion recognition task itself, we utilise a Long Short-Term Memory - Recurrent Neural Network (LSTM-RNN) architecture with a self-attention layer. When utilising audio-only features and testing with non-native German's speaking German we achieve at best a concordance correlation coefficient (CCC) of 0.640 and 0.491 for arousal and valence, respectively - demonstrating a strong effect for this task from non-native speakers, as well as promise for the suitability of deep learning for continuous emotion recognition in the context of public speaking. Alice Baird, Shahin Amiriparian, Manuel Milling, Björn W. Schuller |
SLT | 1 |
| 2020 | Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition ModelsabstractThe development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this study, we aim to train deep models to defending against non-targeted white-box adversarial attacks. Adversarial data is first generated from the real data using the fast gradient sign method. Then in the research field of speech emotion recognition, adversarial-based training is employed as a method for protecting against adversarial attack. We then train deep convolutional models with both real and adversarial data, and compare the performances of two adversarial training procedures - namely, vanilla adversarial training, and similarity-based adversarial training. In our experiments, through the use of adversarial data augmentation, both of the considered adversarial training procedures can improve the performance when validated on the real data. Additionally, the similarity-based adversarial training learns a more robust model when working with adversarial data. Finally, the considered VGG-16 model performs the best across all models, for both real and generated data. Zhao Ren, Alice Baird, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 2 |
| 2020 | Stargan for Emotional Speech Conversion: Validated by Data Augmentation of End-To-End Emotion RecognitionabstractIn this paper, we propose an adversarial network implementation for speech emotion conversion as a data augmentation method, validated by a multi-class speech affect recognition task. In our setting, we do not assume the availability of parallel data, and we additionally make it a priority to exploit as much as possible the available training data by adopting a cycle-consistent, class-conditional generative adversarial network with an auxiliary domain classifier. Our generated samples are valuable for data augmentation, achieving a corresponding 2% and 6% absolute increase in Micro- and MacroF1 compared to the baseline in a 3-class classification paradigm using a deep, end-to-end network. We finally perform a human perception evaluation of the samples, through which we conclude that our samples are indicative of their target emotion, albeit showing a tendency for confusion in cases where the emotional attribute of valence and arousal are inconsistent. Georgios Rizos, Alice Baird, Max Elliott, Björn W. Schuller |
ICASSP | 2 |
| 2020 | An Evaluation of the Effect of Anxiety on Speech - Computational Prediction of Anxiety from Sustained VowelsabstractThe current level of global uncertainty is having an implicit effect on those with a diagnosed anxiety disorder.Anxiety can impact vocal qualities, particularly as physical symptoms of anxiety include muscle tension and shortness of breath.To this end, in this study, we explore the effect of anxiety on speech -focusing on four classes of sustained vowels (sad, smiling, comfortable, and powerful) -via feature analysis and a series of regression experiments.We extract three well-known acoustic feature sets and evaluate the efficacy of machine learning for prediction of anxiety based on the Beck Anxiety Inventory (BAI) score.Of note, utilising a support vector regressor, we find that the effects of anxiety in speech appear to be stronger at higher BAI levels.Significant differences (p < 0.05) between test predictions of Low and High-BAI groupings support this.Furthermore, when utilising a High-BAI grouping for the prediction of standardised BAI, significantly higher results are obtained for smiling sustained vowels, of up to 0.646 Spearman's Correlation Coefficient (ρ), and up to 0.592 ρ with all sustained vowels.A significantly stronger (Cohens d of 1.718) result than all data combined without grouping, which achieves at best 0.234 ρ. Alice Baird, Nicholas Cummins, Sebastian Schnieder, Jarek Krajewski, Björn W. Schuller |
INTERSPEECH | 1 |
| 2020 | Deep Attentive End-to-End Continuous Breath Sensing from SpeechabstractModelling of the breath signal is of high interest to both \nhealthcare professionals and computer scientists, as a source \nof diagnosis-related information, or a means for curating higher \nquality datasets in speech analysis research. The formation of \na breath signal gold standard is, however, not a straightforward \ntask, as it requires specialised equipment, human annotation \nbudget, and even then, it corresponds to lab recording settings, \nthat are not reproducible in-the-wild. Herein, we explore deep \nlearning based methodologies, as an automatic way to predict a \ncontinuous-time breath signal by solely analysing spontaneous \nspeech. We address two task formulations, those of continuousvalued signal prediction, as well as inhalation event prediction, \nthat are of great use in various healthcare and Automatic Speech \nRecognition applications, and showcase results that outperform \ncurrent baselines. Most importantly, we also perform an initial \nexploration into explaining which parts of the input audio signal \nare important with respect to the prediction. Alexis Deighton MacIntyre, Georgios Rizos, Anton Batliner, Alice Baird, Shahin Amiriparian, Antonia F. de C. Hamilton, Björn W. Schuller |
INTERSPEECH | 4 |
| 2020 | The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly Emotion, Breathing & MasksabstractThe INTERSPEECH 2020 Computational Paralinguistics Challenge addresses three different problems for the first time in a research competition under well-defined conditions: In the Elderly Emotion Sub-Challenge, arousal and valence in the speech of elderly individuals have to be modelled as a 3-class problem; in the Breathing Sub-Challenge, breathing has to be assessed as a regression problem; and in the Mask Sub-Challenge, speech without and with a surgical mask has to be told apart.We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AUDEEP toolkit, and deep feature extraction from pre-trained CNNs using the DEEP SPECTRUM toolkit; in addition, we partially add deep end-to-end sequential modelling, and, for the first time in the challenge, linguistic analysis. Björn W. Schuller, Anton Batliner, Christian Bergler, Eva-Maria Messner, Antonia F. de C. Hamilton, Shahin Amiriparian, Alice Baird, Georgios Rizos, Maximilian Schmitt, Lukas Stappen, Harald Baumeister, Alexis Deighton MacIntyre, Simone Hantke |
INTERSPEECH | 7 |
| 2020 | An Evolutionary-based Generative Approach for Audio Data AugmentationabstractIn this paper, we introduce a novel framework to augment raw audio data for machine learning classification tasks. For the first part of our framework, we employ a generative adversarial network (GAN) to create new variants of the audio samples that are already existing in our source dataset for the classification task. In the second step, we then utilize an evolutionary algorithm to search the input domain space of the previously trained GAN, with respect to predefined characteristics of the generated audio. This way we are able to generate audio in a controlled manner that contributes to an improvement in classification performance of the original task. To validate our approach, we chose to test it on the task of soundscape classification. We show that our approach leads to a substantial improvement in classification results when compared to a training routine without data augmentation and training with uncontrolled data augmentation with GANs. Silvan Mertes, Alice Baird, Dominik Schiller, Björn W. Schuller, Elisabeth André |
MMSP | 2 |
| 2020 | Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS - The Heart Sounds Shenzhen CorpusabstractAuscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49.7% unweighted average recall on a three category task (normal, mild, moderate/severe). Fengquan Dong, Kun Qian 0003, Zhao Ren, Alice Baird, Zhenyu Dai, Florian Metze, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Audiovisual Analysis for Recognising Frustration during Game-Play: Introducing the Multimodal Game Frustration DatabaseabstractAutomatic recognition of frustration, by analysing facial and vocal expressions, can help user experience designers to identify interaction obstacles. To encourage the development of automated systems such as these, we present a novel audiovisual database: the Multimodal Game Frustration Database (MGFD), consisting of ca. 5 hours of audiovisual data, collected from 67 Chinese students speaking in English. For data collection, we developed ‘Crazy Trophy’, a Wizard-of-Oz voice activated web-game designed with a variety of usability problems and aimed to induce increasing amounts of frustration. We also present a baseline for binary multimodal frustration classification (frustration vs no-frustration). For this, we compare the performance of a conventional method, Support Vector Machine classifier, and a state-of-the-art method utilising Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN), extracting both audio (Mel-frequency Cepstral Coefficients) and video (facial action units) features. Using LSTM-RNN and a feature-based multi-model fusion strategy, the best result acheived for the baseline was 60.3 % UAR. To enable further research in this area, the game (‘Crazy Trophy’), the database (MGFD), and the partitioning considered in the presented baseline, are made accessible to the research community. Meishu Song, Zijiang Yang 0007, Alice Baird, Emilia Parada-Cabaleiro, Zixing Zhang 0001, Ziping Zhao 0001, Björn W. Schuller |
ACII | 3 |
| 2019 | Performance Analysis of Unimodal and Multimodal Models in Valence-Based Empathy RecognitionabstractThe human ability to empathise is a core aspect of successful interpersonal relationships. In this regard, human-robot interaction can be improved through the automatic perception of empathy, among other human attributes, allowing robots to affectively adapt their actions to interactants' feelings in any given situation. This paper presents our contribution to the generalised track of the One-Minute Gradual (OMG) Empathy Prediction Challenge by describing our approach to predict a listener's valence during semi-scripted actor-listener interactions. We extract visual and acoustic features from the interactions and feed them into a bidirectional long short-term memory network to capture the time-dependencies of the valence-based empathy during the interactions. Generalised and personalised unimodal and multimodal valence-based empathy models are then trained to assess the impact of each modality on the system performance. Furthermore, we analyse if intra-subject dependencies on empathy perception affect the system performance. We assess the models by computing the concordance correlation coefficient (CCC) between the predicted and self-annotated valence scores. The results support the suitability of employing multimodal data to recognise participants' valence-based empathy during the interactions, and highlight the subject-dependency of empathy. In particular, we obtained our best result with a personalised multimodal model, which achieved a CCC of 0.11 on the test set. Adria Mallol-Ragolta, Maximilian Schmitt, Alice Baird, Nicholas Cummins, Björn W. Schuller |
FG | 3 |
| 2019 | Audio-based Recognition of Bipolar Disorder Utilising Capsule NetworksabstractBipolar disorder (BD) is an acute mood condition, in which states can drastically shift from one extreme to another, considerably impacting an individual's wellbeing. Automatic recognition of a BD diagnosis can help patients to obtain medical treatment at an earlier stage and therefore have a better overall prognosis. With this in mind, in this study, we utilise a Capsule Neural Network (CapsNet) for audio-based classification of patients who were suffering from BD after a mania episode into three classes of Remission, Hypomania, and Mania. The CapsNet attempts to address the limitations of Convolutional Neural Networks (CNNs) by considering vital spatial hierarchies between the extracted images from audio files. We develop a framework around the CapsNet in order to analyse and classify audio signals. First, we create a spectrogram from short segments of speech recordings from individuals with a bipolar diagnosis. We then train the CapsNet on the spectrograms with 32 low- level and three high-level capsules, each for one of the BD classes. These capsules attempt both to form a meaningful representation of the input data and to learn the correct BD class. The output of each capsule represents an activity vector. The length of this vector encodes the presence of the corresponding type of BD in the input, and its orientation represents the properties of this specific instance of BD. We show that using our CapsNet framework, it is possible to achieve competitive results for the aforementioned task by reaching a UAR of 46.2 % and 45.5 % on the development and test partitions, respectively. Furthermore, the efficacy of our approach is compared with a sequence to sequence autoencoder and a CNN-based neural network. Shahin Amiriparian, Arsany Awad, Maurice Gerczuk, Lukas Stappen, Alice Baird, Sandra Ottl, Björn W. Schuller |
IJCNN | 5 |
| 2019 | Using Speech to Predict Sequentially Measured Cortisol Levels During a Trier Social Stress TestabstractThe effect of stress on the human body is substantial, potentially resulting in serious health implications.Furthermore, with modern stressors seemingly on the increase, there is an abundance of contributing factors which lead to a diagnosis of acute stress.However, observing biological stress reactions usually includes costly and time consuming sequential fluidbased samples to determine the degree of biological stress.On the contrary, a speech monitoring approach would allow for a non-invasive indication of stress.To evaluate the efficacy of the speech signal as a marker of stress, we explored, for the first time, the relationship between sequential cortisol samples and speech-based features.Utilising a novel corpus of 43 individuals undergoing a standardised Trier Social Stress Test (TSST), we extract a variety of feature sets and observe a correlation between speech and sequential cortisol measurements.For prediction of mean cortisol levels from speech, results show that for the entire TSST oral presentation, handcrafted COMPARE features achieve best results of 0.244 root mean square error [0 ;1] for the sample 20 minutes after the TSST.Correlation also increases at minute 20, with a Spearman's correlation coefficient of 0.421, and Cohen's d of 0.883 between the baseline and minute 20 cortisol predictions. Alice Baird, Shahin Amiriparian, Nicholas Cummins, Sarah Sturmbauer, Johanna Janson, Eva-Maria Messner, Harald Baumeister, Nicolas Rohleder, Björn W. Schuller |
INTERSPEECH | 1 |
| 2019 | Sincerity in Acted Speech: Presenting the Sincere Apology Corpus and ResultsabstractThe ability to discern an individual's level of sincerity varies from person to person and across cultures.Sincerity is typically a key indication of personality traits such as trustworthiness, and portraying sincerity can be integral to an abundance of scenarios, e. g. , when apologising.Speech signals are one important factor when discerning sincerity and, with more modern interactions occurring remotely, automatic approaches for the recognition of sincerity from speech are beneficial during both interpersonal and professional scenarios.In this study we present details of the Sincere Apology Corpus (SINA-C ).Annotated by 22 individuals for their perception of sincerity, SINA-C is an English acted-speech corpus of 32 speakers, apologising in multiple ways.To provide an updated baseline for the corpus, various machine learning experiments are conducted.Finding that extracting deep data-representations (utilising the DEEP SPECTRUM toolkit) from the speech signals is best suited.Classification results on the binary (sincere / not sincere) task are at best 79.2 % Unweighted Average Recall and for regression, in regards to the degree of sincerity, a Root Mean Square Error of 0.395 from the standardised range [-1.51; 1.72] is obtained. Alice Baird, Eduardo Coutinho, Julia Hirschberg, Björn W. Schuller |
INTERSPEECH | 1 |
| 2019 | Predicting Biological Signals from Speech: Introducing a Novel Multimodal Dataset and ResultsabstractIn recent years, diagnosis and awareness of mental health conditions, e. g., chronic stress, have been increasing globally. Biological signals can be an effective way to monitor such conditions, yet acquisition can be cumbersome and invasive. Alternatively, acoustic features offer non-invasive and efficient monitoring of an array of health and wellbeing characteristics. This study presents the BioSpeech Database (BioS-DB), a novel database of audio and biological signals - blood volume pulse (BVP) and skin conductance (SC) - from 55 individuals speaking aloud in front of others, whilst having their emotional state annotated in real time. Through a variation of conventional and state-of-the-art approaches, initial experiments have shown for the first time that acoustic features can be applied for the task of BVP prediction. Notably, using deep representations of audio and a sequence-to-sequence auto-encoders with a GRU-RNN as a time-dependent regressor achieved at best 0.075 and 0.123 RMSE for [0; 1] normalised BVP and SC, respectively. Alice Baird, Shahin Amiriparian, Miriam Berschneider, Maximilian Schmitt, Björn W. Schuller |
MMSP | 1 |
| 2019 | Can Deep Generative Audio be Emotional? Towards an Approach for Personalised Emotional Audio GenerationabstractThe ability for sound to evoke states of emotion is well known across fields of research, with clinical and holistic practitioners utilising audio to create listener experiences which target specific needs. Neural network-based generative models have in recent years shown promise for generating high-fidelity based on a raw audio input. With this in mind, this study utilises the WaveNet generative model to explore the ability of such networks to retain the emotionality of raw audio speech inputs. We train various models on 2-classes (happy and sad) of an emotional speech corpus containing 68 native Italian speakers. When classifying the combined original and generated audio, hand-crafted feature sets achieve at best 75.5 % unweighted average recall, a 2 percent point improvement over the original only audio features. Additionally, from a two-tailed test on the predictions, we find that the audio features from the original speech concatenated with the generated audio features provides significantly different test result compared to the baseline. Both findings indicating promise for emotion-based audio generation. Alice Baird, Shahin Amiriparian, Björn W. Schuller |
MMSP | 1 |
| 2019 | The ASC-Inclusion Perceptual Serious Gaming Platform for Autistic Childrenabstract“Serious games” are becoming extremely relevant to individuals who have specific needs, such as children with an autism spectrum condition (ASC). Often, individuals with an ASC have difficulties in interpreting verbal and nonverbal communication cues during social interactions. The ASC-Inclusion EU-FP7 funded project aims to provide children who have an ASC with a platform to learn emotion expression and recognition, through play in the virtual world. In particular, the ASC-Inclusion platform focuses on the expression of emotion via facial, vocal, and bodily gestures. The platform combines multiple analysis tools, using onboard microphone and webcam capabilities. The platform utilizes these capabilities via training games, text-based communication, animations, video, and audio clips. This paper introduces current findings and evaluations of the ASC-Inclusion platform and provides detailed description for the different modalities. Erik Marchi, Tadas Baltrusaitis, Andra Adams, Marwa Mahmoud, Ofer Golan, Shimrit Fridenson-Hayo, Shahar Tal, Shai Newman, Noga Meir-Goren, Antonio Camurri, Stefano Piana, Björn W. Schuller, Sven Bölte, Tevfik Metin Sezgin, Nese Alyüz, Agnieszka Rynkiewicz, Aurelie Baranger, Alice Baird, Simon Baron-Cohen, Amandine Lassalle, Helen O'Reilly, Delia Pigat, Peter Robinson 0001, Ian Davies |
IEEE Trans. Games | 18 |
| 2018 | Recognition of Echolalic Autistic Child Vocalisations Utilising Convolutional Recurrent Neural NetworksabstractAutism spectrum conditions (ASC) are a set of neurodevelopmental conditions partly characterised by difficulties with communication.Individuals with ASC can show a variety of atypical speech behaviours, including echolalia or the 'echoing' of another's speech.We herein introduce a new dataset of 15 Serbian ASC children in a human-robot interaction scenario, annotated for the presence of echolalia amongst other ASC vocal behaviours.From this, we propose a four-class classification problem and investigate the suitability of applying a 2D convolutional neural network augmented with a recurrent neural network with bidirectional long short-term memory cells to solve the proposed task of echolalia recognition.In this approach, log Mel-spectrograms are first generated from the audio recordings and then fed as input into the convolutional layers to extract high-level spectral features.The subsequent recurrent layers are applied to learn the long-term temporal context from the obtained features.Finally, we use a feed forward neural network with softmax activation to classify the dataset.To evaluate the performance of our deep learning approach, we use leave-onesubject-out cross-validation.Key results presented indicate the suitability of our approach by achieving a classification accuracy of 83.5 % unweighted average recall. Shahin Amiriparian, Alice Baird, Sahib Julka, Alyssa Alcorn, Sandra Ottl, Suncica Petrovic, Eloise Ainger, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 2 |
| 2018 | The Perception and Analysis of the Likeability and Human Likeness of Synthesized SpeechabstractThe synthesized voice has become an ever present aspect of daily life.Heard through our smart-devices and from public announcements, engineers continue in an endeavour to achieve naturalness in such voices.Yet, the degree to which these methods can produce likeable, human like voices, has not been fully evaluated.With recent advancements in synthetic speech technology suggesting that human like imitation is more obtainable, this study asked 25 listeners to evaluate both the likeability and human likeness of a corpus of 13 German male voices, produced via 5 synthesis approaches (from formant to hybrid unit selection, deep neural network systems), and 1 Human control.Results show that unlike visual artificially intelligent elements -as posed by the concept of the Uncanny Valley -likeability consistently improves along with human likeness for the synthesized voice, with recent methods achieving substantially closer results to human speech than older methods.A small scale acoustic analysis shows that the F0 of hybrid systems correlates less closely to human speech with a higher standard deviation for F0.This analysis suggests that limited variance in F0 is linked to a reduction in human likeness, resulting in lower likeability for conventional synthetic speech methods. Alice Baird, Emilia Parada-Cabaleiro, Simone Hantke, Felix Burkhardt, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 1 |
| 2018 | Categorical vs Dimensional Perception of Italian Emotional SpeechabstractCulture and measurement strategies are influential factors when evaluating the perception of emotion in speech.However, multilingual databases suitable for such a study are missing, and there is no agreement on the most suitable emotional model.To address this gap, we present EmoFilm, a new multilingual emotional speech corpus, consisting of 1115 English, Spanish, and Italian emotional utterances extracted from 43 films and 207 speakers.We have performed a within-culture categorical vs dimensional perceptual evaluation, employing 225 native Italian listeners, who evaluated the Italian section of the database with the emotional states of anger, sadness, happiness, fear, and contempt.The aim of this study is to assess whether the emotional model (categorical or dimensional), taken as reference for measurement, influences a listener's perception of emotional speech, and-to what extent-both models are complementary or not.We show that the measurement strategy chosen does influence a listener's response, especially for some emotions, e. g., contempt.The confusion patterns typical of a categorical evaluation are not always mirrored by the dimensional assessment. Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, Björn W. Schuller |
INTERSPEECH | 4 |
| 2018 | The INTERSPEECH 2018 Computational Paralinguistics Challenge: Atypical & Self-Assessed Affect, Crying & Heart BeatsabstractThe INTERSPEECH 2018 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Atypical Affect Sub-Challenge, four basic emotions annotated in the speech of handicapped subjects have to be classified; in the Self-Assessed Affect Sub-Challenge, valence scores given by the speakers themselves are used for a three-class classification problem; in the Crying Sub-Challenge, three types of infant vocalisations have to be told apart; and in the Heart Beats Sub-Challenge, three different types of heart beats have to be determined.We describe the Sub-Challenges, their conditions, and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by end-to-end learning, the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit for the first time in the challenge series. Björn W. Schuller, Stefan Steidl, Anton Batliner, Peter B. Marschik, Harald Baumeister, Fengquan Dong, Simone Hantke, Florian B. Pokorny, Eva-Maria Rathner, Katrin D. Bartl-Pokorny, Christa Einspieler, Dajie Zhang, Alice Baird, Shahin Amiriparian, Kun Qian 0003, Zhao Ren, Maximilian Schmitt, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 13 |
| 2017 | Stimulation of psychological listener experiences by semi-automatically composed electroacoustic environmentsabstractThis work represents the first steps in an almost completely unexplored field, in which electroacoustic composition, based on several signal processing techniques including additive synthesis, pitch extraction and bandpass filtering, is exploited to produce unique sonic environments for stimulating targeted listener experiences. We propose three semi-automatically composed electroacoustic environments and evaluate their psychological potential considering four areas: creativity, emotion, self-perception and mental associations. This empirical study uses a cross-modal perceptual test, completed by 100 listeners. Results presented indicate that electroacoustic music can successfully evoke specific colour connections and individual self-perception. Additionally, we show that synthesised sound based on bio-signals, such as a heart beat, can promote concrete thought and negative emotional states. Our future goal is to fully-automate electroacoustic music composition environments for use in therapy, education and entertainment to promote and encourage human well-being. Emilia Parada-Cabaleiro, Alice Baird, Nicholas Cummins, Björn W. Schuller |
ICME | 2 |
| 2017 | Snore Sound Classification Using Image-Based Deep Spectrum FeaturesabstractIn this paper, we propose a method for automatically detecting various types of snore sounds using image classification convolutional neural network (CNN) descriptors extracted from audio file spectrograms.The descriptors, denoted as deep spectrum features, are derived from forwarding spectrograms through very deep task-independent pre-trained CNNs.Specifically, activations of fully connected layers from two common image classification CNNs, AlexNet and VGG19, are used as feature vectors.Moreover, we investigate the impact of differing spectrogram colour maps and two CNN architectures on the performance of the system.Results presented indicate that deep spectrum features extracted from the activations of the second fully connected layer of AlexNet using a viridis colour map are well suited to the task.This feature space, when combined with a support vector classifier, outperforms the more conventional knowledge-based features of 6 373 acoustic functionals used in the INTERSPEECH ComParE 2017 Snoring sub-challenge baseline system.In comparison to the baseline, unweighted average recall is increased from 40.6 % to 44.8 % on the development partition, and from 58.5 % to 67.0 % on the test partition. Shahin Amiriparian, Maurice Gerczuk, Sandra Ottl, Nicholas Cummins, Michael Freitag 0003, Sergey Pugachevskiy, Alice Baird, Björn W. Schuller |
INTERSPEECH | 7 |
| 2017 | Automatic Classification of Autistic Child Vocalisations: A Novel Database and ResultsabstractHumanoid robots have in recent years shown great promise for supporting the educational needs of children on the autism spectrum.To further improve the efficacy of such interactions, user-adaptation strategies based on the individual needs of a child are required.In this regard, the proposed study assesses the suitability of a range of speech-based classification approaches for automatic detection of autism severity according to the commonly used Social Responsiveness Scale™ second edition (SRS-2).Autism is characterised by socialisation limitations including child language and communication ability.When compared to neurotypical children of the same age these can be a strong indication of severity.This study introduces a novel dataset of 803 utterances recorded from 14 autistic children aged between 4 -10 years, during Wizard-of-Oz interactions with a humanoid robot.Our results demonstrate the suitability of support vector machines (SVMs) which use acoustic feature sets from multiple Interspeech COMPARE challenges.We also evaluate deep spectrum features, extracted via an image classification convolutional neural network (CNN) from the spectrogram of autistic speech instances.At best, by using SVMs on the acoustic feature sets, we achieved a UAR of 73.7 % for the proposed 3-class task. Alice Baird, Shahin Amiriparian, Nicholas Cummins, Alyssa Alcorn, Anton Batliner, Sergey Pugachevskiy, Michael Freitag 0003, Maurice Gerczuk, Björn W. Schuller |
INTERSPEECH | 1 |
| 2017 | The Perception of Emotions in Noisified Nonsense SpeechabstractNoise pollution is part of our daily life, affecting millions of people, particularly those living in urban environments.Noise alters our perception and decreases our ability to understand others.Considering this, speech perception in background noise has been extensively studied, showing that especially white noise can damage listener perception.However, the perception of emotions in noisified speech has not been explored with as much depth.In the present study, we use artificial background noise conditions, by applying noise to a subset of the GEMEP corpus (emotions expressed in nonsense speech).Noises were at varying intensities and 'colours'; white, pink, and brownian.The categorical and dimensional perceptual test was completed by 26 listeners.The results indicate that background noise conditions influence the perception of emotion in speechpink noise most, brownian least.Worsened perception invokes higher confusion, especially with sadness, an emotion with less pronounced prosodic characteristics.Yet, all this does not lead to a break-down of the 'cognitive-emotional space' in a Nonmetric MultiDimensional Scaling representation.The gender of speakers and the cultural background of listeners do not seem to play a role. Emilia Parada-Cabaleiro, Alice Baird, Anton Batliner, Nicholas Cummins, Simone Hantke, Björn W. Schuller |
INTERSPEECH | 2 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: Deception, Sincerity & Native LanguageabstractThe INTERSPEECH 2016 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: classification of deceptive vs. non-deceptive speech, the estimation of the degree of sincerity, and the identification of the native language out of eleven L1 classes of English L2 speakers.In this paper, we describe these sub-challenges, their conditions, the baseline feature extraction and classifiers, and the resulting baselines, as provided to the participants. Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |
| 2016 | The Deception Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |
| 2016 | The Sincerity Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |
| 2016 | The Native Language Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: A Summary of Results
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |
| 2016 | Discussion
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 6 |