EDBT 2026 Demo / reviewers in the wild / expert
Nicholas Cummins
dblp:79/10648
· DBLP profile ↗
80ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0002-1178-917XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 13 first-author · 17 since 2021Artificial intelligence and machine learning · 57 · 7 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAKE: Context-Aware Kid Emotion In-the-Wild Dataset
Sornsiri Poovongsaroj, Zhuo Zeng, Nicholas Cummins, Oya Çeliktutan |
FG | 3 |
| 2025 | Towards the Objective Characterisation of Major Depressive Disorder Using Speech Data from a 12-week Observational Study with Daily Measurements
Robert Lewis 0001, Szymon Fedor, Nelson Hidalgo Julia, Joshua Curtiss, Jiyeon Kim, Noah Jones, David Mischoulon, Thomas F. Quatieri, Nicholas Cummins, Paola Pedrelli, Rosalind W. Picard |
INTERSPEECH | 9 |
| 2025 | Meta-Learning Approaches for Speaker-Dependent Voice Fatigue ModelsabstractSpeaker-dependent modelling can substantially improve performance in speech-based health monitoring applications. While mixed-effect models are commonly used for such speaker adaptation, they require computationally expensive retraining for each new observation, making them impractical in a production environment. We reformulate this task as a meta-learning problem and explore three approaches of increasing complexity: ensemble-based distance models, prototypical networks, and transformer-based sequence models. Using pre-trained speech embeddings, we evaluate these methods on a large longitudinal dataset of shift workers (N=1,185, 10,286 recordings), predicting time since sleep from speech as a function of fatigue, a symptom commonly associated with ill-health. Our results demonstrate that all meta-learning approaches tested outperformed both cross-sectional and conventional mixed-effects models, with a transformer-based method achieving the strongest performance. Roseline Polle, Agnes Norbury, Alexandra Livia Georgescu, Nicholas Cummins, Stefano Goria |
INTERSPEECH | 4 |
| 2025 | Speech Reference Intervals: An Assessment of Feasibility in Depression Symptom Severity PredictionabstractMajor Depressive Disorder (MDD) is a prevalent mental disorder. Combining speech features and machine learning has promise for predicting MDD, but interpretability is crucial for clinical applications. Reference intervals (RIs) represent a typical range for a speech feature in a population. RIs could increase interpretability and help clinicians identify deviations from norms. They could also replace conventional speech features in machine learning models. However, no work has yet assessed the feasibility of speech RIs in MDD. We generated and compared RIs from three reference datasets varying in size, elicitation prompt, and health information. We then calculated deviations from each RI set for people with MDD to compare performance on a depression symptom severity prediction task. Our RI-based models trained with demographic data performed similarly to each other and equivalent models using conventional features or demographics only, demonstrating the value of RI-derived features. Lauren L. White, Ewan Carr, Judith Dineley, Catarina Botelho, Pauline Conde, Faith Matcham, Carolin Oetzmann, Amos Folarin, George Fairs, Agnes Norbury, Stefano Goria, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Alberto Abad, Isabel Trancoso, Nicholas Cummins |
INTERSPEECH | 19 |
| 2025 | Challenges and practical guidelines for atypical speech data collection, annotation, usage and sharing: A multi-project perspectiveabstractContains fulltext : 325867.pdf (Publisher’s version ) (Open Access) Zhengjun Yue, Mara Barberis, Tanvina Patel, Judith Dineley, Willemijn Doedens, Lottie Stipdonk, Elke De Witte, Erfan Loweimi, Hugo Van hamme, Djaina Satoer, Marina B. Ruiter, Laureano Moro-Velázquez, Nicholas Cummins, Odette Scharenborg |
INTERSPEECH | 14 |
| 2024 | Longitudinal Modeling of Depression Shifts Using Speech and LanguageabstractSpeech analysis can provide a potential non-invasive and objective means of assessing and monitoring an individual’s mental health. Most studies to date have focused on cross-sectional analysis and have not explored the benefits of speech analysis as a longitudinal monitoring tool that can assist in the management of chronic conditions such as major depressive disorder (MDD). Objectively monitoring for shifts in depression symptom severity levels over time presents a notable challenge, which we address through an automated approach using longitudinal English and Spanish speech samples collected from a clinical population. We employ time–frequency representations and linguistic embeddings to enhance the early recognition of alterations in depression levels in individuals with MDD. We investigate the suitability of using siamese-based training for modeling these changes, intending to enable personalized and adaptive interventions. Paula Andrea Pérez-Toro, Judith Dineley, Agnieszka Kaczkowska, Pauline Conde, Yuezhou Zhang 0001, Faith Matcham, Sara Siddi, Josep Maria Haro, Stuart Bruce, Til Wykes, Raquel Bailón, Srinivasan Vairavan, Richard J. B. Dobson, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave, Vaibhav A. Narayan, Nicholas Cummins |
ICASSP | 18 |
| 2024 | Variability of speech timing features across repeated recordings: a comparison of open-source extraction techniquesabstractVariations in speech timing features have been reliably linked to symptoms of various health conditions, demonstrating clinical potential. However, replication challenges hinder their translation; extracted speech features are susceptible to methodological variations in the recording and processing pipeline. Investigating this, we compared exemplar timing features extracted via three different techniques from recordings of healthy speech. Our results show that features extracted via an intensity-based method differ from those produced by forced alignment. Different extraction methods also led to differing estimates of within-speaker feature variability over time in an analysis of recordings repeated systematically over three sessions in one day (n=26) and in one week (n=28). Our findings highlight the importance of feature extraction in study design and interpretation, and the need for consistent, accurate extraction techniques for clinical research. Index Terms: speech timing, feature extraction, reproducibility, longitudinal monitoring Judith Dineley, Ewan Carr, Lauren L. White, Catriona Lucas, Zahia Rahman, Tian Pan 0004, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 11 |
| 2024 | Revealing Confounding Biases: A Novel Benchmarking Approach for Aggregate-Level Performance Metrics in Health Assessments
Stefano Goria, Roseline Polle, Salvatore Fara, Nicholas Cummins |
INTERSPEECH | 4 |
| 2023 | Classifying depression symptom severity: Assessment of speech representations in personalized and generalized machine learning modelsabstractThere is an urgent need for new methods that improve the management and treatment of Major Depressive Disorder (MDD). Speech has long been regarded as a promising digital marker in this regard, with many works highlighting that speech changes associated with MDD can be captured through machine learning models. Typically, findings are based on cross-sectional data, with little work exploring the advantages of personalization in building more robust and reliable models. This work assesses the strengths of different combinations of speech representations and machine learning models, in personalized and generalized settings in a two-class depression severity classification paradigm. Key results on a longitudinal dataset highlight the benefits of personalization. Our strongest performing model set-up utilized self-supervised learning features and convolutional neural network (CNN) and long short-term memory (LSTM) back-end. Edward L. Campbell, Judith Dineley, Pauline Conde, Faith Matcham, Katie M. White, Carolin Oetzmann, Sara Simblett, Stuart Bruce, Amos Folarin, Til Wykes, Srinivasan Vairavan, Richard J. B. Dobson, Laura Docío Fernández, Carmen García-Mateo, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins |
INTERSPEECH | 17 |
| 2023 | Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speechabstractSpeech is promising as an objective, convenient tool to monitor health remotely over time using mobile devices. Numerous paralinguistic features have been demonstrated to contain salient information related to an individual’s health. However, mobile device specification and acoustic environments vary widely, risking the reliability of the extracted features. In an initial step towards quantifying these effects, we report the variability of 13 exemplar paralinguistic features commonly reported in the speech-health literature and extracted from the speech of 42 healthy volunteers recorded consecutively in rooms with low and high reverberation with one budget and two higher-end smartphones, and a condenser microphone. Our results show reverberation has a clear effect on several features, in particular voice quality markers. They point to new research directions investigating how best to record and process in-the-wild speech for reliable longitudinal health state assessment. Judith Dineley, Ewan Carr, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 7 |
| 2023 | Bayesian Networks for the robust and unbiased prediction of depression and its symptoms utilizing speech and multimodal dataabstractPredicting the presence of major depressive disorder (MDD) using speech is highly non-trivial. The heterogeneous clinical profile of MDD means that any given speech pattern may be associated with a unique combination of depressive symptoms. Conventional discriminative machine learning models may lack the complexity to robustly model this heterogeneity. Bayesian networks, however, are well-suited to such a scenario. They provide further advantages over standard discriminative modeling by offering the possibility to (i) fuse with other data streams; (ii) incorporate expert opinion into the models; (iii) generate explainable model predictions, inform about the uncertainty of predictions, and (iv) handle missing data. In this study, we apply a Bayesian framework to capture the relationships between depression, depression symptoms, and features derived from speech, facial expression and cognitive game data. Presented results also highlight our model is not subject to demographic biases. Salvatore Fara, Orlaith Hickey, Alexandra Livia Georgescu, Stefano Goria, Emilia Molimpakis, Nicholas Cummins |
INTERSPEECH | 6 |
| 2022 | Speech and the n-Back task as a lens into depression. How combining both may allow us to isolate different core symptoms of depressionabstractEmbedded in any speech signal is a rich combination of cognitive, neuromuscular and physiological information. This richness makes speech a powerful signal in relation to a range of different health conditions, including major depressive disorders (MDD). One pivotal issue in speech-depression research is the assumption that depressive severity is the dominant measurable effect. However, given the heterogeneous clinical profile of MDD, it may actually be the case that speech alterations are more strongly associated with subsets of key depression symptoms. This paper presents strong evidence in support of this argument. First, we present a novel large, cross-sectional, multimodal dataset collected at Thymia. We then present a set of machine learning experiments that demonstrate that combining speech with features from an n-Back working memory assessment improves classifier performance when predicting the popular eight-item Patient Health Questionnaire depression scale (PHQ-8). Finally, we present a set of experiments that highlight the association between different speech and n-Back markers at the PHQ-8 item level. Specifically, we observe that somatic and psychomotor symptoms are more strongly associated with n-Back performance scores, whilst the other items: anhedonia, depressed mood, change in appetite, feelings of worthlessness and trouble concentrating are more strongly associated with speech changes. Salvatore Fara, Stefano Goria, Emilia Molimpakis, Nicholas Cummins |
INTERSPEECH | 4 |
| 2022 | Automatic Detection of Expressed Emotion from Five-Minute Speech Samples: Challenges and OpportunitiesabstractWe present a novel feasibility study on the automatic recognition of Expressed Emotion (EE), a family environment concept based on caregivers speaking freely about their relative/family member.We describe an automated approach for determining the degree of warmth, a key component of EE, from acoustic and text features acquired from a sample of 37 recorded interviews.These recordings, collected over 20 years ago, are derived from a nationally representative birth cohort of 2,232 British twin children and were manually coded for EE.We outline the core steps of extracting usable information from recordings with highly variable audio quality and assess the efficacy of four machine learning approaches trained with different combinations of acoustic and text features.Despite the challenges of working with this legacy data, we demonstrated that the degree of warmth can be predicted with an F1-score of 61.5%.In this paper, we summarise our learning and provide recommendations for future work using real-world speech samples. Bahman Mirheidari, André Bittar, Nicholas Cummins, Johnny Downs, Helen L. Fisher, Heidi Christensen |
INTERSPEECH | 3 |
| 2022 | Fitbeat: COVID-19 estimation based on wristband heart rate using a contrastive convolutional auto-encoder
Shuo Liu 0012, Jing Han 0010, Estela Laporta Puyal, Spyridon Kontaxis, Shaoxiong Sun, Patrick Locatelli, Judith Dineley, Florian B. Pokorny, Gloria Dalla Costa, Letizia Leocani, Ana Isabel Guerrero, Carlos Nos, Ana Zabalza, Per Soelberg Sørensen, Mathias Buron, Melinda Magyari, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Callum L. Stewart, Amos Folarin, Richard J. B. Dobson, Raquel Bailón, Srinivasan Vairavan, Nicholas Cummins, Vaibhav A. Narayan, Matthew Hotopf, Giancarlo Comi, Björn W. Schuller |
Pattern Recognit. | 25 |
| 2022 | Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding PrototypesabstractSpeech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
IEEE Trans. Multim. | 3 |
| 2021 | Hierarchical Attention-Based Temporal Convolutional Networks for Eeg-Based Emotion RecognitionabstractEEG-based emotion recognition is an effective way to infer the inner emotional state of human beings. Recently, deep learning methods, particularly long short-term memory recurrent neural networks (LSTM-RNNs), have made encouraging progress for in the field of emotion recognition. However, the LSTM-RNNs are time-consuming and have difficulty avoiding the problem of exploding/vanishing gradients when during training. In addition, EEG-based emotion recognition often suffers due to the existence of silent and emotional irrelevant frames from intra-channel. Not all channels carry the same emotional discriminative information. In order to tackle these problems, a hierarchical attention-based temporal convolutional networks (HATCN) for efficient EEG-based emotion recognition is proposed. Firstly, a spectrogram representation is generated from raw EEG signals in each channel to capture their time and frequency information. Secondly, temporal convolutional networks (TCNs) are utilised to automatically learn more robust/intrinsic long-term dynamic characters in emotion response. Next, a hierarchical attention mechanism is investigated that aggregates the emotional information at both the frame and channel level. The experimental results on the DEAP dataset show that our method achieves an average recognition accuracy of 0.716 and an F1-score of 0.642 over four emotional dimensions and outperforms other state-of-the-art methods in a user-independent scenario. Ziping Zhao 0001, Nicholas Cummins, Björn W. Schuller |
ICASSP | 4 |
| 2021 | Socially Informed AI for Healthcare: Understanding and Generating Multimodal Nonverbal CuesabstractAdvances in the areas of face and gesture analysis, computational paralinguistics, multimodal interaction, and human-computer interaction have all played a major role in shaping research into assistive technologies over the last decade. This has resulted in a breadth of practical applications ranging from diagnosis and treatment tools to social companion technologies. From an analytical perspective, nonverbal cues provide understanding into the assessment of mental health and wellbeing (i.e., detecting depression and pain) and the detection of developmental and neurological conditions such as autism, dementia, and schizophrenia. From both a synthesis and generative perspective, it is necessary that assistive technologies, either disembodied or embodied, are capable of generating engaging, interactive behaviours and interventions that are personalised and adapted to user’s needs, profiles, and preferences. While nonverbal cues play an essential role, there are still many key issues to overcome, which affect both the development and the deployment of multimodal technologies in real-world settings. The key aim of this multidisciplinary workshop is to foster cross-pollination by bringing together computer scientists and social psychologists to discuss innovative ideas, challenges and opportunities for understanding and generating multimodal nonverbal cues within the scope of healthcare applications1. Oya Çeliktutan, Alexandra Livia Georgescu, Nicholas Cummins |
ICMI | 3 |
| 2021 | Remote Smartphone-Based Speech Collection: Acceptance and Barriers in Individuals with Major Depressive DisorderabstractThe ease of in-the-wild speech recording using smartphones has sparked considerable interest in the combined application of speech, remote measurement technology (RMT) and advanced analytics as a research and healthcare tool. For this to be realised, the acceptability of remote speech collection to the user must be established, in addition to feasibility from an analytical perspective. To understand the acceptance, facilitators, and barriers of smartphone-based speech recording, we invited 384 individuals with major depressive disorder (MDD) from the Remote Assessment of Disease and Relapse - Central Nervous System (RADAR-CNS) research programme in Spain and the UK to complete a survey on their experiences recording their speech. In this analysis, we demonstrate that study participants were more comfortable completing a scripted speech task than a free speech task. For both speech tasks, we found depression severity and country to be significant predictors of comfort. Not seeing smartphone notifications of the scheduled speech tasks, low mood and forgetfulness were the most commonly reported obstacles to providing speech recordings. Judith Dineley, Grace Lavelle, Daniel Leightley, Faith Matcham, Sara Siddi, Maria Teresa Peñarrubia-María, Katie M. White, Alina Ivan, Carolin Oetzmann, Sara Simblett, Erin Dawe-Lane, Stuart Bruce, Daniel Stahl, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Amos Folarin, Josep Maria Haro, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Björn W. Schuller, Nicholas Cummins |
Interspeech | 24 |
| 2021 | Representation transfer learning from deep end-to-end speech recognition networks for the classification of health states from speechabstractRepresentation transfer learning has been widely used across a range of machine learning tasks. One such notable approach seen in the speech literature is the use of Convolutional Neural Networks, pre-trained for image classification tasks, to extract features from spectrograms of speech signals. Interestingly, despite the strong performance of such approaches, there have been minimal research efforts exploring the suitability of using speech-specific networks to perform feature extraction. In this regard, a novel feature representation learning framework is presented herein. This approach is comprising the use of Automatic Speech Recognition (ASR) deep neural networks as feature extractors, the fusion of several extracted feature representations using Compact Bilinear Pooling (CBP), and finally inference via a specially optimised Recurrent Neural Network (RNN) classifier. To determine the usefulness of these feature representations, they are comprehensively tested on two representative speech-health classification tasks, namely the food-type being eaten and speaker intoxication. Key results indicate the promise of the extracted features, demonstrating comparable results to other state-of-the-art approaches in the literature. Benjamin Sertolli, Zhao Ren, Björn W. Schuller, Nicholas Cummins |
Comput. Speech Lang. | 4 |
| 2021 | Combining a parallel 2D CNN with a self-attention Dilated Residual Network for CTC-based discrete speech emotion recognition
Ziping Zhao 0001, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Neural Networks | 4 |
| 2021 | Guided Generative Adversarial Neural Network for Representation Learning and Audio Generation Using Fewer Labelled Audio DataabstractThe Generation power of Generative Adversarial Neural Networks (GANs) has shown great promise to learn representations from unlabelled data while guided by a small amount of labelled data. We aim to utilise the generation power of GANs to learn Audio Representations. Most existing studies are, however, focused on images. Some studies use GANs for speech generation, but they are conditioned on text or acoustic features, limiting their use for other audio, such as instruments, and even for speech where transcripts are limited. This paper proposes a novel GAN-based model that we named Guided Generative Adversarial Neural Network (GGAN), which can learn powerful representations and generate good-quality samples using a small amount of labelled data as guidance. Experimental results based on a speech [Speech Command Dataset (S09)] and a non-speech [Musical Instrument Sound dataset (Nsyth)] dataset demonstrate that using only 5% of labelled data as guidance, GGAN learns significantly better representations than the state-of-the-art models. Kazi Nazmul Haque, Rajib Rana, Jiajun Liu 0013, John H. L. Hansen, Nicholas Cummins, Carlos Busso, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Predictable Robots for Autistic Children - Variance in Robot Behaviour, Idiosyncrasies in Autistic Children's Characteristics, and Child-Robot EngagementabstractPredictability is important to autistic individuals, and robots have been suggested to meet this need as they can be programmed to be predictable, as well as elicit social interaction. The effectiveness of robot-assisted interventions designed for social skill learning presumably depends on the interplay between robot predictability, engagement in learning, and the individual differences between different autistic children. To better understand this interplay, we report on a study where 24 autistic children participated in a robot-assisted intervention. We manipulated the variance in the robot’s behaviour as a way to vary predictability, and measured the children’s behavioural engagement, visual attention, as well as their individual factors. We found that the children will continue engaging in the activity behaviourally, but may start to pay less visual attention over time to activity-relevant locations when the robot is less predictable. Instead, they increasingly start to look away from the activity. Ultimately, this could negatively influence learning, in particular for tasks with a visual component. Furthermore, severity of autistic features and expressive language ability had a significant impact on behavioural engagement. We consider our results as preliminary evidence that robot predictability is an important factor for keeping children in a state where learning can occur. Bob Schadenberg, Dennis Reidsma, Vanessa Evers, Daniel P. Davison, Jamy Li, Dirk Heylen, Carlos Neves 0004, Paulo Alvito, Jie Shen 0008, Maja Pantic, Björn W. Schuller, Nicholas Cummins, Vlad Olaru, Cristian Sminchisescu, Snezana Babovic, Suncica Petrovic, Aurelie Baranger, Alria Williams, Alyssa Alcorn, Elizabeth Pellicano |
ACM Trans. Comput. Hum. Interact. | 12 |
| 2021 | Self-attention transfer networks for speech emotion recognitionabstractA crucial element of human–machine interaction, the automatic detection of emotional states from human speech has long been regarded as a challenging task for machine learning models. One vital challenge in speech emotion recognition (SER) is how to learn robust and discriminative representations from speech. Meanwhile, although machine learning methods have been widely applied in SER research, the inadequate amount of available annotated data has become a bottleneck that impedes the extended application of techniques (e.g., deep neural networks). To address this issue, we present a deep learning method that combines knowledge transfer and self-attention for SER tasks. Here, we apply the log-Mel spectrogram with deltas and delta-deltas as input. Moreover, given that emotions are time-dependent, we apply Temporal Convolutional Neural Networks (TCNs) to model the variations in emotions. We further introduce an attention transfer mechanism, which is based on a self-attention algorithm in order to learn long-term dependencies. The Self-Attention Transfer Network (SATN) in our proposed approach, takes advantage of attention autoencoders to learn attention from a source task, and then from speech recognition, followed by transferring this knowledge into SER. Evaluation built on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) demonstrates the effectiveness of the novel model. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Shihuang Sun, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 4 |
| 2020 | A Curriculum Learning Approach for Pain Intensity Recognition from Facial ExpressionsabstractThe high prevalence of chronic pain in society raises the need to develop new digital tools that can automatically and objectively assess pain intensity in individuals. These tools can contribute to an optimisation of clinical resources, as they offer cost-effective solutions for early detection, continuous monitoring, and treatment personalisation by utilising Artificial Intelligence techniques. In this work, we present our contribution to the Pain Intensity Estimation from Facial Expressions task of the EMOPAIN 2020 Challenge. Specifically, we compare the performance of Recurrent Neural Networks trained with standard or Curriculum Learning (CL) approaches to predict the pain intensity level of individuals reported in an 11-point scale from facial expressions. The results obtained using the test partition support the use of CL-based approaches in the automatic prediction of pain from facial features. The best model trained using a CL approach achieved a Concordance Correlation Coefficient (CCC) of 0.196 in the test partition, while the model trained using a standard approach, without CL, achieved a CCC of 0.174. In terms of CCC, these results respectively represent an improvement of 0.136 and 0.114 on the best results of the baseline system reported by the Challenge organisers using the test partition. Adria Mallol-Ragolta, Shuo Liu 0012, Nicholas Cummins, Björn W. Schuller |
FG | 3 |
| 2020 | Hierarchical Attention Transfer Networks for Depression Assessment from SpeechabstractA growing area of mental health research is the search for speech-based objective markers for conditions such as depression. However, when combined with machine learning, this search can be challenging due to a limited amount of annotated training data. In this paper, we propose a novel crosstask approach which transfers attention mechanisms from speech recognition to aid depression severity measurement. This transfer is applied in a two-level hierarchical network which mirrors the natural hierarchical structure of speech. Experiments based on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) dataset, as used in the 2017 Audio/Visual Emotion Challenge, demonstrate the effectiveness of our Hierarchical Attention Transfer Network. On the development set, the proposed approach achieves a root mean square error (RMSE) of 3.85, and a mean absolute error (MAE) of 2.99, on a Patient Health Questionnaire (PHQ)-8 scale [0], [24], while on the test set, it achieves an RMSE of 5.66 and an MAE of 4.28. To the best of our knowledge, these scores represent the best-known speech-only results to date on this corpus. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
ICASSP | 4 |
| 2020 | Hierarchical Component-attention Based Speaker Turn Embedding for Emotion RecognitionabstractTraditional discrete-time Speech Emotion Recognition (SER) modelling techniques typically assume that an entire speaker chunk or turn is indicative of its corresponding label. An alternative approach is to assume emotional saliency varies over the course of a speaker turn and use modelling techniques capable of identifying and utilising the most emotionally salient segments, such as those with higher emotional intensity. This strategy has the potential to improve the accuracy of SER systems. Towards this goal, we developed a novel hierarchical recurrent neural network model that produces turn level embeddings for SER. Specifically, we apply two levels of attention to learn to identify salient emotional words in a turn as well as the more informative frames within these words. In a set of experiments on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, we demonstrate that component-attention is more effective within our hierarchical framework than both standard soft-attention and conventional local-attention. Our best network, a hierarchical component-attention network with an attention scope of seven, achieved an Unweighted Average Recall (UAR) of 65.0 % and a Weighted Average Recall (WAR) of 66.1 %, outperforming other baseline attention approaches on the IEMOCAP database. Shuo Liu 0012, Jinlong Jiao, Ziping Zhao 0001, Judith Dineley, Nicholas Cummins, Björn W. Schuller |
IJCNN | 5 |
| 2020 | Squeeze for Sneeze: Compact Neural Networks for Cold and Flu RecognitionabstractIn digital health applications, speech offers advantages over other physiological signals, in that it can be easily collected, transmitted, and stored using mobile and Internet of Things (IoT) technologies.However, to take full advantage of this positioning, speech-based machine learning models need to be deployed on devices that can have considerable memory and power constraints.These constraints are particularly apparent when attempting to deploy deep learning models, as they require substantial amounts of memory and data movement operations.Herein, we test the suitability of pruning and quantisation as two methods to compress the overall size of neural networks trained for a health-driven speech classification task. Key results presented on the Upper Respiratory Tract InfectionCorpus indicate that pruning, then quantising a network can reduce the number of operational weights by almost 90 %.They also demonstrate the overall size of the network can be reduced by almost 95 %, as measured in MB, without affecting overall recognition performance. Merlin Albes, Zhao Ren, Björn W. Schuller, Nicholas Cummins |
INTERSPEECH | 4 |
| 2020 | An Evaluation of the Effect of Anxiety on Speech - Computational Prediction of Anxiety from Sustained VowelsabstractThe current level of global uncertainty is having an implicit effect on those with a diagnosed anxiety disorder.Anxiety can impact vocal qualities, particularly as physical symptoms of anxiety include muscle tension and shortness of breath.To this end, in this study, we explore the effect of anxiety on speech -focusing on four classes of sustained vowels (sad, smiling, comfortable, and powerful) -via feature analysis and a series of regression experiments.We extract three well-known acoustic feature sets and evaluate the efficacy of machine learning for prediction of anxiety based on the Beck Anxiety Inventory (BAI) score.Of note, utilising a support vector regressor, we find that the effects of anxiety in speech appear to be stronger at higher BAI levels.Significant differences (p < 0.05) between test predictions of Low and High-BAI groupings support this.Furthermore, when utilising a High-BAI grouping for the prediction of standardised BAI, significantly higher results are obtained for smiling sustained vowels, of up to 0.646 Spearman's Correlation Coefficient (ρ), and up to 0.592 ρ with all sustained vowels.A significantly stronger (Cohens d of 1.718) result than all data combined without grouping, which achieves at best 0.234 ρ. Alice Baird, Nicholas Cummins, Sebastian Schnieder, Jarek Krajewski, Björn W. Schuller |
INTERSPEECH | 2 |
| 2020 | A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia RecognitionabstractContains fulltext : 228158.pdf (Publisher’s version ) (Open Access) Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä |
INTERSPEECH | 1 |
| 2020 | An Investigation of Cross-Cultural Semi-Supervised Learning for Continuous Affect RecognitionabstractOne of the keys for supervised learning techniques to succeed resides in the access to vast amounts of labelled training data. The process of data collection, however, is expensive, time- consuming, and application dependent. In the current digital era, data can be collected continuously. This continuity renders data annotation into an endless task, which potentially, in problems such as emotion recognition, requires annotators with different cultural backgrounds. Herein, we study the impact of utilising data from different cultures in a semi-supervised learning ap- proach to label training material for the automatic recognition of arousal and valence. Specifically, we compare the performance of culture-specific affect recognition models trained with man- ual or cross-cultural automatic annotations. The experiments performed in this work use the dataset released for the Cross- cultural Emotion Sub-challenge of the Audio/Visual Emotion Challenge (AVEC) 2019. The results obtained convey that the cultures used for training impact on the system performance. Furthermore, in most of the scenarios assessed, affect recogni- tion models trained with hybrid solutions, combining manual and automatic annotations, surpass the baseline model, which was exclusively trained with manual annotations. Adria Mallol-Ragolta, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 2 |
| 2020 | Enhancing Transferability of Black-Box Adversarial Attacks via Lifelong Learning for Speech Emotion Recognition ModelsabstractWell-designed adversarial examples can easily fool deep speech emotion recognition models into misclassifications.The transferability of adversarial attacks is a crucial evaluation indicator when generating adversarial examples to fool a new target model or multiple models.Herein, we propose a method to improve the transferability of black-box adversarial attacks using lifelong learning.First, black-box adversarial examples are generated by an atrous Convolutional Neural Network (CNN) model.This initial model is trained to attack a CNN target model.Then, we adapt the trained atrous CNN attacker to a new CNN target model using lifelong learning.We use this paradigm, as it enables multi-task sequential learning, which saves more memory space than conventional multi-task learning.We verify this property on an emotional speech database, by demonstrating that the updated atrous CNN model can attack all target models which have been learnt, and can better attack a new target model than an attack model trained on one target model only. Zhao Ren, Jing Han 0010, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | Hybrid Network Feature Extraction for Depression Assessment from SpeechabstractA fast-growing area of mental health research is the search for speech-based objective markers for conditions such as depression.One vital challenge in the development of speech-based depression severity assessment systems is the extraction of depression-relevant features from speech signals.In order to deliver more comprehensive feature representation, we herein explore the benefits of a hybrid network that encodes depressionrelated characteristics in speech for the task of depression severity assessment.The proposed network leverages self-attention networks (SAN) trained on low-level acoustic features and deep convolutional neural networks (DCNN) trained on 3D Log-Mel spectrograms.The feature representations learnt in the SAN and DCNN are concatenated and average pooling is exploited to aggregate complementary segment-level features.Finally, support vector regression is applied to predict a speaker's Beck Depression Inventory-II score.Experiments based on a subset of the Audio-Visual Depressive Language Corpus, as used in the 2013 and 2014 Audio/Visual Emotion Challenges, demonstrate the effectiveness of our proposed hybrid approach. Ziping Zhao 0001, Nicholas Cummins, Bin Liu 0041, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | I see it in your eyes: Training the shallowest-possible CNN to recognise emotions and pain from muted web-assisted in-the-wild video-chats in real-time
Vedhas Pandit, Maximilian Schmitt, Nicholas Cummins, Björn W. Schuller |
Inf. Process. Manag. | 3 |
| 2020 | Generalized Two-Stage Rank Regression Framework for Depression Score Prediction from SpeechabstractThis paper introduces a novel speech-based depression score prediction paradigm, the 2-stage ranking prediction framework, and highlights the benefits it brings to depression prediction. Conventional regression approaches aim to discern a single functional relationship between speech features and depression scores, making an implicit assumption about the existence of a single fixed relationship between the features and scores. However, as the relationship between severity of depression and the clinical score may vary over the range of the assessment scale, this style of analysis may not be suited to depression prediction. The proposed framework on the other hand, imposes a series of partitions on the feature space, with each partition corresponding to a distinct predefined range of depression scores, and predicts the score based on measures of membership to each partition. This approach provides additional flexibility by allowing different rankings to be learnt for different depression scores, and relaxes assumptions made by conventional regression approaches. Results demonstrate the framework's suitability for depression score prediction: different 2-stage implementations, based on heterogeneous feature extraction and modelling approaches, produce state-of-the-art results on the AVEC-2013 dataset. It is also demonstrated that, unlike fusion of conventional regression systems, the fusion of two-stage systems consistently improves prediction performance. Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, James R. Williamson, Thomas F. Quatieri, Jarek Krajewski |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | "Are You Playing a Shooter Again?!" Deep Representation Learning for Audio-Based Video Game Genre RecognitionabstractIn this paper, we present a novel computer audition task: audio-based video game genre classification. The aim of this study is threefold: 1) to check the feasibility of the proposed task; 2) to introduce a new corpus: The Game Genre by Audio + Multimodal Extracts (G2 AME), collected entirely from social multimedia; and 3) to compare the efficacy of various acoustic feature spaces to classify the G2 AME corpus into six game genres using a linear support vector machine classifier. For the classification we extract three different feature representations from the game audio files: 1) Knowledge-based acoustic features; 2) DEEP SPECTRUM features; and 3) quantized DEEP SPECTRUM features using Bag-of-Audio-Words. The DEEP SPECTRUM features are a deep-learning-based representation derived from forwarding the visual representations of the audio instances, in particular spectrograms, mel-spectrograms, chromagrams, and their deltas through deep task-independent pretrained CNNs. Specifically, activations of fully connected layers from three common image classification CNNs, GoogLeNet, AlexNet, and VGG16 are used as feature vectors. Results for the six-genre classification problem indicate the suitability of our deep learning approach for this task. Our best method achieves an accuracy of up to 66.9% unweighted average recall using tenfold cross-validation. Shahin Amiriparian, Nicholas Cummins, Maurice Gerczuk, Sergey Pugachevskiy, Sandra Ottl, Björn W. Schuller |
IEEE Trans. Games | 2 |
| 2019 | I Know How you Feel Now, and Here's why!: Demystifying Time-Continuous High Resolution Text-Based Affect Predictions in the WildabstractAffective computing 'in the wild' is of huge relevance to the healthcare field, like it is for many industries today. Applications of direct relevance are patient monitoring (e.g., emotional state, depression and pain monitoring), health information mining, diagnosis and opinion mining (e.g., from medical reports and drug reviews). The prevalence of the text modality in the medical field for various reasons - e.g., privacy laws, high costs and prohibitory memory requirements for audio and video data - has made the text modality the most popular. Deviating away from traditionally a classification task at a sample-level, the promising baseline results for the Audio/Visual Emotion Challenge (AVEC) 2017 make a strong case for the suitability of text data for a 'time-continuous' affect estimation. For the very first time, we present insights into the inner workings of deep learning, 'in the wild' affect-predicting, time-continuous regression model. We compute relevance of the sparse text-based bag-of-words features (BoTW) of the AVEC 2017 challenge in estimating the three affect labels, viz. arousal, valence and liking, by using a layerwise relevance propagation method(LRP). Interestingly, the trained models are found to rely more on adjectives and adverbs such as 'schlecht', 'gut', 'genau' with positive or negative connotations, and action descriptors such asand- quite analogous to the human perception of emotion expression. Vedhas Pandit, Maximilian Schmitt, Nicholas Cummins, Björn W. Schuller |
CBMS | 3 |
| 2019 | Performance Analysis of Unimodal and Multimodal Models in Valence-Based Empathy RecognitionabstractThe human ability to empathise is a core aspect of successful interpersonal relationships. In this regard, human-robot interaction can be improved through the automatic perception of empathy, among other human attributes, allowing robots to affectively adapt their actions to interactants' feelings in any given situation. This paper presents our contribution to the generalised track of the One-Minute Gradual (OMG) Empathy Prediction Challenge by describing our approach to predict a listener's valence during semi-scripted actor-listener interactions. We extract visual and acoustic features from the interactions and feed them into a bidirectional long short-term memory network to capture the time-dependencies of the valence-based empathy during the interactions. Generalised and personalised unimodal and multimodal valence-based empathy models are then trained to assess the impact of each modality on the system performance. Furthermore, we analyse if intra-subject dependencies on empathy perception affect the system performance. We assess the models by computing the concordance correlation coefficient (CCC) between the predicted and self-annotated valence scores. The results support the suitability of employing multimodal data to recognise participants' valence-based empathy during the interactions, and highlight the subject-dependency of empathy. In particular, we obtained our best result with a personalised multimodal model, which achieved a CCC of 0.11 on the test set. Adria Mallol-Ragolta, Maximilian Schmitt, Alice Baird, Nicholas Cummins, Björn W. Schuller |
FG | 4 |
| 2019 | Context Modelling Using Hierarchical Attention Networks for Sentiment and Self-assessed Emotion Detection in Spoken NarrativesabstractAutomatic detection of sentiment and affect in personal narratives through word usage has the potential to assist in the automated detection of change in psychotherapy. Such a tool could, for instance, provide an efficient, objective measure of the time a person has been in a positive or negative state-of-mind. Towards this goal, we propose and develop a hierarchical attention model for the tasks of sentiment (positive and negative) and self-assessed affect detection in transcripts of personal narratives. We also perform a qualitative analysis of the word attentions learnt by our sentiment analysis model. In a key result, our attention model achieved an un-weighted average recall (UAR) of 91.0 % in a binary sentiment detection task on the test partition of the Ulm State-of-Mind in Speech (USoMS) corpus. We also achieved UARs of 73.7 % and 68.6 % in the 3-class tasks of arousal and valence detection respectively. Finally, our qualitative analysis associates colloquial reinforcements with positive sentiments, and uncertain phrasing with negative sentiments. Lukas Stappen, Nicholas Cummins, Eva-Maria Messner, Harald Baumeister, Judith Dineley, Björn W. Schuller |
ICASSP | 2 |
| 2019 | Analysing and Inferring of Intimacy Based on fNIRS Signals and Peripheral Physiological SignalsabstractIntimacy refers to a relatively long-lasting affinity relationship between individuals, which involves complex neuronal activities and physiological changes in the body. Recent advancements in the field of neuroimaging have demonstrated that functional near-infrared spectroscopy (fNIRS) has excellent potential for intimate relationship analysis. Signals such as fNIRS and physiological signals are increasingly utilised in this regard due to their consistency and complementarity. In this paper, first, we apply fNIRS and physiological database collected from 26 subjects when viewing lover, friend and stranger pictures to analyse and infer the intimacy. Then, the time domain information from both the fNIRS and physiological signals are utilised to exploit the representation of intimacy by General Linear Model (GLM) and Complex Brain Network Analysis (CBNA) methods. Based on these two methods, the intimacy can be analysed with different brain activation patterns. Finally, different machine learning techniques are utilised to predict the intimate relationship. The results demonstrate that multi-modal features are more efficient for intimacy research. Moreover, the average classification accuracy of ensemble learning is 98.72% whereas for KNN it is 91.03%. Ziping Zhao 0001, Li Gu, Nicholas Cummins, Björn W. Schuller |
IJCNN | 5 |
| 2019 | Using Speech to Predict Sequentially Measured Cortisol Levels During a Trier Social Stress TestabstractThe effect of stress on the human body is substantial, potentially resulting in serious health implications.Furthermore, with modern stressors seemingly on the increase, there is an abundance of contributing factors which lead to a diagnosis of acute stress.However, observing biological stress reactions usually includes costly and time consuming sequential fluidbased samples to determine the degree of biological stress.On the contrary, a speech monitoring approach would allow for a non-invasive indication of stress.To evaluate the efficacy of the speech signal as a marker of stress, we explored, for the first time, the relationship between sequential cortisol samples and speech-based features.Utilising a novel corpus of 43 individuals undergoing a standardised Trier Social Stress Test (TSST), we extract a variety of feature sets and observe a correlation between speech and sequential cortisol measurements.For prediction of mean cortisol levels from speech, results show that for the entire TSST oral presentation, handcrafted COMPARE features achieve best results of 0.244 root mean square error [0 ;1] for the sample 20 minutes after the TSST.Correlation also increases at minute 20, with a Spearman's correlation coefficient of 0.421, and Cohen's d of 0.883 between the baseline and minute 20 cortisol predictions. Alice Baird, Shahin Amiriparian, Nicholas Cummins, Sarah Sturmbauer, Johanna Janson, Eva-Maria Messner, Harald Baumeister, Nicolas Rohleder, Björn W. Schuller |
INTERSPEECH | 3 |
| 2019 | A Hierarchical Attention Network-Based Approach for Depression Detection from Transcribed Clinical InterviewsabstractThe high prevalence of depression in society has given rise to a need for new digital tools that can aid its early detection. Among other effects, depression impacts the use of language. Seeking to exploit this, this work focuses on the detection of depressed and non-depressed individuals through the analysis of linguistic information extracted from transcripts of clinical interviews with a virtual agent. Specifically, we investigated the advantages of employing hierarchical attention-based networks for this task. Using Global Vectors (GloVe) pretrained word embedding models to extract low-level representations of the words, we compared hierarchical local-global attention networks and hierarchical contextual attention networks. We performed our experiments on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WoZ) dataset, which contains audio, visual, and linguistic information acquired from participants during a clinical session. Our results using the DAIC-WoZ test set indicate that hierarchical contextual attention networks are the most suitable configuration to detect depression from transcripts. The configuration achieves an Unweighted Average Recall (UAR) of .66 using the test set, surpassing our baseline, a Recurrent Neural Network that does not use attention. Adria Mallol-Ragolta, Ziping Zhao 0001, Lukas Stappen, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 4 |
| 2019 | Continuous Emotion Recognition in Speech - Do We Need Recurrence?abstractEmotion recognition in speech is a meaningful task in affective computing and human-computer interaction.As human emotion is a frequently changing state, it is usually represented as a densely sampled time series of emotional dimensions, typically arousal and valence.For this, recurrent neural network (RNN) architectures are employed by default when it comes to modelling the contours with deep learning approaches.However, the amount of temporal context required is questionable, and it has not yet been clarified whether the consideration of long-term dependencies is actually beneficial.In this contribution, we demonstrate that RNNs are not necessary to accomplish the task of time-continuous emotion recognition.Indeed, results gained indicate that deep neural networks incorporating less complex convolutional layers can provide more accurate models.We highlight the pros and cons of recurrent and nonrecurrent approaches and evaluate our methods on the public SEWA database, which was used as a benchmark in the 2017 and 2018 editions of the Audio-Visual Emotion Challenge. Maximilian Schmitt, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 2 |
| 2019 | Autonomous Emotion Learning in Speech: A View of Zero-Shot Speech Emotion RecognitionabstractConventionally, speech emotion recognition is achieved using passive learning approaches.Differing from such approaches, we herein propose and develop a dynamic method of autonomous emotion learning based on zero-shot learning.The proposed methodology employs emotional dimensions as the attributes in the zero-shot learning paradigm, resulting in two phases of learning, namely attribute learning and label learning.Attribute learning connects the paralinguistic features and attributes utilising speech with known emotional labels, while label learning aims at defining unseen emotions through the attributes.The experimental results achieved on the CINEMO corpus indicate that zero-shot learning is a useful technique for autonomous speech-based emotion learning, achieving accuracies considerably better than chance level and an attribute-based gold-standard setup.Furthermore, different emotion recognition tasks, emotional attributes, and employed approaches strongly influence system performance. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
INTERSPEECH | 3 |
| 2019 | Attention-Enhanced Connectionist Temporal Classification for Discrete Speech Emotion RecognitionabstractDiscrete speech emotion recognition (SER), the assignment of a single emotion label to an entire speech utterance, is typically performed as a sequence-to-label task.This approach, however, is limited, in that it can result in models that do not capture temporal changes in the speech signal, including those indicative of a particular emotion.One potential solution to overcome this limitation is to model SER as a sequence-to-sequence task instead.In this regard, we have developed an attention-based bidirectional long short-term memory (BLSTM) neural network in combination with a connectionist temporal classification (CTC) objective function (Attention-BLSTM-CTC) for SER.We also assessed the benefits of incorporating two contemporary attention mechanisms, namely component attention and quantum attention, into the CTC framework.To the best of the authors' knowledge, this is the first time that such a hybrid architecture has been employed for SER.We demonstrated the effectiveness of our approach on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and FAU-Aibo Emotion corpora.The experimental results demonstrate that our proposed model outperforms current state-of-the-art approaches. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
INTERSPEECH | 4 |
| 2019 | AVEC'19: Audio/Visual Emotion Challenge and WorkshopabstractThe ninth Audio-Visual Emotion Challenge and workshop AVEC 2019 was held in conjunction with ACM Multimedia'19. This year, the AVEC series addressed major novelties with three distinct tasks: State-of-Mind Sub-challenge (SoMS), Detecting Depression with Artificial Intelligence Sub-challenge (DDS), and Cross-cultural Emotion Sub-challenge (CES). The SoMS was based on a novel dataset (USoM corpus) that includes self-reported mood (10-point Likert scale) after the narrative of personal stories (two positive and two negative). The DDS was based on a large extension of the DAIC-WOZ corpus (c.f. AVEC 2016) that includes new recordings of patients suffering from depression with the virtual agent conducting the interview being, this time, wholly driven by AI, i.e., without any human intervention. The CES was based on the SEWA dataset (c.f. AVEC 2018) that has been extended with the inclusion of new participants in order to investigate how emotion knowledge of Western European cultures (German, Hungarian) can be transferred to the Chinese culture. In this summary, we mainly describe participation and conditions of the AVEC Challenge. Fabien Ringeval, Björn W. Schuller, Michel F. Valstar, Nicholas Cummins, Roddy Cowie, Maja Pantic |
ACM Multimedia | 4 |
| 2019 | From Speech to Facial Activity: Towards Cross-modal Sequence-to-Sequence Attention NetworksabstractMultimodal data sources offer the possibility to capture and model interactions between modalities, leading to an improved understanding of underlying relationships. In this regard, the work presented in this paper explores the relationship between facial muscle movements and speech signals. Specifically, we explore the efficacy of different sequence-to-sequence neural network architectures for the task of predicting Facial Action Coding System Action Units (AUs) from one of two acoustic feature representations extracted from speech signals, namely the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPs) or the Interspeech Computational Paralinguistics Challenge features set (ComParE). Furthermore, these architectures were enhanced by two different attention mechanisms (intra- and inter-attention) and various state-of-the-art network settings to improve prediction performance. Results indicate that a sequence-to-sequence model with inter-attention can achieve on average an Unweighted Average Recall (UAR) of 65.9 % for AU onset, 67.8 % for AU apex (both eGeMAPs), 79.7 % for AU offset and 65.3 % for AU occurrence (both ComParE) detection over all AUs. Lukas Stappen, Vincent Karas, Nicholas Cummins, Fabien Ringeval, Klaus R. Scherer, Björn W. Schuller |
MMSP | 3 |
| 2018 | Multimodal Bag-of-Words for Cross Domains Sentiment AnalysisabstractThe advantages of using cross domain data when performing text-based sentiment analysis have been established; however, similar findings have yet to be observed when performing multimodal sentiment analysis. A potential reason for this is that systems based on feature extracted from speech and facial features are susceptible to confounding effecting caused by different recording conditions associated with data collected in different locations. In this regard, we herein explore different Bag-of-Words paradigms to aid sentiment detection by providing training material from an additional dataset. Key results presented indicate that using a Bag-of-Words extraction paradigm that takes into account information from both the test domain and the out of domain datasets yields gains in system performance. Nicholas Cummins, Shahin Amiriparian, Sandra Ottl, Maurice Gerczuk, Maximilian Schmitt, Björn W. Schuller |
ICASSP | 1 |
| 2018 | What is my Dog Trying to Tell Me? the Automatic Recognition of the Context and Perceived Emotion of Dog BarksabstractA wide range of research disciplines are deeply interested in the measurement of animal emotions, including evolutionary zoology, affective neuroscience and comparative psychology. However, only a few studies have investigated the effect of phenomena such as emotion on the acoustic parameters of (non-human) mammalian species. In this contribution, we explore if commonly used affective computing-based acoustic feature sets can be used to classify either the context, the emotion, or predict the emotional intensity of dog bark sequences. This comparison study includes an in-depth analysis of obtainable classification performances. Results presented indicate that the tested feature representations are suitable for the proposed recognition tasks. Of particular note are results that demonstrate machine learning-based acoustic analysis can achieve above human level performance when classifying the context of a dog bark. Simone Hantke, Nicholas Cummins, Björn W. Schuller |
ICASSP | 2 |
| 2018 | Deep End-to-End Representation Learning for Food Type Recognition from SpeechabstractThe use of Convolutional Neural Networks (CNN) pre-trained for a particular task, as a feature extractor for an alternate task, is a standard practice in many image classification paradigms. However, to date there have been comparatively few works exploring this technique for speech classification tasks. Herein, we utilise a pre-trained end-to-end Automatic Speech Recognition CNN as a feature extractor for the task of food-type recognition from speech. Furthermore, we also explore the benefits of Compact Bilinear Pooling for combining multiple feature representations extracted from the CNN. Key results presented indicate the suitability of this approach. When combined with a Recurrent Neural Network classifier, our strongest system achieves, for a seven-class food-type classification task an unweighted average recall of 73.3% on the test set of the iHEARu-EAT database. Benjamin Sertolli, Nicholas Cummins, Abdulkadir Sengür, Björn W. Schuller |
ICMI | 2 |
| 2018 | Bag-of-Deep-Features: Noise-Robust Deep Feature Representations for Audio AnalysisabstractIn the era of deep learning, research into the classification of various components of the acoustic environment, especially in-the-wild recordings, is gaining in popularity. This is due in part to the increasing computational capacities and the expanding amount of real-world data available on social multimedia. However, the noisy nature of this data can add an additional complexity to the already complex deep learning systems. Herein, we tackle this issue by quantising deep feature representations of various in-the-wild audio data sets. The aim of this paper is twofold: 1) to assess the feasibility of the proposed feature quantisation task, and 2) to compare the efficacy of various feature spaces extracted from different fully connected deep neural networks to classify six real-world audio corpora. For the classification, we extract two feature sets: i) DEEP SPECTRUM features which are derived from forwarding the visual representations of the audio instances, in particular mel-spectrograms through very deep task-independent pre-trained Convolutional Neural Networks (CNNs), and ii) Bag-of-Deep-Features (BODF) which is the quantisation of the DEEP SPECTRUM features. Using BODF, we show the suitability of quantising the deep representations for noisy in-the-wild audio data. Finally, we analyse the effect of early and late fusion of the CNN features and models on the classification results. Shahin Amiriparian, Maurice Gerczuk, Sandra Ottl, Nicholas Cummins, Sergey Pugachevskiy, Björn W. Schuller |
IJCNN | 4 |
| 2018 | Recognition of Echolalic Autistic Child Vocalisations Utilising Convolutional Recurrent Neural NetworksabstractAutism spectrum conditions (ASC) are a set of neurodevelopmental conditions partly characterised by difficulties with communication.Individuals with ASC can show a variety of atypical speech behaviours, including echolalia or the 'echoing' of another's speech.We herein introduce a new dataset of 15 Serbian ASC children in a human-robot interaction scenario, annotated for the presence of echolalia amongst other ASC vocal behaviours.From this, we propose a four-class classification problem and investigate the suitability of applying a 2D convolutional neural network augmented with a recurrent neural network with bidirectional long short-term memory cells to solve the proposed task of echolalia recognition.In this approach, log Mel-spectrograms are first generated from the audio recordings and then fed as input into the convolutional layers to extract high-level spectral features.The subsequent recurrent layers are applied to learn the long-term temporal context from the obtained features.Finally, we use a feed forward neural network with softmax activation to classify the dataset.To evaluate the performance of our deep learning approach, we use leave-onesubject-out cross-validation.Key results presented indicate the suitability of our approach by achieving a classification accuracy of 83.5 % unweighted average recall. Shahin Amiriparian, Alice Baird, Sahib Julka, Alyssa Alcorn, Sandra Ottl, Suncica Petrovic, Eloise Ainger, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 8 |
| 2018 | The Perception and Analysis of the Likeability and Human Likeness of Synthesized SpeechabstractThe synthesized voice has become an ever present aspect of daily life.Heard through our smart-devices and from public announcements, engineers continue in an endeavour to achieve naturalness in such voices.Yet, the degree to which these methods can produce likeable, human like voices, has not been fully evaluated.With recent advancements in synthetic speech technology suggesting that human like imitation is more obtainable, this study asked 25 listeners to evaluate both the likeability and human likeness of a corpus of 13 German male voices, produced via 5 synthesis approaches (from formant to hybrid unit selection, deep neural network systems), and 1 Human control.Results show that unlike visual artificially intelligent elements -as posed by the concept of the Uncanny Valley -likeability consistently improves along with human likeness for the synthesized voice, with recent methods achieving substantially closer results to human speech than older methods.A small scale acoustic analysis shows that the F0 of hybrid systems correlates less closely to human speech with a higher standard deviation for F0.This analysis suggests that limited variance in F0 is linked to a reduction in human likeness, resulting in lower likeability for conventional synthetic speech methods. Alice Baird, Emilia Parada-Cabaleiro, Simone Hantke, Felix Burkhardt, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 5 |
| 2018 | How Did You like 2017? Detection of Language Markers of Depression and Narcissism in Personal NarrativesabstractLanguage analyses reveals crucial information about an individual's current state of mind.Maladaptive psychological functioning appears in cognition, emotional experience and behaviour.In the time of the internet of things, a vast number of text and speech is available; subsequently, the interest in the automated detection of psychological functioning via language is rising.The current study indicates that depression and narcissism can be predicted through word use in personal narratives.Both conditions are characterised by an altered word count regarding anxiety and we (LIWC-based).While depressive individuals use less social words and more anxietyrelated words, narcissists do the opposite.This might reflect the verbal correlate of the cognitive triad in depression.In contrast, narcissists' word use mirrors their excommunicated anxiety of being an undesired self and their inability to reach long-term goals due to a lack of impulse control.The automated recognition of mental state through word use could improve early detection of mental disease, monitoring of disease course, delivery of tailored interventions and evaluation of therapy outcome. Eva-Maria Rathner, Julia Djamali, Yannik Terhorst, Björn W. Schuller, Nicholas Cummins, Gudrun Salamon, Christina Hunger-Schoppe, Harald Baumeister |
INTERSPEECH | 5 |
| 2018 | State of Mind: Classification through Self-reported Affect and Word Use in SpeechabstractHuman state-of-mind (SOM; e.g.: perception, cognition, attention) constantly shifts due to internal and external demands.Mental health is influenced by the habitual use of either adaptive or maladaptive SOM.Therefore, the training of conscious regulation of SOM could be promising in self-help (e-and m-health), blended care and psychotherapy.The presented study indicates that SOM can be influenced by telling personal narratives.Furthermore, SOM and narrative sentiment (positive vs. negative) can be predicted through word use.Such results lay the groundwork for the development of applications that analyse text and speech for: i) the early detection of mental health; ii) the early detection of maladaptive changes in emotion dynamics; (iii) the use of personal narratives to improve emotion regulation skills; iv) the distribution of tailored interventions; and finally, v) the evaluation of therapy outcome. Eva-Maria Rathner, Yannik Terhorst, Nicholas Cummins, Björn W. Schuller, Harald Baumeister |
INTERSPEECH | 3 |
| 2017 | CAST a database: Rapid targeted large-scale big data acquisition via small-world modelling of social media platformsabstractThe adage that there is no data like more data is not new in affective computing; however, with recent advances in deep learning technologies, such as end-to-end learning, the need for extracting big data is greater than ever. Multimedia resources available on social media represent a wealth of data more than large enough to satisfy this need. However, an often prohibitive amount of effort has been required to source and label such instances. As a solution, we introduce Cost-efficient Audio-visual Acquisition via Social-media Small-world Targeting (CAS2T) for efficient large-scale big data collection from online social media platforms. Our system is based on a unique combination of small-world modelling, unsupervised audio analysis, and semi-supervised active learning. Such an approach facilitates rapid training on entirely new tasks sourced in their entirety from social multimedia. We demonstrate the high capability of our methodology via collection of original datasets containing a range of naturalistic, in-the-wild examples of human behaviours. Shahin Amiriparian, Sergey Pugachevskiy, Nicholas Cummins, Simone Hantke, Jouni Pohjalainen, Gil Keren, Björn W. Schuller |
ACII | 3 |
| 2017 | Enhancing Speech-Based Depression Detection Through Gender Dependent Vowel-Level Formant Features
Nicholas Cummins, Bogdan Vlasenko, Hesam Sagha, Björn W. Schuller |
AIME | 1 |
| 2017 | Stimulation of psychological listener experiences by semi-automatically composed electroacoustic environmentsabstractThis work represents the first steps in an almost completely unexplored field, in which electroacoustic composition, based on several signal processing techniques including additive synthesis, pitch extraction and bandpass filtering, is exploited to produce unique sonic environments for stimulating targeted listener experiences. We propose three semi-automatically composed electroacoustic environments and evaluate their psychological potential considering four areas: creativity, emotion, self-perception and mental associations. This empirical study uses a cross-modal perceptual test, completed by 100 listeners. Results presented indicate that electroacoustic music can successfully evoke specific colour connections and individual self-perception. Additionally, we show that synthesised sound based on bio-signals, such as a heart beat, can promote concrete thought and negative emotional states. Our future goal is to fully-automate electroacoustic music composition environments for use in therapy, education and entertainment to promote and encourage human well-being. Emilia Parada-Cabaleiro, Alice Baird, Nicholas Cummins, Björn W. Schuller |
ICME | 3 |
| 2017 | Snore Sound Classification Using Image-Based Deep Spectrum FeaturesabstractIn this paper, we propose a method for automatically detecting various types of snore sounds using image classification convolutional neural network (CNN) descriptors extracted from audio file spectrograms.The descriptors, denoted as deep spectrum features, are derived from forwarding spectrograms through very deep task-independent pre-trained CNNs.Specifically, activations of fully connected layers from two common image classification CNNs, AlexNet and VGG19, are used as feature vectors.Moreover, we investigate the impact of differing spectrogram colour maps and two CNN architectures on the performance of the system.Results presented indicate that deep spectrum features extracted from the activations of the second fully connected layer of AlexNet using a viridis colour map are well suited to the task.This feature space, when combined with a support vector classifier, outperforms the more conventional knowledge-based features of 6 373 acoustic functionals used in the INTERSPEECH ComParE 2017 Snoring sub-challenge baseline system.In comparison to the baseline, unweighted average recall is increased from 40.6 % to 44.8 % on the development partition, and from 58.5 % to 67.0 % on the test partition. Shahin Amiriparian, Maurice Gerczuk, Sandra Ottl, Nicholas Cummins, Michael Freitag 0003, Sergey Pugachevskiy, Alice Baird, Björn W. Schuller |
INTERSPEECH | 4 |
| 2017 | Automatic Classification of Autistic Child Vocalisations: A Novel Database and ResultsabstractHumanoid robots have in recent years shown great promise for supporting the educational needs of children on the autism spectrum.To further improve the efficacy of such interactions, user-adaptation strategies based on the individual needs of a child are required.In this regard, the proposed study assesses the suitability of a range of speech-based classification approaches for automatic detection of autism severity according to the commonly used Social Responsiveness Scale™ second edition (SRS-2).Autism is characterised by socialisation limitations including child language and communication ability.When compared to neurotypical children of the same age these can be a strong indication of severity.This study introduces a novel dataset of 803 utterances recorded from 14 autistic children aged between 4 -10 years, during Wizard-of-Oz interactions with a humanoid robot.Our results demonstrate the suitability of support vector machines (SVMs) which use acoustic feature sets from multiple Interspeech COMPARE challenges.We also evaluate deep spectrum features, extracted via an image classification convolutional neural network (CNN) from the spectrogram of autistic speech instances.At best, by using SVMs on the acoustic feature sets, we achieved a UAR of 73.7 % for the proposed 3-class task. Alice Baird, Shahin Amiriparian, Nicholas Cummins, Alyssa Alcorn, Anton Batliner, Sergey Pugachevskiy, Michael Freitag 0003, Maurice Gerczuk, Björn W. Schuller |
INTERSPEECH | 3 |
| 2017 | An 'End-to-Evolution' Hybrid Approach for Snore Sound ClassificationabstractWhilst snoring itself is usually not harmful to a person's health, it can be an indication of Obstructive Sleep Apnoea (OSA), a serious sleep-related disorder.As a result, studies into using snoring as acoustic based marker of OSA are gaining in popularity.Motivated by this, the INTERSPEECH 2017 ComParE Snoring sub-challenge requires classification from which areas in the upper airways different snoring sounds originate.This paper explores a hybrid approach combining evolutionary feature selection based on competitive swarm optimisation and deep convolutional neural networks (CNN).Feature selection is applied to novel deep spectrum features extracted directly from spectrograms using pre-trained image classification CNN.Key results presented demonstrate that our hybrid approach can substantially increase the performance of a linear support vector machine on a set of low-level features extracted from the Snoring sub-challenge data.Even without subset selection, the deep spectrum features are sufficient to outperform the challenge baseline, and competitive swarm optimisation further improves system performance.In comparison to the challenge baseline, unweighted average recall is increased from 40.6 % to 57.6 % on the development partition, and from 58.5 % to 66.5 % on the test partition, using 2 246 of the 4 096 deep spectrum features. Michael Freitag 0003, Shahin Amiriparian, Nicholas Cummins, Maurice Gerczuk, Björn W. Schuller |
INTERSPEECH | 3 |
| 2017 | "Did you laugh enough today?" - Deep Neural Networks for Mobile and Wearable Laughter Trackers
Gerhard Hagerer, Nicholas Cummins, Florian Eyben, Björn W. Schuller |
INTERSPEECH | 2 |
| 2017 | Emotional Speech of Mentally and Physically Disabled Individuals: Introducing the EmotAsS Database and First Findings
Simone Hantke, Hesam Sagha, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 3 |
| 2017 | The Perception of Emotions in Noisified Nonsense SpeechabstractNoise pollution is part of our daily life, affecting millions of people, particularly those living in urban environments.Noise alters our perception and decreases our ability to understand others.Considering this, speech perception in background noise has been extensively studied, showing that especially white noise can damage listener perception.However, the perception of emotions in noisified speech has not been explored with as much depth.In the present study, we use artificial background noise conditions, by applying noise to a subset of the GEMEP corpus (emotions expressed in nonsense speech).Noises were at varying intensities and 'colours'; white, pink, and brownian.The categorical and dimensional perceptual test was completed by 26 listeners.The results indicate that background noise conditions influence the perception of emotion in speechpink noise most, brownian least.Worsened perception invokes higher confusion, especially with sadness, an emotion with less pronounced prosodic characteristics.Yet, all this does not lead to a break-down of the 'cognitive-emotional space' in a Nonmetric MultiDimensional Scaling representation.The gender of speakers and the cultural background of listeners do not seem to play a role. Emilia Parada-Cabaleiro, Alice Baird, Anton Batliner, Nicholas Cummins, Simone Hantke, Björn W. Schuller |
INTERSPEECH | 4 |
| 2017 | Earlier Identification of Children with Autism Spectrum Disorder: An Automatic Vocalisation-Based ApproachabstractAutism spectrum disorder (ASD) is a neurodevelopmental disorder usually diagnosed in or beyond toddlerhood.ASD is defined by repetitive and restricted behaviours, and deficits in social communication.The early speech-language development of individuals with ASD has been characterised as delayed.However, little is known about ASD-related characteristics of pre-linguistic vocalisations at the feature level.In this study, we examined pre-linguistic vocalisations of 10-month-old individuals later diagnosed with ASD and a matched control group of typically developing individuals (N = 20).We segmented 684 vocalisations from parent-child interaction recordings.All vocalisations were annotated and signal-analytically decomposed.We analysed ASD-related vocalisation specificities on the basis of a standardised set (eGeMAPS) of 88 acoustic features selected for clinical speech analysis applications.54 features showed evidence for a differentiation between vocalisations of individuals later diagnosed with ASD and controls.In addition, we evaluated the feasibility of automated, vocalisation-based identification of individuals later diagnosed with ASD.We compared linear kernel support vector machines and a 1-layer bidirectional long short-term memory neural network.Both classification approaches achieved an accuracy of 75% for subject-wise identification in a subject-independent 3-fold cross-validation scheme.Our promising results may be an important contribution en-route to facilitate earlier identification of ASD. Florian B. Pokorny, Björn W. Schuller, Peter B. Marschik, Raymond Brueckner, Pär Nyström, Nicholas Cummins, Sven Bölte, Christa Einspieler, Terje Falck-Ytter |
INTERSPEECH | 6 |
| 2017 | Implementing Gender-Dependent Vowel-Level Analysis for Boosting Speech-Based Depression RecognitionabstractLIDIAP Bogdan Vlasenko, Hesam Sagha, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 3 |
| 2017 | An Image-based Deep Spectrum Feature Representation for the Recognition of Emotional SpeechabstractThe outputs of the higher layers of deep pre-trained convolutional neural networks (CNNs) have consistently been shown to provide a rich representation of an image for use in recognition tasks. This study explores the suitability of such an approach for speech-based emotion recognition tasks. First, we detail a new acoustic feature representation, denoted as deep spectrum features, derived from feeding spectrograms through a very deep image classification CNN and forming a feature vector from the activations of the last fully connected layer. We then compare the performance of our novel features with standardised brute-force and bag-of-audio-words (BoAW) acoustic feature representations for 2- and 5-class speech-based emotion recognition in clean, noisy and denoised conditions. The presented results show that image-based approaches are a promising avenue of research for speech-based recognition tasks. Key results indicate that deep-spectrum features are comparable in performance with the other tested acoustic feature representations in matched for noise type train-test conditions; however, the BoAW paradigm is better suited to cross-noise-type train-test conditions. Nicholas Cummins, Shahin Amiriparian, Gerhard Hagerer, Anton Batliner, Stefan Steidl, Björn W. Schuller |
ACM Multimedia | 1 |
| 2017 | Strength modelling for real-worldautomatic continuous affect recognition from audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Nicholas Cummins, Fabien Ringeval, Björn W. Schuller |
Image Vis. Comput. | 3 |
| 2017 | auDeep: Unsupervised Learning of Representations from Audio with Deep Recurrent Neural Networks
Michael Freitag 0003, Shahin Amiriparian, Sergey Pugachevskiy, Nicholas Cummins, Björn W. Schuller |
J. Mach. Learn. Res. | 4 |
| 2017 | A Two-Dimensional Framework of Multiple Kernel Subspace Learning for Recognizing Emotion in SpeechabstractAs a highly active topic in computational paralinguistics, speech emotion recognition (SER) aims to explore ideal representations for emotional factors in speech. In order to improve the performance of SER, multiple kernel learning (MKL) dimensionality reduction has been utilized to obtain effective information for recognizing emotions. However, the solution of MKL usually provides only one nonnegative mapping direction for multiple kernels; this may lead to loss of valuable information. To address this issue, we propose a two-dimensional framework for multiple kernel subspace learning. This framework provides more linear combinations on the basis of MKL without nonnegative constraints, which preserves more information in the learning procedures. It also leverages both of MKL and two-dimensional subspace learning, combining them into a unified structure. To apply the framework to SER, we also propose an algorithm, namely generalised multiple kernel discriminant analysis (GMKDA), by employing discriminant embedding graphs in this framework. GMKDA takes advantage of the additional mapping directions for multiple kernels in the proposed framework. In order to evaluate the performance of the proposed algorithm a wide range of experiments is carried out on several key emotional corpora. These experimental results demonstrate that the proposed methods can achieve better performance compared with some conventional and subspace learning methods in dealing with SER. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | An Investigation of Emotional Speech in Depression Classification
Brian Stasak, Julien Epps, Nicholas Cummins, Roland Göcke |
INTERSPEECH | 3 |
| 2015 | Weighted pairwise Gaussian likelihood regression for depression score predictionabstractThis paper presents a technique in which feature vectors are mapped onto ordinal ranges of clinical depression scores using weighted pairwise Gaussians. The position of a test vector with respect to these partitions is used to perform depression score prediction. Results found on a set of spectral and formant based speech characteristics indicate the potential of this technique for performing depression score prediction. Key results on the AVEC 2013 development set indicate that the inclusion of weights and Bayesian adaptation improves system performance by 16.5% - 18.5% when compared to using an unweighted non-adapted system. Fusing results from Bayesian adapted models corresponding to different feature spaces offers up to 8% further improvement. Further, fusion consistently improves performance on both the AVEC 2013 development and test set, in contrast to conventional regressor fusion. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski |
ICASSP | 1 |
| 2015 | Relevance vector machine for depression prediction
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski |
INTERSPEECH | 1 |
| 2015 | Analysis of acoustic space variability in speech affected by depression
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Sebastian Schnieder, Jarek Krajewski |
Speech Commun. | 1 |
| 2015 | A review of depression and suicide risk assessment using speech analysis
Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, Thomas F. Quatieri |
Speech Commun. | 1 |
| 2014 | Variability compensation in small data: Oversampled extraction of i-vectors for the classification of depressed speechabstractVariations in the acoustic space due to changes in speaker mental state are potentially overshadowed by variability due to speaker identity and phonetic content. Using the Audio/Visual Emotion Challenge and Workshop 2013 Depression Dataset we explore the suitability of i-vectors for reducing these latter sources of variability for distinguishing between low or high levels of speaker depression. In addition we investigate whether supervised variability compensation methods such as Linear Discriminant Analysis (LDA), and Within Class Covariance Normalisation (WCCN), applied in the i-vector domain, could be used to compensate for speaker and phonetic variability. Classification results show that i-vectors formed using an over-sampling methodology outperform a baseline set by KL-means supervectors. However the effect of these two compensation methods does not appear to improve system accuracy. Visualisations afforded by the t-Distributed Stochastic Neighbour Embedding (t-SNE) technique suggest that despite the application of these techniques, speaker variability is still a strong confounding effect. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski |
ICASSP | 1 |
| 2014 | Probabilistic acoustic volume analysis for speech affected by depressionabstractAlterations in speech motor control in depressed individuals have been found to manifest as a reduction in spectral variability. In this paper we present a novel method for measuring acoustic volume a model-based measure that is reflective of this decrease in spectral variability and assess the ability of features resulting from this measure for indexing a speaker’s level of depression. A Monte Carlo approximation that enables the computation of this measure is also outlined in this paper. Results found using the AVEC 2013 Challenge Dataset indicate there is a statistically significant reduction in acoustic variation with increasing levels of speaker depression, and using features designed to capture this change it is possible to outperform a range of conventional spectral measures when predicting a speaker’s level of depression. Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski |
INTERSPEECH | 1 |
| 2013 | Spectro-temporal analysis of speech affected by depression and psychomotor retardationabstractTo enhance current diagnostic methods used when assessing a depressed individual, an objective screening mechanism, ideally based on non-intrusive behavioral signals, is needed. Given the clinical description of depression speech as `dull, monotonous and flat' and promising previous results from spectral features, we hypothesize that the effects of depression on speech are embedded in spectro-temporal events. To test this hypothesis we explore different methodologies, based on the modulation spectrum, for extracting long-term spectro-temporal information from speech and assess their suitability as a clinical marker of depression. Results indicate that: depressive speech information is captured in the modulation spectrum, long-term spectro-temporal information is important in depressed speech identification and there are potential differences in the effects that depression and psychomotor retardation have on speech production mechanisms. Nicholas Cummins, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 1 |
| 2013 | Modeling spectral variability for the classification of depressed speechabstractQuantifying how the spectral content of speech relates to changes in mental state may be crucial in building an objective speech-based depression classification system with clinical utility. This paper investigates the hypothesis that important depression based information can be captured within the covariance structure of a Gaussian Mixture Model (GMM) of recorded speech. Significant negative correlations found between a speaker’s average weighted variance- a GMM-based indicator of speaker variability- and their level of depression support this hypothesis. Further evidence is provided by the comparison of classification accuracies from seven different GMM-UBM systems, each formed by varying different parameter combinations during MAP adaption. This analysis shows that variance-only adaptation either outperforms or matches the de facto standard mean-only adaptation when classifying both the presence and severity of depression. This result is perhaps the first of its kind seen in GMM-UBM speech classification. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Michael Breakspear, Roland Göcke |
INTERSPEECH | 1 |
| 2012 | A Comparison of Classification Paradigms for Speaker Likeability DeterminationabstractIn this paper we investigate the performance of different classification paradigms, testing each with a range of acoustic features, to find a system that is well suited to speaker likeability classification. We introduce a Sparse Representation Classifier for paralinguistic classification and explore the role of training data selection for a GMM classifier. Results demonstrate that (1) Single dimensional features of pitch direction, shimmer and spectral roll-off were the most suitable features found when testing on the development set but we were unable to reproduce their performance in the final classification task, (2) Using UBM training data selection increased accuracy of MFCC's and (3) Sparse Representation showed promise as a paralinguistic classifier with results comparable to that of SVM. Nicholas Cummins, Julien Epps, Jia Min Karen Kua |
INTERSPEECH | 1 |
| 2011 | An Investigation of Depressed Speech Detection: Features and NormalizationabstractIn recent years, the problem of automatic detection of mental illness from the speech signal has gained some initial interest, however questions remaining include how speech segments should be selected, what features provide good discrimination, and what benefits feature normalization might bring given the speaker-specific nature of mental disorders. In this paper, these questions are addressed empirically using classifier configurations employed in emotion recognition from speech, evaluated on a 47-speaker depressed/neutral read sentence speech database. Results demonstrate that (1) detailed spectral features are well suited to the task, (2) speaker normalization provides benefits mainly for less detailed features, and (3) dynamic information appears to provide little benefit. Classification accuracy using a combination of MFCC and formant based features approached 80 % for this database. Index Terms: mental state recognition, depressed speech, feature comparison, MFCC, Gaussian mixture models Nicholas Cummins, Julien Epps, Michael Breakspear, Roland Göcke |
INTERSPEECH | 1 |