Florian Eyben

dblp:97/6615 · DBLP profile ↗
← Back
92ranked-venue papers
16as first author
14since 2021 · last 2025
0009-0003-0330-8545ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 53 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author
YearPublicationVenuePosition
2025 EmoDB 2.0: A Database of Emotional Speech in a World that is not Black or White but Grey
Felix Burkhardt, Oliver Schrüfer, Uwe D. Reichel, Hagen Wierstorf, Anna Derington, Florian Eyben, Björn W. Schuller
INTERSPEECH6
2025 Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
abstract
Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or tone of voice. Those models are mainly evaluated in terms of correlation or recall, and always show some errors in their predictions. The errors manifest themselves in model behaviour, which can be very different along different dimensions even if the same recall or correlation is achieved by the model. This paper introduces a testing framework to investigate behaviour of speech emotion recognition models, by requiring different metrics to reach a certain threshold in order to pass a test. The test metrics can be grouped in terms of correctness, fairness, and robustness. It also provides a method for automatically specifying test thresholds for fairness tests, based on the datasets used, and recommendations on how to select the remaining test thresholds. We evaluated a xLSTM-based and nine transformer-based acoustic foundation models against a convolutional baseline model, testing their performance on arousal, valence, dominance, and emotional category classification. The test results highlight, that models with high correlation or recall might rely on shortcuts – such as text sentiment –, and differ in terms of fairness.
Anna Derington, Hagen Wierstorf, Ali Gürcan Özkil, Florian Eyben, Felix Burkhardt, Björn W. Schuller
IEEE Trans. Affect. Comput.4
2024 A Comparative Analysis of Federated Learning for Speech-Based Cognitive Decline Detection
abstract
Speech-based machine learning models that can distinguish between a healthy cognitive state and different stages of cognitive decline would enable a more appropriate and timely treatment of patients.However, their development is often hampered by data scarcity.Federated Learning (FL) is a potential solution that could enable entities with limited voice recordings to collectively build effective models.Motivated by this, we compare centralised, local, and federated learning for building speechbased models to discern Alzheimer's Disease, Mild Cognitive Impairment, and a healthy state.For a more realistic evaluation, we use three independently collected datasets to simulate healthcare institutions employing these strategies.Our initial analysis shows that FL may not be the best solution in every scenario, as performance improvements are not guaranteed even with small amounts of available data, and further research is needed to determine the conditions under which it is beneficial.
Stefan Kalabakov, Monica González Machorro, Florian Eyben, Björn W. Schuller, Bert Arnrich
INTERSPEECH3
2024 Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
Oliver Schrüfer, Manuel Milling, Felix Burkhardt, Florian Eyben, Björn W. Schuller
INTERSPEECH4
2023 Multimodal Recognition of Valence, Arousal and Dominance via Late-Fusion of Text, Audio and Facial Expressions
abstract
We present an approach for the prediction of valence, arousal, and dominance of people communicating via text/audio/video streams for a translation from and to sign languages.The approach consists of the fusion of the output of three CNN-based models dedicated to the analysis of text, audio, and facial expressions.Our experiments show that any combination of two or three modalities increases prediction performance for valence and arousal.
Fabrizio Nunnari, Annette Rios, Uwe D. Reichel, Chirag Bhuvaneshwara, Panagiotis Paraskevas Filntisis, Petros Maragos, Felix Burkhardt, Florian Eyben, Björn W. Schuller, Sarah Ebling
ESANN8
2023 Masking Speech Contents by Random Splicing: is Emotional Expression Preserved?
abstract
We discuss the influence of random splicing on the perception of emotional expression in speech signals. Random splicing is the randomized reconstruction of short audio snippets with the aim to obfuscate the speech contents. A part of the German parliament recordings has been random spliced and both versions – the original and the scrambled ones – manually labeled with respect to the arousal, valence and dominance dimensions. Additionally, we run a state-of-the-art transformer-based pre-trained emotional model on the data. We find sufficiently high correlation for the annotations and predictions of emotional dimensions between both sample versions to be confident that machine learners can be trained with random spliced data.
Felix Burkhardt, Anna Derington, Matthias Kahlau, Klaus R. Scherer, Florian Eyben, Björn W. Schuller
ICASSP5
2023 Nkululeko: Machine Learning Experiments on Speaker Characteristics Without Programming
Felix Burkhardt, Florian Eyben, Björn W. Schuller
INTERSPEECH2
2023 Towards Supporting an Early Diagnosis of Multiple Sclerosis using Vocal Features
abstract
Multiple sclerosis (MS) is a neuroinflammatory disease that affects millions of people worldwide. Since dysarthria is prominent in people with MS (pwMS), this paper aims to identify acoustic features that differ between people with MS and healthy controls (HC). Additionally, we develop automatic classification methods to distinguish between pwMS and HC. In this work, we present a new dataset of a German-speaking cohort which contains 39 patients with low disability of relapsing MS and 16 HC. Findings suggest that certain interpretable speech features could be useful in diagnosing MS, and that machine learning methods could potentially support fast and unobtrusive screening in clinical practice. The study emphasises the importance of analysing free speech compared to read speech.
Monica González Machorro, Pascal Hecker, Uwe D. Reichel, Helly N. Hammer, Robert Hoepner, Lisa Pedrotti, Alisha Zmutt, Hesam Sagha, Johan van Beek, Florian Eyben, Dagmar Schuller, Björn W. Schuller, Bert Arnrich
INTERSPEECH10
2023 Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence Gap
abstract
Recent advances in transformer-based architectures have shown promise in several machine learning tasks. In the audio domain, such architectures have been successfully utilised in the field of speech emotion recognition (SER). However, existing works have not evaluated the influence of model size and pre-training data on downstream performance, and have shown limited attention to generalisation, robustness, fairness, and efficiency. The present contribution conducts a thorough analysis of these aspects on several pre-trained variants of wav2vec 2.0 and HuBERT that we fine-tuned on the dimensions arousal, dominance, and valence of MSP-Podcast, while additionally using IEMOCAP and MOSI to test cross-corpus generalisation. To the best of our knowledge, we obtain the top performance for valence prediction without use of explicit linguistic information, with a concordance correlation coefficient (CCC) of. 638 on MSP-Podcast. Our investigations reveal that transformer-based architectures are more robust compared to a CNN-based baseline and fair with respect to gender groups, but not towards individual speakers. Finally, we show that their success on valence is based on implicit linguistic information, which explains why they perform on-par with recent multimodal approaches that explicitly utilise textual information. To make our findings reproducible, we release the best performing model to the community.
Johannes Wagner 0001, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, Björn W. Schuller
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Quantifying Cognitive Load from Voice using Transformer-Based Models and a Cross-Dataset Evaluation
abstract
Cognitive load is frequently induced in laboratory setups to measure responses to stress, and its impact on voice has been studied in the field of computational paralinguistics. One dataset on this topic was provided in the Computational Paralinguistics Challenge (ComParE) 2014, and therefore offers great comparability. Recently, transformer-based deep learning architectures established a new state-of-the-art and are finding their way gradually into the audio domain. In this context, we investigate the performance of popular transformer architectures in the audio domain on the ComParE 2014 dataset, and the impact of different pre-training and fine-tuning setups on these models. Further, we recorded a small custom dataset, designed to be comparable with the ComParE 2014 one, to assess cross-corpus model generalisability. We find that the transformer models outperform the challenge baseline, the challenge winner, and more recent deep learning approaches. Models based on the ‘large’ architecture perform well on the task at hand, while models based on the ‘base’ architecture perform at chance level. Fine-tuning on related domains (such as ASR or emotion), before fine-tuning on the targets, yields no higher performance compared to models pre-trained only in a self-supervised manner. The generalisability of the models between datasets is more intricate than expected, as seen in an unexpected low performance on the small custom dataset, and we discuss potential ‘hidden’ underlying discrepancies between the datasets. In summary, transformer-based architectures outperform previous attempts to quantify cognitive load from voice. This is promising, in particular for healthcare-related problems in computational paralinguistics applications, since datasets are sparse in that realm.
Pascal Hecker, Arpita Kappattanavar, Maximilian Schmitt, Sidratul Moontaha, Johannes Wagner 0001, Florian Eyben, Björn W. Schuller, Bert Arnrich
ICMLA6
2022 Probing speech emotion recognition transformers for linguistic knowledge
abstract
Large, pre-trained neural networks consisting of self-attention layers (transformers) have recently achieved state-of-the-art results on several speech emotion recognition (SER) datasets. These models are typically pre-trained in self-supervised manner with the goal to improve automatic speech recognition performance -- and thus, to understand linguistic information. In this work, we investigate the extent in which this information is exploited during SER fine-tuning. Using a reproducible methodology based on open-source tools, we synthesise prosodically neutral speech utterances while varying the sentiment of the text. Valence predictions of the transformer model are very reactive to positive and negative sentiment content, as well as negations, but not to intensifiers or reducers, while none of those linguistic features impact arousal or dominance. These findings show that transformers can successfully leverage linguistic information to improve their valence predictions, and that linguistic analysis should be included in their testing.
Andreas Triantafyllopoulos, Johannes Wagner 0001, Hagen Wierstorf, Maximilian Schmitt, Uwe D. Reichel, Florian Eyben, Felix Burkhardt, Björn W. Schuller
INTERSPEECH6
2022 Nkululeko: A Tool For Rapid Speaker Characteristics Detection
abstract
We present advancements with a software tool called Nkululeko, that lets users perform (semi-) supervised machine learning experiments in the speaker characteristics domain. It is based on audformat, a format for speech database metadata description. Due to an interface based on configurable templates, it supports best practise and very fast setup of experiments without the need to be proficient in the underlying language: Python. The paper explains the handling of Nkululeko and presents two typical experiments: comparing the expert acoustic features with artificial neural net embeddings for emotion classification and speaker age regression.
Felix Burkhardt, Johannes Wagner 0001, Hagen Wierstorf, Florian Eyben, Björn W. Schuller
LREC4
2022 A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine Learning
abstract
Since several decades emotional databases have been recorded by various laboratories. Many of them contain acted portrays of Darwin’s famous “big four” basic emotions. In this paper, we investigate in how far a selection of them are comparable by two approaches: on the one hand modeling similarity as performance in cross database machine learning experiments and on the other by analyzing a manually picked set of four acoustic features that represent different phonetic areas. It is interesting to see in how far specific databases (we added a synthetic one) perform well as a training set for others while some do not. Generally speaking, we found indications for both similarity as well as specificiality across languages.
Felix Burkhardt, Anabell Hacker, Uwe D. Reichel, Hagen Wierstorf, Florian Eyben, Björn W. Schuller
LREC5
2021 Speaking Corona? Human and Machine Recognition of COVID-19 from Voice
abstract
With the COVID-19 pandemic, several research teams have reported successful advances in automated recognition of COVID-19 by voice. Resulting voice-based screening tools for COVID-19 could support large-scale testing efforts. While capabilities of machines on this task are progressing, we approach the so far unexplored aspect whether human raters can distinguish COVID-19 positive and negative tested speakers from voice samples, and compare their performance to a machine learning baseline. To account for the challenging symptom similarity between COVID-19 and other respiratory diseases, we use a carefully balanced dataset of voice samples, in which COVID-19 positive and negative tested speakers are matched by their symptoms alongside COVID-19 negative speakers without symptoms. Both human raters and the machine struggle to reliably identify COVID-19 positive speakers in our dataset. These results indicate that particular attention should be paid to the distribution of symptoms across all speakers of a dataset when assessing the capabilities of existing systems. The identification of acoustic aspects of COVID-19-related symptom manifestations might be the key for a reliable voice-based COVID-19 detection in the future by both trained human raters and machine learning models. Copyright ©2021 ISCA.
Pascal Hecker, Florian B. Pokorny, Katrin D. Bartl-Pokorny, Uwe D. Reichel, Zhao Ren, Simone Hantke, Florian Eyben, Dagmar Schuller, Bert Arnrich, Björn W. Schuller
Interspeech7
2020 Exploiting time-frequency patterns with LSTM-RNNs for low-bitrate audio restoration
Björn W. Schuller, Florian Eyben, Dagmar Schuller, Zixing Zhang 0001, Holly Francois, Eunmi Oh
Neural Comput. Appl.3
2019 Affective and behavioural computing: Lessons learnt from the First Computational Paralinguistics Challenge
Björn W. Schuller, Felix Weninger, Yue Zhang 0014, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli, Klaus R. Scherer, Mohamed Chetouani, Marcello Mortillaro
Comput. Speech Lang.7
2017 Automatic multi-lingual arousal detection from voice applied to real product testing applications
abstract
A method is presented which applies Long Short-Term Memory Recurrent Neural Networks on real market-research voice recordings in order to automatically predict emotional arousal from speech. While most previous work has dealt with evaluations of algorithms within the same speech corpus, the novelty of this paper lies in an extensive evaluation across corpora and languages. The approach is evaluated on seven large data sets collected in real tests of TV commercials and new product concepts across four languages. We observe excellent performance within and between the different corpora when compared against the gold standard of arousal ratings by human annotators. Even in the cross-language validation the models show good performance which almost reaches human rater agreement.
Florian Eyben, Matthias Unfried, Gerhard Hagerer, Björn W. Schuller
ICASSP1
2017 Seeking the SuperStar: Automatic assessment of perceived singing quality
abstract
The quality of the singing voice is an important aspect of subjective, aesthetic perception of music. In this contribution, we propose a method to automatically assess perceived singing quality. We classify monophonic vocal recordings without accompaniment into one of three classes of singing quality. Unprocessed private and non-commercial recordings from a social media website are utilised. In addition to the user ratings given on the website, we let both subjects with and without a musical background annotate the samples. Building on musicological foundations, we define and extract acoustic parameters describing the quality of the sound, musical expression and intonation of the singing. Besides features which are already established in the field of Music Information Retrieval, such as loudness and mel-frequency cepstral coefficients, we propose and employ new types of features which are specific to intonation. For automatic classification by supervised machine learning methods, models predicting the subjective ratings and the user ratings on the social media website are learnt. We perform an exhaustive evaluation of both different classifiers and combinations of features. We show that the performance of automatic classification is close to that of human evaluators. Utilising support vector machines, an accuracy of classification of 55.4 %, based on the subjective ratings, and of 84.7 %, based on the user ratings of the social media website, are achieved.
Johanna Bohm, Florian Eyben, Maximilian Schmitt, Harald Kosch, Björn W. Schuller
IJCNN2
2017 "Did you laugh enough today?" - Deep Neural Networks for Mobile and Wearable Laughter Trackers
Gerhard Hagerer, Nicholas Cummins, Florian Eyben, Björn W. Schuller
INTERSPEECH3
2017 A Paralinguistic Approach To Speaker Diarisation: Using Age, Gender, Voice Likability and Personality Traits
abstract
In this work, we present a new view on automatic speaker diarisation, i.e., assessing "who speaks when", based on the recognition of speaker traits such as age, gender, voice likability, and personality. Traditionally, speaker diarisation is accomplished using low-level audio descriptors (e.g., cepstral or spectral features), neglecting the fact that speakers can be well discriminated by humans according to various perceived characteristics. Thus, we advocate a novel paralinguistic approach that combines speaker diarisation with speaker characterisation by automatically identifying the speakers according to their individual traits. In a three-tier processing flow, speaker segmentation by voice activity detection (VAD) is initially performed to detect speaker turns. Next, speaker attributes are predicted using pre-trained paralinguistic models. To tag the speakers, clustering algorithms are applied to the predicted traits. We evaluate our methods against state-of-the-art open source and commercial systems on a corpus of realistic, spontaneous dyadic conversations recorded in the wild from three different cultures (Chinese, English, German). Our results provide clear evidence that using paralinguistic features for speaker diarisation is a promising avenue of research.
Yue Zhang 0014, Felix Weninger, Boqing Liu, Maximilian Schmitt, Florian Eyben, Björn W. Schuller
ACM Multimedia5
2016 Real-Time Tracking of Speakers' Emotions, States, and Traits on Mobile Platforms
Erik Marchi, Florian Eyben, Gerhard Hagerer, Björn W. Schuller
INTERSPEECH2
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.1
2015 Real-time robust recognition of speakers' emotions and characteristics on mobile platforms
abstract
We demonstrate audEERING's sensAI technology running natively on low-resource mobile devices applied to emotion analytics and speaker characterisation tasks. A showcase application for the Android platform is provided, where au-dEERING's highly noise robust voice activity detection based on Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN) is combined with our core emotion recognition and speaker characterisation engine natively on the mobile device. This eliminates the need for network connectivity and allows to perform robust speaker state and trait recognition efficiently in real-time without network transmission lags. Real-time factors are benchmarked for a popular mobile device to demonstrate the efficiency, and average response times are compared to a server based approach. The output of the emotion analysis is visualized graphically in the arousal and valence space alongside the emotion category and further speaker characteristics.
Florian Eyben, Bernd Huber, Erik Marchi, Dagmar Schuller, Björn W. Schuller
ACII1
2015 iHEARu-PLAY: Introducing a game for crowdsourced data collection for affective computing
abstract
We introduce iHEARu-PLAY, a web-based multi-player game for crowdsourced database collection and — most important — labelling. Existing databases (with speech and video content) can be added to the game and labelling tasks can be defined via a web-interface. The primary purpose of iHEARu-PLAY is multi-label, holistic annotation of multi-modal affective speech databases. Players perform labelling (or prompted recording) tasks and are rewarded with scores and prizes, which are computed based on the “correctness” of their annotations, e.g., the agreement with a pre-defined gold standard or with the other players. iHEARu-PLAY is implemented with the open source high-level Python Web framework Django and can be installed on Unix and Windows platforms. Its modular architecture allows for easy integration of custom extensions: New gaming components can be added as plugins in order to support new databases and modalities. Label categories for each database are individually selectable and editable. Audio, image and video annotation are currently supported. iHEARu-PLAY will be available to the research community as a ready-to-use web-service. Researchers can add their own databases, optionally post rewards, and receive annotation results in the end. General users can register to play the game, have fun, compete with other players, and at the same time support science.
Simone Hantke, Florian Eyben, Tobias Appel, Björn W. Schuller
ACII2
2015 Context-sensitive learning for enhanced audiovisual emotion classification (Extended abstract)
abstract
Human emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs) and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations.
Angeliki Metallinou, Athanasios Katsamanis, Martin Wöllmer, Florian Eyben, Björn W. Schuller, Shri Narayanan
ACII4
2015 Building autonomous sensitive artificial listeners (Extended abstract)
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and non-verbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and non-verbal behaviours required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and non-verbal behaviour, since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on non-verbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling etc. We first summarise three prototype versions of the SAL scenario, in which the behaviour of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analysing and synthesising the respective behaviours. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behaviour, dialogue management, and synthesis of speaker and listener behaviour of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse, and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
ACII4
2015 Cross-corpus acoustic emotion recognition: Variances and strategies (Extended abstract)
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: acted data is often used rather than spontaneous data, results are reported on pre-selected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. A considerably more realistic impression can be gathered by inter-set evaluation: we therefore show results employing six standard databases in a cross-corpora evaluation experiment. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter- to intra-corpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
ACII3
2015 A novel approach for automatic acoustic novelty detection using a denoising autoencoder with bidirectional LSTM neural networks
abstract
Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel unsupervised approach based on a denoising autoencoder. In our approach auditory spectral features are processed by a denoising autoencoder with bidirectional Long Short-Term Memory recurrent neural networks. We use the reconstruction error between the input and the output of the autoencoder as activation signal to detect novel events. The autoencoder is trained on a public database which contains recordings of typical in-home situations such as talking, watching television, playing and eating. The evaluation was performed on more than 260 different abnormal events. We compare results with state-of-the-art methods and we conclude that our novel approach significantly outperforms existing methods by achieving up to 93.4% F-Measure.
Erik Marchi, Fabio Vesperini, Florian Eyben, Stefano Squartini, Björn W. Schuller
ICASSP3
2015 Non-linear prediction with LSTM recurrent neural networks for acoustic novelty detection
abstract
Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel approach based on non-linear predictive denoising autoencoders. In our approach, auditory spectral features of the next short-term frame are predicted from the previous frames by means of Long-Short Term Memory (LSTM) recurrent denoising autoencoders. We show that this yields an effective generative model for audio. The reconstruction error between the input and the output of the autoencoder is used as activation signal to detect novel events. The autoencoder is trained on a public database which contains recordings of typical in-home situations such as talking, watching television, playing and eating. The evaluation was performed on more than 260 different abnormal events. We compare results with state-of-the-art methods and we conclude that our novel approach significantly outperforms existing methods by achieving up to 94.4% F-Measure.
Erik Marchi, Fabio Vesperini, Felix Weninger, Florian Eyben, Stefano Squartini, Björn W. Schuller
IJCNN4
2015 Does my speech rock? automatic assessment of public speaking skills
abstract
In this paper, we introduce results for the task of Automatic Public Speech Assessment (APSA).Given the comparably sparse work carried out on this task up to this point, a novel database was required for training and evaluation of machine learning models.As a basis, the freely available oral presentations of the ICASSP conference in 2011 were selected due to their transcription including non-verbal vocalisations.The data was specifically labelled in terms of the perceived oratory ability of the speakers by five raters according to a 5-point Public Speaking Skill Rating Likert scale.We investigate the feasibility of speaker-independent APSA using different standardised acoustic feature sets computed per fixed chunk of an oral presentation in a series of ternary classification and continuous regression experiments.Further, we compare the relevance of different feature groups related to fluency (speech/hesitation rate), prosody, voice quality and a variety of spectral features.Our results demonstrate that oratory speaking skills can be reliably assessed using suprasegmental audio features, with prosodic ones being particularly suited.
Lucas Azaïs, Adrien Payan, Tianjiao Sun, Guillaume Vidal, Tina Zhang, Eduardo Coutinho, Florian Eyben, Björn W. Schuller
INTERSPEECH7
2015 A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
Comput. Speech Lang.9
2015 Prediction of asynchronous dimensional emotion ratings from audiovisual and physiological data
Fabien Ringeval, Florian Eyben, Eleni Kroupi, Anil Yüce, Jean-Philippe Thiran, Touradj Ebrahimi, Denis Lalanne, Björn W. Schuller
Pattern Recognit. Lett.2
2014 A frequency-weighted post-filtering transform for compensation of the over-smoothing effect in HMM-based speech synthesis
abstract
Over-smoothing is one of the major sources of quality degradation in statistical parametric speech synthesis. Many methods have been proposed to compensate over-smoothing with the speech parameter generation algorithm considering Global Variance (GV) being one of the most successfull. This paper models over-smoothing as a radial relocation of poles and zeros of the spectral envelope towards the origin of the z-plane and uses radial scaling to enhance spectral peaks and to deepen spectral valeys. The radial scaling technique is improved by introducing over-emphasis, spectral-tilt compensation and frequency weighting. Listening test results indicate that the proposed method is 11%-13% more preferable than GV while it has less algorithmic delay (only 5 ms) and computational complexity.
Florian Eyben, Yannis Agiomyrgiannakis
ICASSP1
2014 CCA based feature selection with application to continuous depression recognition from acoustic speech features
abstract
In this study we make use of Canonical Correlation Analysis (CCA) based feature selection for continuous depression recognition from speech. Besides its common use in multi-modal/multi-view feature extraction, CCA can be easily employed as a feature selector. We introduce several novel ways of CCA based filter (ranking) methods, showing their relations to previous work. We test the suitability of proposed methods on the AVEC 2013 dataset under the ACM MM 2013 Challenge protocol. Using 17% of features, we obtained a relative improvement of 30% on the challenge's test-set baseline Root Mean Square Error.
Heysem Kaya, Florian Eyben, Albert Ali Salah, Björn W. Schuller
ICASSP2
2014 Multi-resolution linear prediction based features for audio onset detection with bidirectional LSTM neural networks
abstract
A plethora of different onset detection methods have been proposed in the recent years. However, few attempts have been made with respect to widely-applicable approaches in order to achieve superior performances over different types of music and with considerable temporal precision. In this paper, we present a multi-resolution approach based on discrete wavelet transform and linear prediction filtering that improves time resolution and performance of onset detection in different musical scenarios. In our approach, wavelet coefficients and forward prediction errors are combined with auditory spectral features and then processed by a bidirectional Long Short-Term Memory recurrent neural network, which acts as reduction function. The network is trained with a large database of onset data covering various genres and onset types. We compare results with state-of-the-art methods on a dataset that includes Bello, Glover and ISMIR 2004 Ballroom sets, and we conclude that our approach significantly outperforms existing methods in terms of F-Measure. For pitched non percussive music an absolute improvement of 7.5% is reported.
Erik Marchi, Giacomo Ferroni, Florian Eyben, Leonardo Gabrielli, Stefano Squartini, Björn W. Schuller
ICASSP3
2014 Single-channel speech separation with memory-enhanced recurrent neural networks
abstract
In this paper we propose the use of Long Short-Term Memory recurrent neural networks for speech enhancement. Networks are trained to predict clean speech as well as noise features from noisy speech features, and a magnitude domain soft mask is constructed from these features. Extensive tests are run on 73 k noisy and reverberated utterances from the Audio-Visual Interest Corpus of spontaneous, emotionally colored speech, degraded by several hours of real noise recordings comprising stationary and non-stationary sources and convolutive noise from the Aachen Room Impulse Response database. In the result, the proposed method is shown to provide superior noise reduction at low signal-to-noise ratios while creating very little artifacts at higher signal-to-noise ratios, thereby outperforming unsupervised magnitude domain spectral subtraction by a large margin in terms of source-distortion ratio.
Felix Weninger, Florian Eyben, Björn W. Schuller
ICASSP2
2014 On-line continuous-time music mood regression with deep recurrent neural networks
abstract
This paper proposes a novel machine learning approach for the task of on-line continuous-time music mood regression, i.e., low-latency prediction of the time-varying arousal and valence in musical pieces. On the front-end, a large set of segmental acoustic features is extracted to model short-term variations. Then, multi-variate regression is performed by deep recurrent neural networks to model longer-range context and capture the time-varying emotional profile of musical pieces appropriately. Evaluation is done on the 2013 MediaEval Challenge corpus consisting of 1000 pieces annotated in continous time and continuous arousal and valence by crowd-sourcing. In the result, recurrent neural networks outperform SVR and feedforward neural networks both in continuous-time and static music mood regression, and achieve an R2of up to .70 and .50 with arousal and valence annotations.
Felix Weninger, Florian Eyben, Björn W. Schuller
ICASSP2
2014 MAPTRAITS 2014 - The First Audio/Visual Mapping Personality Traits Challenge - An Introduction: Perceived Personality and Social Dimensions
abstract
The Audio/Visual Mapping Personality Challenge and Workshop (MAPTRAITS) is a competition event that is organised to facilitate the development of signal processing and machine learning techniques for the automatic analysis of personality traits and social dimensions. MAPTRAITS includes two sub-challenges, the continuous space-time sub-challenge and the quantised space-time sub-challenge. The continuous sub-challenge evaluated how systems predict the variation of perceived personality traits and social dimensions in time, whereas the quantised challenge evaluated the ability of systems to predict the overall perceived traits and dimensions in shorter video clips. To analyse the effect of audio and visual modalities on personality perception, we compared systems under three different settings: visual-only, audio-only and audio-visual. With MAPTRAITS we aimed at improving the knowledge on the automatic analysis of personality traits and social dimensions by producing a benchmarking protocol and encouraging the participation of various research groups from different backgrounds.
Oya Çeliktutan, Florian Eyben, Evangelos Sariyanidi, Hatice Gunes, Björn W. Schuller
ICMI2
2014 Emotion Recognition in the Wild: Incorporating Voice and Lip Activity in Multimodal Decision-Level Fusion
abstract
In this paper, we investigate the relevance of using voice and lip activity to improve performance of audiovisual emotion recognition in unconstrained settings, as part of the 2014 Emotion Recognition in the Wild Challenge (EmotiW14). Indeed, the dataset provided by the organisers contains movie excerpts with highly challenging variability in terms of audiovisual content; e.g., speech and/or face of the subject expressing the emotion can be absent in the data. We therefore propose to tackle this issue by incorporating both voice and lip activity as additional features in a decision-level fusion. Results obtained on the blind test set show that the decision-level fusion can improve the best mono-modal approach, and that the addition of both voice and lip activity in the feature set leads to the best performance (UAR=35.27%), with an absolute improvement of 5.36% over the baseline.
Fabien Ringeval, Shahin Amiriparian, Florian Eyben, Klaus R. Scherer, Björn W. Schuller
ICMI3
2014 Audio onset detection: A wavelet packet based approach with recurrent neural networks
abstract
This paper concerns the exploitation of multi-resolution time-frequency features via Wavelet Packet Transform to improve audio onset detection. In our approach, Wavelet Packet Energy Coefficients (WPEC) and Auditory Spectral Features (ASF) are processed by Bidirectional Long Short-Term Memory (BLSTM) recurrent neural network that yields the onsets location. The combination of the two feature sets, together with the BLSTM based detector, form an advanced energy-based approach that takes advantage from the multi-resolution analysis given by the wavelet decomposition of the audio input signal. The neural network is trained with a large database of onset data covering various genres and onset types. Due to its data-driven nature, our approach does not require the onset detection method and its parameters to be tuned to a particular type of music. We show a comparison with other types and sizes of recurrent neural networks and we compare results with state-of-the-art methods on the whole onset dataset. We conclude that our approach significantly increase performance in terms of F-measure without any music genres or onset type constraints.
Erik Marchi, Giacomo Ferroni, Florian Eyben, Stefano Squartini, Björn W. Schuller
IJCNN3
2014 The INTERSPEECH 2014 computational paralinguistics challenge: cognitive & physical load
abstract
The INTERSPEECH 2014 Computational Paralinguistics Challenge provides for the first time a unified test-bed for the automatic recognition of speakers’ cognitive and physical load in speech. In this paper, we describe these two Sub-Challenges, their conditions, baseline results and experimental procedures, as well as the COMPARE baseline features generated with the openSMILE toolkit and provided to the participants in the Challenge.
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julien Epps, Florian Eyben, Fabien Ringeval, Erik Marchi, Yue Zhang 0014
INTERSPEECH5
2014 The Munich Biovoice Corpus: Effects of Physical Exercising, Heart Rate, and Skin Conductance on Human Speech Production
Björn W. Schuller, Felix Friedmann, Florian Eyben
LREC3
2014 Emotional Analysis of Music: A Comparison of Methods
abstract
Music as a form of art is intentionally composed to be emotionally expressive. The emotional features of music are invaluable for music indexing and recommendation. In this paper we present a cross-comparison of automatic emotional analysis of music. We created a public dataset of Creative Commons licensed songs. Using valence and arousal model, the songs were annotated both in terms of the emotions that were expressed by the whole excerpt and dynamically with 1 Hz temporal resolution. Each song received 10 annotations on Amazon Mechanical Turk and the annotations were averaged to form a ground truth. Four different systems from three teams and the organizers were employed to tackle this problem in an open challenge. We compare their performances and discuss the best practices. While the effect of a larger feature set was not very apparent in the static emotion estimation, the combination of a comprehensive feature set and a recurrent neural network that models temporal dependencies has largely outperformed the other proposed methods for dynamic music emotion estimation.
Mohammad Soleymani 0001, Anna Aljanaki, Yi-Hsuan Yang, Michael N. Caro, Florian Eyben, Konstantin Markov, Björn W. Schuller, Remco C. Veltkamp, Felix Weninger, Frans Wiering
ACM Multimedia5
2014 Medium-term speaker states - A review on intoxication, sleepiness and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Florian Schiel, Jarek Krajewski, Felix Weninger, Florian Eyben
Comput. Speech Lang.7
2014 Autoencoder-based Unsupervised Domain Adaptation for Speech Emotion Recognition
abstract
With the availability of speech data obtained from different devices and varied acquisition conditions, we are often faced with scenarios, where the intrinsic discrepancy between the training and the test data has an adverse impact on affective speech analysis. To address this issue, this letter introduces an Adaptive Denoising Autoencoder based on an unsupervised domain adaptation method, where prior knowledge learned from a target set is used to regularize the training on a source set. Our goal is to achieve a matched feature space representation for the target and source sets while ensuring target domain knowledge transfer. The method has been successfully evaluated on the 2009 INTERSPEECH Emotion Challenge's FAU Aibo Emotion Corpus as target corpus and two other publicly available speech emotion corpora as sources. The experimental results show that our method significantly improves over the baseline performance and outperforms related feature domain adaptation methods.
Zixing Zhang 0001, Florian Eyben, Björn W. Schuller
IEEE Signal Process. Lett.3
2013 Real-life voice activity detection with LSTM Recurrent Neural Networks and an application to Hollywood movies
abstract
A novel, data-driven approach to voice activity detection is presented. The approach is based on Long Short-Term Memory Recurrent Neural Networks trained on standard RASTA-PLP frontend features. To approximate real-life scenarios, large amounts of noisy speech instances are mixed by using both read and spontaneous speech from the TIMIT and Buckeye corpora, and adding real long term recordings of diverse noise types. The approach is evaluated on unseen synthetically mixed test data as well as a real-life test set consisting of four full-length Hollywood movies. A frame-wise Equal Error Rate (EER) of 33.2% is obtained for the four movies and an EER of 9.6% is obtained for the synthetic test data at a peak SNR of 0 dB, clearly outperforming three state-of-the-art reference algorithms under the same conditions.
Florian Eyben, Felix Weninger, Stefano Squartini, Björn W. Schuller
ICASSP1
2013 Automatic recognition of physiological parameters in the human voice: Heart rate and skin conductance
abstract
We show that high pulse/low pulse, heart rate and skin conductance recognition can reach good accuracies using classification on a large group of 4k audio features extracted from sustained vowels and breathing periods. A database containing audio, heart rate and skin conductance recordings from 19 subjects is established for evaluation of audio-based bio-signal recognition. On this database in speaker-dependent testing, heart rate and skin conductance can be determined with a correlation coefficient of .861/.960 and mean absolute error of 8.1 BPM/88.2 μMhO for regression based on sustained vowels recorded from a room microphone. Using the same set-up, a high pulse/low pulse classification can reach an unweighted accuracy of 82.7%. The results are largely independent from microphone type and the two bio-signals can be determined from breathing periods as well. Performance does, however, degrade in speaker-independent setting.
Björn W. Schuller, Felix Friedmann, Florian Eyben
ICASSP3
2013 Affect recognition in real-life acoustic conditions - a new perspective on feature selection
abstract
Automatic emotion recognition and computational paralinguistics have matured to some robustness under controlled laboratory settings, however, the accuracies are degraded in real-life conditions such as the presence of noise and reverberation.In this paper we take a look at the relevance of acoustic features for expression of valence, arousal, and interest conveyed by a speaker's voice.Experiments are conducted on the GEMEP and TUM AVIC databases.To simulate realistically degraded conditions the audio is corrupted with real room impulse responses and real-life noise recordings.Features well correlated with the target (emotion) over a wide range of acoustic conditions are analysed and an interpretation is given.Classification results in matched and mismatched settings with multi-condition training are provided to validate the benefit of the feature selection method.Our proposed way of selecting features over a range of noise types considerably boosts the generalisation ability of the classifiers.
Florian Eyben, Felix Weninger, Björn W. Schuller
INTERSPEECH1
2013 Using linguistic information to detect overlapping speech
abstract
Overlapping speech is still a major cause of error in many speech processing applications, currently without any satisfactory solution. This paper considers the problem of detecting segments of overlapping speech within meeting recordings. Using an HMM-based framework recordings are segmented into intervals containing non-speech, speech and overlapping speech. New to this contribution is the use of linguistic information, where spoken content is used to improve overlap detection. Using language models for speech and overlap, an overlap score is created for every spoken word and used as an additional feature within the HMM framework. Experiments conducted on the AMI corpus demonstrate the potential of the proposed linguistic features.
Jürgen T. Geiger, Florian Eyben, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH2
2013 Detecting overlapping speech with long short-term memory recurrent neural networks
abstract
Detecting segments of overlapping speech (when two or more speakers are active at the same time) is a challenging problem.Previously, mostly HMM-based systems have been used for overlap detection, employing various different audio features.In this work, we propose a novel overlap detection system using Long Short-Term Memory (LSTM) recurrent neural networks.LSTMs are used to generate framewise overlap predictions which are applied for overlap detection.Furthermore, a tandem HMM-LSTM system is obtained by adding LSTM predictions to the HMM feature set.Experiments with the AMI corpus show that overlap detection performance of LSTMs is comparable to HMMs.The combination of HMMs and LSTMs improves overlap detection by achieving higher recall.
Jürgen T. Geiger, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH2
2013 The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism
abstract
International audience
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim
INTERSPEECH9
2013 Recent developments in openSMILE, the munich open-source multimedia feature extractor
abstract
We present recent developments in the openSMILE feature extraction toolkit. Version 2.0 now unites feature extraction paradigms from speech, music, and general sound events with basic video features for multi-modal processing. Descriptors from audio and video can be processed jointly in a single framework allowing for time synchronization of parameters, on-line incremental processing as well as off-line and batch processing, and the extraction of statistical functionals (feature summaries), such as moments, peaks, regression parameters, etc. Postprocessing of the features includes statistical classifiers such as support vector machine models or file export for popular toolkits such as Weka or HTK. Available low-level descriptors include popular speech, music and video features including Mel-frequency and similar cepstral and spectral coefficients, Chroma, CENS, auditory model based loudness, voice quality, local binary pattern, color, and optical flow histograms. Besides, voice activity detection, pitch tracking and face detection are supported. openSMILE is implemented in C++, using standard open source libraries for on-line audio and video input. It is fast, runs on Unix and Windows platforms, and has a modular, component based architecture which makes extensions via plug-ins easy. openSMILE 2.0 is distributed under a research license and can be downloaded from http://opensmile.sourceforge.net/.
Florian Eyben, Felix Weninger, Florian Groß, Björn W. Schuller
ACM Multimedia1
2013 LSTM-Modeling of continuous emotions in an audiovisual affect recognition framework
Martin Wöllmer, Moritz Kaiser, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
Image Vis. Comput.3
2012 Unsupervised clustering of emotion and voice styles for expressive TTS
abstract
Current text-to-speech synthesis (TTS) systems are often perceived as lacking expressiveness, limiting the ability to fully convey information. This paper describes initial investigations into improving expressiveness for statistical speech synthesis systems. Rather than using hand-crafted definitions of expressive classes, an unsupervised clustering approach is described which is scalable to large quantities of training data. To incorporate this “expression cluster” information into an HMM-TTS system two approaches are described: cluster questions in the decision tree construction; and average expression speech synthesis (AESS) using cluster-based linear transform adaptation. The performance of the approaches was evaluated on audiobook data in which the reader exhibits a wide range of expressiveness. A subjective listening test showed that synthesising with AESS results in speech that better reflects the expressiveness of human speech than a baseline expression-independent system.
Florian Eyben, Sabine Buchholz, Norbert Braunschweiler, Javier Latorre, Vincent Wan, Mark J. F. Gales, Kate M. Knill
ICASSP1
2012 Audiovisual vocal outburst classification in noisy acoustic conditions
abstract
In this study, we investigate an audiovisual approach for classification of vocal outbursts (non-linguistic vocalisations) in noisy conditions using Long Short-Term Memory (LSTM) Recurrent Neural Networks and Support Vector Machines. Fusion of geometric shape features and acoustic low-level descriptors is performed on the feature level. Three different types of acoustic noise are considered: babble, office and street noise. Experiments are conducted on every noise type to asses the benefit of the fusion in each case. As database for evaluations serves the INTERSPEECH 2010 Paralinguistic Challenge's Audiovisual Interest Corpus of human-to-human natural conversation. The results show that even when training is performed on noise corrupted audio which matches the test conditions the addition of visual features is still beneficial.
Florian Eyben, Stavros Petridis, Björn W. Schuller, Maja Pantic
ICASSP1
2012 Robust feature extraction for automatic recognition of vibrato singing in recorded polyphonic music
abstract
We address the robustness of features for fully automatic recognition of vibrato, which is usually defined as a periodic oscillation of the pitch (F0) of the singing voice, in recorded polyphonic music. Using an evaluation database covering jazz, pop and opera music, we show that the extraction of pitch is challenging in the presence of instrumental accompaniment, leading to unsatisfactory classification accuracy (61.1 %) if only the F0 frequency spectrum is used as features. To alleviate, we investigate alternative functionals of F0, alternative low-level features besides F0, and extraction of vocals by monaural source separation. Finally, we propose to use inter-quartile ranges of F0 delta regression coefficients as features which are highly robust against pitch extraction errors, reaching up to 86.9% accuracy in real-life conditions without any signal enhancement.
Felix Weninger, Noam Amir, Ofer Amir, Irit Ronen, Florian Eyben, Björn W. Schuller
ICASSP5
2012 Improving generalisation and robustness of acoustic affect recognition
abstract
Emotion recognition in real-life conditions faces several challenging factors, which most studies on emotion recognition do not consider. Such factors include background noise, varying recording levels, and acoustic properties of the environment, for example. This paper presents a systematic evaluation of the influence of background noise of various types and SNRs, as well as recording level variations on the performance of automatic emotion recognition from speech. Both, natural and spontaneous as well as acted/prototypical emotions are considered. Besides the well known influence of additive noise, a significant influence of the recording level on the recognition performance is observed. Multi-condition learning with various noise types and recording levels is proposed as a way to increase robustness of methods based on standard acoustic feature sets and commonly used classifiers. It is compared to matched conditions learning and is found to be almost on par for many settings.
Florian Eyben, Björn W. Schuller, Gerhard Rigoll
ICMI1
2012 Preserving actual dynamic trend of emotion in dimensional speech emotion recognition
abstract
In this paper, we use the concept of dynamic trend of emotion to describe how a human's emotion changes over time, which is believed to be important for understanding one's stance toward current topic in interactions. However, the importance of this concept - to our best knowledge - has not been paid enough attention before in the field of speech emotion recognition (SER). Inspired by this, this paper aims to evoke researchers' attention on this concept and makes a primary effort on the research of predicting correct dynamic trend of emotion in the process of SER. Specifically, we propose a novel algorithm named Order Preserving Network (OPNet) to this end. First, as the key issue for OPNet construction, we propose employing a probabilistic method to define an emotion trend-sensitive loss function. Then, a nonlinear neural network is trained using the gradient descent as optimization algorithm to minimize the constructed loss function. We validated the prediction performance of OPNet on the VAM corpus, by mean linear error as well as a rank correlation coefficient γ as measures. Comparing to k-Nearest Neighbor and support vector regression, the proposed OPNet performs better on the preservation of actual dynamic trend of emotion.
Wenjing Han, Haifeng Li 0001, Florian Eyben, Lin Ma 0003, Jiayin Sun, Björn W. Schuller
ICMI3
2012 AVEC 2012: the continuous audio/visual emotion challenge
abstract
We present the second Audio-Visual Emotion recognition Challenge and workshop (AVEC 2012), which aims to bring together researchers from the audio and video analysis communities around the topic of emotion recognition. The goal of the challenge is to recognise four continuously valued affective dimensions: arousal, expectancy, power, and valence. There are two sub-challenges: in the Fully Continuous Sub-Challenge participants have to predict the values of the four dimensions at every moment during the recordings, while for the Word-Level Sub-Challenge a single prediction has to be given per word uttered by the user. This paper presents the challenge guidelines, the common data used, and the performance of the baseline system on the two tasks.
Björn W. Schuller, Michel F. Valstar, Florian Eyben, Roddy Cowie, Maja Pantic
ICMI3
2012 The INTERSPEECH 2012 Speaker Trait Challenge
abstract
LIDIAP
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
INTERSPEECH9
2012 Temporal and Situational Context Modeling for Improved Dominance Recognition in Meetings
abstract
We present and evaluate a novel approach towards automatically detecting a speaker's level of dominance in a meeting scenario.Since previous studies reveal that audio appears to be the most important modality for dominance recognition, we focus on the analysis of the speech signals recorded in multiparty meetings.Unlike recently published techniques which concentrate on frame-level hidden Markov modeling, we propose a recognition framework operating on segmental data and investigate context modeling on three different levels to explore possible performance gains.First, we apply a set of statistical functionals to capture large-scale feature-level context within a speech segment.Second, we consider bidirectional Long Short-Term Memory recurrent neural networks for long-range temporal context modeling between segments.Finally, we evaluate the benefit of situational context incorporation by simultaneously modeling speech of all meeting participants.Overall, our approach leads to a remarkable increase of recognition accuracy when compared to hidden Markov modeling.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH2
2012 Context-Sensitive Learning for Enhanced Audiovisual Emotion Classification
abstract
Human emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore, the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs), and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations. Context-sensitive approaches outperform those without context for classification tasks such as discrimination between valence levels or between clusters in the valence-activation space. The analysis of emotional transitions in our database sheds light into the flow of affective expressions, revealing potentially useful patterns.
Angeliki Metallinou, Martin Wöllmer, Athanasios Katsamanis, Florian Eyben, Björn W. Schuller, Shri Narayanan
IEEE Trans. Affect. Comput.4
2012 Building Autonomous Sensitive Artificial Listeners
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and nonverbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and nonverbal behaviors required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and nonverbal behavior since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on nonverbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling, etc. We first report on three prototype versions of the SAL scenario in which the behavior of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analyzing and synthesizing the respective behaviors. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behavior, dialogue management, and synthesis of speaker and listener behavior of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
IEEE Trans. Affect. Comput.4
2012 A multitask approach to continuous five-dimensional affect sensing in natural speech
abstract
Automatic affect recognition is important for the ability of future technical systems to interact with us socially in an intelligent way by understanding our current affective state. In recent years there has been a shift in the field of affect recognition from “in the lab” experiments with acted data to “in the wild” experiments with spontaneous and naturalistic data. Two major issues thereby are the proper segmentation of the input and adequate description and modeling of affective states. The first issue is crucial for responsive, real-time systems such as virtual agents and robots, where the latency of the analysis must be as small as possible. To address this issue we introduce a novel method of incremental segmentation to be used in combination with supra-segmental modeling. For modeling of continuous affective states we use Long Short-Term Memory Recurrent Neural Networks, with which we can show an improvement in performance over standard recurrent neural networks and feed-forward neural networks as well as Support Vector Regression. For experiments we use the SEMAINE database, which contains recordings of spontaneous and natural human to Wizard-of-Oz conversations. The recordings are annotated continuously in time and magnitude with FeelTrace for five affective dimensions, namely activation, expectation, intensity, power/dominance, and valence. To exploit dependencies between the five affective dimensions we investigate multitask learning of all five dimensions augmented with inter-rater standard deviation. We can show improvements for multitask over single-task modeling. Correlation coefficients of up to 0.81 are obtained for the activation dimension and up to 0.58 for the valence dimension. The performance for the remaining dimensions were found to be in between that for activation and valence.
Florian Eyben, Martin Wöllmer, Björn W. Schuller
ACM Trans. Interact. Intell. Syst.1
2011 AVEC 2011-The First International Audio/Visual Emotion Challenge
Björn W. Schuller, Michel F. Valstar, Florian Eyben, Gary McKeown, Roddy Cowie, Maja Pantic
ACII (2)3
2011 String-based audiovisual fusion of behavioural events for the assessment of dimensional affect
abstract
The automatic assessment of affect is mostly based on feature-level approaches, such as distances between facial points or prosodic and spectral information when it comes to audiovisual analysis. However, it is known and intuitive that behavioural events such as smiles, head shakes or laughter and sighs also bear highly relevant information regarding a subject's affective display. Accordingly, we propose a novel string-based prediction approach to fuse such events and to predict human affect in a continuous dimensional space. Extensive analysis and evaluation has been conducted using the newly released SEMAINE database of human-to-agent communication. For a thorough understanding of the obtained results, we provide additional benchmarks by more conventional feature-level modelling, and compare these and the string-based approach to fusion of signal-based features and string-based events. Our experimental results show that the proposed string-based approach is the best performing approach for automatic prediction of Valence and Expectation dimensions, and improves prediction performance for the other dimensions when combined with at least acoustic signal-based features.
Florian Eyben, Martin Wöllmer, Michel F. Valstar, Hatice Gunes, Björn W. Schuller, Maja Pantic
FG1
2011 Come and have an emotional workout with sensitive artificial listeners!
abstract
This demonstration aims to showcase the recently completed SEMAINE system. The SEMAINE system is a publicly available, fully autonomous Sensitive Artificial Listeners (SAL) system that consists of virtual dialog partners based on audiovisual analysis and synthesis (see http://semaine.opendfki.de/wiki). The system runs in real-time, and combines incremental analysis of user behavior, dialog management, and synthesis of speaker and listener behavior of a SAL character, displayed as a virtual agent. The SAL characters intend to engage the user in a conversation by paying attention to the user's emotions and nonverbal expressions. The characters have their own emotionally defined personality. During an interaction, the characters attempt to create an emotional workout for the user by drawing her/him towards their dominant emotion, through a combination of verbal and nonverbal expressions.
Marc Schröder 0001, Sathish Pammi, Hatice Gunes, Maja Pantic, Michel F. Valstar, Roddy Cowie, Gary McKeown, Dirk Heylen, Mark ter Maat, Florian Eyben, Björn W. Schuller, Martin Wöllmer, Elisabetta Bevacqua, Catherine Pelachaud, Etienne de Sevin
FG10
2011 Audiovisual classification of vocal outbursts in human conversation using Long-Short-Term Memory networks
abstract
We investigate classification of non-linguistic vocalisations with a novel audiovisual approach and Long Short-Term Memory (LSTM) Recurrent Neural Networks as highly successful dynamic sequence classifiers. As database of evaluation serves this year's Paralinguistic Challenge's Audiovisual Interest Corpus of human-to-human natural conversation. For video-based analysis we compare shape and appearance based features. These are fused in an early manner with typical audio descriptors. The results show significant improvements of LSTM networks over a static approach based on Support Vector Machines. More important, we can show a significant gain in performance when fusing audio and visual shape features.
Florian Eyben, Stavros Petridis, Björn W. Schuller, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic
ICASSP1
2011 Syllabification of conversational speech using Bidirectional Long-Short-Term Memory Neural Networks
abstract
Segmentation of speech signals is a crucial task in many types of speech analysis. We present a novel approach at segmentation on a syllable level, using a Bidirectional Long-Short-Term Memory Neural Network. It performs estimation of syllable nucleus positions based on regression of perceptually motivated input features to a smooth target function. Peak selection is performed to attain valid nuclei positions. Performance of the model is evaluated on the levels of both syllables and the vowel segments making up the syllable nuclei. The general applicability of the approach is illustrated by good results for two common databases-Switchboard and TIMIT-for both read and spontaneous speech, and a favourable comparison with other published results.
Christian Landsiedel, Jens Edlund, Florian Eyben, Daniel Neiberg, Björn W. Schuller
ICASSP3
2011 Deep neural networks for acoustic emotion recognition: Raising the benchmarks
abstract
Deep Neural Networks (DNNs) denote multilayer artificial neural networks with more than one hidden layer and millions of free parameters. We propose a Generalized Discriminant Analysis (GerDA) based on DNNs to learn discriminative features of low dimension optimized with respect to a fast classification from a large set of acoustic features for emotion recognition. On nine frequently used emotional speech corpora, we compare the performance of GerDA features and their subsequent linear classification with previously reported benchmarks obtained using the same set of acoustic features classified by Support Vector Machines (SVMs). Our results impressively show that low-dimensional GerDA features capture hidden information from the acoustic features leading to a significantly raised unweighted average recall and considerably raised weighted average recall.
André Stuhlsatz, Christine Meyer, Florian Eyben, Thomas Zielke, Hans-Günter Meier, Björn W. Schuller
ICASSP3
2011 Combining monaural source separation with Long Short-Term Memory for increased robustness in vocalist gender recognition
abstract
We present a novel and unique combination of algorithms to detect the gender of the leading vocalist in recorded popular music. Building on our previous successful approach that enhanced the harmonic parts by means of Non-Negative Matrix Factorization (NMF) for increased accuracy, we integrate on the one hand a new source separation algorithm specifically tailored to extracting the leading voice from monaural recordings. On the other hand, we introduce Bidirectional Long Short-Term Memory Recurrent Neural Networks (BLSTM-RNNs) as context-sensitive classifiers for this scenario, which have lately led to great success in Music Information Retrieval tasks. Through a combination of leading voice separation and BLSTM networks, as opposed to a baseline approach using Hidden Naive Bayes on the original recordings, the accuracy of simultaneous detection of vocal presence and vocalist gender on beat level is improved by up to 10% absolute. Furthermore, using this technique we achieve 91.6% accuracy in determining the gender of the predominant vocalist on song level, which is 4% absolute above our previous best result.
Felix Weninger, Jean-Louis Durrieu, Florian Eyben, Gaël Richard, Björn W. Schuller
ICASSP3
2011 A multi-stream ASR framework for BLSTM modeling of conversational speech
abstract
We propose a novel multi-stream framework for continuous conversational speech recognition which employs bidirectional Long Short-Term Memory (BLSTM) networks for phoneme prediction. The BLSTM architecture allows recurrent neural nets to model long range context, which led to improved ASR performance when combined with conventional triphone modeling in a Tandem system. In this paper, we extend the principle of joint BLSTM and triphone modeling to a multi-stream system which uses MFCC features and BLSTM predictions as observations originating from two independent data streams. Using the COSINE database, we show that this technique prevails over a recently proposed single-stream Tandem system as well as over a conventional HMM recognizer.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
ICASSP2
2011 Acoustic-Linguistic Recognition of Interest in Speech with Bottleneck-BLSTM Nets
abstract
This paper proposes a novel technique for speech-based interest recognition in natural conversations. We introduce a fully automatic system that exploits the principle of bidirectional Long Short-Term Memory (BLSTM) as well as the structure of socalled bottleneck networks. BLSTM nets are able to model a self-learned amount of context information, which was shown to be beneficial for affect recognition applications, while bottleneck networks allow for efficient feature compression within neural networks. In addition to acoustic features, our technique considers linguistic information obtained from a multi-stream BLSTM-HMM speech recognizer. Evaluations on the TUM AVIC corpus reveal that the bottleneck-BLSTM method prevails over all approaches that have been proposed for the Interspeech 2010 Paralinguistic Challenge task. Index Terms: affective computing, interest recognition, recurrent neural networks
Martin Wöllmer, Felix Weninger, Florian Eyben, Björn W. Schuller
INTERSPEECH3
2010 Late fusion of individual engines for improved recognition of negative emotion in speech - learning vs. democratic vote
abstract
The fusion of multiple recognition engines is known to be able to outperform individual ones, given sufficient independence of methods, models, and knowledge sources. We therefore investigate late fusion of different speech-based recognizers of emotion. Two generally different streams of information are considered: acoustics and linguistics fed by state-of-the-art automatic speech recognition. A total of five emotion recognition engines from different sites that provide heterogeneous output information are integrated by either simple democratic vote or learning `which predictor to trust when'. We are able to significantly outperform the best individual engine by fusion, and the so far best reported result on the recently introduced Emotion Challenge task.
Björn W. Schuller, Florian Metze, Stefan Steidl, Anton Batliner, Florian Eyben, Tim Polzehl
ICASSP5
2010 Spoken term detection with Connectionist Temporal Classification: A novel hybrid CTC-DBN decoder
abstract
This paper proposes a novel system for robust keyword detection in continuous speech. Our decoder is composed of a bidirectional Long Short-Term Memory recurrent neural network using a Connectionist Temporal Classification (CTC) output layer, and a Dynamic Bayesian Network (DBN). The CTC network exploits bidirectional context information to reliably identify phonemes, whereas the DBN is able to discriminate between keywords and arbitrary speech while explicitly modeling substitutions, deletions, and insertions in the CTC phoneme output string. Our technique is vocabulary independent and does not require an explicit garbage model. Experiments show that our system architecture prevails over a standard Hidden Markov Model approach.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
ICASSP2
2010 Emotion recognition using imperfect speech recognition
abstract
This paper investigates the use of speech-to-text methods for assigning an emotion class to a given speech utterance. Previous work shows that an emotion extracted from text can convey complementary evidence to the information extracted by classifiers based on spectral, or other non-linguistic features. As speech-to-text usually presents significantly more computational effort, in this study we investigate the degree of speech-to-text accuracy needed for reliable detection of emotions from an automatically generated transcription of an utterance. We evaluate the use of hypotheses in both training and testing, and compare several classification approaches on the same task. Our results show that emotion recognition performance stays roughly constant as long as word accuracy doesn't fall below a reasonable value, making the use of speech-to-text viable for training of emotion classifiers based on linguistics.
Florian Metze, Anton Batliner, Florian Eyben, Tim Polzehl, Björn W. Schuller, Stefan Steidl
INTERSPEECH3
2010 Recognition of spontaneous conversational speech using long short-term memory phoneme predictions
abstract
We present a novel continuous speech recognition framework designed to unite the principles of triphone and Long Short-Term Memory (LSTM) modeling. The LSTM principle allows a recurrent neural network to store and to retrieve information over long time periods, which was shown to be well-suited for the modeling of co-articulation effects in human speech. Our system uses a bidirectional LSTM network to generate a phoneme prediction feature that is observed by a triphone-based large-vocabulary continuous speech recognition (LVCSR) decoder, together with conventional MFCC features. We evaluate both, phoneme prediction error rates of various network architectures and the word recognition performance of our Tandem approach using the COSINE database- a large corpus of conversational and noisy speech, and show that incorporating LSTM phoneme predictions in to an LVCSR system leads to significantly higher word accuracies.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH2
2010 Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional LSTM modeling
abstract
In this paper, we apply a context-sensitive technique for multimodal emotion recognition based on feature-level fusion of acoustic and visual cues. We use bidirectional Long Short-Term Memory (BLSTM) networks which, unlike most other emotion recognition approaches, exploit long-range contextual information for modeling the evolution of emotion within a conversation. We focus on recognizing dimensional emotional labels, which enables us to classify both prototypical and nonprototypical emotional expressions contained in a large audiovisual database. Subject-independent experiments on various classification tasks reveal that the BLSTM network approach generally prevails over standard classification techniques such as Hidden Markov Models or Support Vector Machines, and achieves F1-measures of the order of 72 %, 65 %, and 55 % for the discrimination of three clusters in emotional space and the distinction between three levels of valence and activation, respectively. Index Terms: emotion recognition, multimodality, long shortterm memory, hidden markov models, context modeling
Martin Wöllmer, Angeliki Metallinou, Florian Eyben, Björn W. Schuller, Shri Narayanan
INTERSPEECH3
2010 Long short-term memory networks for noise robust speech recognition
abstract
In this paper we introduce a novel hybrid model architecture for speech recognition and investigate its noise robustness on the Aurora 2 database.Our model is composed of a bidirectional Long Short-Term Memory (BLSTM) recurrent neural net exploiting long-range context information for phoneme prediction and a Dynamic Bayesian Network (DBN) for decoding.The DBN is able to learn pronunciation variants as well as typical phoneme confusions of the BLSTM predictor in order to compensate signal disturbances.Unlike conventional Hidden Markov Model (HMM) systems, the proposed architecture is not based on Gaussian mixture modeling.Even without any feature enhancement, our BLSTM-DBN system outperforms a baseline HMM recognizer by up to 18 %.
Martin Wöllmer, Florian Eyben, Björn W. Schuller
INTERSPEECH3
2010 Opensmile: the munich versatile and fast open-source audio feature extractor
abstract
We introduce the openSMILE feature extraction toolkit, which unites feature extraction algorithms from the speech processing and the Music Information Retrieval communities. Audio low-level descriptors such as CHROMA and CENS features, loudness, Mel-frequency cepstral coefficients, perceptual linear predictive cepstral coefficients, linear predictive coefficients, line spectral frequencies, fundamental frequency, and formant frequencies are supported. Delta regression and various statistical functionals can be applied to the low-level descriptors. openSMILE is implemented in C++ with no third-party dependencies for the core functionality. It is fast, runs on Unix and Windows platforms, and has a modular, component based architecture which makes extensions via plug-ins easy. It supports on-line incremental processing for all implemented features as well as off-line and batch processing. Numeric compatibility with future versions is ensured by means of unit tests. openSMILE can be downloaded from http://opensmile.sourceforge.net/.
Florian Eyben, Martin Wöllmer, Björn W. Schuller
ACM Multimedia1
2010 Cross-Corpus Acoustic Emotion Recognition: Variances and Strategies
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: Acted data is often used rather than spontaneous data, results are reported on preselected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. Even speaker disjunctive evaluation can give only a little insight into the generalization ability of today's emotion recognition engines since training and test data used for system development usually tend to be similar as far as recording conditions, noise overlay, language, and types of emotions are concerned. A considerably more realistic impression can be gathered by interset evaluation: We therefore show results employing six standard databases in a cross-corpora evaluation experiment which could also be helpful for learning about chances to add resources for training and overcoming the typical sparseness in the field. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter to intracorpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
IEEE Trans. Affect. Comput.3
2009 From speech to letters - using a novel neural network architecture for grapheme based ASR
abstract
Main-stream automatic speech recognition systems are based on modelling acoustic sub-word units such as phonemes. Phonemisation dictionaries and language model based decoding techniques are applied to transform the phoneme hypothesis into orthographic transcriptions. Direct modelling of graphemes as sub-word units using HMM has not been successful. We investigate a novel ASR approach using Bidirectional Long Short-Term Memory Recurrent Neural Networks and Connectionist Temporal Classification, which is capable of transcribing graphemes directly and yields results highly competitive with phoneme transcription. In design of such a grapheme based speech recognition system phonemisation dictionaries are no longer required. All that is needed is text transcribed on the sentence level, which greatly simplifies the training procedure. The novel approach is evaluated extensively on the Wall Street Journal 1 corpus.
Florian Eyben, Martin Wöllmer, Björn W. Schuller, Alex Graves
ASRU1
2009 Acoustic emotion recognition: A benchmark comparison of performances
abstract
In the light of the first challenge on emotion recognition from speech we provide the largest-to-date benchmark comparison under equal conditions on nine standard corpora in the field using the two pre-dominant paradigms: modeling on a frame-level by means of hidden Markov models and supra-segmental modeling by systematic feature brute-forcing. Investigated corpora are the ABC, AVIC, DES, EMO-DB, eNTERFACE, SAL, SmartKom, SUSAS, and VAM databases. To provide better comparability among sets, we additionally cluster each database's emotions into binary valence and arousal discrimination tasks. In the result large differences are found among corpora that mostly stem from naturalistic emotions and spontaneous speech vs. more prototypical events. Further, supra-segmental modeling proves significantly beneficial on average when several classes are addressed at a time.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Gerhard Rigoll, Andreas Wendemuth
ASRU3
2009 Robust vocabulary independent keyword spotting with graphical models
abstract
This paper introduces a novel graphical model architecture for robust and vocabulary independent keyword spotting which does not require the training of an explicit garbage model. We show how a graphical model structure for phoneme recognition can be extended to a keyword spotter that is robust with respect to phoneme recognition errors. We use a hidden garbage variable together with the concept of switching parents to model keywords as well as arbitrary speech. This implies that keywords can be added to the vocabulary without having to re-train the model. Thereby the design of our model architecture is optimised to reliably detect keywords rather than to decode keyword phoneme sequences as arbitrary speech, while offering a parameter to adjust the operating point on the receiver operating characteristics curve. Experiments on the TIMIT corpus reveal that our graphical model outperforms a comparable hidden Markov model based keyword spotter that uses conventional garbage modelling.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
ASRU2
2009 Robust discriminative keyword spotting for emotionally colored spontaneous speech using bidirectional LSTM networks
abstract
In this paper we propose a new technique for robust keyword spotting that uses bidirectional long short-term memory (BLSTM) recurrent neural nets to incorporate contextual information in speech decoding. Our approach overcomes the drawbacks of generative HMM modeling by applying a discriminative learning procedure that non-linearly maps speech features into an abstract vector space. By incorporating the outputs of a BLSTM network into the speech features, it is able to make use of past and future context for phoneme predictions. The robustness of the approach is evaluated on a keyword spotting task using the HUMAINE sensitive artificial listener (SAL) database, which contains accented, spontaneous, and emotionally colored speech. The test is particularly stringent because the system is not trained on the SAL database, but only on the TIMIT corpus of read speech. We show that our method prevails over a discriminative keyword spotter without BLSTM-enhanced feature functions, which in turn has been proven to outperform HMM-based techniques.
Martin Wöllmer, Florian Eyben, Joseph Keshet, Alex Graves, Björn W. Schuller, Gerhard Rigoll
ICASSP2
2009 Data-driven clustering in emotional space for affect recognition using discriminatively trained LSTM networks
abstract
In today's affective databases speech turns are often labelled on a continuous scale for emotional dimensions such as valence or arousal to better express the diversity of human affect.However, applications like virtual agents usually map the detected emotional user state to rough classes in order to reduce the multiplicity of emotion dependent system responses.Since these classes often do not optimally reflect emotions that typically occur in a given application, this paper investigates data-driven clustering of emotional space to find class divisions that better match the training data and the area of application.Thereby we consider the Belfast Sensitive Artificial Listener database and TV talkshow data from the VAM corpus.We show that a discriminatively trained Long Short-Term Memory (LSTM) recurrent neural net that explicitly learns clusters in emotional space and additionally models context information outperforms both, Support Vector Machines and a Regression-LSTM net.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Ellen Douglas-Cowie, Roddy Cowie
INTERSPEECH2
2009 Robust in-car spelling recognition - a tandem BLSTM-HMM approach
abstract
As an intuitive hands-free input modality automatic spelling recognition is especially useful for in-car human-machine interfaces.However, for today's speech recognition engines it is extremely challenging to cope with similar sounding spelling speech sequences in the presence of noises such as the driving noise inside a car.Thus, we propose a novel Tandem spelling recogniser, combining a Hidden Markov Model (HMM) with a discriminatively trained bidirectional Long Short-Term Memory (BLSTM) recurrent neural net.The BLSTM network captures long-range temporal dependencies to learn the properties of in-car noise, which makes the Tandem BLSTM-HMM robust with respect to speech signal disturbances at extremely low signal-to-noise ratios and mismatches between training and test noise conditions.Experiments considering various driving conditions reveal that our Tandem recogniser outperforms a conventional HMM by up to 33%.
Martin Wöllmer, Florian Eyben, Björn W. Schuller, Tobias Moosmayr, Nhu Nguyen-Thien
INTERSPEECH2
2009 A multidimensional dynamic time warping algorithm for efficient multimodal fusion of asynchronous data streams
Martin Wöllmer, Marc A. Al-Hames, Florian Eyben, Björn W. Schuller, Gerhard Rigoll
Neurocomputing3
2009 Being bored? Recognising natural interest by extensive audiovisual integration for real-life application
Björn W. Schuller, Ronald Müller, Florian Eyben, Jürgen Gast, Benedikt Hörnler, Martin Wöllmer, Gerhard Rigoll, Anja Höthker, Hitoshi Konosu
Image Vis. Comput.3
2008 Abandoning emotion classes - towards continuous emotion recognition with modelling of long-range dependencies
abstract
Class based emotion recognition from speech, as performed in most works up to now, entails many restrictions for practical applications. Human emotion is a continuum and an automatic emotion recognition system must be able to recognise it as such. We present a novel approach for continuous emotion recognition based on Long Short-Term Memory Recurrent Neural Networks which include modelling of long-range dependencies between observations and thus outperform techniques like Support-Vector Regression. Transferring the innovative concept of additionally modelling emotional history to the classification of discrete levels for the emotional dimensions “valence ” and “activation ” we also apply Conditional Random Fields which prevail over the commonly used Support-Vector Machines. Experiments conducted on data that was recorded while humans interacted with a Sensitive Artificial Listener prove that for activation the derived classifiers perform as well as human annotators.
Martin Wöllmer, Florian Eyben, Stephan Reiter, Björn W. Schuller, Cate Cox, Ellen Douglas-Cowie, Roddy Cowie
INTERSPEECH2
2007 Fast and Robust Meter and Tempo Recognition for the Automatic Discrimination of Ballroom Dance Styles
abstract
Fast and robust recognition of a song's meter, and quarter note tempo is crucial in many music information retrieval tasks dealing especially with large databases or real-time musical stream processing. We therefore introduce a novel approach that is capable of extracting musical meter features and tempo in beats per minute. The method is extendable in order to return the locations of beat onsets suitable for example for beat synchronization or musical audio segmentation. We use a simplified psychoacoustic model to split the input into audible frequency bands and two phase comb filtering on those bands to find the quarter note tempo and metrical structure. Based on these features we discriminate the nine classic ballroom dance styles and duple or triple meter by support-vector-machines as exemplary application. Test-runs are carried out on a public ballroom dance music database containing 1.8 k titles and the public MTV-Europe Most Wanted 1981-2000 to demonstrate the high effectiveness for popular music with respect to meter, tempo and ballroom dance style recognition.
Björn W. Schuller, Florian Eyben, Gerhard Rigoll
ICASSP (1)2
2007 Wearable Assistance for the Ballroom-Dance Hobbyist - Holistic Rhythm Analysis and Dance-Style Classification
abstract
Automated retrieval of high level information from ballroom dance music is challenging, but has many practical applications. These include, for example, a fully automatic ballroom dance D.J., robots capable of performing ballroom dances, or wearable dance-assistance, as considered herein. It is necessary, for such a system, to retrieve information about the song's quarter note tempo, meter and beat positions. Further, the system must be able to discriminate between the nine standard and Latin ballroom dances. In this paper we present a model that combines all these requirements in one holistic approach. The polyphonic input is processed by a simplified psychoacoustic model. Tatum, tempo and meter features are extracted using resonant filters. The filter output is used for beat tracking. The extracted features are used for a ballroom dance-style classification by support-vector-machines. To show the high effectiveness regarding dance-style recognition and beat tracking, test-runs are carried out on a database containing 1.8k titles.
Florian Eyben, Björn W. Schuller, Stephan Reiter, Gerhard Rigoll
ICME1