Bogdan Vlasenko

dblp:78/2841 · DBLP profile ↗
← Back
31ranked-venue papers
16as first author
6since 2021 · last 2025
0000-0003-2248-6200ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 9 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode Prediction
Bogdan Vlasenko, Mathew Magimai-Doss
INTERSPEECH1
2024 Comparing data-Driven and Handcrafted Features for Dimensional Emotion Recognition
abstract
Speech Emotion Recognition (SER) has garnered significant attention over the past two decades. In the early stages of SER technology, ’brute force’-based techniques led to a significant expansion in knowledge-based acoustic feature representation (FR) for modeling sparse emotional data. However, as deep learning techniques have become more powerful, their direct application has been limited by the scarcity of well-annotated emotional data. As a result, pretrained neural embeddings on large speech corpora have gained popularity for SER tasks. These embeddings leverage existing transfer learning methods suitable for general-purpose self-supervised learning (SSL) representations. Recent studies on downstream SSL techniques for dimensional SER have shown promising results. In this research, we aim to evaluate the emotion-discriminative characteristics of neural embeddings in general cases (out-of-domain) and when fine-tuned for SER (in-domain). Given that most SSL techniques are pre-trained primarily on English speech, we plan to use speech emotion corpora in both language-matched and mismatched conditions. We will assess the discriminative characteristics of both handcrafted and standalone neural embeddings as FRs.
Bogdan Vlasenko, Sargam Vyas, Mathew Magimai-Doss
ICASSP1
2024 Cross-transfer Knowledge between Speech and Text Encoders to Evaluate Customer Satisfaction
L. Felipe Parra-Gallego, Tilak Purohit, Bogdan Vlasenko, Juan Rafael Orozco-Arroyave, Mathew Magimai-Doss
INTERSPEECH3
2023 Towards Learning Emotion Information from Short Segments of Speech
abstract
Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this approach has been successful, there is a growing interest in modelling speech emotion information at the short segment level, at around 250ms-500ms (e.g. the 2021-22 MuSe Challenges). This paper investigates both hand-crafted feature-based and end-to-end raw waveform DNN approaches for modelling speech emotion information in such short segments. Through experimental studies on IEMOCAP corpus, we demonstrate that the end-to-end raw waveform modelling approach is more effective than using hand-crafted features for short-segment level modelling. Furthermore, through relevance signal-based analysis of the trained neural networks, we observe that the top performing end-to-end approach tends to emphasize cepstral information instead of spectral information (such as flux and harmonicity).
Tilak Purohit, Sarthak Yadav, Bogdan Vlasenko, S. Pavankumar Dubagunta, Mathew Magimai-Doss
ICASSP3
2023 Implicit phonetic information modeling for speech emotion recognition
Tilak Purohit, Bogdan Vlasenko, Mathew Magimai-Doss
INTERSPEECH2
2022 Modeling of Pre-Trained Neural Network Embeddings Learned From Raw Waveform for COVID-19 Infection Detection
abstract
COVID-19 is a respiratory system disorder that can disrupt the function of lungs. Effects of dysfunctional respiratory mechanism can reflect upon other modalities which function in close coupling. Audio signals result from modulation of respiration through speech production system, and hence acoustic information can be modeled for detection of COVID-19. In that direction, this paper is addressing the second DiCOVA challenge that deals with COVID-19 detection based on speech, cough and breathing. We investigate modeling of (a) ComParE LLD representations derived at frame- and turn-level resolutions and (b) neural representations obtained from pre-trained neural networks trained to recognize phones and estimate breathing patterns. On Track 1, the ComParE LLD representations yield a best performance of 78.05% area under the curve (AUC). Experimental studies on Track 2 and Track 3 demonstrate that neural representations tend to yield better detection than ComParE LLD representations. Late fusion of different utterance level representations of neural embeddings yielded a best performance of 80.64% AUC.
Zohreh Mostaani, RaviShankar Prasad, Bogdan Vlasenko, Mathew Magimai-Doss
ICASSP3
2019 Learning Voice Source Related Information for Depression Detection
abstract
During depression neurophysiological changes can occur, which may affect laryngeal control i.e. behaviour of the vocal folds. Characterising these changes in a precise manner from speech signals is a non trivial task, as this typically involves reliable separation of the voice source information from them. In this paper, by exploiting the abilities of CNNs to learn task-relevant information from the input raw signals, we investigate several methods to model voice source related information for depression detection. Specifically, we investigate modelling of low pass filtered speech signals, linear prediction residual signals, homomorphically filtered voice source signals and zero frequency filtered signals to learn voice source related information for depression detection. Our investigations show that subsegmental level modelling of linear prediction residual signals or zero frequency filtered signals leads to systems better than the state-of-the-art low level descriptor based systems and deep learning based systems modelling the vocal tract system information.
S. Pavankumar Dubagunta, Bogdan Vlasenko, Mathew Magimai-Doss
ICASSP2
2018 Implementing Fusion Techniques for the Classification of Paralinguistic Information
abstract
This work tests several classification techniques and acoustic features and further combines them using late fusion to classify paralinguistic information for the ComParE 2018 challenge. We use Multiple Linear Regression (MLR) with Ordinary Least Squares (OLS) analysis to select the most informative features for Self-Assessed Affect (SSA) sub-Challenge. We also propose to use raw-waveform convolutional neural networks (CNN) in the context of three paralinguistic sub-challenges. By using combined evaluation split for estimating codebook, we obtain better representation for Bag-of-Audio-Words approach. We preprocess the speech to vocalized segments to improve classification performance. For fusion of our leading classification techniques, we use weighted late fusion approach applied for confidence scores. We use two mismatched evaluation phases by exchanging the training and development sets, and this estimates the optimal fusion weight. Weighted late fusion provides better performance on development sets in comparison with baseline techniques. Raw-waveform techniques perform comparable to the baseline.
Bogdan Vlasenko, Jilt Sebastian, Pavan Kumar D. S., Mathew Magimai-Doss
INTERSPEECH1
2017 Enhancing Speech-Based Depression Detection Through Gender Dependent Vowel-Level Formant Features
Nicholas Cummins, Bogdan Vlasenko, Hesam Sagha, Björn W. Schuller
AIME2
2017 Implementing Gender-Dependent Vowel-Level Analysis for Boosting Speech-Based Depression Recognition
abstract
LIDIAP
Bogdan Vlasenko, Hesam Sagha, Nicholas Cummins, Björn W. Schuller
INTERSPEECH1
2015 Cross-corpus acoustic emotion recognition: Variances and strategies (Extended abstract)
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: acted data is often used rather than spontaneous data, results are reported on pre-selected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. A considerably more realistic impression can be gathered by inter-set evaluation: we therefore show results employing six standard databases in a cross-corpora evaluation experiment. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter- to intra-corpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
ACII2
2015 Exploring dataset similarities using PCA-based feature selection
abstract
In emotion recognition from speech, several well-established corpora are used to date for the development of classification engines. The data is annotated differently, and the community in the field uses a variety of feature extraction schemes. The aim of this paper is to investigate promising features for individual corpora and then compare the results for proposing optimal features across data sets, introducing a new ranking method. Further, this enables us to present a method for automatic identification of groups of corpora with similar characteristics. This answers an urgent question in classifier development, namely whether data from different corpora is similar enough to jointly be used as training material, overcoming shortage of material in matching domains. We compare the results of this method with manual groupings of corpora. We consider the established emotional speech corpora AVIC, ABC, DES, EMO-DB, ENTERFACE, SAL, SMARTKOM, SUSAS and VAM, however our approach is general.
Ingo Siegert, Ronald Böck, Andreas Wendemuth, Bogdan Vlasenko
ACII4
2015 Annotators' agreement and spontaneous emotion classification performance
abstract
The combination of various types of data can significantly in-crease the amount of emotional material for training of more reliable real-life emotion classifiers. There are two well-known schemes of annotation utilized for emotional speech: multi-dimensional and categories-based. Multi-dimensional anno-tation is usually applied for labeling spontaneous emotional events, and categorial-based annotation is used for specifica-tion of the acted ”full blown ” emotional chunks. In order to simulate real-life conditions we used a cross-corpora evalua-tion strategy for datasets with different schemes of emotional annotation. Emotional models were trained on acted material from the EMO-DB (categories based annotation) dataset and evaluated on spontaneous data from the VAM dataset (multi-dimensional annotation). The best emotion classification per-formance was obtained on real-life emotional instances with the most intense arousal labels provided by a majority voting strategy (out of 17 annotators). We find that the correspond-ing spontaneous speech samples containing the most intensive emotional content are comparable with acted instances. The im-portance of employing a larger number of emotional annotators was finally addressed in our article. Index Terms: emotion recognition, cross-corpora evaluation, phoneme-level emotional models, turn-level emotional models, emotional intensity 1.
Bogdan Vlasenko, Andreas Wendemuth
INTERSPEECH1
2014 Location of an emotionally neutral region in valence-arousal space: Two-class vs. three-class cross corpora emotion recognition evaluations
abstract
There are two main emotion annotation techniques: multidimensional and categories based. In order to conduct experiments on emotional data annotated with different techniques, two-classes emotion mapping strategies (e.g. high-vs. low-arousal) are commonly used. The ”affective computing” community could not specify a location of emotionally neutral area in multi-dimensional emotional space (e.g. valence-arousal-dominance (VAD)). Nonetheless, in the current research a neutral state is added to the standard two-classes emotion classification task. Within experiments a possible location of a neutral arousal region in valence-arousal space was determined. We employed general and phonetic pattern dependent emotion classification techniques for cross-corpora experiments. Emotional models were trained on the VAM dataset (multi-dimensional annotation) and evaluated them on the EMO-DB dataset (categories based annotation).
Bogdan Vlasenko, Andreas Wendemuth
ICME1
2014 Modeling phonetic pattern variability in favor of the creation of robust emotion classifiers for real-life applications
Bogdan Vlasenko, Dmytro Prylipko, Ronald Böck, Andreas Wendemuth
Comput. Speech Lang.1
2013 Parameter Optimization Issues for Cross-corpora Emotion Classification
abstract
As speech based emotion recognition has matured to a degree where it becomes applicable within real-life conditions, it is time for a realistic view on obtainable performances. Most state-of-the-art emotion recognition methods are based on turn- and frame-level analysis independent of phonetic transcription. True speaker disjoint partitioning of training and test sets is still less common than simple cross-validation. Even speaker disjoint experiments can give only little insight into the generalization ability of modern emotion recognition engines since training and test sets used for system development usually tend to be similar as far as acoustic channel, noise overlay, and language are concerned. A considerably more realistic impression can be gathered by cross-corpora evaluation. Tuning of the emotion classification engine (feature set optimization and normalization, selection of a classification technique and corresponding parameter configuration) is an important issue of realistic evaluations. In the ideal case, an optimal classifier configuration estimated on training data should provide an outstanding recognition performance on unseen data. We therefore compare cross-corpora classification performances of optimized and non-optimized general and phonetic-pattern dependent classifiers.
Bogdan Vlasenko, David Philippou-Hübner, Andreas Wendemuth
ACII1
2013 Determining the Smallest Emotional Unit for Level of Arousal Classification
abstract
Most state-of-the-art emotion recognition methods are based on turn- and frame-level analysis independent from phonetic transcription. Currently "affective computing" community could not specify the smallest emotional standard unit which can be easily classified and determined by any "advanced" and "non-advanced" listener. It is known that, acoustic modeling on the smallest phonetic unit (phoneme) started a new era in automatic speech recognition: switch from speaker dependent isolated word recognition to speaker independent continuous speech recognition. In or current research we showed that phoneme can be used as as smallest unit for high and low arousal emotion classification task. We trained our classifications models on the VAM dataset material and evaluated them on speech samples from the DES dataset. For our experiments we employed two different emotion classification approaches: general (phonetic pattern independent) and phoneme-based (phonetic pattern dependent). Both classification approaches used MFFC features extracted on the frame level. Our experimental results impressively show that the proposed phoneme-based classification technique could increase emotion classification performance by about 9.68% absolute (15.98% relative). We showed that phoneme-level emotion models trained on "natural" emotions could provide impressive classification performance on dataset with acted affective content.
Bogdan Vlasenko, Andreas Wendemuth
ACII1
2012 The Performance of the Speaking Rate Parameter in Emotion Recognition from Speech
abstract
The speaking rate is a quite obvious prosodic characteristic of speech and humans can easily estimate how fast an interlocutor is talking. Further, different emotional dispositions of a person are strongly expressed in his/her speaking rate. In this paper we investigate the performance gain originating from the use of the speaking rate parameter in emotion recognition from speech. The speaking rates are determined by applying a broad phonetic class recognizer. The classifier is trained on cepstral features extracted on the emotionally neutral RM1 speech corpus and provides low average recognition errors of one phoneme/second. We present the results of an empirical approach on the emotionally expressive Emo-DB corpus applying a neural network classifier and prove the significant influence of the speaking rate in emotion classification. The performances of Multi-Layer Perceptrons trained on cepstral turn-level features are analyzed with respect to the presence and absence of the speaking rate feature. An increase of accuracy up to 3.7% in certain emotion categories is reported.
David Philippou-Hübner, Bogdan Vlasenko, Ronald Böck, Andreas Wendemuth
ICME2
2011 Appropriate emotional labelling of non-acted speech using basic emotions, geneva emotion wheel and self assessment manikins
abstract
In emotion recognition from speech, a good transcription and annotation of given material is crucial. Moreover, the question of how to find good emotional labels for new data material is a basic issue. It is not only the question of which emotion labels to choose, it is also a matter of how labellers can cope with annotation methods. In this paper, we present our investigations for emotional labelling with three different methods (Basic Emotions, Geneva Emotion Wheel and Self Assessment Manikins) and compare them in terms of emotion coverage and usability. We show that emotion labels derived from Geneva Emotion Wheel or Self Assessment Manikins fulfill our requirements, but Basic Emotions are not feasible for emotion labelling from spontaneous speech.
Ingo Siegert, Ronald Böck, Bogdan Vlasenko, David Philippou-Hübner, Andreas Wendemuth
ICME3
2011 Vowels formants analysis allows straightforward detection of high arousal emotions
abstract
Recently, automatic emotion recognition from speech has achieved growing interest within the human-machine interaction research community. Most part of emotion recognition methods use context independent frame-level analysis or turn-level analysis. In this article, we introduce context dependent vowel level analysis applied for emotion classification. An average first formant value extracted on vowel level has been used as unidimensional acoustic feature vector. The Neyman-Pearson criterion has been used for classification purpose. Our classifier is able to detect high-arousal emotions with small error rates. Within our research we proved that the smallest emotional unit should be the vowel instead of the word. We find out that using vowel level analysis can be an important issue during developing a robust emotion classifier. Also, our research can be useful for developing robust affective speech recognition methods and high quality emotional speech synthesis systems.
Bogdan Vlasenko, David Philippou-Hübner, Dmytro Prylipko, Ronald Böck, Ingo Siegert, Andreas Wendemuth
ICME1
2011 Vowels Formants Analysis Allows Straightforward Detection of High Arousal Acted and Spontaneous Emotions
Bogdan Vlasenko, Dmytro Prylipko, David Philippou-Hübner, Andreas Wendemuth
INTERSPEECH1
2010 Determining optimal features for emotion recognition from speech by applying an evolutionary algorithm
abstract
Abstract The automated recognition of emotions from speech is a chal-lenging issue. In order to build an emotion recognizer well de-fined features and optimized parameter sets are essential. Thispaper will show how an optimal parameter set for HMM-basedrecognizerscanbefoundbyapplyinganevolutionaryalgorithmon standard features in automated speech recognition. For this,we compared different signal features, as well as several archi-tectures of HMMs. The system was evaluated on a non-acteddatabase and its performance was compared to a baseline sys-tem. We present an optimal feature set for the public part of theSmartKom database.Index Terms: Emotion Recognition, Evolutionary Algorithms,Feature Optimization, Hidden-Markov Models 1. Introduction The interaction between men and machines using language isnowadays becoming more and more self-evident, but machinesstill lack of many human abilities which would considerablysimplify communication and would also help to increase theacceptance of such systems. For some time, research activitiesalso focus stronger on the emotional aspect of speech. Exploit-ing information about the emotional state of a user, machinescan be enabled to adapt their dialog strategy online, depend-ing on the user’s emotions and hence react in a more appro-priate and empathic manner. As emotion recognition in manyapplications goes hand in hand with automated speech recog-nition (ASR) it would be favorable to make use of the samefeatures or even a subset thereof. Especially small devices likesmart phones or PDAs, which do not provide huge computa-tionalpowerwouldbenefitfromsuchasparsefeatureapproach.Parallel research of other groups bases on pooling together(high level) features including the application of brute forcemethods in order to fully exploit the feature space (compare[1]). This paper however describes an evolutionary strategy(ES) of finding an optimal sparse feature set, given not morethan the common acoustic features used in ASR. Applying anES has the advantage of self adaptation of its parameters andis able to find optimal parameter constellations in high dimen-sional search spaces and further gives insights into the rele-vance of each parameter in terms of the model’s accuracy. Es-pecially in case of many parameters with unknown relation-ships ES quickly avoids wasting time on generating and test-ingunsuitableparametercombinationsastheevolutionaryforceminimizes the probability of the evolvement of such combina-tions effectively. In ASR Mel-Frequency-Cepstral-Coefficients(MFCCs) have established as a basic feature in order to trainphoneme based recognizers. As we do not want additional pa-rameters to be extracted from the speech signal we concentrateonly on MFCCs, which have also proven to perform well inemotion recognition during the Emotion Challenge within In-terspeech 2009 (compare [2]).Thispaperisstructuredasfollows: InSection2wedescribethespontaneous database and the emotions we want to recognize.Section 3 describes, which parameters we investigated and inwhich range they were allowed to change during the evolutionprocess. Section4introducestheevolutionaryalgorithm,showshow fitness is measured and the population evolves over time.TheresultsarepresentedinSection5andcomparedtoourbase-line recognizer. Finally Section 6 summarizes our findings andgives an outlook.
David Philippou-Hübner, Bogdan Vlasenko, Tobias Grosser, Andreas Wendemuth
INTERSPEECH2
2010 Cross-Corpus Acoustic Emotion Recognition: Variances and Strategies
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: Acted data is often used rather than spontaneous data, results are reported on preselected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. Even speaker disjunctive evaluation can give only a little insight into the generalization ability of today's emotion recognition engines since training and test data used for system development usually tend to be similar as far as recording conditions, noise overlay, language, and types of emotions are concerned. A considerably more realistic impression can be gathered by interset evaluation: We therefore show results employing six standard databases in a cross-corpora evaluation experiment which could also be helpful for learning about chances to add resources for training and overcoming the typical sparseness in the field. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter to intracorpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
IEEE Trans. Affect. Comput.2
2009 Acoustic emotion recognition: A benchmark comparison of performances
abstract
In the light of the first challenge on emotion recognition from speech we provide the largest-to-date benchmark comparison under equal conditions on nine standard corpora in the field using the two pre-dominant paradigms: modeling on a frame-level by means of hidden Markov models and supra-segmental modeling by systematic feature brute-forcing. Investigated corpora are the ABC, AVIC, DES, EMO-DB, eNTERFACE, SAL, SmartKom, SUSAS, and VAM databases. To provide better comparability among sets, we additionally cluster each database's emotions into binary valence and arousal discrimination tasks. In the result large differences are found among corpora that mostly stem from naturalistic emotions and spontaneous speech vs. more prototypical events. Further, supra-segmental modeling proves significantly beneficial on average when several classes are addressed at a time.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Gerhard Rigoll, Andreas Wendemuth
ASRU2
2009 Heading toward to the natural way of human-machine interaction: the nimitek project
abstract
Spoken human-machine interaction supported by state-of-the-art dialog systems is becoming a standard technology. A lot of effort was invested for this kind of artificial communication interface. But still the spoken dialog systems (SDS) are not able to provide for the user a natural way of communication. Because, existing automated dialog system do not dedicate enough attention to problems in the interaction related to affected user behavior. This paper addresses some aspects of design and implementation of user behavior models in dialog systems aimed to provide naturalness of human-machine interaction. We discuss a viable integration technique of speech based emotion classification in SDS for robust affected automatic speech recognition and user emotion correlated dialog strategy. First of all, we describe existing methods of emotion recognition within speech and affected speech adapted ASR methods. Second, we introduce an approach to achieve emotion adaptive dialog management in human-machine interaction. A multimodal human-machine interaction system with integrated user behavior model is created within the project ldquoNeurobiologically Inspired, Multimodal Intention Recognition for Technical Communication Systemsrdquo (NIMITEK). Currently NIMITEK provides a technical demonstrator to study these principles in a dedicated prototypical task, namely solving the game Towers of Hanoi. In this paper, we will describe the general approach NIMITEK takes to emotional man-machine interactions.
Bogdan Vlasenko, Andreas Wendemuth
ICME1
2009 Processing affected speech within human machine interaction
abstract
Spoken dialog systems (SDS) integrated into human-machine interaction interfaces is becoming a standard technology. Current state-of-the-art SDS, usually, is not able to provide for the user a natural way of communication. Existing automated dialog systems do not dedicate enough attention to problems in the interaction related to affected user behavior. As a result, Automatic Speech Recognition (ASR) engines are not able to recognize affected speech and dialog strategy does not make use of the user’s emotional state. This paper addresses some aspects of processing affected speech withinnatural human-machine interaction. First of all, we propose an affected speech adapted ASR engine. Second, we describe our methods of emotion recognition within speech and present our results of emotion classification within Interspeech 2009 Emotion Challenge. Third, we test affected speech adapted speech recognition models and introduce an approach to achieve emotion adaptive dialog management in human-machine interaction.
Bogdan Vlasenko, Andreas Wendemuth
INTERSPEECH1
2008 Combining speech recognition and acoustic word emotion models for robust text-independent emotion recognition
abstract
Recognition of emotion in speech usually uses acoustic models that ignore the spoken content. Likewise one general model per emotion is trained independent of the phonetic structure. Given sufficient data, this approach seemingly works well enough. Yet, this paper tries to answer the question whether acoustic emotion recognition strongly depends on phonetic content, and if models tailored for the spoken unit can lead to higher accuracies. We therefore investigate phoneme-, and word-models by use of a large prosodic, spectral, and voice quality feature space and Support Vector Machines (SVM). Experiments also take the necessity of ASR into account to select appropriate unit- models. Test-runs on the well-known EMO-DB database facing speaker-independence demonstrate superiority of word emotion models over today's common general models provided sufficient occurrences in the training corpus.
Björn W. Schuller, Bogdan Vlasenko, Dejan Arsic, Gerhard Rigoll, Andreas Wendemuth
ICME2
2008 Balancing spoken content adaptation and unit length in the recognition of emotion and interest
abstract
Recognition and detection of non-lexical or paralinguistic cues from speech usually uses one general model per event (emotional state, level of interest).Commonly this model is trained independent of the phonetic structure.Given sufficient data, this approach seemingly works well enough.Yet, this paper addresses the question on which phonetic level there is the onset of emotions and level of interest.We therefore compare phoneme-, word-and sentence-level analysis for emotional sentence classification by use of a large prosodic, spectral, and voice quality feature space for SVM and MFCC for HMM/GMM.Experiments also take the necessity of ASR into account to select appropriate unit-models.In experiments on the well-known public EMO-DB database, and the SUSAS and AVIC spontaneous interest corpora, we found that the emotion recognition by sentence level analysis shows the best results.We discuss the implications of these types of analysis on the design of robust emotion and interest recognition of usable human-machine interfaces (HMI).
Bogdan Vlasenko, Björn W. Schuller, Kinfe Tadesse Mengistu, Gerhard Rigoll, Andreas Wendemuth
INTERSPEECH1
2007 Frame vs. Turn-Level: Emotion Recognition from Speech Considering Static and Dynamic Processing
Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll
ACII1
2007 Comparing one and two-stage acoustic modeling in the recognition of emotion in speech
abstract
In the search for a standard unit for use in recognition of emotion in speech, a whole turn, that is the full section of speech by one person in a conversation, is common. Within applications such turns often seem favorable. Yet, high effectiveness of sub-turn entities is known. In this respect a two-stage approach is investigated to provide higher temporal resolution by chunking of speech-turns according to acoustic properties, and multi-instance learning for turn-mapping after individual chunk analysis. For chunking fast pre-segmentation into emotionally quasi-stationary segments by one-pass Viterbi beam search with token passing basing on MFCC is used. Chunk analysis is realized by brute-force large feature space construction with subsequent subset selection, SVM classification, and speaker normalization. Extensive tests reveal differences compared to one-stage processing. Alternatively, syllables are used for chunking.
Björn W. Schuller, Bogdan Vlasenko, Ricardo Minguez, Gerhard Rigoll, Andreas Wendemuth
ASRU2
2007 Combining frame and turn-level information for robust recognition of emotions within speech
abstract
Current approaches to the recognition of emotion within speech usually use statistic feature information obtained by application of functionals on turn- or chunk levels. Yet, it is well known that thereby important information on temporal sub-layers as the frame-level is lost. We therefore investigate the benefits of integration of such information within turn-level feature space. For frame-level analysis we use GMM for classification and 39 MFCC and energy features with CMS. In a subsequent step output scores are fed forward into a 1.4k large-feature-space turn-level SVM emotion recognition engine. Thereby we use a variety of Low-Level-Descriptors and functionals to cover prosodic, speech quality, and articulatory aspects. Extensive testruns are carried out on the public databases EMO-DB and SUSAS. Speaker-independent analysis is faced by speaker normalization. Overall results highly emphasize the benefits of feature integration on diverse time scales.
Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll
INTERSPEECH1