Andreas Wendemuth

dblp:89/3836 · DBLP profile ↗
← Back
60ranked-venue papers
11as first author
4since 2021 · last 2026
0000-0001-6917-8198ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-authorHuman-computer interaction and ubiquitous computing · 12 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorTheory of computation · 2 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Intent Recognition in Speech-to-Text Processing in the Context of Natural Interaction with Cognitive Assistive Systems
Behnam Ensan, Magnus Jung, Matthias Busch, Andreas Wendemuth
LREC4
2026 A Poisson limit for the number of sub-matrices of random binary matrices satisfying the majority rule
abstract
We consider m × n random binary matrices, m = m ( n ) . For arbitrary odd integer k we investigate the asymptotic distribution of the random number of sub-matrices of size k × n for which the number of ones in every column satisfies the majority rule. We discuss possible impacts of our result and give examples of applications.
Italo Simonelli, Andreas Wendemuth
Discret. Appl. Math.2
2023 Multiparty Dialogic Processes of Goal and Strategy Formation in Hybrid Teams
Andreas Wendemuth, Stefan Kopp
CHIRA (1)1
2022 On Emotions as Features for Speech Overlaps Classification
abstract
Although being a frequently occurring phenomenon in spoken communication, speech overlaps did not obtain the deserved attention in research so far—in both Human-Human Interaction (HHI) and Human-Computer Interaction (HCI). It is common knowledge that overlaps can figure as a competitive, rude interruption as well as a cooperative, convenient feedback signal giving important insight on the course of the interaction—but how are they related to the internal state of the overlapping speaker or the overlapped speaker? In this paper, we investigate dyadic human-human interactions and focus on the relations between the emotional changes occurring around overlaps in both interaction participants. Further to an in-depth statistical analysis of the changes in control and valence levels with respect to the nature of the overlap, we also present a classification approach based on features derived from such emotional changes surrounding an overlap and compare the classification performance of these features to classic acoustic features. We show that the automatic classification of competitive and cooperative overlaps using the changes in valence and control levels of the overlapping speaker outperforms common approaches employing acoustic and linguistic features.
Olga Egorow, Andreas Wendemuth
IEEE Trans. Affect. Comput.2
2019 Employing Bottleneck and Convolutional Features for Speech-Based Physical Load Detection on Limited Data Amounts
Olga Egorow, Tarik Mrech, Norman Weißkirchen, Andreas Wendemuth
INTERSPEECH4
2019 Towards cognitive systems for assisted cooperative processes of goal finding and strategy change
abstract
In the future, people and intelligent technical systems will cooperate in situations in which the goals, means, or actions are not completely prespecified, but develop in the course of a process that entails also the finding of new goals and strategies. This position paper presents an account of how these processes can be described such that they enable support of human decision-making and action coordination in such open, under-specified scenarios. We discuss how technical cognitive systems can assist these processes by bringing together techniques for multi-modal processing, information retrieval, situated action planning and autonomous action generation, with novel capabilities of recognizing and anticipating task-related (cognitive-intentional, procedural, affective) states of the actors, and for cooperative goal refinement and action coordination among the actors. A foremost requirement is to automatically provide markers for the indication of necessary strategy changes that reconFigure the space of actions, where it is to be expected that such strategy changes may require explanation and mediation. To that end, cognitive systems and robots must be endowed with new kinds of models of (explicit or implicit) cooperative processes.
Andreas Wendemuth, Stefan Kopp
SMC1
2018 Recognizing Behavioral Factors while Driving: A Real-World Multimodal Corpus to Monitor the Driver's Affective State
Alicia Flores Lotz, Klas Ihme, Audrey Charnoz, Pantelis Maroudis, Ivan Dmitriev, Andreas Wendemuth
LREC6
2018 Significance of Feature Differences in the Distinction of Mental-Load
abstract
The use of voice as a main form of communication is becoming more relevant for modern human-machine interactions. Based on former works concerned primarily with the understanding of speech-to-text and the recognition of emotional information through voice, modern applications strive to advance the level of information that can be gathered by recording and analyzing the voice of users. This can include the general disposition of the speaker, beyond the direct emotional state but also the general state of behavior, with which systems might anticipate the most likely actions the user might take and prepare for them. In this paper we investigate how the general state of the mind and situation of the user influence his/her voice and if it is possible to measure this behavior in a reproducible manner.
Norman Weißkirchen, Ronald Böck, Andreas Wendemuth, Andreas Nürnberger
SMC3
2018 Intention-Based Anticipatory Interactive Systems
abstract
Intention-based, anticipatory, interactive systems (IAIS) represent a new class of user-centered assistance systems. IAIS uses actions and system intentions derived from signal data, and the affective state of the user. By anticipating the further action of the user, solutions are interactively negotiated. The active roles of humans and systems change strategically, which requires behavioral models, which in turn can be specified by life sciences' results. Deployed human-machine-systems and lab tests have the goal of understanding of the situated interaction. This supports integration of assistance systems for Industry 4.0 and in the context of demographic change. We provide a definition of IAIS, discuss the underlying requirements, goals and challenges and provide a brief review of the state-of-the-art in the involved research areas.
Andreas Wendemuth, Ronald Böck, Andreas Nürnberger, Ayoub Al-Hamadi, André Brechmann, Frank W. Ohl
SMC1
2018 Using a PCA-based dataset similarity measure to improve cross-corpus emotion recognition
Ingo Siegert, Ronald Böck, Andreas Wendemuth
Comput. Speech Lang.3
2017 Emotional Features for Speech Overlaps Classification
Olga Egorow, Andreas Wendemuth
INTERSPEECH2
2017 Towards a Sensor Failure-Dependent Performance Adaptation Using the Validity Concept
Juliane Höbel, Georg Jäger, Sebastian Zug, Andreas Wendemuth
SAFECOMP4
2015 Cross-corpus acoustic emotion recognition: Variances and strategies (Extended abstract)
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: acted data is often used rather than spontaneous data, results are reported on pre-selected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. A considerably more realistic impression can be gathered by inter-set evaluation: we therefore show results employing six standard databases in a cross-corpora evaluation experiment. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter- to intra-corpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
ACII6
2015 Exploring dataset similarities using PCA-based feature selection
abstract
In emotion recognition from speech, several well-established corpora are used to date for the development of classification engines. The data is annotated differently, and the community in the field uses a variety of feature extraction schemes. The aim of this paper is to investigate promising features for individual corpora and then compare the results for proposing optimal features across data sets, introducing a new ranking method. Further, this enables us to present a method for automatic identification of groups of corpora with similar characteristics. This answers an urgent question in classifier development, namely whether data from different corpora is similar enough to jointly be used as training material, overcoming shortage of material in matching domains. We compare the results of this method with manual groupings of corpora. We consider the established emotional speech corpora AVIC, ABC, DES, EMO-DB, ENTERFACE, SAL, SMARTKOM, SUSAS and VAM, however our approach is general.
Ingo Siegert, Ronald Böck, Andreas Wendemuth, Bogdan Vlasenko
ACII3
2015 NaLMC: A Database on Non-acted and Acted Emotional Sequences in HCI
abstract
We report on the investigation on acted and non-acted emotional speech and the resulting Non-/acted LAST MINUTE corpus (NaLMC) database. The database consists of newly recorded acted emotional speech samples which were designed to allow the direct comparison of acted and non-acted emotional speech. The non-acted samples are taken from the LAST MINUTE corpus (LMC) [1]. Furthermore, emotional labels were added to selected passages of the LMC and a self-rating of the LMC recordings was performed. Although the main objective of the NaLMC database is to allow the comparative analysis of acted and non-acted emotional speech, both audio and video signals were recorded to allow multimodal investigations.
Kim Hartmann, Julia Krüger, Jörg Frommer, Andreas Wendemuth
ICMI4
2015 Annotators' agreement and spontaneous emotion classification performance
abstract
The combination of various types of data can significantly in-crease the amount of emotional material for training of more reliable real-life emotion classifiers. There are two well-known schemes of annotation utilized for emotional speech: multi-dimensional and categories-based. Multi-dimensional anno-tation is usually applied for labeling spontaneous emotional events, and categorial-based annotation is used for specifica-tion of the acted ”full blown ” emotional chunks. In order to simulate real-life conditions we used a cross-corpora evalua-tion strategy for datasets with different schemes of emotional annotation. Emotional models were trained on acted material from the EMO-DB (categories based annotation) dataset and evaluated on spontaneous data from the VAM dataset (multi-dimensional annotation). The best emotion classification per-formance was obtained on real-life emotional instances with the most intense arousal labels provided by a majority voting strategy (out of 17 annotators). We find that the correspond-ing spontaneous speech samples containing the most intensive emotional content are comparable with acted instances. The im-portance of employing a larger number of emotional annotators was finally addressed in our article. Index Terms: emotion recognition, cross-corpora evaluation, phoneme-level emotional models, turn-level emotional models, emotional intensity 1.
Bogdan Vlasenko, Andreas Wendemuth
INTERSPEECH2
2014 Location of an emotionally neutral region in valence-arousal space: Two-class vs. three-class cross corpora emotion recognition evaluations
abstract
There are two main emotion annotation techniques: multidimensional and categories based. In order to conduct experiments on emotional data annotated with different techniques, two-classes emotion mapping strategies (e.g. high-vs. low-arousal) are commonly used. The ”affective computing” community could not specify a location of emotionally neutral area in multi-dimensional emotional space (e.g. valence-arousal-dominance (VAD)). Nonetheless, in the current research a neutral state is added to the standard two-classes emotion classification task. Within experiments a possible location of a neutral arousal region in valence-arousal space was determined. We employed general and phonetic pattern dependent emotion classification techniques for cross-corpora experiments. Emotional models were trained on the VAM dataset (multi-dimensional annotation) and evaluated them on the EMO-DB dataset (categories based annotation).
Bogdan Vlasenko, Andreas Wendemuth
ICME2
2014 Application of image processing methods to filled pauses detection from spontaneous speech
Dmytro Prylipko, Olga Egorow, Ingo Siegert, Andreas Wendemuth
INTERSPEECH4
2014 Modeling phonetic pattern variability in favor of the creation of robust emotion classifiers for real-life applications
Bogdan Vlasenko, Dmytro Prylipko, Ronald Böck, Andreas Wendemuth
Comput. Speech Lang.4
2014 Learning long-term dependencies in segmented-memory recurrent neural networks with backpropagation of error
Stefan Glüge, Ronald Böck, Günther Palm, Andreas Wendemuth
Neurocomputing4
2013 Annotation and Classification of Changes of Involvement in Group Conversation
abstract
The detection of involvement in a conversation is important to assess the level humans are participating in either a human-human or human-computer interaction. Especially, detecting changes in a group's involvement in a multi-party interaction is of interest to distinguish several constellations in the group itself. This information can further be used in situations where technical support of meetings is favoured, for instance, focusing a camera, switching microphones, etc. Moreover, this information could also help to improve the performance of technical systems applied in human-machine interaction. In this paper, we concentrate on video material given by the Table Talk corpus. Therefore, we introduce a way of annotating and classifying changes of involvement and discuss the reliability of the annotation. Further, we present classification results based on video features using Multi-Layer Networks.
Ronald Böck, Stefan Glüge, Ingo Siegert, Andreas Wendemuth
ACII4
2013 Parameter Optimization Issues for Cross-corpora Emotion Classification
abstract
As speech based emotion recognition has matured to a degree where it becomes applicable within real-life conditions, it is time for a realistic view on obtainable performances. Most state-of-the-art emotion recognition methods are based on turn- and frame-level analysis independent of phonetic transcription. True speaker disjoint partitioning of training and test sets is still less common than simple cross-validation. Even speaker disjoint experiments can give only little insight into the generalization ability of modern emotion recognition engines since training and test sets used for system development usually tend to be similar as far as acoustic channel, noise overlay, and language are concerned. A considerably more realistic impression can be gathered by cross-corpora evaluation. Tuning of the emotion classification engine (feature set optimization and normalization, selection of a classification technique and corresponding parameter configuration) is an important issue of realistic evaluations. In the ideal case, an optimal classifier configuration estimated on training data should provide an outstanding recognition performance on unseen data. We therefore compare cross-corpora classification performances of optimized and non-optimized general and phonetic-pattern dependent classifiers.
Bogdan Vlasenko, David Philippou-Hübner, Andreas Wendemuth
ACII3
2013 Determining the Smallest Emotional Unit for Level of Arousal Classification
abstract
Most state-of-the-art emotion recognition methods are based on turn- and frame-level analysis independent from phonetic transcription. Currently "affective computing" community could not specify the smallest emotional standard unit which can be easily classified and determined by any "advanced" and "non-advanced" listener. It is known that, acoustic modeling on the smallest phonetic unit (phoneme) started a new era in automatic speech recognition: switch from speaker dependent isolated word recognition to speaker independent continuous speech recognition. In or current research we showed that phoneme can be used as as smallest unit for high and low arousal emotion classification task. We trained our classifications models on the VAM dataset material and evaluated them on speech samples from the DES dataset. For our experiments we employed two different emotion classification approaches: general (phonetic pattern independent) and phoneme-based (phonetic pattern dependent). Both classification approaches used MFFC features extracted on the frame level. Our experimental results impressively show that the proposed phoneme-based classification technique could increase emotion classification performance by about 9.68% absolute (15.98% relative). We showed that phoneme-level emotion models trained on "natural" emotions could provide impressive classification performance on dataset with acted affective content.
Bogdan Vlasenko, Andreas Wendemuth
ACII2
2013 Auto-encoder pre-training of segmented-memory recurrent neural networks
Stefan Glüge, Ronald Böck, Andreas Wendemuth
ESANN3
2012 Fine-tuning HMMS for nonverbal vocalizations in spontaneous speech: A multicorpus perspective
abstract
Phenomena like filled pauses, laughter, breathing, hesitation, etc. play significant role in everyday human-to-human conversation and have a significant influence on speech recognition accuracy [1]. Because of their nature (e. g. long duration), they should be modeled with different number of emitting states and Gaussian mixtures. In this paper we address this question and try to determine the most suitable method for finding these parameters: we provide an examination of two methods for optimization of hidden Markov model (HMM) configurations for better classification and recognition of nonverbal vocalizations within speech. Experiments were conducted on three conversational databases: TUM AVIC, Verbmobil, and SmartKom. These experiments show that with HMMs configurations tailored to a particular database we can achieve 1-3% improvement in speech recognition accuracy with comparison to a baseline topology. An in-depth analysis of discussed methods is provided.
Dmytro Prylipko, Björn W. Schuller, Andreas Wendemuth
ICASSP3
2012 The Performance of the Speaking Rate Parameter in Emotion Recognition from Speech
abstract
The speaking rate is a quite obvious prosodic characteristic of speech and humans can easily estimate how fast an interlocutor is talking. Further, different emotional dispositions of a person are strongly expressed in his/her speaking rate. In this paper we investigate the performance gain originating from the use of the speaking rate parameter in emotion recognition from speech. The speaking rates are determined by applying a broad phonetic class recognizer. The classifier is trained on cepstral features extracted on the emotionally neutral RM1 speech corpus and provides low average recognition errors of one phoneme/second. We present the results of an empirical approach on the emotionally expressive Emo-DB corpus applying a neural network classifier and prove the significant influence of the speaking rate in emotion classification. The performances of Multi-Layer Perceptrons trained on cepstral turn-level features are analyzed with respect to the presence and absence of the speaking rate feature. An increase of accuracy up to 3.7% in certain emotion categories is reported.
David Philippou-Hübner, Bogdan Vlasenko, Ronald Böck, Andreas Wendemuth
ICME4
2012 Extension of Backpropagation through Time for Segmented-memory Recurrent Neural Networks
Stefan Glüge, Ronald Böck, Andreas Wendemuth
IJCCI3
2012 Towards Emotion and Affect Detection in the Multimodal LAST MINUTE Corpus
Jörg Frommer, Bernd Michaelis, Dietmar F. Rösner, Andreas Wendemuth, Rafael Friesen, Matthias Haase, Manuela Kunze, Rico Andrich, Julia Lange, Axel Panning, Ingo Siegert
LREC4
2012 Majority Decisions in Overlapping Committees and Asymptotic Size of Dichotomies
abstract
In this paper we settle an open combinatorial conjecture in artificial neural networks: we show that the bound on the number of dichotomies given by Mitchinson and Durbin [Biological Cybernetics, 60 (1989), pp. 345--365] is tight, and that their structural asymptotics remain unchanged with varying required success probability. In our proof we use Rényi's graph-sieves inequalities, and we derive a contracted version and a sharp bound for a triple-indexed sum of binomial coefficients with dependent indices.
Andreas Wendemuth, Italo Simonelli
SIAM J. Discret. Math.1
2011 ikannotate - A Tool for Labelling, Transcription, and Annotation of Emotionally Coloured Speech
Ronald Böck, Ingo Siegert, Matthias Haase, Julia Lange, Andreas Wendemuth
ACII (1)5
2011 Appropriate emotional labelling of non-acted speech using basic emotions, geneva emotion wheel and self assessment manikins
abstract
In emotion recognition from speech, a good transcription and annotation of given material is crucial. Moreover, the question of how to find good emotional labels for new data material is a basic issue. It is not only the question of which emotion labels to choose, it is also a matter of how labellers can cope with annotation methods. In this paper, we present our investigations for emotional labelling with three different methods (Basic Emotions, Geneva Emotion Wheel and Self Assessment Manikins) and compare them in terms of emotion coverage and usability. We show that emotion labels derived from Geneva Emotion Wheel or Self Assessment Manikins fulfill our requirements, but Basic Emotions are not feasible for emotion labelling from spontaneous speech.
Ingo Siegert, Ronald Böck, Bogdan Vlasenko, David Philippou-Hübner, Andreas Wendemuth
ICME5
2011 Vowels formants analysis allows straightforward detection of high arousal emotions
abstract
Recently, automatic emotion recognition from speech has achieved growing interest within the human-machine interaction research community. Most part of emotion recognition methods use context independent frame-level analysis or turn-level analysis. In this article, we introduce context dependent vowel level analysis applied for emotion classification. An average first formant value extracted on vowel level has been used as unidimensional acoustic feature vector. The Neyman-Pearson criterion has been used for classification purpose. Our classifier is able to detect high-arousal emotions with small error rates. Within our research we proved that the smallest emotional unit should be the vowel instead of the word. We find out that using vowel level analysis can be an important issue during developing a robust emotion classifier. Also, our research can be useful for developing robust affective speech recognition methods and high quality emotional speech synthesis systems.
Bogdan Vlasenko, David Philippou-Hübner, Dmytro Prylipko, Ronald Böck, Ingo Siegert, Andreas Wendemuth
ICME6
2011 Vowels Formants Analysis Allows Straightforward Detection of High Arousal Acted and Spontaneous Emotions
Bogdan Vlasenko, Dmytro Prylipko, David Philippou-Hübner, Andreas Wendemuth
INTERSPEECH4
2010 Determining optimal features for emotion recognition from speech by applying an evolutionary algorithm
abstract
Abstract The automated recognition of emotions from speech is a chal-lenging issue. In order to build an emotion recognizer well de-fined features and optimized parameter sets are essential. Thispaper will show how an optimal parameter set for HMM-basedrecognizerscanbefoundbyapplyinganevolutionaryalgorithmon standard features in automated speech recognition. For this,we compared different signal features, as well as several archi-tectures of HMMs. The system was evaluated on a non-acteddatabase and its performance was compared to a baseline sys-tem. We present an optimal feature set for the public part of theSmartKom database.Index Terms: Emotion Recognition, Evolutionary Algorithms,Feature Optimization, Hidden-Markov Models 1. Introduction The interaction between men and machines using language isnowadays becoming more and more self-evident, but machinesstill lack of many human abilities which would considerablysimplify communication and would also help to increase theacceptance of such systems. For some time, research activitiesalso focus stronger on the emotional aspect of speech. Exploit-ing information about the emotional state of a user, machinescan be enabled to adapt their dialog strategy online, depend-ing on the user’s emotions and hence react in a more appro-priate and empathic manner. As emotion recognition in manyapplications goes hand in hand with automated speech recog-nition (ASR) it would be favorable to make use of the samefeatures or even a subset thereof. Especially small devices likesmart phones or PDAs, which do not provide huge computa-tionalpowerwouldbenefitfromsuchasparsefeatureapproach.Parallel research of other groups bases on pooling together(high level) features including the application of brute forcemethods in order to fully exploit the feature space (compare[1]). This paper however describes an evolutionary strategy(ES) of finding an optimal sparse feature set, given not morethan the common acoustic features used in ASR. Applying anES has the advantage of self adaptation of its parameters andis able to find optimal parameter constellations in high dimen-sional search spaces and further gives insights into the rele-vance of each parameter in terms of the model’s accuracy. Es-pecially in case of many parameters with unknown relation-ships ES quickly avoids wasting time on generating and test-ingunsuitableparametercombinationsastheevolutionaryforceminimizes the probability of the evolvement of such combina-tions effectively. In ASR Mel-Frequency-Cepstral-Coefficients(MFCCs) have established as a basic feature in order to trainphoneme based recognizers. As we do not want additional pa-rameters to be extracted from the speech signal we concentrateonly on MFCCs, which have also proven to perform well inemotion recognition during the Emotion Challenge within In-terspeech 2009 (compare [2]).Thispaperisstructuredasfollows: InSection2wedescribethespontaneous database and the emotions we want to recognize.Section 3 describes, which parameters we investigated and inwhich range they were allowed to change during the evolutionprocess. Section4introducestheevolutionaryalgorithm,showshow fitness is measured and the population evolves over time.TheresultsarepresentedinSection5andcomparedtoourbase-line recognizer. Finally Section 6 summarizes our findings andgives an outlook.
David Philippou-Hübner, Bogdan Vlasenko, Tobias Grosser, Andreas Wendemuth
INTERSPEECH4
2010 Cross-Corpus Acoustic Emotion Recognition: Variances and Strategies
abstract
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: Acted data is often used rather than spontaneous data, results are reported on preselected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. Even speaker disjunctive evaluation can give only a little insight into the generalization ability of today's emotion recognition engines since training and test data used for system development usually tend to be similar as far as recording conditions, noise overlay, language, and types of emotions are concerned. A considerably more realistic impression can be gathered by interset evaluation: We therefore show results employing six standard databases in a cross-corpora evaluation experiment which could also be helpful for learning about chances to add resources for training and overcoming the typical sparseness in the field. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter to intracorpus testing.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll
IEEE Trans. Affect. Comput.6
2009 Acoustic emotion recognition: A benchmark comparison of performances
abstract
In the light of the first challenge on emotion recognition from speech we provide the largest-to-date benchmark comparison under equal conditions on nine standard corpora in the field using the two pre-dominant paradigms: modeling on a frame-level by means of hidden Markov models and supra-segmental modeling by systematic feature brute-forcing. Investigated corpora are the ABC, AVIC, DES, EMO-DB, eNTERFACE, SAL, SmartKom, SUSAS, and VAM databases. To provide better comparability among sets, we additionally cluster each database's emotions into binary valence and arousal discrimination tasks. In the result large differences are found among corpora that mostly stem from naturalistic emotions and spontaneous speech vs. more prototypical events. Further, supra-segmental modeling proves significantly beneficial on average when several classes are addressed at a time.
Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Gerhard Rigoll, Andreas Wendemuth
ASRU5
2009 Heading toward to the natural way of human-machine interaction: the nimitek project
abstract
Spoken human-machine interaction supported by state-of-the-art dialog systems is becoming a standard technology. A lot of effort was invested for this kind of artificial communication interface. But still the spoken dialog systems (SDS) are not able to provide for the user a natural way of communication. Because, existing automated dialog system do not dedicate enough attention to problems in the interaction related to affected user behavior. This paper addresses some aspects of design and implementation of user behavior models in dialog systems aimed to provide naturalness of human-machine interaction. We discuss a viable integration technique of speech based emotion classification in SDS for robust affected automatic speech recognition and user emotion correlated dialog strategy. First of all, we describe existing methods of emotion recognition within speech and affected speech adapted ASR methods. Second, we introduce an approach to achieve emotion adaptive dialog management in human-machine interaction. A multimodal human-machine interaction system with integrated user behavior model is created within the project ldquoNeurobiologically Inspired, Multimodal Intention Recognition for Technical Communication Systemsrdquo (NIMITEK). Currently NIMITEK provides a technical demonstrator to study these principles in a dedicated prototypical task, namely solving the game Towers of Hanoi. In this paper, we will describe the general approach NIMITEK takes to emotional man-machine interactions.
Bogdan Vlasenko, Andreas Wendemuth
ICME2
2009 Processing affected speech within human machine interaction
abstract
Spoken dialog systems (SDS) integrated into human-machine interaction interfaces is becoming a standard technology. Current state-of-the-art SDS, usually, is not able to provide for the user a natural way of communication. Existing automated dialog systems do not dedicate enough attention to problems in the interaction related to affected user behavior. As a result, Automatic Speech Recognition (ASR) engines are not able to recognize affected speech and dialog strategy does not make use of the user’s emotional state. This paper addresses some aspects of processing affected speech withinnatural human-machine interaction. First of all, we propose an affected speech adapted ASR engine. Second, we describe our methods of emotion recognition within speech and present our results of emotion classification within Interspeech 2009 Emotion Challenge. Third, we test affected speech adapted speech recognition models and introduce an approach to achieve emotion adaptive dialog management in human-machine interaction.
Bogdan Vlasenko, Andreas Wendemuth
INTERSPEECH2
2008 Combining speech recognition and acoustic word emotion models for robust text-independent emotion recognition
abstract
Recognition of emotion in speech usually uses acoustic models that ignore the spoken content. Likewise one general model per emotion is trained independent of the phonetic structure. Given sufficient data, this approach seemingly works well enough. Yet, this paper tries to answer the question whether acoustic emotion recognition strongly depends on phonetic content, and if models tailored for the spoken unit can lead to higher accuracies. We therefore investigate phoneme-, and word-models by use of a large prosodic, spectral, and voice quality feature space and Support Vector Machines (SVM). Experiments also take the necessity of ASR into account to select appropriate unit- models. Test-runs on the well-known EMO-DB database facing speaker-independence demonstrate superiority of word emotion models over today's common general models provided sufficient occurrences in the training corpus.
Björn W. Schuller, Bogdan Vlasenko, Dejan Arsic, Gerhard Rigoll, Andreas Wendemuth
ICME5
2008 Making the Lipschitz Classifier Practical via Semi-infinite Programming
abstract
This paper presents a new implementable algorithm for solving the Lipschitz classifier that is a generalization of the maximum margin concept from Hilbert to Banach spaces. In contrast to the support vector machine approach, our algorithm is free to use any finite family of continuously differentiable functions which linearly compose the decision function. Nevertheless, robustness properties are maintained due to a maximizing margin. To obtain a useful algorithm, the inherent difficult problem is formulated in a convex semi-infinite program. Using this new formulation, we develop a duality result enabling us to solve the original problem iteratively as a finite sequence of constrained quadratic programming problems over a convex hull of matrices. We compare the performance of the Lipschitz classifier algorithm with state-of-the-art machine learning methodologies using a benchmark data set as well as a data set randomly generated from Gaussian mixtures.
André Stuhlsatz, Hans-Günter Meier, Andreas Wendemuth
ICMLA3
2008 Balancing spoken content adaptation and unit length in the recognition of emotion and interest
abstract
Recognition and detection of non-lexical or paralinguistic cues from speech usually uses one general model per event (emotional state, level of interest).Commonly this model is trained independent of the phonetic structure.Given sufficient data, this approach seemingly works well enough.Yet, this paper addresses the question on which phonetic level there is the onset of emotions and level of interest.We therefore compare phoneme-, word-and sentence-level analysis for emotional sentence classification by use of a large prosodic, spectral, and voice quality feature space for SVM and MFCC for HMM/GMM.Experiments also take the necessity of ASR into account to select appropriate unit-models.In experiments on the well-known public EMO-DB database, and the SUSAS and AVIC spontaneous interest corpora, we found that the emotion recognition by sentence level analysis shows the best results.We discuss the implications of these types of analysis on the design of robust emotion and interest recognition of usable human-machine interfaces (HMI).
Bogdan Vlasenko, Björn W. Schuller, Kinfe Tadesse Mengistu, Gerhard Rigoll, Andreas Wendemuth
INTERSPEECH5
2008 Hierarchical HMM-based semantic concept labeling model
abstract
An utterance can be conceived as a hidden sequence of semantic concepts expressed in words or phrases. The problem of understanding the meaning underlying a spoken utterance in a dialog system can be partly solved by decoding the hidden sequence of semantic concepts from the observed sequence of words. In this paper, we describe a hierarchical HMM-based semantic concept labeling model trained on semantically unlabeled data. The hierarchical model is compared with a flat concept based model in terms of performance, ambiguity resolution ability and expressive power of the output. It is shown that the proposed method outperforms the flat-concept model in these points.
Kinfe Tadesse Mengistu, Mirko Hannemann, Tobias Baum, Andreas Wendemuth
SLT4
2007 Frame vs. Turn-Level: Emotion Recognition from Speech Considering Static and Dynamic Processing
Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll
ACII3
2007 Comparing one and two-stage acoustic modeling in the recognition of emotion in speech
abstract
In the search for a standard unit for use in recognition of emotion in speech, a whole turn, that is the full section of speech by one person in a conversation, is common. Within applications such turns often seem favorable. Yet, high effectiveness of sub-turn entities is known. In this respect a two-stage approach is investigated to provide higher temporal resolution by chunking of speech-turns according to acoustic properties, and multi-instance learning for turn-mapping after individual chunk analysis. For chunking fast pre-segmentation into emotionally quasi-stationary segments by one-pass Viterbi beam search with token passing basing on MFCC is used. Chunk analysis is realized by brute-force large feature space construction with subsequent subset selection, SVM classification, and speaker normalization. Extensive tests reveal differences compared to one-stage processing. Alternatively, syllables are used for chunking.
Björn W. Schuller, Bogdan Vlasenko, Ricardo Minguez, Gerhard Rigoll, Andreas Wendemuth
ASRU5
2007 Updates for Nonlinear Discriminants
Edin Andelic, Martin Schafföner, Marcel Katz, Sven Krüger 0002, Andreas Wendemuth
IJCAI5
2007 Dynamics of Temporal Difference Learning
Andreas Wendemuth
IJCAI1
2007 Combining frame and turn-level information for robust recognition of emotions within speech
abstract
Current approaches to the recognition of emotion within speech usually use statistic feature information obtained by application of functionals on turn- or chunk levels. Yet, it is well known that thereby important information on temporal sub-layers as the frame-level is lost. We therefore investigate the benefits of integration of such information within turn-level feature space. For frame-level analysis we use GMM for classification and 39 MFCC and energy features with CMS. In a subsequent step output scores are fed forward into a 1.4k large-feature-space turn-level SVM emotion recognition engine. Thereby we use a variety of Low-Level-Descriptors and functionals to cover prosodic, speech quality, and articulatory aspects. Extensive testruns are carried out on the public databases EMO-DB and SUSAS. Speaker-independent analysis is faced by speaker normalization. Overall results highly emphasize the benefits of feature integration on diverse time scales.
Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll
INTERSPEECH3
2006 Limited Training Data Robust Speech Recognition Using Kernel-Based Acoustic Models
abstract
Contemporary automatic speech recognition uses hidden-Markov-models (HMMs) to model the temporal structure of speech where one HMM is used for each phonetic unit. The states of the HMMs are associated with state-conditional probability density functions (PDFs) which are typically realized using mixtures of Gaussian PDFs (GMMs). Training of GMMs is error-prone especially if training data size is limited. This paper evaluates two new methods of modeling state-conditional PDFs using probabilistically interpreted support vector machines and kernel Fisher discriminants. Extensive experiments on the RMI (P. Price et al., 1988) corpus yield substantially improved recognition rates compared to traditional GMMs. Due to their generalization ability, our new methods reduce the word error rate by up to 13% using the complete training set and up to 33% when the training set size is reduced
Martin Schafföner, Sven Krüger 0002, Edin Andelic, Marcel Katz, Andreas Wendemuth
ICASSP (1)5
2006 Kernel Least-Squares Models Using Updates of the Pseudoinverse
abstract
Sparse nonlinear classification and regression models in reproducing kernel Hilbert spaces (RKHSs) are considered. The use of Mercer kernels and the square loss function gives rise to an overdetermined linear least-squares problem in the corresponding RKHS. When we apply a greedy forward selection scheme, the least-squares problem may be solved by an order-recursive update of the pseudoinverse in each iteration step. The computational time is linear with respect to the number of the selected training samples.
Edin Andelic, Martin Schafföner, Marcel Katz, Sven Krüger 0002, Andreas Wendemuth
Neural Comput.5
2005 Speech recognition with support vector machines in a hybrid system
abstract
While the temporal dynamics of speech can be represented very efficiently by Hidden Markov Models (HMMs), the classification of speech into single speech units (phonemes) is usually done with Gaussian mixture models which do not discriminate well. Here, we use Support Vector Machines (SVMs) for classification by integrating this method in a HMM-based speech recognition system. In this hybrid SVM/HMM system we translate the outputs of the SVM classifiers into conditional probabilities and use them as emission probabilities in a HMM-based decoder. SVMs are very appealing due to their association with statistical learning theory. They have already shown very good classification results in other fields of pattern recognition. We train and test our hybrid system on the DARPA Resource Management (RM1) corpus. Our results show better performance than HMM-based decoder using Gaussian mixtures. 1.
Sven Krüger 0002, Martin Schafföner, Marcel Katz, Edin Andelic, Andreas Wendemuth
INTERSPEECH5
2003 Improved robustness of automatic speech recognition using a new class definition in linear discriminant analysis
abstract
This work discusses the improvements which can be expected when applying linear feature-space transformations based on Linear Discriminant Analysis (LDA) within automatic speechrecognition (ASR). It is shown that different factors influence the effectiveness of LDA-transformations. Most importantly, increasing the number of LDA-classes by using time-aligned states of Hidden-Markov-Models instead of phonemes is necessary to obtain improvements predictably. An extension of LDA is presented, which utilises the elementary Gaussian components of the mixture probability-density functions of the Hidden-Markov-Models' states to define actual Gaussian LDAclasses. Experimental results on the TIMIT and WSJCAM0 recognition task are given, where relative improvements of the error-rate of 3.2% and 3.9%, respectively, were obtained.
Martin Schafföner, Marcel Katz, Sven Krüger 0002, Andreas Wendemuth
INTERSPEECH4
2002 Large vocabulary continuous speech recognition of Broadcast News - The Philips/RWTH approach
Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Hermann Ney, Michael Pitz, Achim Sixtus
Speech Commun.6
2001 Modeling uncertainty of data observation
abstract
An approach is presented both theoretically and experimentally which overcomes a number of existing conceptual and performance problems in density estimation. The theoretical approach shows methods for incorporating or estimating uncertainties into speech recognition. In the maximum mutual information (MMI) and maximum likelihood (ML) case, precise formulae are given for estimation of densities for uncertainty variances small compared to the curvature of the posteriors. For implementation, the theoretical formulae are presented in such a way that the additional computation effort goes linearly with the number of densities. Experiments on car digits show relative improvements in word error rate of at most 4.8% relative. Uncertainty modelling is shown to help remedy effects of the sparse data problem in density estimation.
Andreas Wendemuth
ICASSP1
1999 Advances in confidence measures for large vocabulary
abstract
This paper addresses the correct choice and combination of confidence measures in large vocabulary speech recognition tasks. We classify single words within continuous as well as large vocabulary utterances into two categories: utterances within the vocabulary which are recognized correctly, and other utterances, namely misrecognized utterances or (less frequent) out-of-vocabulary (OOV). To this end, we investigate the classification error rate (CER) of several classes of confidence measures and transformations. In particular, we employed data-independent and data-dependent measures. The transformations we investigated include mapping to single confidence measures and linear combinations of these measures. These combinations are computed by means of neural networks trained with Bayes-optimal, and with Gardner-Derrida-optimal criteria. Compared to a recognition system without confidence measures, the selection of (various combinations of) confidence measures, the selection of suitable neural network architectures and training methods, continuously improves the CER.
Andreas Wendemuth, Georg Rose, Hans J. G. A. Dolfing
ICASSP1
1999 Optimal training parameters in multilayer feedforward networks
abstract
We present a systematic investigation of the training behavior for multilayer feedforward neural networks. Usually learning is governed by three metaparameters, which are learning rate, momentum and offset. We apply a (nearly) exhaustive search method to find optimal parameter sets throughout the complete sequence of training cycles, regarding the training process as a finite state network in the space of metaparameter configurations. Minimization of training time is achieved by methods of dynamic programming. A detailed analysis is given for the choice of error criteria and necessary widths and prunings of network 'beams' in search space. It is shown for a representative set of training patterns, that the number of network training iterations is largely independent of both, the metaparameter initialization and the random weight initialization. Training is twice as fast as with conventional metaparameter adaptation strategies, such as RPROP or local fuzzy inference.
Andreas Wendemuth, Michael Gerke 0001
IJCNN1
1999 The philips/RWTH system for transcription of broadcast news
abstract
This paper contains a description of the Philips/RWTH 1998 HUB4 system which has been build in a joint eort of Philips Research Laboratories Aachen and Aachen University o f T echnology.We will focus our discussion on recent improvements compared to the original 1997 HUB4 system and evaluate them on the HUB4'97 evaluation data.The paper will deal with 1. a rough system overview including feature extraction, acoustic training, audio stream segmentation, and decoding 2. log-linear interpolation of distance-language models, 3. and the integration of various acoustic and language models via Discriminative Model Combination (DMC).The performance of the described system is 23% (relative) better than the performance of the 1997 Philips HUB4 system.A w ord error rate of 17.9% was achieved on the 1997 HUB4 evaluation set, compared to 23.5% using the original 1997 system.
Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Michael Pitz, Achim Sixtus
EUROSPEECH6
1998 Combination of confidence measures in isolated word recognition
abstract
In the context of command-and-control applications, we exploit confidence measures in order to classify single-word utterances into two categories: utterances within the vocabulary which are recognized correctly, and other utterances, namely out-ofvocabulary (OOV) or misrecognized utterances.
Hans J. G. A. Dolfing, Andreas Wendemuth
ICSLP2
1995 Stabilities in optimal cluster separation networks
Andreas Wendemuth
Neural Networks1
1994 Storage Capacity Bounds in Multilayer Neural Networks
abstract
General analytic expressions are given for the lower and upper storage capacity bounds of multilayer neural networks which have variable weights between input layer and first hidden layer, and a fixed output function implemented between first hidden and output layer. The special cases of committee and parity machines as well as the limiting cases of networks with minimum and maximum storage capacities are discussed. The results are compared with replica calculations and simulations. An explanation is given as to why the latter have a storage capacity just slightly above the lower limit and how this can be improved.
Andreas Wendemuth
Int. J. Neural Syst.1
1993 Fast Learning Of Biased Patterns In Neural Networks
abstract
Usual neural network gradient descent training algorithms require training times of the same order as the number of neurons N if the patterns are biased. In this paper, modified algorithms are presented which require training times equal to those in unbiased cases which are of order 1. Exact convergence proofs are given. Gain parameters which produce minimal learning times in large networks are computed by replica methods. It is demonstrated how these modified algorithms are applied in order to produce four types of solutions to the learning problem: 1. A solution with all internal fields equal to the desired output, 2. The Adaline (or pseudo-inverse) solution, 3. The perceptron of optimal stability without threshold and 4. The perceptron of optimal stability with threshold.
Andreas Wendemuth, David Sherrington
Int. J. Neural Syst.1