VLDB 2026 Research / reviewers in the wild / expert
Felix Burkhardt
dblp:24/368
· DBLP profile ↗
41ranked-venue papers
19as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 16 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EmoDB 2.0: A Database of Emotional Speech in a World that is not Black or White but Grey
Felix Burkhardt, Oliver Schrüfer, Uwe D. Reichel, Hagen Wierstorf, Anna Derington, Florian Eyben, Björn W. Schuller |
INTERSPEECH | 1 |
| 2025 | Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition ModelsabstractMachine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or tone of voice. Those models are mainly evaluated in terms of correlation or recall, and always show some errors in their predictions. The errors manifest themselves in model behaviour, which can be very different along different dimensions even if the same recall or correlation is achieved by the model. This paper introduces a testing framework to investigate behaviour of speech emotion recognition models, by requiring different metrics to reach a certain threshold in order to pass a test. The test metrics can be grouped in terms of correctness, fairness, and robustness. It also provides a method for automatically specifying test thresholds for fairness tests, based on the datasets used, and recommendations on how to select the remaining test thresholds. We evaluated a xLSTM-based and nine transformer-based acoustic foundation models against a convolutional baseline model, testing their performance on arousal, valence, dominance, and emotional category classification. The test results highlight, that models with high correlation or recall might rely on shortcuts – such as text sentiment –, and differ in terms of fairness. Anna Derington, Hagen Wierstorf, Ali Gürcan Özkil, Florian Eyben, Felix Burkhardt, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | A Single-Ended High-Voltage-Compliant 11-bit Current-Steering Digital-to-Analog Converter for Adaptive Noise Cancellation in Power Over Data Line NetworksabstractAutomotive Ethernet is considered to be the backbone of future in-vehicle data communication. One main feature is its ability to simultaneously transmit data and energy via power over data lines (PoDL). This article proposes the design of a single-ended high-voltage (HV)-compliant 11-bit current-steering digital-to-analog converter (DAC). The converter is tailored for the utilization as digitally controlled current source in an adaptive noise-cancellation filter for PoDL networks. Designed in an HV-compliant 180-nm bipolar complementary metal-oxide-semiconductor (BiCMOS) semiconductor technology, the DAC features a monolithically combined topology of two identical 10-bit low-voltage (LV) current-steering DACs supplied at 1.8 V and two complementary HV-compliant output current stages. Main design features of the segmented LV DAC are the utilization of single-ended current cells with an optimized switching logic, proposed to enhance the cells transient performance and energy efficiency. Furthermore, a newly derived$Q^{4}$asymmetric rotated walk switching scheme is investigated. At a maximum output voltage of 60 V, the proposed DAC can deliver a bidirectional output current with the amplitudes of up to 500 mA. The proposed DAC exhibits the highest voltage compliance combined with the highest output current compared with related works. It also features the second highest resolution. Operated at a sample rate of 10 MS/s with a resolution of 11 bit, a spurious-free dynamic range (SFDR) of 57.8 dB could be measured for a synthesized single tone at 100 kHz, as well as a maximum integral nonlinearity (INL) error of 1.61 LSB and a differential nonlinearity (DNL) error of 1.05 LSB. Felix Burkhardt, Florian Protze, Frank Ellinger |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
Oliver Schrüfer, Manuel Milling, Felix Burkhardt, Florian Eyben, Björn W. Schuller |
INTERSPEECH | 3 |
| 2024 | A High-Speed Dynamic Element Matching Decoder With Integrated Background Calibration ControlabstractA dynamic element matching (DEM) decoder with integrated mismatch calibration control for high-speed current-steering digital-to-analog converters (CS-DACs) and CSDAC- based direct digital frequency synthesizers (DDFSs) is studied and presented. The DEM algorithm achieves very good averaging of mismatch-induced errors in the succeeding CS-DAC. It features a minimum element transition rate, therefore opimizing the power dissipation and ensuring minimal glitch energy at the output. Due to the chosen network-based architecture, with only a few modifications of the hardware, the decoder allows the integration of a comprehensive current source mismatch calibration that can be fully operated in the background and even in parallel to the regular DEM operation. A proof-ofconcept hardware implementation of the presented decoder was fabricated in a 22-nm FD-SOI CMOS process and characterized in a high-speed DDFS system with a sampling rate of 5 GHz. Measurements reveal a significant improvement in the spurious free dynamic range (SFDR) and signal-to-noise-and-distortion ratio (SNDR) when the calibration and DEM are enabled. Compared to the state-of-the-art (SoA), the presented DDFS achieves one of the best figures of merit. Tobias Schirmer, Simon Buhr, Felix Burkhardt, Florian Protze, Frank Ellinger |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Multimodal Recognition of Valence, Arousal and Dominance via Late-Fusion of Text, Audio and Facial ExpressionsabstractWe present an approach for the prediction of valence, arousal, and dominance of people communicating via text/audio/video streams for a translation from and to sign languages.The approach consists of the fusion of the output of three CNN-based models dedicated to the analysis of text, audio, and facial expressions.Our experiments show that any combination of two or three modalities increases prediction performance for valence and arousal. Fabrizio Nunnari, Annette Rios, Uwe D. Reichel, Chirag Bhuvaneshwara, Panagiotis Paraskevas Filntisis, Petros Maragos, Felix Burkhardt, Florian Eyben, Björn W. Schuller, Sarah Ebling |
ESANN | 7 |
| 2023 | Masking Speech Contents by Random Splicing: is Emotional Expression Preserved?abstractWe discuss the influence of random splicing on the perception of emotional expression in speech signals. Random splicing is the randomized reconstruction of short audio snippets with the aim to obfuscate the speech contents. A part of the German parliament recordings has been random spliced and both versions – the original and the scrambled ones – manually labeled with respect to the arousal, valence and dominance dimensions. Additionally, we run a state-of-the-art transformer-based pre-trained emotional model on the data. We find sufficiently high correlation for the annotations and predictions of emotional dimensions between both sample versions to be confident that machine learners can be trained with random spliced data. Felix Burkhardt, Anna Derington, Matthias Kahlau, Klaus R. Scherer, Florian Eyben, Björn W. Schuller |
ICASSP | 1 |
| 2023 | Nkululeko: Machine Learning Experiments on Speaker Characteristics Without Programming
Felix Burkhardt, Florian Eyben, Björn W. Schuller |
INTERSPEECH | 1 |
| 2023 | Ethical Awareness in Paralinguistics: A Taxonomy of ApplicationsabstractSince the end of the last century, the automatic processing of paralinguistics has been investigated widely and put into practice in many applications, on wearables, smartphones, and computers. In this contribution, we address ethical awareness for paralinguistic applications, by establishing taxonomies for data representations, system designs for and a typology of applications, and users/test sets and subject areas. These are related to an “ethical grid” consisting of the most relevant ethical cornerstones, based on principalism. The characteristics of and the interdependencies between these taxonomies are described and exemplified. This makes it possible to assess more or less critical “ethical constellations.” To the best of our knowledge, this is the first attempt of its kind. Anton Batliner, Michael Neumann 0001, Felix Burkhardt, Alice Baird, Sarina Meyer, Ngoc Thang Vu, Björn W. Schuller |
Int. J. Hum. Comput. Interact. | 3 |
| 2023 | Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence GapabstractRecent advances in transformer-based architectures have shown promise in several machine learning tasks. In the audio domain, such architectures have been successfully utilised in the field of speech emotion recognition (SER). However, existing works have not evaluated the influence of model size and pre-training data on downstream performance, and have shown limited attention to generalisation, robustness, fairness, and efficiency. The present contribution conducts a thorough analysis of these aspects on several pre-trained variants of wav2vec 2.0 and HuBERT that we fine-tuned on the dimensions arousal, dominance, and valence of MSP-Podcast, while additionally using IEMOCAP and MOSI to test cross-corpus generalisation. To the best of our knowledge, we obtain the top performance for valence prediction without use of explicit linguistic information, with a concordance correlation coefficient (CCC) of. 638 on MSP-Podcast. Our investigations reveal that transformer-based architectures are more robust compared to a CNN-based baseline and fair with respect to gender groups, but not towards individual speakers. Finally, we show that their success on valence is based on implicit linguistic information, which explains why they perform on-par with recent multimodal approaches that explicitly utilise textual information. To make our findings reproducible, we release the best performing model to the community. Johannes Wagner 0001, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, Björn W. Schuller |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Guest Editorial: Special Issue on Affective Speech and Language Synthesis, Generation, and ConversionabstractThe papers in this special section focus on affective speech and language synthesis, generation, and conversion. As an inseparable and crucial part of spoken language, emotions play a substantial role in human-human and human-technology conversation. They convey information about a person’s needs, how one feels about the objectives of a conversation, the trustworthiness of one’s verbal communication, and more. Accordingly, substantial efforts have been made to generate affective text and speech for conversational AI, artificial storytelling, and machine translation. Similarly, there is a push for converting the affect in text and speech, ideally, in real-time and fully preserving intelligibility, e. g., to hide one’s emotion, for creative applications and in entertainment, or even to augment training data for affect analyzing AI. Shahin Amiriparian, Björn W. Schuller, Nabiha Asghar, Heiga Zen, Felix Burkhardt |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Probing speech emotion recognition transformers for linguistic knowledgeabstractLarge, pre-trained neural networks consisting of self-attention layers (transformers) have recently achieved state-of-the-art results on several speech emotion recognition (SER) datasets. These models are typically pre-trained in self-supervised manner with the goal to improve automatic speech recognition performance -- and thus, to understand linguistic information. In this work, we investigate the extent in which this information is exploited during SER fine-tuning. Using a reproducible methodology based on open-source tools, we synthesise prosodically neutral speech utterances while varying the sentiment of the text. Valence predictions of the transformer model are very reactive to positive and negative sentiment content, as well as negations, but not to intensifiers or reducers, while none of those linguistic features impact arousal or dominance. These findings show that transformers can successfully leverage linguistic information to improve their valence predictions, and that linguistic analysis should be included in their testing. Andreas Triantafyllopoulos, Johannes Wagner 0001, Hagen Wierstorf, Maximilian Schmitt, Uwe D. Reichel, Florian Eyben, Felix Burkhardt, Björn W. Schuller |
INTERSPEECH | 7 |
| 2022 | Nkululeko: A Tool For Rapid Speaker Characteristics DetectionabstractWe present advancements with a software tool called Nkululeko, that lets users perform (semi-) supervised machine learning experiments in the speaker characteristics domain. It is based on audformat, a format for speech database metadata description. Due to an interface based on configurable templates, it supports best practise and very fast setup of experiments without the need to be proficient in the underlying language: Python. The paper explains the handling of Nkululeko and presents two typical experiments: comparing the expert acoustic features with artificial neural net embeddings for emotion classification and speaker age regression. Felix Burkhardt, Johannes Wagner 0001, Hagen Wierstorf, Florian Eyben, Björn W. Schuller |
LREC | 1 |
| 2022 | A Comparative Cross Language View On Acted Databases Portraying Basic Emotions Utilising Machine LearningabstractSince several decades emotional databases have been recorded by various laboratories. Many of them contain acted portrays of Darwin’s famous “big four” basic emotions. In this paper, we investigate in how far a selection of them are comparable by two approaches: on the one hand modeling similarity as performance in cross database machine learning experiments and on the other by analyzing a manually picked set of four acoustic features that represent different phonetic areas. It is interesting to see in how far specific databases (we added a synthetic one) perform well as a training set for others while some do not. Generally speaking, we found indications for both similarity as well as specificiality across languages. Felix Burkhardt, Anabell Hacker, Uwe D. Reichel, Hagen Wierstorf, Florian Eyben, Björn W. Schuller |
LREC | 1 |
| 2018 | The Perception and Analysis of the Likeability and Human Likeness of Synthesized SpeechabstractThe synthesized voice has become an ever present aspect of daily life.Heard through our smart-devices and from public announcements, engineers continue in an endeavour to achieve naturalness in such voices.Yet, the degree to which these methods can produce likeable, human like voices, has not been fully evaluated.With recent advancements in synthetic speech technology suggesting that human like imitation is more obtainable, this study asked 25 listeners to evaluate both the likeability and human likeness of a corpus of 13 German male voices, produced via 5 synthesis approaches (from formant to hybrid unit selection, deep neural network systems), and 1 Human control.Results show that unlike visual artificially intelligent elements -as posed by the concept of the Uncanny Valley -likeability consistently improves along with human likeness for the synthesized voice, with recent methods achieving substantially closer results to human speech than older methods.A small scale acoustic analysis shows that the F0 of hybrid systems correlates less closely to human speech with a higher standard deviation for F0.This analysis suggests that limited variance in F0 is linked to a reduction in human likeness, resulting in lower likeability for conventional synthetic speech methods. Alice Baird, Emilia Parada-Cabaleiro, Simone Hantke, Felix Burkhardt, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 4 |
| 2016 | A Taxonomy of Specific Problem Classes in Text-to-Speech Synthesis: Comparing Commercial and Open Source Performance
Felix Burkhardt, Uwe D. Reichel |
LREC | 1 |
| 2015 | A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
Comput. Speech Lang. | 6 |
| 2015 | Introduction
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son |
Comput. Speech Lang. | 5 |
| 2013 | Voice search in mobile applications with the rootvole framework
Felix Burkhardt |
INTERSPEECH | 1 |
| 2013 | Voice search in mobile applications and the use of linked open data
Felix Burkhardt, Hans Ulrich Nägeli |
INTERSPEECH | 1 |
| 2013 | Paralinguistics in speech and language - State-of-the-art and the challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller 0014, Shri Narayanan |
Comput. Speech Lang. | 4 |
| 2012 | The INTERSPEECH 2012 Speaker Trait ChallengeabstractLIDIAP Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
INTERSPEECH | 6 |
| 2012 | Is 'not bad' good enough? Aspects of unknown voices' likabilityabstractFrom the DTAG likability database the 30 most and 30 least lik-able ones have been selected for further phonetic analysis and expert questionnaire. In contrast to dislikable speakers, likable ones exhibit almost no perceivable accent, command style or disfluencies and were rated very high on all six questionnaire scales. Dislikable speakers are only moderately rated. However, dislikable speakers display also lower pronunciation precision, different amounts of jitter, lower articulation rate and higher pitch, indicating that just the absence of negatively perceived characterizations is not enough to become a likable speaker. Index Terms: likability, speaker traits, speaking style 1. Benjamin Weiss 0001, Felix Burkhardt |
INTERSPEECH | 2 |
| 2012 | Fast Labeling and Transcription with the Speechalyzer Toolkit
Felix Burkhardt |
LREC | 1 |
| 2012 | "You Seem Aggressive!" Monitoring Anger in a Practical Application
Felix Burkhardt |
LREC | 1 |
| 2011 | EmotionML - An Upcoming Standard for Representing Emotions and Related States
Marc Schröder 0001, Paolo Baggia, Felix Burkhardt, Catherine Pelachaud, Christian Peter, Enrico Zovato |
ACII (1) | 3 |
| 2011 | An Affective Spoken Storyteller
Felix Burkhardt |
INTERSPEECH | 1 |
| 2011 | "Would You Buy a Car from Me?" - On the Likability of Telephone VoicesabstractWe researched how “likable ” or “pleasant ” a speaker appears based on a subset of the “Agender ” database which was recently introduced at the 2010 Interspeech Paralinguistic Challenge. 32 participants rated the stimuli according to their likability on a seven point scale. An Anova showed that the samples rated are significantly different although the inter-rater agreement is not very high. Experiments with automatic regression and clas-sification by REPTree ensemble learning resulted in a cross-correlation of up to.378 with the evaluator weighted estimator, and 67.6 % accuracy in binary classification (likable / not lik-able). Analysis of individual acoustic feature groups reveals that for this data, auditory spectral features seem to contribute the most to reliable automatic likability analysis. Index Terms: speaker traits, likability, classification 1. Felix Burkhardt, Björn W. Schuller, Benjamin Weiss 0001, Felix Weninger |
INTERSPEECH | 1 |
| 2010 | Learning with synthesized speech for automatic emotion recognitionabstractData sparseness is an ever dominating problem in automatic emotion recognition. Using artificially generated speech for training or adapting models could potentially ease this: though less natural than human speech, one could synthesize the exact spoken content in different emotional nuances - of many speakers and even in different languages. To investigate chances, the phonemisation components Txt2Pho and openMary are used with Emofilt and Mbrola for emotional speech synthesis. Analysis is realized with our Munich open Emotion and Affect Recognition toolkit. As test set we gently limit to the acted Berlin and eNTERFACE databases for the moment. In the result synthesized speech can indeed be used for the recognition of human emotional speech. Björn W. Schuller, Felix Burkhardt |
ICASSP | 2 |
| 2010 | Automatic speaker age and gender recognition in the car for tailoring dialog and mobile servicesabstractCar manufacturers are faced with a new challenge. While a new generation of “digital natives ” becomes a new customer group, the problem of aging society is still increasing. This em-phasizes the need of providing flexible in-car dialog that take into account the specific needs and preferences of the respective user (group). Along the lines of this year’s Interspeech motto “Spoken Language Processing for All”, we address the ques-tion how we find out which group the current user belongs to. Michael Feld, Felix Burkhardt, Christian Müller 0014 |
INTERSPEECH | 2 |
| 2010 | The INTERSPEECH 2010 paralinguistic challengeabstractMost paralinguistic analysis tasks are lacking agreed-upon evaluation procedures and comparability, in contrast to more ‘traditional ’ disciplines in speech analysis. The INTERSPEECH 2010 Paralinguistic Challenge shall help overcome the usually low compatibility of results, by addressing three selected subchallenges. In the Age Sub-Challenge, the age of speakers has to be determined in four groups. In the Gender Sub-Challenge, a three-class classification task has to be solved and finally, the Affect Sub-Challenge asks for speakers ’ interest in ordinal representation. This paper introduces the conditions, the Challenge corpora “aGender ” and “TUM AVIC ” and standard feature sets that may be used. Further, baseline results are given. Björn W. Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller 0014, Shri Narayanan |
INTERSPEECH | 4 |
| 2010 | Voice attributes affecting likability perceptionabstractRatings of voices ’ likability were collected in two successive studies. A single scale seems to be sufficient for assessing such ratings. Based on limited but controlled data, spectral parame-ters as well as f0 and articulation rate correlate with the ratings obtained. An automatic classification confirms the relevance of spectral features for the perception of likability. As a sim-ple method of collecting more data for further studies, the sin-gle scale was validated within the bounds of the small data set. Both the spectral parameters and items from a comprehensive questionnaire indicate the relevance of timbre for the likability perception. Index Terms: likability, voice, timbre, speaking style 1. Benjamin Weiss 0001, Felix Burkhardt |
INTERSPEECH | 2 |
| 2010 | A Database of Age and Gender Annotated Telephone Speech
Felix Burkhardt, Martin Eckert, Wiebke Johannsen, Joachim Stegmann |
LREC | 1 |
| 2009 | Detecting real life angerabstractAcoustic anger detection in voice portals can help to enhance human computer interaction. A comprehensive voice portal data collection has been carried out and gives new insight on the nature of real life data. Manual labeling revealed a high percentage of non-classifiable data. Experiments with a statistical classifier indicate that, in contrast to pitch and energy related features, duration measures do not play an important role for this data while cepstral information does. Also in a direct comparison between Gaussian Mixture Models and Support Vector Machines the latter gave better results. Felix Burkhardt, Tim Polzehl, Joachim Stegmann, Florian Metze, Richard Huber |
ICASSP | 1 |
| 2009 | Rule-based voice quality variation with formant synthesisabstractWe describe an approach to simulate different phonation types, following John Laver’s terminology, by means of a hybrid (rule-based and unit concatenating) formant synthesizer. Different voice qualities were generated by following hints from the lit-erature and applying the revised KLGLOTT88 model. Within a listener perception experiment, we show that the phonation types get distinguished by the listeners and lead to emotional impression as predicted by literature. The synthesis system and its source code, as well as audio samples can be downloaded at Felix Burkhardt |
INTERSPEECH | 1 |
| 2008 | Age and gender recognition for telephone applications based on GMM supervectors and support vector machinesabstractThis paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans. Tobias Bocklet, Andreas K. Maier, Josef G. Bauer, Felix Burkhardt, Elmar Nöth |
ICASSP | 4 |
| 2007 | Comparison of Four Approaches to Age and Gender Recognition for Telephone ApplicationsabstractThis paper presents a comparative study of four different approaches to automatic age and gender classification using seven classes on a telephony speech task and also compares the results with human performance on the same data. The automatic approaches compared are based on (1) a parallel phone recognizer, derived from an automatic language identification system; (2) a system using dynamic Bayesian networks to combine several prosodic features; (3) a system based solely on linear prediction analysis; and (4) Gaussian mixture models based on MFCCs for separate recognition of age and gender. On average, the parallel phone recognizer performs as well as Human listeners do, while loosing performance on short utterances. The system based on prosodic features however shows very little dependence on the length of the utterance. Florian Metze, Jitendra Ajmera, Roman Englert, Udo Bub, Felix Burkhardt, Joachim Stegmann, Christian Müller 0014, Richard Huber, Bernt Andrassy, Josef G. Bauer, Bernhard Littel |
ICASSP (4) | 5 |
| 2007 | Combining short-term cepstral and long-term pitch features for automatic recognition of speaker ageabstractThe most successful systems in previous comparative studies on speaker age recognition used short-term cepstral features modeled with Gaussian Mixture Models (GMMs) or applied multiple phone recognizers trained with the data of speakers of the respective class. Acoustic analyses, however, indicate that certain features such as pitch extracted from a longer span of speech correlate clearly with the speaker age although the systems based on those features have been inferior to the before mentioned approaches. In this paper, three novel systems combining short-term cepstral features and long-term features for speaker age recognition are compared to each other. A system combining GMMs using frame-based MFCCs and SupportVector-Machines using long-term pitch performs best. The results indicate that the combination of the two feature types is a promising approach, which corresponds to findings in related fields like speaker recognition. Christian Müller 0014, Felix Burkhardt |
INTERSPEECH | 2 |
| 2006 | Detecting anger in automated voice portal dialogsabstractAnger detection is a topic that is gaining more and more attention with voice portal carriers, as it can be useful for quality measurement and emotion-aware dialog strategies. In the context of a prototype voice portal we describe methods to search for training data, report on the performance of the prosodic classifier under real world conditions and explore the use of dialog information for anger prediction. The results show that, although significantly worse than under laboratory conditions, anger detection still works well above chance level and can be used to enhance real world voice-portal usability. Index Terms: Emotion Recognition, Voice Portal, Speech Classification, Dialogue System. Felix Burkhardt, Jitendra Ajmera, Roman Englert, Joachim Stegmann, Winslow Burleson |
INTERSPEECH | 1 |
| 2005 | Emofilt: the simulation of emotional speech by prosody-transformationabstractEmofilt is a software program intended to simulate emotional arousal with speech synthesis based on the free-for-noncommercial-use MBROLA synthesis engine. It acts as a transformer between the phonetisation and the speech-generation component. Originally developed at the Technical University of Berlin it was recently revived as an open-source project written in Java (http://emofilt.sourceforge.net). Emofilt’s languagedependent modules are controlled by external XML-files and it is as multilingual as MBROLA which currently supports 35 languages. It might be used for research, teaching or to implement applications that include the simulation of emotional speech. Felix Burkhardt |
INTERSPEECH | 1 |
| 2005 | A database of German emotional speechabstractThe article describes a database of emotional speech. Ten actors (5 female and 5 male) simulated the emotions, producing 10 German utterances (5 short and 5 longer sentences) which could be used in everyday communication and are interpretable in all applied emotions. The recordings were taken in an anechoic chamber with high-quality recording equipment. In addition to the sound electro-glottograms were recorded. The speech material comprises about 800 sentences (seven emotions * ten actors * ten sentences + some second versions). The complete database was evaluated in a perception test regarding the recognisability of emotions and their naturalness. Utterances recognised better than 80 % and judged as natural by more than 60 % of the listeners were phonetically labelled in a narrow transcription with special markers for voice-quality, phonatory and articulatory settings and articulatory features. The database can be accessed by the public via the internet Felix Burkhardt, Astrid Paeschke, M. Rolfes, Walter F. Sendlmeier, Benjamin Weiss 0001 |
INTERSPEECH | 1 |