EDBT 2026 Demo / reviewers in the wild / expert
Anton Batliner
dblp:88/6105
· DBLP profile ↗
112ranked-venue papers
22as first author
13since 2021 · last 2024
0000-0002-0946-7497ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 86 · 16 first-author · 7 since 2021Artificial intelligence and machine learning · 84 · 16 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Sustained Vowels for Pre- vs Post-Treatment COPD ClassificationabstractChronic obstructive pulmonary disease (COPD) is a serious inflammatory lung disease affecting millions of people around the world.Due to an obstructed airflow from the lungs, it also becomes manifest in patients' vocal behaviour.Of particular importance is the detection of an exacerbation episode, which marks an acute phase and often requires hospitalisation and treatment.Previous work has shown that it is possible to distinguish between a pre-and a post-treatment state using automatic analysis of read speech.In this contribution, we examine whether sustained vowels can provide a complementary lens for telling apart these two states.Using a cohort of 50 patients, we show that the inclusion of sustained vowels can improve performance to up to 79% unweighted average recall, from a 71% baseline using read speech.We further identify and interpret the most important acoustic features that characterise the manifestation of COPD in sustained vowels. Andreas Triantafyllopoulos, Anton Batliner, Wolfgang Mayr, Markus Fendler, Florian B. Pokorny, Maurice Gerczuk, Shahin Amiriparian, Thomas M. Berghaus, Björn W. Schuller |
INTERSPEECH | 2 |
| 2024 | INTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition
Andreas Triantafyllopoulos, Anton Batliner, Simon David Noel Rampp, Manuel Milling, Björn W. Schuller |
INTERSPEECH | 2 |
| 2023 | The MASCFLICHT Corpus: Face Mask Type and Coverage Area Recognition from Speech
Adria Mallol-Ragolta, Nils Urbach, Shuo Liu 0012, Anton Batliner, Björn W. Schuller |
INTERSPEECH | 4 |
| 2023 | The ACM Multimedia 2023 Computational Paralinguistics Challenge: Emotion Share & RequestsabstractThe ACM Multimedia 2023 Computational Paralinguistics Challenge addresses two different problems for the first time in a research competition under well-defined conditions: In the Emotion Share Sub-Challenge, a regression on speech has to be made; and in the Requests Sub-Challenges, requests and complaints need to be detected. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' ComPaRE features, the auDeep toolkit, and deep feature extraction from pre-trained CNNs using the DeepSpectRum toolkit; in addition, wav2vec2 models are used. Björn W. Schuller, Anton Batliner, Shahin Amiriparian, Alexander Barnhill, Maurice Gerczuk, Andreas Triantafyllopoulos, Alice Baird, Panagiotis Tzirakis, Chris Gagne 0001, Alan Cowen, Nikola Lackovic, Marie-José Caraty, Claude Montacié |
ACM Multimedia | 2 |
| 2023 | Classification of stuttering - The ComParE challenge and beyond
Sebastian P. Bayerl, Maurice Gerczuk, Anton Batliner, Christian Bergler, Shahin Amiriparian, Björn W. Schuller, Elmar Nöth, Korbinian Riedhammer |
Comput. Speech Lang. | 3 |
| 2023 | Ethical Awareness in Paralinguistics: A Taxonomy of ApplicationsabstractSince the end of the last century, the automatic processing of paralinguistics has been investigated widely and put into practice in many applications, on wearables, smartphones, and computers. In this contribution, we address ethical awareness for paralinguistic applications, by establishing taxonomies for data representations, system designs for and a typology of applications, and users/test sets and subject areas. These are related to an “ethical grid” consisting of the most relevant ethical cornerstones, based on principalism. The characteristics of and the interdependencies between these taxonomies are described and exemplified. This makes it possible to assess more or less critical “ethical constellations.” To the best of our knowledge, this is the first attempt of its kind. Anton Batliner, Michael Neumann 0001, Felix Burkhardt, Alice Baird, Sarina Meyer, Ngoc Thang Vu, Björn W. Schuller |
Int. J. Hum. Comput. Interact. | 1 |
| 2022 | Distinguishing between pre- and post-treatment in the speech of patients with chronic obstructive pulmonary diseaseabstractChronic obstructive pulmonary disease (COPD) causes lung inflammation and airflow blockage leading to a variety of respiratory symptoms; it is also a leading cause of death and affects millions of individuals around the world. Patients often require treatment and hospitalisation, while no cure is currently available. As COPD predominantly affects the respiratory system, speech and non-linguistic vocalisations present a major avenue for measuring the effect of treatment. In this work, we present results on a new COPD dataset of 20 patients, showing that, by employing personalisation through speaker-level feature normalisation, we can distinguish between pre- and post-treatment speech with an unweighted average recall (UAR) of up to 82% in (nested) leave-one-speaker-out cross-validation. We further identify the most important features and link them to pathological voice properties, thus enabling an auditory interpretation of treatment effects. Monitoring tools based on such approaches may help objectivise the clinical status of COPD patients and facilitate personalised treatment plans. Andreas Triantafyllopoulos, Markus Fendler, Anton Batliner, Maurice Gerczuk, Shahin Amiriparian, Thomas M. Berghaus, Björn W. Schuller |
INTERSPEECH | 3 |
| 2022 | The ACM Multimedia 2022 Computational Paralinguistics Challenge: Vocalisations, Stuttering, Activity, & MosquitoesabstractThe ACM Multimedia 2022 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Vocalisations and Stuttering Sub-Challenges, a classification on human non-verbal vocalisations and speech has to be made; the Activity Sub-Challenge aims at beyond-audio human activity recognition from smartwatch sensor data; and in the Mosquitoes Sub-Challenge, mosquitoes need to be detected. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' ComParE and BoAW features, the auDeep toolkit, and deep feature extraction from pre-trained CNNs using the DeepSpectrum toolkit; in addition, we add end-to-end sequential modelling, and a log-mel-128-BNN. Björn W. Schuller, Anton Batliner, Shahin Amiriparian, Christian Bergler, Maurice Gerczuk, Natalie Holz, Pauline Larrouy-Maestri, Sebastian P. Bayerl, Korbinian Riedhammer, Adria Mallol-Ragolta, Maria Pateraki, Harry Coppock, Ivan Kiskin, Marianne Sinka, Stephen J. Roberts |
ACM Multimedia | 2 |
| 2022 | The phonetic footprint of Parkinson's disease
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Anton Batliner, Elmar Nöth |
Comput. Speech Lang. | 6 |
| 2022 | AI-Based human audio processing for COVID-19: A comprehensive overview
Gauri Deshpande, Anton Batliner, Björn W. Schuller |
Pattern Recognit. | 2 |
| 2022 | Face mask recognition from audio: The MASC database and an overview on the mask challenge
Mostafa M. Mohamed, Mina A. Nessiem, Anton Batliner, Christian Bergler, Simone Hantke, Maximilian Schmitt, Alice Baird, Adria Mallol-Ragolta, Vincent Karas, Shahin Amiriparian, Björn W. Schuller |
Pattern Recognit. | 3 |
| 2022 | Ethics and Good Practice in Computational ParalinguisticsabstractWith the advent of ‘heavy Artificial Intelligence’ – big data, deep learning, and ubiquitous use of the internet, ethical considerations are widely dealt with in public discussions and governmental bodies. Within Computational Paralinguistics with its manifold topics and possible applications (modelling of long-term, medium-term, and short-term traits and states such as personality, emotion, or speech pathology), we have not yet seen that many contributions. In this article, we try to set the scene by (1) giving a short overview of ethics and privacy, (2) describing the field of Computational Paralinguistics, its history and exemplary use cases, as well as (de-)anonymisation and peculiarities of speech and text data, and (3) proposing rules for good practice in the field, such as choosing the right performance measure, and accounting for representativity and interpretability. Anton Batliner, Simone Hantke, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & PrimatesabstractThe INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the COVID-19 Cough and COVID-19 Speech Sub-Challenges, a binary classification on COVID-19 infection has to be made based on coughing sounds and speech; in the Escalation SubChallenge, a three-way assessment of the level of escalation in a dialogue is featured; and in the Primates Sub-Challenge, four species vs background need to be classified. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AuDeep toolkit, and deep feature extraction from pre-trained CNNs using the Deep Spectrum toolkit; in addition, we add deep end-to-end sequential modelling, and partially linguistic analysis. Björn W. Schuller, Anton Batliner, Christian Bergler, Cecilia Mascolo, Jing Han 0010, Iulia Lefter, Heysem Kaya, Shahin Amiriparian, Alice Baird, Lukas Stappen, Sandra Ottl, Maurice Gerczuk, Panagiotis Tzirakis, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Léon J. M. Rothkrantz, Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp |
Interspeech | 2 |
| 2020 | Deep Attentive End-to-End Continuous Breath Sensing from SpeechabstractModelling of the breath signal is of high interest to both \nhealthcare professionals and computer scientists, as a source \nof diagnosis-related information, or a means for curating higher \nquality datasets in speech analysis research. The formation of \na breath signal gold standard is, however, not a straightforward \ntask, as it requires specialised equipment, human annotation \nbudget, and even then, it corresponds to lab recording settings, \nthat are not reproducible in-the-wild. Herein, we explore deep \nlearning based methodologies, as an automatic way to predict a \ncontinuous-time breath signal by solely analysing spontaneous \nspeech. We address two task formulations, those of continuousvalued signal prediction, as well as inhalation event prediction, \nthat are of great use in various healthcare and Automatic Speech \nRecognition applications, and showcase results that outperform \ncurrent baselines. Most importantly, we also perform an initial \nexploration into explaining which parts of the input audio signal \nare important with respect to the prediction. Alexis Deighton MacIntyre, Georgios Rizos, Anton Batliner, Alice Baird, Shahin Amiriparian, Antonia F. de C. Hamilton, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly Emotion, Breathing & MasksabstractThe INTERSPEECH 2020 Computational Paralinguistics Challenge addresses three different problems for the first time in a research competition under well-defined conditions: In the Elderly Emotion Sub-Challenge, arousal and valence in the speech of elderly individuals have to be modelled as a 3-class problem; in the Breathing Sub-Challenge, breathing has to be assessed as a regression problem; and in the Mask Sub-Challenge, speech without and with a surgical mask has to be told apart.We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AUDEEP toolkit, and deep feature extraction from pre-trained CNNs using the DEEP SPECTRUM toolkit; in addition, we partially add deep end-to-end sequential modelling, and, for the first time in the challenge, linguistic analysis. Björn W. Schuller, Anton Batliner, Christian Bergler, Eva-Maria Messner, Antonia F. de C. Hamilton, Shahin Amiriparian, Alice Baird, Georgios Rizos, Maximilian Schmitt, Lukas Stappen, Harald Baumeister, Alexis Deighton MacIntyre, Simone Hantke |
INTERSPEECH | 2 |
| 2019 | The INTERSPEECH 2019 Computational Paralinguistics Challenge: Styrian Dialects, Continuous Sleepiness, Baby Sounds & Orca ActivityabstractThe INTERSPEECH 2019 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Styrian Dialects Sub-Challenge, three types of Austrian-German dialects have to be classified; in the Continuous Sleepiness Sub-Challenge, the sleepiness of a speaker has to be assessed as regression problem; in the Baby Sound Sub-Challenge, five types of infant sounds have to be classified; and in the Orca Activity Sub-Challenge, orca sounds have to be detected.We describe the Sub-Challenges and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit. Björn W. Schuller, Anton Batliner, Christian Bergler, Florian B. Pokorny, Jarek Krajewski, Margaret Cychosz, Ralf Vollmann, Sonja-Dana Roelen, Sebastian Schnieder, Elika Bergelson, Alejandrina Cristià, Amanda Seidl, Anne S. Warlaumont, Lisa Yankowitz, Elmar Nöth, Shahin Amiriparian, Simone Hantke, Maximilian Schmitt |
INTERSPEECH | 2 |
| 2019 | Affective and behavioural computing: Lessons learnt from the First Computational Paralinguistics Challenge
Björn W. Schuller, Felix Weninger, Yue Zhang 0014, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli, Klaus R. Scherer, Mohamed Chetouani, Marcello Mortillaro |
Comput. Speech Lang. | 5 |
| 2018 | Categorical vs Dimensional Perception of Italian Emotional SpeechabstractCulture and measurement strategies are influential factors when evaluating the perception of emotion in speech.However, multilingual databases suitable for such a study are missing, and there is no agreement on the most suitable emotional model.To address this gap, we present EmoFilm, a new multilingual emotional speech corpus, consisting of 1115 English, Spanish, and Italian emotional utterances extracted from 43 films and 207 speakers.We have performed a within-culture categorical vs dimensional perceptual evaluation, employing 225 native Italian listeners, who evaluated the Italian section of the database with the emotional states of anger, sadness, happiness, fear, and contempt.The aim of this study is to assess whether the emotional model (categorical or dimensional), taken as reference for measurement, influences a listener's perception of emotional speech, and-to what extent-both models are complementary or not.We show that the measurement strategy chosen does influence a listener's response, especially for some emotions, e. g., contempt.The confusion patterns typical of a categorical evaluation are not always mirrored by the dimensional assessment. Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, Björn W. Schuller |
INTERSPEECH | 3 |
| 2018 | The INTERSPEECH 2018 Computational Paralinguistics Challenge: Atypical & Self-Assessed Affect, Crying & Heart BeatsabstractThe INTERSPEECH 2018 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Atypical Affect Sub-Challenge, four basic emotions annotated in the speech of handicapped subjects have to be classified; in the Self-Assessed Affect Sub-Challenge, valence scores given by the speakers themselves are used for a three-class classification problem; in the Crying Sub-Challenge, three types of infant vocalisations have to be told apart; and in the Heart Beats Sub-Challenge, three different types of heart beats have to be determined.We describe the Sub-Challenges, their conditions, and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by end-to-end learning, the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit for the first time in the challenge series. Björn W. Schuller, Stefan Steidl, Anton Batliner, Peter B. Marschik, Harald Baumeister, Fengquan Dong, Simone Hantke, Florian B. Pokorny, Eva-Maria Rathner, Katrin D. Bartl-Pokorny, Christa Einspieler, Dajie Zhang, Alice Baird, Shahin Amiriparian, Kun Qian 0003, Zhao Ren, Maximilian Schmitt, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 3 |
| 2017 | Automatic Classification of Autistic Child Vocalisations: A Novel Database and ResultsabstractHumanoid robots have in recent years shown great promise for supporting the educational needs of children on the autism spectrum.To further improve the efficacy of such interactions, user-adaptation strategies based on the individual needs of a child are required.In this regard, the proposed study assesses the suitability of a range of speech-based classification approaches for automatic detection of autism severity according to the commonly used Social Responsiveness Scale™ second edition (SRS-2).Autism is characterised by socialisation limitations including child language and communication ability.When compared to neurotypical children of the same age these can be a strong indication of severity.This study introduces a novel dataset of 803 utterances recorded from 14 autistic children aged between 4 -10 years, during Wizard-of-Oz interactions with a humanoid robot.Our results demonstrate the suitability of support vector machines (SVMs) which use acoustic feature sets from multiple Interspeech COMPARE challenges.We also evaluate deep spectrum features, extracted via an image classification convolutional neural network (CNN) from the spectrogram of autistic speech instances.At best, by using SVMs on the acoustic feature sets, we achieved a UAR of 73.7 % for the proposed 3-class task. Alice Baird, Shahin Amiriparian, Nicholas Cummins, Alyssa Alcorn, Anton Batliner, Sergey Pugachevskiy, Michael Freitag 0003, Maurice Gerczuk, Björn W. Schuller |
INTERSPEECH | 5 |
| 2017 | Description of the Munich-Passau Snore Sound Corpus (MPSSC)
Christoph Janott, Anton Batliner |
INTERSPEECH | 2 |
| 2017 | Description of the Upper Respiratory Tract Infection Corpus (URTIC)
Jarek Krajewski, Sebastian Schnieder, Anton Batliner |
INTERSPEECH | 3 |
| 2017 | The Perception of Emotions in Noisified Nonsense SpeechabstractNoise pollution is part of our daily life, affecting millions of people, particularly those living in urban environments.Noise alters our perception and decreases our ability to understand others.Considering this, speech perception in background noise has been extensively studied, showing that especially white noise can damage listener perception.However, the perception of emotions in noisified speech has not been explored with as much depth.In the present study, we use artificial background noise conditions, by applying noise to a subset of the GEMEP corpus (emotions expressed in nonsense speech).Noises were at varying intensities and 'colours'; white, pink, and brownian.The categorical and dimensional perceptual test was completed by 26 listeners.The results indicate that background noise conditions influence the perception of emotion in speechpink noise most, brownian least.Worsened perception invokes higher confusion, especially with sadness, an emotion with less pronounced prosodic characteristics.Yet, all this does not lead to a break-down of the 'cognitive-emotional space' in a Nonmetric MultiDimensional Scaling representation.The gender of speakers and the cultural background of listeners do not seem to play a role. Emilia Parada-Cabaleiro, Alice Baird, Anton Batliner, Nicholas Cummins, Simone Hantke, Björn W. Schuller |
INTERSPEECH | 3 |
| 2017 | Discussion
Björn W. Schuller, Anton Batliner |
INTERSPEECH | 2 |
| 2017 | The INTERSPEECH 2017 Computational Paralinguistics Challenge: Addressee, Cold & SnoringabstractThe INTERSPEECH 2017 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: In the Addressee sub-challenge, it has to be determined whether speech produced by an adult is directed towards another adult or towards a child; in the Cold sub-challenge, speech under cold has to be told apart from ‘healthy’ speech; and in the Snoring subchallenge, four different types of snoring have to be classified. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, which include data-learnt feature representations by end-to-end learning with convolutional and recurrent neural networks, and bag-of-audiowords for the first time in the challenge series Björn W. Schuller, Stefan Steidl, Anton Batliner, Elika Bergelson, Jarek Krajewski, Christoph Janott, Andrei Amatuni, Marisa Casillas, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont, Guillermo Hidalgo, Sebastian Schnieder, Clemens Heiser, Winfried Hohenhorst, Michael Herzog, Maximilian Schmitt, Kun Qian 0003, Yue Zhang 0014, George Trigeorgis, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 3 |
| 2017 | An Image-based Deep Spectrum Feature Representation for the Recognition of Emotional SpeechabstractThe outputs of the higher layers of deep pre-trained convolutional neural networks (CNNs) have consistently been shown to provide a rich representation of an image for use in recognition tasks. This study explores the suitability of such an approach for speech-based emotion recognition tasks. First, we detail a new acoustic feature representation, denoted as deep spectrum features, derived from feeding spectrograms through a very deep image classification CNN and forming a feature vector from the activations of the last fully connected layer. We then compare the performance of our novel features with standardised brute-force and bag-of-audio-words (BoAW) acoustic feature representations for 2- and 5-class speech-based emotion recognition in clean, noisy and denoised conditions. The presented results show that image-based approaches are a promising avenue of research for speech-based recognition tasks. Key results indicate that deep-spectrum features are comparable in performance with the other tested acoustic feature representations in matched for noise type train-test conditions; however, the BoAW paradigm is better suited to cross-noise-type train-test conditions. Nicholas Cummins, Shahin Amiriparian, Gerhard Hagerer, Anton Batliner, Stefan Steidl, Björn W. Schuller |
ACM Multimedia | 4 |
| 2016 | Language proficiency assessment of English L2 speakers based on joint analysis of prosody and native languageabstractIn this work, we present an in-depth analysis of the interdependency between the non-native prosody and the native language (L1) of English L2 speakers, as separately investigated in the Degree of Nativeness Task and the Native Language Task of the INTERSPEECH 2015 and 2016 Computational Paralinguistics ChallengE (ComParE). To this end, we propose a multi-task learning scheme based on auxiliary attributes for jointly learning the tasks of L1 classification and prosody score regression. The effectiveness of this scheme is demonstrated in extensive experimental runs, comparing various standardised feature sets of prosodic, cepstral, spectral, and voice quality descriptors, as well as automatic feature selection. In the result, we show that the prediction of both prosody score and L1 can be improved by considering both tasks in a holistic way. In particular, we achieve an 11% relative gain in regression performance (Spearman's correlation coefficient) on prosody scores, when comparing the best multi- and single-task learning results. Yue Zhang 0014, Felix Weninger, Anton Batliner, Florian Hönig, Björn W. Schuller |
ICMI | 3 |
| 2016 | Combining Semantic Word Classes and Sub-Word Unit Speech Recognition for Robust OOV DetectionabstractOut-of-vocabulary words (OOVs) are often the main reason for the failure of tasks like automated voice searches or humanmachine dialogs.This is especially true if rare but task-relevant content words, e.g.person or location names, are not in the recognizer's vocabulary.Since applications like spoken dialog systems use the result of the speech recognizer to extract a semantic representation of a user utterance, the detection of OOVs as well as their (semantic) word class can support to manage a dialog successfully.In this paper we suggest to combine two wellknown approaches in the context of OOV detection: semantic word classes and OOV models based on sub-word units.With our system, which builds upon the widely used Kaldi speech recognition toolkit, we show on two different data sets that -compared to other methods -such a combination improves OOV detection performance for open word classes at a given false alarm rate.Another result of our approach is a reduction of the word error rate (WER). Axel Horndasch, Anton Batliner, Caroline Kaufhold, Elmar Nöth |
INTERSPEECH | 2 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: Deception, Sincerity & Native LanguageabstractThe INTERSPEECH 2016 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: classification of deceptive vs. non-deceptive speech, the estimation of the degree of sincerity, and the identification of the native language out of eleven L1 classes of English L2 speakers.In this paper, we describe these sub-challenges, their conditions, the baseline feature extraction and classifiers, and the resulting baselines, as provided to the participants. Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | The Deception Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | The Sincerity Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | The Native Language Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: A Summary of Results
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | Discussion
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 3 |
| 2016 | Assessing the Prosody of Non-Native Speakers of English: Measures and Feature Sets
Eduardo Coutinho, Florian Hönig, Yue Zhang 0014, Simone Hantke, Anton Batliner, Elmar Nöth, Björn W. Schuller |
LREC | 5 |
| 2015 | The eating condition sub-challenge: the data
Anton Batliner |
INTERSPEECH | 1 |
| 2015 | Wrapping up: the story of the compare challenges, what we learned and where to go
Anton Batliner |
INTERSPEECH | 1 |
| 2015 | The INTERSPEECH 2015 computational paralinguistics challenge: nativeness, parkinson's & eating conditionabstractThe INTERSPEECH 2015 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: the estimation of the degree of nativeness, the neurological state of patients with Parkinson’s condition, and the eating conditions of speakers, i. e., whether and which food type they are eating in a seven-class problem. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, as provided to the participants. Index Terms: Computational Paralinguistics, Challenge, Degree of Nativeness, Parkinson’s Condition, Eating Condition Björn W. Schuller, Stefan Steidl, Anton Batliner, Simone Hantke, Florian Hönig, Juan Rafael Orozco-Arroyave, Elmar Nöth, Yue Zhang 0014, Felix Weninger |
INTERSPEECH | 3 |
| 2015 | Visual comparison of speaker groups
Sebastian Wankerl, Florian Hönig, Anton Batliner, Juan Rafael Orozco-Arroyave, Elmar Nöth |
INTERSPEECH | 3 |
| 2015 | A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
Comput. Speech Lang. | 3 |
| 2015 | Introduction
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son |
Comput. Speech Lang. | 3 |
| 2014 | Are men more sleepy than women or does it only look like - Automatic analysis of sleepy speechabstractThe degree of sleepiness in the Sleepy Language Corpus from the Interspeech 2011 Speaker State Challenge is predicted with regression and a very large feature vector. Most notable is the great gender difference which can mainly be attributed to females showing their sleepiness less than males do. Florian Hönig, Anton Batliner, Tobias Bocklet, Georg Stemmer, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski |
ICASSP | 2 |
| 2014 | Automatic modelling of depressed speech: relevant features and relevance of genderabstractDepression is an affective disorder characterised by psychomotor retardation; in speech, this shows up in reduction of pitch (variation, range), loudness, and tempo, and in voice qualities different from those of typical modal speech.A similar reduction can be observed in sleepy speech (relaxation).In this paper, we employ a small group of acoustic features modelling prosody and spectrum that have been proven successful in the modelling of sleepy speech, enriched with voice quality features, for the modelling of depressed speech within a regression approach.This knowledge-based approach is complemented by and compared with brute-forcing and automatic feature selection.We further discuss gender differences and the contributions of (groups of) features both for the modelling of depression and across depression and sleepiness. Florian Hönig, Anton Batliner, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski |
INTERSPEECH | 2 |
| 2014 | The INTERSPEECH 2014 computational paralinguistics challenge: cognitive & physical loadabstractThe INTERSPEECH 2014 Computational Paralinguistics Challenge provides for the first time a unified test-bed for the automatic recognition of speakers’ cognitive and physical load in speech. In this paper, we describe these two Sub-Challenges, their conditions, baseline results and experimental procedures, as well as the COMPARE baseline features generated with the openSMILE toolkit and provided to the participants in the Challenge. Björn W. Schuller, Stefan Steidl, Anton Batliner, Julien Epps, Florian Eyben, Fabien Ringeval, Erik Marchi, Yue Zhang 0014 |
INTERSPEECH | 3 |
| 2014 | Introduction to the Special Issue on Broadening the View on Speaker Analysis
Björn W. Schuller, Stefan Steidl, Anton Batliner, Florian Schiel, Jarek Krajewski |
Comput. Speech Lang. | 3 |
| 2014 | Medium-term speaker states - A review on intoxication, sleepiness and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Florian Schiel, Jarek Krajewski, Felix Weninger, Florian Eyben |
Comput. Speech Lang. | 3 |
| 2013 | The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autismabstractInternational audience Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim |
INTERSPEECH | 3 |
| 2013 | Introduction to the special issue on Paralinguistics in Naturalistic Speech and Language
Björn W. Schuller, Stefan Steidl, Anton Batliner |
Comput. Speech Lang. | 3 |
| 2013 | Paralinguistics in speech and language - State-of-the-art and the challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller 0014, Shri Narayanan |
Comput. Speech Lang. | 3 |
| 2012 | The Automatic Assessment of Non-native Prosody: Combining Classical Prosodic Analysis with Acoustic ModellingabstractIn earlier studies, we employed a large prosodic feature vector to assess the quality of L2 learner's utterances with respect to sentence melody and rhythm.In this paper, we combine these features with two standard approaches in paralinguistic analysis: (1) features derived from a Gaussian Mixture Model used as Universal Background Model (GMM-UBM), and (2) openSMILE, an open-source toolkit for extracting acoustic features.We evaluate our approach with English speech from 94 non-native speakers perceptually scored by 62 native labellers.GMM-UBM or openSMILE modelling alone yields lower performance than our prosodic feature vector; however, adding information from the GMM-UBM modelling or openSMILE by late fusion improves results. Florian Hönig, Tobias Bocklet, Korbinian Riedhammer, Anton Batliner, Elmar Nöth |
INTERSPEECH | 4 |
| 2012 | The INTERSPEECH 2012 Speaker Trait ChallengeabstractLIDIAP Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
INTERSPEECH | 3 |
| 2012 | Applying multiple classifiers and non-linear dynamics features for detecting sleepiness from speech
Jarek Krajewski, Sebastian Schnieder, David Sommer, Anton Batliner, Björn W. Schuller |
Neurocomputing | 4 |
| 2012 | Guest Editorial: Special Section on Naturalistic Affect Resources for System Building and EvaluationabstractThe papers in this special section focus on the deployment of naturalistic affect resources for systems design and analysis. Björn W. Schuller, Ellen Douglas-Cowie, Anton Batliner |
IEEE Trans. Affect. Comput. | 3 |
| 2012 | The Voice of Leadership: Models and Performances of Automatic Analysis in Online SpeechesabstractWe introduce the automatic determination of leadership emergence by acoustic and linguistic features in online speeches. Full realism is provided by the varying and challenging acoustic conditions of the presented YouTube corpus of online available speeches labeled by 10 raters and by processing that includes Long Short-Term Memory-based robust voice activity detection (VAD) and automatic speech recognition (ASR) prior to feature extraction. We discuss cluster-preserving scaling of 10 original dimensions for discrete and continuous task modeling, ground truth establishment, and appropriate feature extraction for this novel speaker trait analysis paradigm. In extensive classification and regression runs, different temporal chunkings and optimal late fusion strategies (LFSs) of feature streams are presented. In the result, achievers, charismatic speakers, and teamplayers can be recognized significantly above chance level, reaching up to 72.5 percent accuracy on unseen test data. Felix Weninger, Jarek Krajewski, Anton Batliner, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2011 | Associating children's non-verbal and verbal behaviour: Body movements, emotions, and laughter in a human-robot interactionabstractIn this article, we associate different types of vocal behaviour denoting emotional user states and laughter with different types of body movements such as gestures, forward bends, or liveliness. Our subjects are German children giving commands to Sony's Aibo robot; the data are fully realistic. The analysis reveals characteristic and significant co-occurrences of body movements and vocal events. Anton Batliner, Stefan Steidl, Elmar Nöth |
ICASSP | 1 |
| 2011 | Does it Groove or does it Stumble - Automatic Classification of Alcoholic Intoxication using Prosodic FeaturesabstractThis paper studies how prosodic features can help in the automatic detection of alcoholic intoxication.We compute features that have recently been proposed to model speech rhythm such as the pair-wise variability index for consonantal and vocalic segments (PVI) and study their aptness for the task.Further, we use a large prosodic feature vector modelling the usual candidates -pitch, intensity, and duration -and apply it onto different units such as words, syllables and stressed syllables to create generalizations of the rhythm features mentioned.The results show that the prosodic features computed are suitable for detecting alcoholic intoxication and add complementary information to state-of-the-art features.The database is the intoxication database provided by the organizers of the 2011 Interspeech Speaker State Challenge. Florian Hönig, Anton Batliner, Elmar Nöth |
INTERSPEECH | 2 |
| 2011 | The INTERSPEECH 2011 Speaker State ChallengeabstractWhile the first open comparative challenges in the field of paralinguistics targeted more 'conventional' phenomena such as emotion, age, and gender, there still exists a multiplicity of not yet covered, but highly relevant speaker states and traits.The INTERSPEECH 2011 Speaker State Challenge thus addresses two new sub-challenges to overcome the usually low compatibility of results: In the Intoxication Sub-Challenge, alcoholisation of speakers has to be determined in two classes; in the Sleepiness Sub-Challenge, another two-class classification task has to be solved.This paper introduces the conditions, the Challenge corpora "Alcohol Language Corpus" and "Sleepy Language Corpus", and a standard feature set that may be used.Further, baseline results are given. Björn W. Schuller, Stefan Steidl, Anton Batliner, Florian Schiel, Jarek Krajewski |
INTERSPEECH | 3 |
| 2011 | Speech-Based Non-Prototypical Affect Recognition for Child-Robot Interaction in Reverberated EnvironmentsabstractWe present a study on the effect of reverberation on acousticlinguistic recognition of non-prototypical emotions during child-robot interaction. Investigating the well-defined Interspeech 2009 Emotion Challenge task of recognizing negative emotions in children’s speech, we focus on the impact of artificial and real reverberation conditions on the quality of linguistic features and on emotion recognition accuracy. To maintain acceptable recognition performance of both, spoken content and affective state, we consider matched and multi-condition training and apply our novel multi-stream automatic speech recognition system which outperforms conventional Hidden Markov Modeling. Depending on the acoustic condition, we obtain unweighted emotion recognition accuracies of between 65.4 % and 70.3 % applying our multi-stream system in combination with the SimpleLogistic algorithm for joint acoustic-linguistic analysis. Index Terms: child-robot interaction, affective computing, acoustic-linguistic emotion recognition, reverberation Martin Wöllmer, Felix Weninger, Stefan Steidl, Anton Batliner, Björn W. Schuller |
INTERSPEECH | 4 |
| 2011 | Whodunnit - Searching for the most important feature types signalling emotion-related user states in speech
Anton Batliner, Stefan Steidl, Björn W. Schuller, Dino Seppi, Thurid Vogt, Johannes Wagner 0001, Laurence Devillers, Laurence Vidrascu, Vered Aharonson, Loïc Kessous, Noam Amir |
Comput. Speech Lang. | 1 |
| 2011 | Introduction to the special issue on sensing emotion and affect - Facing realism in speech processing
Björn W. Schuller, Anton Batliner, Stefan Steidl |
Speech Commun. | 2 |
| 2011 | Recognising realistic emotions and affect in speech: State of the art and lessons learnt from the first challenge
Björn W. Schuller, Anton Batliner, Stefan Steidl, Dino Seppi |
Speech Commun. | 2 |
| 2010 | Late fusion of individual engines for improved recognition of negative emotion in speech - learning vs. democratic voteabstractThe fusion of multiple recognition engines is known to be able to outperform individual ones, given sufficient independence of methods, models, and knowledge sources. We therefore investigate late fusion of different speech-based recognizers of emotion. Two generally different streams of information are considered: acoustics and linguistics fed by state-of-the-art automatic speech recognition. A total of five emotion recognition engines from different sites that provide heterogeneous output information are integrated by either simple democratic vote or learning `which predictor to trust when'. We are able to significantly outperform the best individual engine by fusion, and the so far best reported result on the recently introduced Emotion Challenge task. Björn W. Schuller, Florian Metze, Stefan Steidl, Anton Batliner, Florian Eyben, Tim Polzehl |
ICASSP | 4 |
| 2010 | Comparing Multiple Classifiers for Speech-Based Detection of Self-Confidence - A Pilot StudyabstractThe aim of this study is to compare several classifiers commonly used within the field of speech emotion recognition (SER) on the speech based detection of self-confidence. A standard acoustic feature set was computed, resulting in 170 features per one-minute speech sample (e.g. fundamental frequency, intensity, formants, MFCCs). In order to identify speech correlates of self-confidence, the lectures of 14 female participants were recorded, resulting in 306 one-minute segments of speech. Five expert raters independently assessed the self-confidence impression. Several classification models (e.g. Random Forest, Support Vector Machine, Naïve Bayes, Multi-Layer Perceptron) and ensemble classifiers (AdaBoost, Bagging, Stacking) were trained. AdaBoost procedures turned out to achieve best performance, both for single models (AdaBoost LR: 75.2% class-wise averaged recognition rate) and for average boosting (59.3%) within speaker-independent settings. Jarek Krajewski, Anton Batliner, Silke Kessel |
ICPR | 2 |
| 2010 | Emotion recognition using imperfect speech recognitionabstractThis paper investigates the use of speech-to-text methods for assigning an emotion class to a given speech utterance. Previous work shows that an emotion extracted from text can convey complementary evidence to the information extracted by classifiers based on spectral, or other non-linguistic features. As speech-to-text usually presents significantly more computational effort, in this study we investigate the degree of speech-to-text accuracy needed for reliable detection of emotions from an automatically generated transcription of an utterance. We evaluate the use of hypotheses in both training and testing, and compare several classification approaches on the same task. Our results show that emotion recognition performance stays roughly constant as long as word accuracy doesn't fall below a reasonable value, making the use of speech-to-text viable for training of emotion classifiers based on linguistics. Florian Metze, Anton Batliner, Florian Eyben, Tim Polzehl, Björn W. Schuller, Stefan Steidl |
INTERSPEECH | 2 |
| 2010 | The INTERSPEECH 2010 paralinguistic challengeabstractMost paralinguistic analysis tasks are lacking agreed-upon evaluation procedures and comparability, in contrast to more ‘traditional ’ disciplines in speech analysis. The INTERSPEECH 2010 Paralinguistic Challenge shall help overcome the usually low compatibility of results, by addressing three selected subchallenges. In the Age Sub-Challenge, the age of speakers has to be determined in four groups. In the Gender Sub-Challenge, a three-class classification task has to be solved and finally, the Affect Sub-Challenge asks for speakers ’ interest in ordinal representation. This paper introduces the conditions, the Challenge corpora “aGender ” and “TUM AVIC ” and standard feature sets that may be used. Further, baseline results are given. Björn W. Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller 0014, Shri Narayanan |
INTERSPEECH | 3 |
| 2009 | Emotion recognition from speech: Putting ASR in the loopabstractThis paper investigates the automatic recognition of emotion from spoken words by vector space modeling vs. string kernels which have not been investigated in this respect, yet. Apart from the spoken content directly, we integrate part-of-speech and higher semantic tagging in our analyses. As opposed to most works in the field, we evaluate the performance with an ASR engine in the loop. Extensive experiments are run on the FAU Aibo Emotion Corpus of 4 k spontaneous emotional child-robot interactions and show surprisingly low performance degradation with real ASR over transcription-based emotion recognition. In the result, bag of words dominate over all other modeling forms based on the spoken content. Björn W. Schuller, Anton Batliner, Stefan Steidl, Dino Seppi |
ICASSP | 2 |
| 2009 | A language-independent feature set for the automatic evaluation of prosodyabstractIn second language learning, the correct use of prosody plays a vital role.Therefore, an automatic method to evaluate the naturalness of the prosody of a speaker is desirable.We present a novel method to model prosody independently of the text and thus independently of the language as well.For this purpose, the voiced and unvoiced speech segments are extracted and a 187-dimensional feature vector is computed for each voiced segment.This approach is compared to word based prosodic features on a German text passage.Both are confronted with the perceptive evaluation of two native speakers of German.The word-based feature set yielded correlations of up to 0.92 while the text-independent feature set yielded 0.88.This is in the same range as the inter-rater correlation with 0.88.Furthermore, the text-independent features were computed for a Japanese translation of the passage which was also rated by two native speakers of Japanese.Again, the correlation between the automatic system and the human perception of the naturalness was high with 0.83 and not significantly lower than the inter-rater correlation of 0.92. Andreas K. Maier, Florian Hönig, Viktor Zeißler, Anton Batliner, Erik Körner, Nobuyuki Yamanaka, Peter Ackermann, Elmar Nöth |
INTERSPEECH | 4 |
| 2009 | The INTERSPEECH 2009 emotion challengeabstractThe last decade has seen a substantial body of literature on the recognition of emotion from speech. However, in comparison to related speech processing tasks such as Automatic Speech and Speaker Recognition, practically no standardised corpora and test-conditions exist to compare performances under exactly the same conditions. Instead a multiplicity of evaluation strategies employed – such as cross-validation or percentage splits without proper instance definition – prevents exact reproducibility. Further, in order to face more realistic scenarios, the community is in desperate need of more spontaneous and less prototypical data. This INTERSPEECH 2009 Emotion Challenge aims at bridging such gaps between excellent research on human emotion recognition from speech and low compatibility of results. The FAU Aibo Emotion Corpus [1] serves as basis with clearly defined test and training partitions incorporating speaker independence and different room acoustics as needed in most reallife settings. This paper introduces the challenge, the corpus, the features, and benchmark results of two popular approaches towards emotion recognition from speech. Index Terms: emotion, challenge, feature types, classification 1. Björn W. Schuller, Stefan Steidl, Anton Batliner |
INTERSPEECH | 3 |
| 2009 | PEAKS - A system for the automatic evaluation of voice and speech disorders
Andreas K. Maier, Tino Haderlein, Ulrich Eysholdt, Frank Rosanowski, Anton Batliner, Maria Schuster, Elmar Nöth |
Speech Commun. | 5 |
| 2008 | Mothers, adults, children, pets - towards the acoustics of intimacyabstractIn this paper, we investigate acoustic features which differentiate the two speech registers neutral and intimate within different constellations of speakers and addressees. Three different types of speakers are considered: mothers addressing their own children or an unknown adult, women with no children addressing an imaginary child or an imaginary adult, and children addressing a pet robot using both intimate and neutral speech. We use a large, systematically generated feature vector, upsampling, and SVM and RF for learning. Results are reported for extensive test-runs facing speaker- independency and using PCA-SFFS vs. SVM-SFFS for feature ranking. Classification performance and most relevant feature types are discussed in detail. Anton Batliner, Björn W. Schuller, Sonja Schaeffler, Stefan Steidl |
ICASSP | 1 |
| 2008 | An Acoustic Framework for Detecting Fatigue in Speech Based Human-Computer-Interaction
Jarek Krajewski, Rainer Wieland, Anton Batliner |
ICCHP | 3 |
| 2008 | Multiple classifier applied on predicting microsleep from speechabstractThe aim of this study is to apply a state-of-the-art speech emotion recognition engine on the detection of microsleep endangered sleepiness states. Current approaches in speech emotion recognition use low-level descriptors and functionals to compute brute-force feature sets. This paper describes a further enrichment of the temporal information, aggregating functionals and utilizing a broad pool of diverse elementary statistics and spectral descriptors. The resulting 45,088 features were applied to speech samples gained from a car simulator based sleep deprivation study. After a correlation-filter based feature subset selection, which was employed on the feature space in an attempt to maximize relevance, several classification models were trained. The best model (Support Vector Machine, dot kernel) achieved 86.1% recognition rate in predicting microsleep endangered sleepiness stages. Jarek Krajewski, Anton Batliner, Rainer Wieland |
ICPR | 2 |
| 2008 | Patterns, prototypes, performance: classifying emotional user statesabstractIn this paper, we report on classification results for emotional user states (4 classes, German database of children interacting with a pet robot).Starting with 5 emotion labels per word, we obtained chunks with different degrees of prototypicality.Six sites computed acoustic and linguistic features independently from each other.A total of 4232 features were pooled together and grouped into 10 low level descriptor types.For each of these groups separately and for all taken together, classification results using Support Vector Machines are reported for 150 features each with the highest individual Information Gain Ratio, for a scale of prototypicality.With both acoustic and linguistic features, we obtained a relative improvement of up to 27.6%, going from low to higher prototypicality. Dino Seppi, Anton Batliner, Björn W. Schuller, Stefan Steidl, Thurid Vogt, Johannes Wagner 0001, Laurence Devillers, Laurence Vidrascu, Noam Amir, Vered Aharonson |
INTERSPEECH | 2 |
| 2008 | Private emotions versus social interaction: a data-driven approach towards analysing emotion in speech
Anton Batliner, Stefan Steidl, Christian Hacker, Elmar Nöth |
User Model. User Adapt. Interact. | 1 |
| 2007 | The HUMAINE Database: Addressing the Collection and Annotation of Naturalistic and Induced Emotional Data
Ellen Douglas-Cowie, Roddy Cowie, Ian Sneddon, Cate Cox, Orla Lowry, Margaret McRorie, Jean-Claude Martin, Laurence Devillers, Sarkis Abrilian, Anton Batliner, Noam Amir, Kostas Karpouzis |
ACII | 10 |
| 2007 | 'You are Sooo Cool, Valentina!' Recognizing Social Attitude in Speech-Based Dialogues with an ECA
Fiorella de Rosis, Anton Batliner, Nicole Novielli, Stefan Steidl |
ACII | 2 |
| 2007 | Towards More Reality in the Recognition of Emotional SpeechabstractAs automatic emotion recognition based on speech matures, new challenges can be faced. We therefore address the major aspects in view of potential applications in the field, to benchmark today's emotion recognition systems and bridge the gap between commercial interest and current performances: acted vs. spontaneous speech, realistic emotions, noise and microphone conditions, and speaker independence. Three different data-sets are used: the Berlin Emotional Speech Database, the Danish Emotional Speech Database, and the spontaneous AIBO Emotion Corpus. By using different feature types such as word- or turn-based statistics, manual versus forced alignment, and optimization techniques we show how to best cope with this demanding task and how noise addition or different microphone positions affect emotion recognition. Björn W. Schuller, Dino Seppi, Anton Batliner, Andreas K. Maier, Stefan Steidl |
ICASSP (4) | 3 |
| 2007 | Automatic scoring of the intelligibility in patients with cancer of the oral cavityabstractAfter surgical treatment of cancer of the oral cavity patients often suffer from functional restrictions such as speech disorders.In this paper we present a novel approach to assess the outcome of the treatment w.r.t. the intelligibility of the patient using the result of an automatic speech recognition system.The word recognition rate was taken as intelligibility score.Compared to four speech experts this method yields results that are as good as the best speech expert compared to the other experts.The correlation between our system and the mean opinion of the experts is .92.Furthermore we show that our system has better performance than the average expert and is more reliable. Andreas K. Maier, Maria Schuster, Anton Batliner, Elmar Nöth, Emeka Nkenke |
INTERSPEECH | 3 |
| 2007 | The relevance of feature type for the automatic classification of emotional user states: low level descriptors and functionalsabstractIn this paper, we report on classification results for emotional user states (4 classes, German database of children interacting with a pet robot).Six sites computed acoustic and linguistic features independently from each other, following in part different strategies.A total of 4244 features were pooled together and grouped into 12 low level descriptor types and 6 functional types.For each of these groups, classification results using Support Vector Machines and Random Forests are reported for the full set of features, and for 150 features each with the highest individual Information Gain Ratio.The performance for the different groups varies mostly between ≈ 50% and ≈ 60%. Björn W. Schuller, Anton Batliner, Dino Seppi, Stefan Steidl, Thurid Vogt, Johannes Wagner 0001, Laurence Devillers, Laurence Vidrascu, Noam Amir, Loïc Kessous, Vered Aharonson |
INTERSPEECH | 2 |
| 2006 | Phoneme-to-grapheme mapping for spoken inquiries to the semantic webabstractAutomatic methods for grapheme-to-phoneme (G2P) and phoneme-to-grapheme (P2G) conversion have become very popular in recent years.Their performance has improved considerably, while at the same time these developments required less input from expert lexicographers.Continuing in this tradition we will present in this paper a data-driven, language-independent approach called MASSIVE 1 with which it is possible to create efficient online modules for automatic symbol mapping.Our framework is solely based on statistical methods for training and run-time and has been optimized for P2G conversion in the context of spoken inquiries to the Semantic Web, an issue researched in the SmartWeb project 2 .MASSIVE systems can be trained using a pronunciation lexicon, the output of a phone recognizer or any other suitable set of corresponding symbol strings.Successful tests have been performed on German and English data sets. Axel Horndasch, Elmar Nöth, Anton Batliner, Volker Warnke |
INTERSPEECH | 3 |
| 2005 | Can you Understand him? Let's Look at his Word Accuracy - Automatic Evaluation of Tracheoesophageal SpeechabstractTracheoesophageal (TE) speech is a possibility to restore the ability to speak after laryngectomy. TE speech often shows low intelligibility. An objective means to determine and quantify the intelligibility does not exist until now and an automation of this procedure is desirable. We used a speech recognizer trained on normal, non-pathologic voices. We compared intelligibility scores for TE speech from five experienced raters with the word accuracy (WA) of our speech recognizer. A correlation coefficient of -0.84 shows that WA can be a good indicator of intelligibility for pathologic voices. An outlook for future work is presented. Maria Schuster, Elmar Nöth, Tino Haderlein, Stefan Steidl, Anton Batliner, Frank Rosanowski |
ICASSP (1) | 5 |
| 2005 | "Of All Things the Measure Is Man" : Automatic Classification of Emotions and Inter-Labeler ConsistencyabstractIn traditional classification problems, the reference needed for training a classifier is given and considered to be absolutely correct. However, this does not apply to all tasks. In emotion recognition in non-acted speech, for instance, one often does not know which emotion was really intended by the speaker. Hence, the data is annotated by a group of human labelers who do not agree on one common class in most cases. Often, similar classes are confused systematically. We propose a new entropy-based method to evaluate classification results taking into account these systematic confusions. We can show that a classifier which achieves a recognition rate of "only" about 60 % on a four-class-problem performs as well as our five human labelers on average. Stefan Steidl, Michael Levit, Anton Batliner, Elmar Nöth, Heinrich Niemann |
ICASSP (1) | 3 |
| 2005 | The PF_STAR children's speech corpusabstractThis paper describes the corpus of recordings of children's speech which was collected as part of the EU FP5 PF STAR project.The corpus contains more than 60 hours of speech, including read and imitated native-language speech in British English, German and Swedish, read and imitated non-nativelanguage English speech from German, Italian and Swedish children, and native-language spontaneous and emotional speech in English and German. 1 This work was conducted as part of EU FP5 PF STAR (Preparing Future Multisensorial Interaction Research) Anton Batliner, Mats Blomberg, Shona D'Arcy, Daniel Elenius, Diego Giuliani, Matteo Gerosa, Christian Hacker, Martin J. Russell, Stefan Steidl |
INTERSPEECH | 1 |
| 2005 | Tales of tuning - prototyping for automatic classification of emotional user statesabstractClassification performance for emotional user states found in the few realistic, spontaneous databases available is as yet not very high.We present a database with emotional children's speech in a human-robot scenario.Baseline classification performance for seven classes is 44.5%, for four classes 59.2%.We discuss possible strategies for tuning, e.g., using only prototypes (based on annotation correspondence or classification scores), or taking into account requirements and feasibility in possible applications (weighting of false alarms or speakerspecific overall frequencies). Anton Batliner, Stefan Steidl, Christian Hacker, Elmar Nöth, Heinrich Niemann |
INTERSPEECH | 1 |
| 2004 | "You Stupid Tin Box" - Children Interacting with the AIBO Robot: A Cross-linguistic Emotional Speech Corpus
Anton Batliner, Christian Hacker, Stefan Steidl, Elmar Nöth, Shona D'Arcy, Martin J. Russell |
LREC | 1 |
| 2003 | We are not amused - but how do you know? user states in a multi-modal dialogue systemabstractFor the multi-modal dialogue system SmartKom, emotional user states in a Wizard-of-Oz experiment as, e.g., joyful, angry, helpless, are annotated holistically and based purely on facial expressions; other phenomena (prosodic peculiarities, offtalk, i.e., speaking aside, etc.) are labelled as well.We present the correlations between these different annotations and report classification results using a large prosodic feature vector.The performance of the user state classification is not yet satisfactory; possible reasons and remedies are discussed. Anton Batliner, Viktor Zeißler, Carmen Frank, Johann Adelhardt, Rui Ping Shi, Elmar Nöth |
INTERSPEECH | 1 |
| 2003 | How to find trouble in communication
Anton Batliner, K. Fischer, Richard Huber, Jörg Spilker, Elmar Nöth |
Speech Commun. | 1 |
| 2002 | On the use of prosody in automatic dialogue understanding
Elmar Nöth, Anton Batliner, Volker Warnke, Manuela Boros, Jan Buckow, Richard Huber, Florian Gallwitz, M. Nutt, Heinrich Niemann |
Speech Commun. | 2 |
| 2001 | Boiling down prosody for the classification of boundaries and accents in German and EnglishabstractIn the focus of this paper is a comparison of the most relevant prosodic features/feature classes for the classification of boundaries and accents in German and in English.Principal components were computed based on a large prosodic feature vector; these principal components were used as predictor variables in a Linear Discriminant analysis as well as in a Classification and Regression Tree.The number of the most relevant principal components was between three and five; for both languages and for boundary and accent classification alike, most important were principal components modelling duration, in combination with energy, followed by pauses and F0. Anton Batliner, Jan Buckow, Richard Huber, Volker Warnke, Elmar Nöth, Heinrich Niemann |
INTERSPEECH | 1 |
| 2001 | Prosodic models, automatic speech understanding, and speech synthesis: towards the common groundabstractAutomatic speech understanding and speech synthesis, two of the major speech processing applications, impose strikingly different constraints and requirements on prosodic models.The prevalent models of prosody and intonation fail to offer a unified solution to these conflicting constraints.As a consequence, prosodic models have been applied only occasionally in end-toend automatic speech understanding systems; in contrast, they have been applied extensively in speech synthesis systems.In this paper we want to discuss the reasons for this state of affairs as well as possible strategies to overcome the shortcomings of the use of prosodic modelling in automatic speech processing. Anton Batliner, Bernd Möbius, Gregor Möhler, Antje Schweitzer, Elmar Nöth |
INTERSPEECH | 1 |
| 2000 | Recognition of emotion in a realistic dialogue scenarioabstractNowadays modern automatic dialogue systems are able to understand complex sentences instead of only a few commands like Stop or No.In a call-center, such a system should be able to determine in a critical phase of the dialogue if the call should be passed over to a human operator.Such a critical phase can be indicated by the customer's vocal expression.Other studies prooved that it is possible to distinguish between anger and neutral speech w i t h prosodic features alone.Subjects in these studies were mostly people acting or simulating emotions like anger.In this paper we use data from a so-called Wizard of O z (WoZ) scenario to get more realistic data instead of simulated anger.As shown below, the classi cation rate for the two classes "emotion" (class E) and "neutral" (class :E) is signi cantly worse for these more realistic data.Furthermore the classi cation results are heavily speaker dependent.Prosody alone might t h us not be sucient and has to be supplemented by the use of other knowledge sources such as the detection of repetitions, reformulations, swear words, and dialogue acts. Richard Huber, Anton Batliner, Jan Buckow, Elmar Nöth, Volker Warnke, Heinrich Niemann |
INTERSPEECH | 2 |
| 2000 | VERBMOBIL: the use of prosody in the linguistic components of a speech understanding systemabstractWe show how prosody can be used in speech understanding systems. This is demonstrated with the VERBMOBIL speech to-speech translation system which, to our knowledge, is the first complete system which successfully uses prosodic information in the linguistic analysis. Prosody is used by computing probabilities for clause boundaries, accentuation, and different types, of sentence mood for each of the word hypotheses computed by the word recognizer. These probabilities guide the search of the linguistic analysis. Disambiguation is already achieved during the analysis and not by a prosodic verification of different linguistic hypotheses. So far, the most useful prosodic information is provided by clause boundaries. These are detected with a recognition rate of 94%. For the parsing of word hypotheses graphs, the use of clause boundary probabilities yields a speed-up of 92% and a 96% reduction of alternative readings. Elmar Nöth, Anton Batliner, Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Automatic annotation and classification of phrase accents in spontaneous speechabstractDuring the last years, we have been working on the automatic classification of boundaries and accents in the German VERBMOBIL (VM) project (human-human communication, appointment scheduling dialogues). A sub-corpus was annotated manually with prosodic boundary and accent labels, and neural networks (NN) trained with a large set of prosodic features were used for automatic classification. The classification of boundaries could be improved markedly with a combination of the NN with a language model (LM) that was trained with manually annotated syntactic-prosodic boundary labels in a much larger sub-corpus. Here we show how a combination of NN with LM along similar lines can be used for an improvement of accent classification as well. For the training of the LM, accents are annotated automatically in the transliteration with the help of a rule--based system that uses part--of--speech (POS) as well as other linguistic /phonological information. 1. INTRODUCTION This research has been condu... Anton Batliner, M. Nutt, Volker Warnke, Elmar Nöth, Jan Buckow, Richard Huber, Heinrich Niemann |
EUROSPEECH | 1 |
| 1999 | Integrating multiple knowledge sources for word hypotheses graph interpretationabstractWe present a n i n tegrated approach for the interpretation of word hypotheses graphs (WHGs) using multiple knowledge sources.Commonly, dierent knowledge sources in speech understanding are applied sequentially.Typically, speech understanding systems, such as the Verbmobil speech-to-speech translation system, rst use a word recognizer to determine word hypotheses, only based on acoustic and language model (LM) information.The resulting word sequences or WHGs are then segmented according to syntactic and/or prosodic information.Finally, these segments are interpreted by a parser or a stochastic process.Thus, it is impossible to use the knowledge of the syntactic-prosodic process, the parser or any other subsequent process to nd the best word sequence.In our new approach w e use acoustic, prosodic and LM information to determine the best word chain, to detect syntactic/prosodic/pragmatic phrase boundaries and to classify dialog acts in one integrated search procedure, based on a WHG or a word lattice. Volker Warnke, Florian Gallwitz, Anton Batliner, Jan Buckow, Richard Huber, Elmar Nöth, A. Höthker |
EUROSPEECH | 3 |
| 1998 | Dovetailing of acoustics and prosody in spontaneous speech recognitionabstractProsody can be applied to improve the performance of spontaneous speech translation systems like VERBMOBIL.In VERB-MOBIL we previously augmented the output of a word recognizer with prosodic information.Here we present a new approach of interleaving word recognition and prosodic processing.While we still use the output of a word recognizer to determine phrase boundaries, we do not wait until the end of the utterance before we start processing.Instead we intercept chunks of word hypotheses during the forward search of the recognizer.Neural networks and language models are used to predict phrase boundaries.Those boundary hypotheses, in turn, are used by the recognizer to cut the stream of incoming speech into syntactic-prosodic phrases.Thus, incremental processing is possible.We investigate which features are suited for incremental prosodic processing and compare them w.r.t.classification performance and efficiency.We show that with a set of features that can be computed efficiently classification results are achieved which are almost as good as those with the previously used computationally more expensive features. Jan Buckow, Anton Batliner, Richard Huber, Elmar Nöth, Volker Warnke, Heinrich Niemann |
ICSLP | 2 |
| 1998 | Integrated recognition of words and phrase boundariesabstractIn this paper we present an integrated approach for recognizing both the word sequence and the syntactic-prosodic structure of a spontaneous utterance. We take into account the fact that a spontaneous utterance is not merely an unstructured sequence of words by incorporating phrase boundary information into the language model and by providing HMMs to model boundaries. This allows for a distinction between word transitions across phrase boundaries and transitions within a phrase. During recognition, the syntactic-prosodic structure of the utterance is determined implicitly. Without any increase in computational effort, this leads to a 4% reduction of word error rate, and, at the same time, syntactic-prosodic boundary labels are provided for subsequent processing. The boundaries are recognized with a precision and recall rate of about 75% each. They can be used to reduce drastically the computational effort for parsing spontaneous utterances. We also present a system architecture to inco... Florian Gallwitz, Anton Batliner, Jan Buckow, Richard Huber, Heinrich Niemann, Elmar Nöth |
ICSLP | 2 |
| 1998 | M = Syntax + Prosody: A syntactic-prosodic labelling scheme for large spontaneous speech databases
Anton Batliner, Ralf Kompe, Andreas Kießling 0001, Marion Mast, Heinrich Niemann, Elmar Nöth |
Speech Commun. | 1 |
| 1997 | Improving parsing of spontaneous speech with the help of prosodic boundariesabstractParsing can be improved in automatic speech understanding if prosodic boundary marking is taken into account, because syntactic boundaries are often marked by prosodic means. Because large databases are needed for the training of statistical models for prosodic boundaries, we developed a labeling scheme for syntactic-prosodic boundaries within the German Verbmobil project (automatic speech-to-speech translation). We compare the results of classifiers (multi-layer perceptrons and language models) trained on these syntactic-prosodic boundary labels with classifiers trained on perceptual-prosodic and purely syntactic labels. Recognition rates of up to 96% were achieved. The turns that we need to parse consist of 20 words on the average and frequently contain sequences of partial sentence equivalents due to restarts, ellipsis, etc. For this material, the boundary scores computed by our classifiers can successfully be integrated into the syntactic parsing of word graphs; currently, they improve the parse time by 92% and reduce the number of parse trees by 96%. This is achieved by introducing a special prosodic syntactic clause boundary (PSCB) symbol into our grammar and guiding the search for the best word chain with the prosodic boundary scores. Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth, Anton Batliner, Stefanie Schachtl, Tobias Ruland, Hans Ulrich Block |
ICASSP | 5 |
| 1997 | Prosodic processing and its use in VERBMOBILabstractWe present the prosody module of the VERBMOBIL speech-to-speech translation system, the world wide first complete system, which successfully uses prosodic information in linguistic analysis. This is achieved by computing probabilities for clause boundaries, accentuation, and different types of sentence mood for each of the word hypotheses computed by the word recognizer. These probabilities guide the search of the linguistic analysis. Disambiguation is already achieved during the analysis and not by a prosodic verification of different linguistic hypotheses. So far, the most useful prosodic information is provided by clause boundaries. These are detected with a recognition rate of 94%. For the parsing of word hypotheses graphs, the use of clause boundary probabilities yields a speed-up of 92% and a 96% reduction of alternative readings. Heinrich Niemann, Elmar Nöth, Andreas Kießling 0001, Ralf Kompe, Anton Batliner |
ICASSP | 5 |
| 1997 | Tempo and its change in spontaneous speechabstractIn this paper, we give a first account of speech tempo and its change in spontaneous speech in a very large data base (Verbmobil, i.e., human-human appointment dialogs). As features representing speech tempo, we computed mean normalized speech duration (speaking rate) and normalized phone duration in different ways. The importance of these features is evaluated with an automatic classification of boundaries and accents where different sets of prosodic features (including also information about F0, energy, pause, etc.) were used. The best results (83% for accents, 88% for boundaries, two classes each) could be achieved when all features were used. For the 2nd issue change of tempo was labelled manually. We present the characterizing feature values for changes from slow to fast and from fast to slow, as well as the results of an automatic classification of change of tempo (72% for three classes). Finally, we discuss the possible function of change of tempo and its use in automatic speech processing. (orig.) Anton Batliner, Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann, Elmar Nöth |
EUROSPEECH | 1 |
| 1996 | Integrating Syntactic and Prosodic Information for the Efficient Detection of Empty Categories
Anton Batliner, Anke Feldhaus, Stefan Geißler, Andreas Kießling 0001, Tibor Kiss, Ralf Kompe, Elmar Nöth |
COLING | 1 |
| 1996 | Prosody, empty categories and parsing - a success storyabstractWe describe a number of experiments that demonstrate the usefulness of prosodic information for a processing module which parses spoken utterances with a feature-based grammar employing empty categories.We show that by requiring certain prosodic properties from those positions in the input, where the presence of an empty category has to be hypothesized, a derivation can be accomplished more eciently.The approach has been implemented in the machine translation project Verbmobil and results in a signicant reduction of the work-load for the parser. Anton Batliner, Anke Feldhaus, Stefan Geißler, Tibor Kiss, Ralf Kompe, Elmar Nöth |
ICSLP | 1 |
| 1996 | Syntactic-prosodic labeling of large spontaneous speech data-bases
Anton Batliner, Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth |
ICSLP | 1 |
| 1996 | Consistency in transcription and labelling of German intonation with GToBIabstractA diverse set of speech data was labelled in three sites by 13 transcribers with differing levels of expertise, using GToBI, a consensus transcription system for German intonation.Overall inter-transcriber-consistency suggests that, with training, labellers can acquire sufficient skill with GToBI for large-scale database labelling. Martine Grice, Matthias Reyelt, Ralf Benzmüller, Jörg Mayer 0001, Anton Batliner |
ICSLP | 5 |
| 1995 | Prosodic scoring of word hypotheses graphs
Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth, Ernst Günter Schukat-Talamazzini, A. Zottmann, Anton Batliner |
EUROSPEECH | 7 |
| 1994 | Automatic classification of prosodically marked phrase boundaries in GermanabstractA large corpus has been created automatically and read by 100 speakers. Phrase boundaries were labeled in the sentences automatically during sentence generation. Perception experiments on a subset of 500 utterances showed a high agreement between the automatically generated boundary markers and the ones perceived by listeners. Gaussian distribution and polynomial classifiers were trained on a set of prosodic features computed from the speech signal using the automatically generated boundary markers. Comparing the classification results with the judgments of the listeners yielded in a recognition rate of 87%. A combination with stochastic language models improved the recognition rate to 90%. We found that the pause and the durational features are most important for the classification, but that the influence of F0 is not neglectable.> Ralf Kompe, Anton Batliner, Andreas Kießling 0001, Ute Kilian, Heinrich Niemann, Elmar Nöth, Peter Regel-Brietzmann |
ICASSP (2) | 2 |
| 1994 | Improving parsing by incorporating 'prosodic clause boundaries into a grammarabstractIn written language, punctuation is used to separate main and subordinate clause.In spoken language, ambiguities arise due to missing punctuation, but clause boundaries are often marked prosodically and can be used instead.We detect PCBs (Prosodically marked Clause Boundaries) by using prosodic features (duration, intonation, energy, and pause information) with a neural network, achieving a recognition rate of 82%.PCBs are integrated into our grammar using a special syntactic category 'break' that can be used in the phrase-structure rules of the grammar in a similar way as punctuation is used in grammars for written language.Whereas punctuation in most cases is obligatory, PCBs are sometimes optional.Moreover, they can in principle occur everywhere in the sentence due e.g. to hesitations or misrecognition.To cope with these problems we tested two different approaches: A slightly modified parser for word chains containing PCBs and a word graph parser that takes the probabilities of PCBs into account.Tests were conducted on a subset of infinitive subordinate clauses from a large speech database containing sentences from the domain of train table inquiries.The average number of syntactic derivations could be reduced by about 70 % even when working on recognized word graphs. Gabriele Bakenecker, Hans Ulrich Block, Anton Batliner, Ralf Kompe, Elmar Nöth, Peter Regel-Brietzmann |
ICSLP | 3 |
| 1994 | Automatic labeling of phrase accents in GermanabstractIn this paper a method for the automatic labeling of phrase accents is described, based on a large text corpus that has been generated automatically and read by 100 speakers.Perception experiments on a subset of 500 utterances show a high agreement between the automatically generated accent labels and the judgment scores obtained.We computed different prosodic feature vectors from the speech signal for each syllable and trained different Gaussian distribution classifiers and artificial neural networks using the automatically generated accent labels.Recognition rates of up to 83% could be achieved for the distinction of accentuated vs. unaccentuated syllables.Similar results could be obtained for the comparison of the listeners judgments with the automatic classification. Andreas Kießling 0001, Ralf Kompe, Anton Batliner, Heinrich Niemann, Elmar Nöth |
ICSLP | 3 |
| 1994 | Prosody takes over: Towards a prosodically guided dialog system
Ralf Kompe, Elmar Nöth, Andreas Kießling 0001, Thomas Kuhn 0002, Marion Mast, Heinrich Niemann, K. Ott, Anton Batliner |
Speech Communication | 8 |
| 1993 | Prosody takes over: a prosodically guided dialog systemabstractIn this paper rst experiments with naive persons using the speech understanding and dialog system EVAR are discussed.The domain of EVAR is train table inquiry.We observed that in real human-human dialogs when the o cer transmits the information the customer very often interrupts.Many of these interruptions are just repetitions of the time of day given by the o cer.The functional role of these interruptions is determined b y prosodic cues only.An important result of the experiments with EVAR is that it is hard to follow the system giving the train connection via speech synthesis.In this case it is even more important than in human-human dialogs that the user has the opportunity to interact during the answer phase.Therefore we extended the dialog m o dule to allow the user to repeat the time of day and we added a p r osody module guiding the continuation of the dialog. Ralf Kompe, Andreas Kießling 0001, Thomas Kuhn 0002, Marion Mast, Heinrich Niemann, Elmar Nöth, K. Ott, Anton Batliner |
EUROSPEECH | 8 |
| 1992 | DP-based determination of F0 contours from speech signalsabstractA new algorithm for the determination of fundamental frequency (F/sub 0/) contours is presented. For each voiced frame appropriate divisors of the frequency with the maximum energy in the spectrum are taken as F/sub 0/ candidates. An F/sub 0/ contour is computed using a dynamic programming (DP) method by minimizing a weighted sum of the difference between consecutive candidates and the distance of the candidates to a predetermined local target value. With this algorithm a coarse error rate of 0.6% on the frame level and of 6.4% on the sentence level is achieved on a German speech database. On the average the difference to the reference is 1.9 Hz. The algorithm outperforms two conventional algorithms tested on the same data.> Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann, Elmar Nöth, Anton Batliner |
ICASSP | 5 |
| 1989 | The prediction of focusabstractWe present results on how focus is marked intonationally in German.Several speakers produced a large corpus of sentences.The corpus was constructed in a way that sentence modality and place of focus could only be differentiated by intona tional means.Acoustic features representing the intonational para meters pitch, duration, and intensity, were extracted manually or automatically.The relevance of these features and the effect of several transformations were tested with statistical methods.Percep tual experiments where the listeners had to judge the naturalness and categories of the utterances were performed as well.By calcula ting average values for the (appropriately transformed) relevant features we found "normal", prototypical cases.We will show that by looking at utterances where all listeners agreed on the naturalness and (intended) categories we arrived at coinciding results.At the same time we found ''unusual" but regular productions. MATERIAL AND PROCEDURESThis paper is concerned with the prediction of focus; focus is the part of an utterance which is semantically most important.On the phonetic surface focus is marked by the focal accent (FA).To be more exact, we will try to predict the phrase that carries the FA. Anton Batliner, Elmar Nöth |
EUROSPEECH | 1 |