Florian Hönig

dblp:68/7247 · also Florian Hoenig · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-8677-3420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
abstract
Early and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases.
Franziska Braun, Christopher Witzl, Florian Hönig, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer
LREC3
2025 On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
abstract
Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting in non-verbatim transcriptions with limited clinical and research value. We propose a parameter-efficient adaptation method to decode dysfluencies and fluency modifications as special tokens within transcriptions, evaluated on simulated (LibriStutter, English) and natural (KSoF, German) stuttered speech datasets. To mitigate ASR performance disparities and bias towards English, we introduce a multi-step fine-tuning strategy with language-adaptive pretraining. Tokenization analysis further highlights the tokenizer’s English-centric bias, which poses challenges for improving performance on German data. Our findings demonstrate the effectiveness of lightweight adaptation techniques for dysfluency-aware ASR while exposing key limitations in multilingual E2E systems.
Kashaf Gulzar, Dominik Wagner 0002, Sebastian P. Bayerl, Florian Hönig, Tobias Bocklet, Korbinian Riedhammer
ASRU4
2024 Infusing Acoustic Pause Context into Text-Based Dementia Assessment
abstract
Speech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the cognitive states of subjects with no cognitive impairment, mild cognitive impairment, and Alzheimer's dementia based on their speech from a clinical assessment. We address three binary classification tasks: Onset, monitoring, and dementia exclusion. The performance is evaluated through experiments on a German Verbal Fluency Test and a Picture Description Test, comparing the model's effectiveness across different speech production contexts. Starting from a textual baseline, we investigate the effect of incorporation of pause information and acoustic context. We show the test should be chosen depending on the task, and similarly, lexical pause information and acoustic cross-attention contribute differently.
Franziska Braun, Sebastian P. Bayerl, Florian Hönig, Hartmut Lehfeld, Thomas Hillemacher, Tobias Bocklet, Korbinian Riedhammer
INTERSPEECH3
2023 A Stutter Seldom Comes Alone - Cross-Corpus Stuttering Detection as a Multi-label Problem
Sebastian P. Bayerl, Dominik Wagner 0002, Ilja Baumann, Florian Hönig, Tobias Bocklet, Elmar Nöth, Korbinian Riedhammer
INTERSPEECH4
2023 Classifying Dementia in the Presence of Depression: A Cross-Corpus Study
abstract
Automated dementia screening enables early detection and intervention, reducing costs to healthcare systems and increasing quality of life for those affected. Depression has shared symptoms with dementia, adding complexity to diagnoses. The research focus so far has been on binary classification of dementia (DEM) and healthy controls (HC) using speech from picture description tests from a single dataset. In this work, we apply established baseline systems to discriminate cognitive impairment in speech from the semantic Verbal Fluency Test and the Boston Naming Test using text, audio and emotion embeddings in a 3-class classification problem (HC vs. MCI vs. DEM). We perform cross-corpus and mixed-corpus experiments on two independently recorded German datasets to investigate generalization to larger populations and different recording conditions. In a detailed error analysis, we look at depression as a secondary diagnosis to understand what our classifiers actually learn.
Franziska Braun, Sebastian P. Bayerl, Paula Andrea Pérez-Toro, Florian Hönig, Hartmut Lehfeld, Thomas Hillemacher, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer
INTERSPEECH4
2023 Automatic Assessment of Alzheimer's across Three Languages Using Speech and Language Features
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Franziska Braun, Florian Hönig, Carlos Tobon 0001, David Aguillón, Francisco Lopera, Liliana Hincapié-Henao, Maria Schuster, Korbinian Riedhammer, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH4
2022 KSoF: The Kassel State of Fluency Dataset - A Therapy Centered Dataset of Stuttering
abstract
Stuttering is a complex speech disorder that negatively affects an individual’s ability to communicate effectively. Persons who stutter (PWS) often suffer considerably under the condition and seek help through therapy. Fluency shaping is a therapy approach where PWSs learn to modify their speech to help them to overcome their stutter. Mastering such speech techniques takes time and practice, even after therapy. Shortly after therapy, success is evaluated highly, but relapse rates are high. To be able to monitor speech behavior over a long time, the ability to detect stuttering events and modifications in speech could help PWSs and speech pathologists to track the level of fluency. Monitoring could create the ability to intervene early by detecting lapses in fluency. To the best of our knowledge, no public dataset is available that contains speech from people who underwent stuttering therapy that changed the style of speaking. This work introduces the Kassel State of Fluency (KSoF), a therapy-based dataset containing over 5500 clips of PWSs. The clips were labeled with six stuttering-related event types: blocks, prolongations, sound repetitions, word repetitions, interjections, and – specific to therapy – speech modifications. The audio was recorded during therapy sessions at the Institut der Kasseler Stottertherapie. The data will be made available for research purposes upon request.
Sebastian P. Bayerl, Alexander W. von Gudenberg, Florian Hönig, Elmar Nöth, Korbinian Riedhammer
LREC3
2020 Surgical Mask Detection with Deep Recurrent Phonetic Models
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Florian Hönig, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH5
2017 Writer Identification Using GMM Supervectors and Exemplar-SVMs
Vincent Christlein, David Bernecker, Florian Hönig, Andreas K. Maier, Elli Angelopoulou
Pattern Recognit.3
2016 Towards an automatic monitoring of the neurological state of Parkinson's patients from speech
abstract
The suitability of articulation measures and speech intelligibility is evaluated to estimate the neurological state of patients with Parkinson's disease (PD). A set of measures recently introduced to model the articulatory capability of PD patients is considered. Additionally, the speech intelligibility in terms of the word accuracy obtained from the Google® speech recognizer is included. Recordings of patients in three different languages are considered: Spanish, German, and Czech. Additionally, the proposed approach is tested on data recently used in the INTERSPEECH 2015 Computational Paralinguistics Challenge. According to the results, it is possible to estimate the neurological state of PD patients from speech with a Spearman's correlation of up to 0.72 with respect to the evaluations performed by neurologist experts.
Juan Rafael Orozco-Arroyave, Juan Camilo Vásquez-Correa, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
ICASSP3
2016 Language proficiency assessment of English L2 speakers based on joint analysis of prosody and native language
abstract
In this work, we present an in-depth analysis of the interdependency between the non-native prosody and the native language (L1) of English L2 speakers, as separately investigated in the Degree of Nativeness Task and the Native Language Task of the INTERSPEECH 2015 and 2016 Computational Paralinguistics ChallengE (ComParE). To this end, we propose a multi-task learning scheme based on auxiliary attributes for jointly learning the tasks of L1 classification and prosody score regression. The effectiveness of this scheme is demonstrated in extensive experimental runs, comparing various standardised feature sets of prosodic, cepstral, spectral, and voice quality descriptors, as well as automatic feature selection. In the result, we show that the prediction of both prosody score and L1 can be improved by considering both tasks in a holistic way. In particular, we achieve an 11% relative gain in regression performance (Spearman's correlation coefficient) on prosody scores, when comparing the best multi- and single-task learning results.
Yue Zhang 0014, Felix Weninger, Anton Batliner, Florian Hönig, Björn W. Schuller
ICMI4
2016 Assessing the Prosody of Non-Native Speakers of English: Measures and Feature Sets
Eduardo Coutinho, Florian Hönig, Yue Zhang 0014, Simone Hantke, Anton Batliner, Elmar Nöth, Björn W. Schuller
LREC2
2015 The degree of nativeness sub-challenge: the data
Florian Hönig
INTERSPEECH1
2015 Voiced/unvoiced transitions in speech as a potential bio-marker to detect parkinson's disease
abstract
Several studies have addressed the automatic classification of speakers with Parkinson’s disease (PD) and healthy controls (HC). Most of the studies are based on speech recordings of sustained vowels, isolated words, and single sentences. Only few investigations have considered read texts and/or sponta-neous speech. This paper addresses two main questions still open regarding the automatic analysis speech in patients with PD, (a) “Is it possible to classify PD patients and HC through running speech signals in multiple languages?”, and (b) “where is the information to discriminate between speech recordings of PD patients and HC? ” In this paper speech recordings of read texts and monologues spoken in three different languages are considered. The energy content of the borders between voiced and unvoiced sounds is modeled. According to the results with read texts it is possible to achieve accuracies ranging from 91% to 98 % depending on the language. With respect to the re-sults on monologues, the accuracies are above 98 % in all of the three languages. The presence of discriminant information in the voiced/unvoiced and unvoiced/voiced transitions is vali-dated here, evidencing the problems of PD patients to stop/start the vocal folds movement during the production of running speech. Index Terms: Parkinson’s disease, dysarthria, hesitation in speech, language and motor planning, energy content, voiced/unvoiced transitions. 1.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH2
2015 The INTERSPEECH 2015 computational paralinguistics challenge: nativeness, parkinson's & eating condition
abstract
The INTERSPEECH 2015 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: the estimation of the degree of nativeness, the neurological state of patients with Parkinson’s condition, and the eating conditions of speakers, i. e., whether and which food type they are eating in a seven-class problem. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, as provided to the participants. Index Terms: Computational Paralinguistics, Challenge, Degree of Nativeness, Parkinson’s Condition, Eating Condition
Björn W. Schuller, Stefan Steidl, Anton Batliner, Simone Hantke, Florian Hönig, Juan Rafael Orozco-Arroyave, Elmar Nöth, Yue Zhang 0014, Felix Weninger
INTERSPEECH5
2015 Visual comparison of speaker groups
Sebastian Wankerl, Florian Hönig, Anton Batliner, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH2
2015 Spectral and cepstral analyses for Parkinson's disease detection in Spanish vowels and words
abstract
Abstract About 1% of people older than 65 years suffer from Parkinson's disease (PD) and 90% of them develop several speech impairments, affecting phonation, articulation, prosody and fluency. Computer‐aided tools for the automatic evaluation of speech can provide useful information to the medical experts to perform a more accurate and objective diagnosis and monitoring of PD patients and can help also to evaluate the correctness and progress of their therapy. Although there are several studies that consider spectral and cepstral information to perform automatic classification of speech of people with PD, so far it is not known which is the most discriminative, spectral or cepstral analysis. In this paper, the discriminant capability of six sets of spectral and cepstral coefficients is evaluated, considering speech recordings of the five Spanish vowels and a total of 24 isolated words. According to the results, linear predictive cepstral coefficients are the most robust and exhibit values of the area under the receiver operating characteristic curve above 0.85 in 6 of the 24 words.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Elmar Nöth
Expert Syst. J. Knowl. Eng.2
2015 Characterization Methods for the Detection of Multiple Voice Disorders: Neurological, Functional, and Laryngeal Diseases
abstract
This paper evaluates the accuracy of different characterization methods for the automatic detection of multiple speech disorders. The speech impairments considered include dysphonia in people with Parkinson's disease (PD), dysphonia diagnosed in patients with different laryngeal pathologies (LP), and hypernasality in children with cleft lip and palate (CLP). Four different methods are applied to analyze the voice signals including noise content measures, spectral-cepstral modeling, nonlinear features, and measurements to quantify the stability of the fundamental frequency. These measures are tested in six databases: three with recordings of PD patients, two with patients with LP, and one with children with CLP. The abnormal vibration of the vocal folds observed in PD patients and in people with LP is modeled using the stability measures with accuracies ranging from 81% to 99% depending on the pathology. The spectral-cepstral features are used in this paper to model the voice spectrum with special emphasis around the first two formants. These measures exhibit accuracies ranging from 95% to 99% in the automatic detection of hypernasal voices, which confirms the presence of changes in the speech spectrum due to hypernasality. Noise measures suitably discriminate between dysphonic and healthy voices in both databases with speakers suffering from LP. The results obtained in this study suggest that it is not suitable to use every kind of features to model all of the voice pathologies; conversely, it is necessary to study the physiology of each impairment to choose the most appropriate set of features.
Juan Rafael Orozco-Arroyave, Elkyn Alexander Belalcázar-Bolaños, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Khaled Daqrouq, Florian Hönig, Elmar Nöth
IEEE J. Biomed. Health Informatics8
2014 Are men more sleepy than women or does it only look like - Automatic analysis of sleepy speech
abstract
The degree of sleepiness in the Sleepy Language Corpus from the Interspeech 2011 Speaker State Challenge is predicted with regression and a very large feature vector. Most notable is the great gender difference which can mainly be attributed to females showing their sleepiness less than males do.
Florian Hönig, Anton Batliner, Tobias Bocklet, Georg Stemmer, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski
ICASSP1
2014 Automatic modelling of depressed speech: relevant features and relevance of gender
abstract
Depression is an affective disorder characterised by psychomotor retardation; in speech, this shows up in reduction of pitch (variation, range), loudness, and tempo, and in voice qualities different from those of typical modal speech.A similar reduction can be observed in sleepy speech (relaxation).In this paper, we employ a small group of acoustic features modelling prosody and spectrum that have been proven successful in the modelling of sleepy speech, enriched with voice quality features, for the modelling of depressed speech within a regression approach.This knowledge-based approach is complemented by and compared with brute-forcing and automatic feature selection.We further discuss gender differences and the contributions of (groups of) features both for the modelling of depression and across depression and sleepiness.
Florian Hönig, Anton Batliner, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski
INTERSPEECH1
2014 Automatic detection of parkinson's disease from words uttered in three different languages
abstract
About 90% of the people with Parkinson’s disease (PD) develop speech impairments such as monopitch, monoloudness, imprecise articulation, and other symptoms. There are several studies addressing the problem of the automatic detection of PD from speech signals in order to develop computer aided tools for the assessment and monitoring of the patients. Recent works have shown that it is possible to detect PD from speech with accuracies above 90%; however, it is still unclear whether it is possible to make the detection independent of the spoken language. This paper addresses the automatic detection of PD considering speech recordings of three languages: German, Spanish and Czech. According to the results it is possible to classify between speech of people with PD and healthy controls (HC) with accuracies ranging from 84% to 99%, depending on the utterance.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH2
2014 Writer identification and verification using GMM supervectors
abstract
This paper proposes a new system for offline writer identification and writer verification. The proposed method uses GMM supervectors to encode the feature distribution of individual writers. Each supervector originates from an individual GMM which has been adapted from a background model via a maximum-a-posteriori step followed by mixing the new statistics with the background model. We show that this approach improves the TOP-1 accuracy of the current best ranked methods evaluated at the ICDAR-2013 competition dataset from 95.1% [13] to 97.1%, and from 97.9% [11] to 99.2% at the CVL dataset, respectively. Additionally, we compare the GMM supervector encoding with other encoding schemes, namely Fisher vectors and Vectors of Locally Aggregated Descriptors.
Vincent Christlein, David Bernecker, Florian Hönig, Elli Angelopoulou
WACV3
2012 The Automatic Assessment of Non-native Prosody: Combining Classical Prosodic Analysis with Acoustic Modelling
abstract
In earlier studies, we employed a large prosodic feature vector to assess the quality of L2 learner's utterances with respect to sentence melody and rhythm.In this paper, we combine these features with two standard approaches in paralinguistic analysis: (1) features derived from a Gaussian Mixture Model used as Universal Background Model (GMM-UBM), and (2) openSMILE, an open-source toolkit for extracting acoustic features.We evaluate our approach with English speech from 94 non-native speakers perceptually scored by 62 native labellers.GMM-UBM or openSMILE modelling alone yields lower performance than our prosodic feature vector; however, adding information from the GMM-UBM modelling or openSMILE by late fusion improves results.
Florian Hönig, Tobias Bocklet, Korbinian Riedhammer, Anton Batliner, Elmar Nöth
INTERSPEECH1
2011 Does it Groove or does it Stumble - Automatic Classification of Alcoholic Intoxication using Prosodic Features
abstract
This paper studies how prosodic features can help in the automatic detection of alcoholic intoxication.We compute features that have recently been proposed to model speech rhythm such as the pair-wise variability index for consonantal and vocalic segments (PVI) and study their aptness for the task.Further, we use a large prosodic feature vector modelling the usual candidates -pitch, intensity, and duration -and apply it onto different units such as words, syllables and stressed syllables to create generalizations of the rhythm features mentioned.The results show that the prosodic features computed are suitable for detecting alcoholic intoxication and add complementary information to state-of-the-art features.The database is the intoxication database provided by the organizers of the 2011 Interspeech Speaker State Challenge.
Florian Hönig, Anton Batliner, Elmar Nöth
INTERSPEECH1
2011 Java Visual Speech Components for Rapid Application Development of GUI Based Speech Processing Applications
abstract
In this paper, we describe a new Java framework for an easy and efficient way of developing new GUI based speech processing applications. Standard components are provided to display the speech signal, the power plot, and the spectrogram. Furthermore, a component to create a new transcription and to display and manipulate an existing transcription is provided, as well as a component to display and manually correct external pitch values. These Swing components can be easily embedded into own Java programs. They can be synchronized to display the same region of the speech file. The object-oriented design provides base classes for rapid development of own components.
Stefan Steidl, Korbinian Riedhammer, Tobias Bocklet, Florian Hönig, Elmar Nöth
INTERSPEECH4
2009 Immersive Painting
Stefan Soutschek, Florian Hönig, Andreas K. Maier, Stefan Steidl, Michael Stürmer, Hellmut Erzigkeit, Joachim Hornegger, Johannes Kornhuber
ArtsIT2
2009 A language-independent feature set for the automatic evaluation of prosody
abstract
In second language learning, the correct use of prosody plays a vital role.Therefore, an automatic method to evaluate the naturalness of the prosody of a speaker is desirable.We present a novel method to model prosody independently of the text and thus independently of the language as well.For this purpose, the voiced and unvoiced speech segments are extracted and a 187-dimensional feature vector is computed for each voiced segment.This approach is compared to word based prosodic features on a German text passage.Both are confronted with the perceptive evaluation of two native speakers of German.The word-based feature set yielded correlations of up to 0.92 while the text-independent feature set yielded 0.88.This is in the same range as the inter-rater correlation with 0.88.Furthermore, the text-independent features were computed for a Japanese translation of the passage which was also rated by two native speakers of Japanese.Again, the correlation between the automatic system and the human perception of the naturalness was high with 0.83 and not significantly lower than the inter-rater correlation of 0.92.
Andreas K. Maier, Florian Hönig, Viktor Zeißler, Anton Batliner, Erik Körner, Nobuyuki Yamanaka, Peter Ackermann, Elmar Nöth
INTERSPEECH2
2008 Classification of perceived running fatigue in digital sports
abstract
This paper presents methods for collecting and analyzing physiological and biomechanical data during recreational runs in order to classify an athletepsilas perceived fatigue state. Heart rate and its variability, running speed and stride frequency, GPS position and shoe heel compression were recorded continuously while runners moved freely outdoors. During their activity the sportsmen answered questions about their fatigue state in five-minute-intervals. Data from 84 one-hour-runs was collected for analysis. The data was analyzed using features computed for each step of the athlete to distinguish three levels of the runnerpsilas fatigue state with an accuracy of 75.3% across multiple study participants and 91.8% in the intraindividual case. The results show that for most participating runners, a heart rate variability periodogram feature and a step duration feature are best suited for classification of the perceived fatigue level. This information can be used to support sportsmen, for example by adapting their equipment to the specific needs of a fatigued athlete.
Björn M. Eskofier, Florian Hönig, Pascal Kuehner
ICPR2
2008 Automatic evaluation of characteristic speech disorders in children with cleft lip and palate
abstract
Abstract This paper discusses the automatic evaluation of speech of chil-dren with cleft lip and palate (CLP). CLP speech shows specialcharacteristics such as hypernasality, backing, and weakeningof plosives. In total ve criteria were subjectively assessed byan experienced speech expert on the phone level. This subjec-tive evaluation was used as a gold standard to train a classi-cation system. The automatic system achieves recognition re-sults on frame, phone, and word level of up to 75.8% CL. Onspeaker level signicant and high correlations between the sub-jective evaluation and the automatic system of up to 0.89 areobtained. Index Terms : pathologic speech, speech assessment, pronun-ciation scoring, children’s speech 1. Introduction Cleft Lip and Palate (CLP) is the most common malformationof the head. It constitutes almost two-thirds of the major facialdefects and almost 80% of all orofacial clefts [1]. Its prevalencediffers in different populations from 1 in 400 to 500 newborns inAsians to 1 in 1500 to 2000 in African Americans. The preva-lence in Caucasians is 1 in 750 to 900 births [2, 3].In clinical practice, articulation disorders are mainly eval-uated by subjective tools. The simplest method is the audi-tive perception, mostly performed by a speech therapist. Pre-vious studies have shown that experience is an important fac-tor that inuences the subjective estimation of speech disorderswhich leads to inaccurate evaluation by persons with only fewyears of experience as speech therapist [4]. Until now, objectivemeans exist only for quantitative measurements of nasal emis-sions [5, 6, 7] and for the detection of secondary voice disorders[8]. But other specic articulation disorders in CLP cannot besufciently quantied.In this paper, we present a new technical procedure for themeasurement and evaluation of specic speech disorders andcompare the results obtained with subjective ratings of an expe-rienced speech therapist.
Andreas K. Maier, Florian Hönig, Christian Hacker, Maria Schuster, Elmar Nöth
INTERSPEECH2
2005 Revising Perceptual Linear Prediction (PLP)
Florian Hönig, Georg Stemmer, Christian Hacker, Fabio Brugnara
INTERSPEECH1