David Suendermann-Oeft

dblp:92/5083 · also David Suendermann, David Sündermann · DBLP profile ↗
← Back
50ranked-venue papers
14as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 11 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author
YearPublicationVenuePosition
2024 Towards Scalable Remote Assessment of Mild Cognitive Impairment Via Multimodal Dialog
Oliver Roesler, Jackson Liscombe, Michael Neumann 0001, Hardik Kothare, Abhishek Hosamath, Lakshmi Arbatti, Doug Habberstad, Christiane Suendermann-Oeft, Meredith Bartlett, Cathy Zhang, Nikhil Sukhdev, Kolja Wilms, Anusha Badathala, Sandrine Istas, Steve Ruhmel, Bryan Hansen, Madeline Hannan, David Henley, Arthur W. Wallace, Ira Shoulson, David Suendermann-Oeft, Vikram Ramanarayanan
INTERSPEECH21
2023 When Words Speak Just as Loudly as Actions: Virtual Agent Based Remote Health Assessment Integrating What Patients Say with What They Do
Vikram Ramanarayanan, David Pautler, Lakshmi Arbatti, Abhishek Hosamath, Michael Neumann 0001, Hardik Kothare, Oliver Roesler, Jackson Liscombe, Andrew Cornish, Doug Habberstad, Vanessa Richter, David Suendermann-Oeft, Ira Shoulson
INTERSPEECH13
2022 Statistical and clinical utility of multimodal dialogue-based speech and facial metrics for Parkinson's disease assessment
Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Oliver Roesler, William Burke, Andrew Exner, Sandy Snyder, Andrew Cornish, Doug Habberstad, David Pautler, David Suendermann-Oeft, Jessica Huber, Vikram Ramanarayanan
INTERSPEECH11
2021 Investigating the Interplay Between Affective, Phonatory and Motoric Subsystems in Autism Spectrum Disorder Using a Multimodal Dialogue Agent
abstract
Abstract We explore the utility of an on-demand multimodal conversational platform in extracting speech and facial metrics in children with Autism Spectrum Disorder (ASD). We investigate the extent to which these metrics correlate with objective clinical measures, particularly as they pertain to the interplay be-tween the affective, phonatory and motoric subsystems. 22 participants diagnosed with ASD engaged with a virtual agent in conversational affect production tasks designed to elicit facial and vocal affect. We found significant correlations between vocal pitch and loudness extracted by our platform during these tasks and accuracy in recognition of facial and vocal affect, as-sessed via the Diagnostic Analysis of Nonverbal Accuracy-2 (DANVA-2) neuropsychological task. We also found significant correlations between jaw kinematic metrics extracted using our platform and motor speed of the dominant hand assessed via a standardised neuropsychological finger tapping task. These findings offer preliminary evidence for the usefulness of these audiovisual analytic metrics and could help us better model the interplay between different physiological subsystems in individuals with ASD.
Hardik Kothare, Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, William Burke, Andrew Cornish, Doug Habberstad, Alaa Sakallah, Sara Markuson, Seemran Kansara, Afik Faerman, Yasmine Bensidi-Slimane, Laura Fry, Saige Portera, David Suendermann-Oeft, David Pautler, Carly Demopoulos
Interspeech16
2021 Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at Scale
abstract
We propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS.
Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan
Interspeech5
2020 Toward Remote Patient Monitoring of Speech, Video, Cognitive and Respiratory Biomarkers Using Multimodal Dialog Technology
Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, David Pautler, Doug Habberstad, Andrew Cornish, Hardik Kothare, Vignesh Murali, Jackson Liscombe, Dirk Schnelle-Walka, Patrick L. Lange, David Suendermann-Oeft
INTERSPEECH12
2019 NEMSI: A Multimodal Dialog System for Screening of Neurological or Mental Conditions
abstract
We present NEMSI, a cloud-based multimodal dialog system designed to have naturalistic interactions with individuals for the purpose of screening neurological or mental conditions. The system has been used by thousands of people capturing audio and video responses to open-ended questions and structured health surveys.
David Suendermann-Oeft, Amanda Robinson 0002, Andrew Cornish, Doug Habberstad, David Pautler, Dirk Schnelle-Walka, Franziska Haller, Jackson Liscombe, Michael Neumann 0001, Mike Merrill, Oliver Roesler, Renko Geffarth
IVA1
2018 Game-based Spoken Dialog Language Learning Applications for Young Students
Keelan Evanini, Veronika Timpe-Laughlin, Eugene Tsuprun, Ian Blood, Jeremy Lee, James V. Bruno, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft
INTERSPEECH9
2018 An Automated Assistant for Medical Scribes
Gregory P. Finley, Erik Edwards, Amanda Robinson 0002, Najmeh Sadoughi, James Fone, Mark Miller 0001, David Suendermann-Oeft, Michael Brenndoerfer, Nico Axtmann
INTERSPEECH7
2018 Toward Scalable Dialog Technology for Conversational Language Learning: Case Study of the TOEFL® MOOC
Vikram Ramanarayanan, David Pautler, Patrick L. Lange, Eugene Tsuprun, Rutuja Ubale, Keelan Evanini, David Suendermann-Oeft
INTERSPEECH7
2018 Leveraging Multimodal Dialog Technology for the Design of Automated and Interactive Student Agents for Teacher Training
abstract
We present a paradigm for interactive teacher training that leverages multimodal dialog technology to puppeteer customdesigned embodied conversational agents (ECAs) in student roles.We used the open-source multimodal dialog system HALEF to implement a small-group classroom math discussion involving Venn diagrams where a human teacher candidate has to interact with two student ECAs whose actions are controlled by the dialog system.Such an automated paradigm has the potential to be extended and scaled to a wide range of interactive simulation scenarios in education, medicine, and business where group interaction training is essential.
David Pautler, Vikram Ramanarayanan, Kirby Cofino, Patrick L. Lange, David Suendermann-Oeft
SIGDIAL Conference5
2017 Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system
abstract
Spoken language understanding (SLU) in dialog systems is generally performed using a natural language understanding (NLU) model based on the hypotheses produced by an automatic speech recognition (ASR) system. However, when new spoken dialog applications are built from scratch in real user environments that often have sub-optimal audio characteristics, ASR performance can suffer due to factors such as the paucity of training data or a mismatch between the training and test data. To address this issue, this paper proposes an ASR-free, end-to-end (E2E) modeling approach to SLU for a cloud-based, modular spoken dialog system (SDS). We evaluate the effectiveness of our approach on crowdsourced data collected from non-native English speakers interacting with a conversational language learning application. Experimental results show that our approach is particularly promising in situations with low ASR accuracy. It can further improve the performance of a sophisticated CNN-based SLU system with more accurate ASR hypotheses by fusing the scores from E2E system, i.e., the overall accuracy of SLU is improved from 85.6% to 86.5%.
Yao Qian, Rutuja Ubale, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft, Keelan Evanini, Eugene Tsuprun
ASRU5
2017 A modular, multimodal open-source virtual interviewer dialog agent
abstract
We present an open-source multimodal dialog system equipped with a virtual human avatar interlocutor. The agent, rigged in Blender and developed in Unity with WebGL support, interfaces with the HALEF open-source cloud-based standard-compliant dialog framework. To demonstrate the capabilities of the system, we designed and implemented a conversational job interview scenario where the avatar plays the role of an interviewer and responds to user input in real-time to provide an immersive user experience.
Kirby Cofino, Vikram Ramanarayanan, Patrick L. Lange, David Pautler, David Suendermann-Oeft, Keelan Evanini
ICMI5
2017 Crowdsourcing ratings of caller engagement in thin-slice videos of human-machine dialog: benefits and pitfalls
abstract
We analyze the efficacy of different crowds of naive human raters in rating engagement during human--machine dialog interactions. Each rater viewed multiple 10 second, thin-slice videos of native and non-native English speakers interacting with a computer-assisted language learning (CALL) system and rated how engaged and disengaged those callers were while interacting with the automated agent. We observe how the crowd's ratings compared to callers' self ratings of engagement, and further study how the distribution of these rating assignments vary as a function of whether the automated system or the caller was speaking. Finally, we discuss the potential applications and pitfalls of such crowdsourced paradigms in designing, developing and analyzing engagement-aware dialog systems.
Vikram Ramanarayanan, Chee Wee Leong, David Suendermann-Oeft, Keelan Evanini
ICMI3
2017 Improving Sub-Phone Modeling for Better Native Language Identification with Non-Native English Speech
Yao Qian, Keelan Evanini, David Suendermann-Oeft, Robert A. Pugh, Patrick L. Lange, Hillary Molloy, Frank K. Soong
INTERSPEECH4
2017 Human and Automated Scoring of Fluency, Pronunciation and Intonation During Human-Machine Spoken Dialog Interactions
Vikram Ramanarayanan, Patrick L. Lange, Keelan Evanini, Hillary Molloy, David Suendermann-Oeft
INTERSPEECH5
2017 Rushing to Judgement: How do Laypeople Rate Caller Engagement in Thin-Slice Videos of Human-Machine Dialog?
Vikram Ramanarayanan, Chee Wee Leong, David Suendermann-Oeft
INTERSPEECH3
2017 Jee haan, I'd like both, por favor: Elicitation of a Code-Switched Corpus of Hindi-English and Spanish-English Human-Machine Dialog
Vikram Ramanarayanan, David Suendermann-Oeft
INTERSPEECH2
2016 Noise and Metadata Sensitive Bottleneck Features for Improving Speaker Recognition with Non-Native Speech Input
Yao Qian, Jidong Tao, David Suendermann-Oeft, Keelan Evanini, Alexei V. Ivanov, Vikram Ramanarayanan
INTERSPEECH3
2016 Self-Adaptive DNN for Improving Spoken Language Proficiency Assessment
Yao Qian, Keelan Evanini, David Suendermann-Oeft
INTERSPEECH4
2016 Speech Ventures
Nicolas Scheffer, Korbinian Riedhammer, Alexandre Lebrun, David Suendermann-Oeft
INTERSPEECH4
2016 LVCSR System on a Hybrid GPU-CPU Embedded Platform for Real-Time Dialog Applications
abstract
We present the implementation of a largevocabulary continuous speech recognition (LVCSR) system on NVIDIA's Tegra K1 hyprid GPU-CPU embedded platform.The system is trained on a standard 1000hour corpus, LibriSpeech, features a trigram WFST-based language model, and achieves state-of-the-art recognition accuracy.The fact that the system is realtime-able and consumes less than 7.5 watts peak makes the system perfectly suitable for fast, but precise, offline spoken dialog applications, such as in robotics, portable gaming devices, or in-car systems.
Alexei V. Ivanov, Patrick L. Lange, David Suendermann-Oeft
SIGDIAL Conference3
2015 Using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speech
abstract
We introduce a new method to grade non-native spoken language tests automatically. Traditional automated response grading approaches use manually engineered time-aggregated features (such as mean length of pauses). We propose to incorporate general time-sequence features (such as pitch) which preserve more information than time-aggregated features and do not require human effort to design. We use a type of recurrent neural network to jointly optimize the learning of high level abstractions from time-sequence features with the time-aggregated features. We first automatically learn high level abstractions from time-sequence features with a Bidirectional Long Short Term Memory (BLSTM) and then combine the high level abstractions with time-aggregated features in a Multilayer Perceptron (MLP)/Linear Regression (LR). We optimize the BLSTM and the MLP/LR jointly. We find such models reach the best performance in terms of correlation with human raters. We also find that when there are limited time-aggregated features available, our model that incorporates time-sequence features improves performance drastically.
Zhou Yu 0005, Vikram Ramanarayanan, David Suendermann-Oeft, Klaus Zechner, Lei Chen 0004, Jidong Tao, Aliaksei Ivanou, Yao Qian
ASRU3
2015 Evaluating Speech, Face, Emotion and Body Movement Time-series Features for Automated Multimodal Presentation Scoring
abstract
We analyze how fusing features obtained from different multimodal data streams such as speech, face, body movement and emotion tracks can be applied to the scoring of multimodal presentations. We compute both time-aggregated and time-series based features from these data streams--the former being statistical functionals and other cumulative features computed over the entire time series, while the latter, dubbed histograms of cooccurrences, capture how different prototypical body posture or facial configurations co-occur within different time-lags of each other over the evolution of the multimodal, multivariate time series. We examine the relative utility of these features, along with curated speech stream features in predicting human-rated scores of multiple aspects of presentation proficiency. We find that different modalities are useful in predicting different aspects, even outperforming a naive human inter-rater agreement baseline for a subset of the aspects analyzed.
Vikram Ramanarayanan, Chee Wee Leong, Lei Chen 0004, Gary Feng, David Suendermann-Oeft
ICMI5
2015 Pronunciation accuracy and intelligibility of non-native speech
abstract
This paper investigates the connection between intelligibility and pronunciation accuracy. We compare which words in non-native English speech are likely to be misrecognized and which words are likely to be marked as pronunciation errors. We found that only 16% of the variability in word-level intelligibility can be explained by the presence of obvious mispronunciations. In some cases, a word remained recognizable or could be identified from the context despite obvious pronunciation errors. In many other cases, the annotators were unable to identify the word when listening to the audio but did not perceive it as mispronounced when presented with its transcription. At the same time, we see high agreement when the results are aggregated across all words from the same speaker.
Anastassia Loukina, Melissa Lopez, Keelan Evanini, David Suendermann-Oeft, Alexei V. Ivanov, Klaus Zechner
INTERSPEECH4
2015 Expert and crowdsourced annotation of pronunciation errors for automatic scoring systems
abstract
This paper evaluates and compares different approaches to collecting judgments about pronunciation accuracy of nonnative speech. We compare the common approach, which requires expert linguists to provide a detailed phonetic transcription of non-native English speech, with word-level judgments collected from multiple naive listeners using a crowdsourcing platform. In both cases we found low agreement between annotators on what words should be marked as errors. We compare the error detection task to a simple transcription task in which the annotators were asked to transcribe the same fragments using standard English spelling. We argue that the transcription task is a simpler and more practical way of collecting annotations which also leads to more valid data for training an automatic scoring system.
Anastassia Loukina, Melissa Lopez, Keelan Evanini, David Suendermann-Oeft, Klaus Zechner
INTERSPEECH4
2015 An analysis of time-aggregated and time-series features for scoring different aspects of multimodal presentation data
abstract
We present a technique for automated assessment of public speaking and presentation proficiency based on the analysis of concurrently recorded speech and motion capture data. With respect to Kinect motion capture data, we examine both timeaggregated as well as time-series based features. While the former is based on statistical functionals of body-part position and/or velocity computed over the entire series, the latter feature set, dubbed histograms of cooccurrences, captures how often different broad postural configurations co-occur within different time lags of each other over the evolution of the multimodal time series. We examine the relative utility of these features, along with curated features derived from the speech stream, in predicting human-rated scores of different aspects of public speaking and presentation proficiency. We further show that these features outperform the human inter-rater agreement baseline for a subset of the analyzed aspects.
Vikram Ramanarayanan, Lei Chen 0004, Chee Wee Leong, Gary Feng, David Suendermann-Oeft
INTERSPEECH5
2015 Automated Speech Recognition Technology for Dialogue Interaction with Non-Native Interlocutors
abstract
Alexei V. Ivanov, Vikram Ramanarayanan, David Suendermann-Oeft, Melissa Lopez, Keelan Evanini, Jidong Tao. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015.
Alexei V. Ivanov, Vikram Ramanarayanan, David Suendermann-Oeft, Melissa Lopez, Keelan Evanini, Jidong Tao
SIGDIAL Conference3
2015 A distributed cloud-based dialog system for conversational application development
abstract
We have previously presented HALEF-an open-source spoken dialog system-that supports telephonic interfaces and has a distributed architecture.In this paper, we extend this infrastructure to be cloud-based, and thus truly distributed and scalable.This cloud-based spoken dialog system can be accessed both via telephone interfaces as well as through web clients with WebRTC/HTML5 integration, allowing in-browser access to potentially multimodal dialog applications.We demonstrate the versatility of the system with two conversation applications in the educational domain.
Vikram Ramanarayanan, David Suendermann-Oeft, Alexei V. Ivanov, Keelan Evanini
SIGDIAL Conference2
2011 Large-Scale Experiments on Data-Driven Design of Commercial Spoken Dialog Systems
abstract
The design of commercial spoken dialog systems is most commonly based on hand-crafting call flows. Voice interaction designers write prompts, predict caller responses, set speech recognition parameters, implement interaction strategies, all based on “best design practices”. Recently, we presented the mathematical framework “Contender” (similar to reinforcement learning) that allows for replacing manual decisions made during system design by data-driven soft decisions made at system run time optimizing the cumulative reward of an application. The current paper reports on the results of 26 Contenders implemented in commercial applications processing a total of about 15 million calls.
David Suendermann-Oeft, Jackson Liscombe, Jonathan Bloom, Grace Li, Roberto Pieraccini
INTERSPEECH1
2010 A Non-parameterised Hierarchical Pole-based Clustering Algorithm (HPoBC)
Amparo Albalate, Steffen Rhinow, David Suendermann-Oeft
ICAART (1)3
2010 On Ambiguity Detection and Postprocessing Schemes using Cluster Ensembles
Amparo Albalate, Aparna Suchindranath, Mehmet Muti Soenmez, David Suendermann-Oeft
ICAART (1)4
2010 Optimize the obvious: Automatic call flow generation
abstract
In commercial spoken dialog systems, call flows are built by call flow designers implementing a predefined business logic. While it may appear obvious from this logic how the call flow has to look like, i.e., which pieces of information have to be gathered from the caller or back-end systems and in which sequence, there are, in fact, strong arguments for automating call flow generation: 1) manual generation is time-consuming 2) manual generation is suboptimal and error-prone 3) automatic generation can react on dynamically changing business logic or external factors such as the distribution of callers and call reasons This paper presents a method for automatically deriving a call flow minimizing the average number of user turns given a business logic and a frequency distribution of call reasons. As an example, we applied the method to a call routing application whose manually built call flow is processing about 4 million calls per month and whose call reason distribution served to measure the impact of the automatic call flow generation.
David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini
ICASSP1
2010 A semi-supervised cluster-and-label approach for utterance classification
Amparo Albalate, Aparna Suchindranath, David Suendermann-Oeft, Wolfgang Minker
INTERSPEECH3
2010 Is it possible to predict task completion in automated troubleshooters?
abstract
Thede online prediction of task success in Interactive Voice Response (IVR) systems is a comparatively new field of research. It helps to identify problemantic calls and enables the dialog system to react before the caller gets overly frustrated. This publication investigates, to which extent it is possible to predict task completion and how existing approaches generalize for long dialogs. We compare the performance of two different modeling techniques: linear modeling and n-gram modeling. We show that n-gram modeling outperforms linear modeling significantly at later prediction points. From a comprehensive set of interaction parameters, we identify the relevant ones using the Information Gain Ratio. New interaction parameters are presented and evaluated. The study is based on 41,422 calls from an automated Internet troubleshooter with an average of 21.4 turns per call.
Alexander Schmitt, Wolfgang Minker, Jackson Liscombe, David Suendermann-Oeft
INTERSPEECH5
2010 Minimally invasive surgery for spoken dialog systems
abstract
We demonstrate three techniques (Escalator, Engager, and EverywhereContender) designed to optimize performance of commercial spoken dialog systems. These techniques have in common that they produce very small or no negative performance impact even during a potential experimental phase. This is because they can either be applied offline to data collected on a deployed system, or they can be incorporated conservatively such that only a low percentage of calls will get affected until the optimal strategy becomes apparent.
David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini
INTERSPEECH1
2010 How to Drink from a Fire Hose: One Person Can Annoscribe One Million Utterances in One Month
David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini
SIGDIAL Conference1
2010 Contender
abstract
Contender (or what the academic community would refer to as a light version of reinforcement learning) is a simple technique to experiment with a number of competing paths in a (commercial) spoken dialog system. By randomly routing certain portions of traffic to individual paths and computing average rewards for each of the routes, the goal is to find out which one performs best. This paper is to do away with common uncertainties on how to set up contender weights, how much data needs to be accumulated to draw reliable conclusions, and how this all relates to the notion of statistical significance.
David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini
SLT1
2009 From rule-based to statistical grammars: Continuous improvement of large-scale spoken dialog systems
abstract
Statistical Spoken Language Understanding grammars (SSLUs) are often used only at the top recognition contexts of modern large-scale spoken dialog systems. We propose to use SSLUs at every recognition context in a dialog system, effectively replacing conventional, manually written grammars. Furthermore, we present a methodology of continuous improvement in which data are collected at every recognition context over an entire dialog system. These data are then used to automatically generate updated context-specific SSLUs at regular intervals and, in so doing, continually improve system performance over time. We have found that SSLUs significantly and consistently outperform even the most carefully designed rule-based grammars in a wide range of contexts in a corpus of over two million utterances collected for a complex call-routing and troubleshooting dialog system.
David Suendermann-Oeft, Keelan Evanini, Jackson Liscombe, Phillip Hunter, Krishna Dayanidhi, Roberto Pieraccini
ICASSP1
2009 Localization of speech recognition in spoken dialog systems: how machine translation can make our lives easier
abstract
The localization of speech recognition for large-scale spoken dialog systems can be a tremendous exercise. Usually, all in- volved grammars have to be translated by a language expert, and new data has to be collected, transcribed, and annotated for statistical utterance classifiers resulting in a time-consuming and expensive undertaking. Often though, a vast number of transcribed and annotated utterances exists for the source lan- guage. In this paper, we propose to use such data and translate it into the target language using machine translation. The trans- lated utterances and their associated (original) annotations are then used to train statistical grammars for all contexts of the target system. As an example, we localize an English spoken dialog system for Internet troubleshooting to Spanish by trans- lating more than 4 million source utterances without any human intervention. In an application of the localized system to more than 10,000 utterances collected on a similar Spanish Internet troubleshooting system, we show that the overall accuracy was only 5.7% worse than that of the English source system. Index Terms: spoken dialog systems, machine translation, lo- calization
David Suendermann-Oeft, Jackson Liscombe, Krishna Dayanidhi, Roberto Pieraccini
INTERSPEECH1
2009 A Handsome Set of Metrics to Measure Utterance Classification Performance in Spoken Dialog Systems
David Suendermann-Oeft, Jackson Liscombe, Krishna Dayanidhi, Roberto Pieraccini
SIGDIAL Conference1
2008 Caller Experience: A method for evaluating dialog systems and its automatic prediction
abstract
In this paper we introduce a subjective metric for evaluating the performance of spoken dialog systems, caller experience (CE). CE is a useful metric for tracking the overall performance of a system in deployment, as well as for isolating individual problematic calls in which the system underperforms. The proposed CE metric differs from most performance evaluation metrics proposed in the past in that it is a) a subjective, qualitative rating of the call, and b) provided by expert, external listeners, not the callers themselves. The results of an experiment in which a set of human experts listened to the same calls three times are presented. The fact that these results show a high level of agreement among different listeners, despite the subjective nature of the task, demonstrates the validity of using CE as a standard metric. Finally, an automated rating system using objective measures is shown to perform at the same high level as the humans. This is an important advance, since it provides a way to reduce the human labor costs associated with producing a reliable CE.
Keelan Evanini, Phillip Hunter, Jackson Liscombe, David Suendermann-Oeft, Krishna Dayanidhi, Roberto Pieraccini
SLT4
2008 C5
abstract
The annotation of hundreds of thousands of utterances for the training of statistical utterance classifiers requires a careful quality assurance procedure to make the data consistent and reliable. In this paper, we present five methods to analyze different aspects of annotated data to ensure their Completeness, Consistency, Correlation, Congruence and to avoid Confusion-collectively referred to as C5.
David Suendermann-Oeft, Jackson Liscombe, Keelan Evanini, Krishna Dayanidhi, Roberto Pieraccini
SLT1
2007 Call classification for automated troubleshooting on large corpora
abstract
This paper compares six algorithms for call classification in the framework of a dialog system for automated troubleshooting. The comparison is carried out on large datasets, each consisting of over 100,000 utterances from two domains: television (TV) and Internet (INT). In spite of the high number of classes (79 for TV and 58 for INT), the best classifier (maximum entropy on word bigrams) achieved more than 77% classification accuracy on the TV dataset and 81% on the INT dataset.
Keelan Evanini, David Suendermann-Oeft, Roberto Pieraccini
ASRU2
2006 Text-Independent Voice Conversion Based on Unit Selection
abstract
So far, most of the voice conversion training procedures are text-dependent, i.e., they are based on parallel training utterances of source and large speaker. Since several applications (e.g. speech-to-speech translation or dubbing) require text-independent training, over the last two years, training techniques that use non-parallel data were proposed In this paper, we present a new approach that applies unit selection to find corresponding time frames in source and target speech. By means of a subjective experiment it is shown that this technique achieves the same performance as the conventional text-dependent training
David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Alan W. Black, Shri Narayanan
ICASSP (1)1
2006 Text-independent cross-language voice conversion
abstract
So far, cross-language voice conversion requires at least one bilingual speaker and parallel speech data to perform the training. This paper shows how these obstacles can be overcome by means of a recently presented text-independent training method based on unit selection. The new method is evaluated in the framework of the European speech-to-speech translation project TC-Star and achieves a performance similar to that of text-dependent intralingual voice conversion.
David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Julia Hirschberg
INTERSPEECH1
2005 A Study on Residual Prediction Techniques for Voice Conversion
abstract
Several well-studied voice conversion techniques use line spectral frequencies as features to represent the spectral envelopes of the processed speech frames. In order to return to the time domain, these features are converted to linear predictive coefficients that serve as coefficients of a filter applied to an unknown residual signal. We compare several residual prediction approaches that have already been proposed in the literature dealing with voice conversion. We also present a novel technique that outperforms the others in terms of voice conversion performance and sound quality.
David Suendermann-Oeft, Antonio Bonafonte, Hermann Ney, Harald Höge
ICASSP (1)1
2005 Evaluation of VTLN-based voice conversion for embedded speech synthesis
abstract
Recently, we demonstrated that vocal tract length normalization (VTLN) can be applied to voice conversion tasks.In particular, when the conversion algorithm is performed in time domain, this technique is very resource-efficient and, consequently, suitable for embedded applications.In this paper, we use VTLNbased voice conversion as a novel feature of a small footprint speech synthesizer running on mobile devices.The characteristics of this feature are investigated by means of extensive subjective tests.
David Suendermann-Oeft, Guntram Strecha, Antonio Bonafonte, Harald Höge, Hermann Ney
INTERSPEECH1
2004 Error Measures and Bayes Decision Rules Revisited with Applications to POS Tagging
Hermann Ney, Maja Popovic, David Suendermann-Oeft
EMNLP3
2004 A first step towards text-independent voice conversion
abstract
So far, all conventional voice conversion approaches are text-dependent, i.e., they need equivalent training utterances of source and target speaker. Since several recently proposed applications call for renouncing this requirement, in this paper, we present an algorithm which finds corresponding time frames within text-independent training data. The performance of this algorithm is tested by means of a voice conversion framework based on linear transformation of the spectral envelope. Experimental results are reported on a Spanish cross-gender corpus utilizing several objective error measures.
Hermann Ney, David Suendermann-Oeft, Antonio Bonafonte, Harald Höge
INTERSPEECH2