VLDB 2026 Research / reviewers in the wild / expert
Elizabeth Shriberg
dblp:53/2688 · also Liz Shriberg
· DBLP profile ↗
132ranked-venue papers
18as first author
6since 2021 · last 2023
0009-0004-3779-4956ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 116 · 18 first-author · 5 since 2021Artificial intelligence and machine learning · 78 · 15 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Probabilistic Performance Bounds for Evaluating Depression Models Given Noisy Self-Report LabelsabstractAdvances in AI for health applications rely on evaluating performance against labeled test data. In the area of mental health, self-report labels from surveys such as the Patient Health Questionnaire (PHQ) for depression, are useful but noisy. This "fuzzy label" problem is not currently reflected in reporting model performance, adding to the challenge of comparing results across diverse corpora, data sizes, metrics, and test label distributions.To address this issue, we develop an approach inspired by Bayes Error to estimate a model’s upper and lower performance bounds. Unlike past work, our approach can be used for both regression and classification. The method starts with a perfect match between target and prediction vectors, then applies label noise to degrade performance. To obtain confidence intervals, we use test-set bootstrapping to produce prediction and target vectors.We present results using voice-based deep learning models that predict depression risk from a conversational speech sample. Models capture both language and acoustic information. For label noise, we introduce results from a corpus in which 5625 unique subjects completed the PHQ-8 twice, separated by a short distraction task. Speech test data come from three real-world corpora encompassing over 3500 total datapoints. The test sets differ in speech elicitation, speech length, and speaker demographics among other factors.Results illustrate how probabilistic performance bounds based on PHQ-8 label noise affect the interpretation and comparison of models over corpora and metrics. Implications for science, technology, and future directions are discussed. Robert Rozanski, Elizabeth Shriberg, Amir Harati, Tomasz Rutowski, Piotr Chlebek, Tulio Goulart |
BIBM | 2 |
| 2022 | Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion RecognitionabstractWe propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more "emotion aware". We generate targets for the sentiment classification using text-to-sentiment model trained on publicly available data. Finally, we fine-tune the acoustic ASR on emotion annotated speech data. We evaluated the proposed approach on MSP-Podcast dataset, where we achieved the best reported concordance correlation coefficient (CCC) of 0.41 for valence prediction. Ayoub Ghriss, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018 |
ICASSP | 4 |
| 2022 | Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and PrimitivesabstractConfidence estimation for Speech Emotion Recognition (SER) is instrumental in improving the reliability in the behavior of downstream applications. In this work we propose (1) a novel confidence metric for SER based on the relationship between emotion primitives: arousal, valence, and dominance (AVD) and emotion categories (ECs), (2) EmoConfidNet - a DNN trained alongside the EC recognizer to predict the proposed confidence metric, and (3) a data filtering technique used to enhance the training of EmoConfidNet and the EC recognizer. For each training sample, we calculate distances from corresponding AVD annotation vectors to centroids of each EC in the AVD space, and define EC confidences as functions of the evaluated distances. EmoConfidNet is trained to predict confidence from the same acoustic representations used to train the EC recognizer. EmoConfidNet outperforms state-of-the-art confidence estimation methods on the MSP-Podcast and IEMOCAP datasets. For a fixed EC recognizer, after we reject the same number of low confidence predictions using EmoConfidNet, we achieve a higher F1 and unweighted average recall (UAR) than when rejecting using other methods. Yang Li 0149, Constantinos Papayiannis, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018 |
ICASSP | 4 |
| 2022 | Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
Tomasz Rutowski, Amir Harati, Elizabeth Shriberg, Piotr Chlebek |
INTERSPEECH | 3 |
| 2021 | Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large CorpusabstractSpeech-based algorithms have gained interest for the management of behavioral health conditions such as depression. We explore a speech-based transfer learning approach that uses a lightweight encoder and that transfers only the encoder weights, enabling a simplified run-time model. Our study uses a large data set containing roughly two orders of magnitude more speakers and sessions than used in prior work. The large data set enables reliable estimation of improvement from transfer learning. Results for the prediction of PHQ-8 labels show up to 27% relative performance gains for binary classification; these gains are statistically significant with a p-value close to zero. Improvements were also found for regression. Additionally, the gain from transfer learning does not appear to require strong source task performance. Results suggest that this approach is flexible and offers promise for efficient implementation. Amir Harati, Elizabeth Shriberg, Tomasz Rutowski, Piotr Chlebek |
ICASSP | 2 |
| 2021 | Cross-Demographic Portability of Deep NLP-Based Depression ModelsabstractDeep learning models are rapidly gaining interest for real-world applications in behavioral health. An important gap in current literature is how well such models generalize over different populations. We study Natural Language Processing (NLP) based models to explore portability over two different corpora highly mismatched in age. The first and larger corpus contains younger speakers. It is used to train an NLP model to predict depression. When testing on unseen speakers from the same age distribution, this model performs at AUC=0.82. We then test this model on the second corpus, which comprises seniors from a retirement community. Despite the large demographic differences in the two corpora, we saw only modest degradation in performance for the senior-corpus data, achieving AUC=0.76. Interestingly, in the senior population, we find AUC=0.81 for the subset of patients whose health state is consistent over time. Implications for demographic portability of speech-based applications are discussed. Tomek Rutowski, Elizabeth Shriberg, Amir Harati, Piotr Chlebek |
SLT | 2 |
| 2019 | Optimizing Speech-Input Length for Speaker-Independent Depression Classification
Tomasz Rutowski, Amir Harati, Elizabeth Shriberg |
INTERSPEECH | 4 |
| 2018 | Crowdsourcing Emotional SpeechabstractWe describe the methodology for the collection and annotation of a large corpus of emotional speech data through crowdsourcing. The corpus offers 187 hours of data from 2,965 subjects. Data includes non-emotional recordings from each subject as well as recordings for five emotions: angry, happy-low-arousal, happy-high-arousal, neutral, and sad. The data consist of spontaneous speech elicited from subjects via a web-based tool. Subjects used their own personal recording equipment, resulting in a data set that contains variation in room acoustics, microphone, etc. This offers the advantage of matching the type of variation one would expect to see when exposing speech technology in the wild in a web-based environment. The annotation scheme covers the quality of emotion expressed through the tone of voice and what was said, along with common audio-quality issues. We discuss lessons learned in the process of the creation of this corpus. Jennifer Smith 0006, Andreas Tsiartas, Valerie Wagner, Elizabeth Shriberg, Nikoletta Bassiou |
ICASSP | 4 |
| 2017 | Analysis and prediction of heart rate using speech features from natural speechabstractInteractive voice technologies can leverage biosignals, such as heart rate (HR), to infer the psychophysiological state of the user. Voice-based detection of HR is attractive because it does not require additional sensors. We predict HR from speech using the SRI BioFrustration Corpus. In contrast to previous studies we use continuous spontaneous speech as input. Results using random forests show modest but significant effects on HR prediction. We further explore the effects on HR of speaking itself, and contrast the effects when interactions induce neutral versus frustrated responses from users. Results reveal that regardless of the user's emotional state, HR tends to increase while the user is engaged in speaking to a dialog system relative to a silent region right before speech, and that this effect is greater when the subject is expressing frustration. We also find that the user's HR does not recover to pre-speaking levels as quickly after frustrated speech as it does after neutral speech. Implications and future directions are discussed. Jennifer Smith 0006, Andreas Tsiartas, Elizabeth Shriberg, Andreas Kathol, Adrian Willoughby, Massimiliano de Zambotti |
ICASSP | 3 |
| 2017 | Sensay analyticstm: A real-time speaker-state platformabstractGrowth in voice-based applications and personalized systems has led to increasing demand for speech- analytics technologies that estimate the state of a speaker from speech. Such systems support a wide range of applications, from more traditional call-center monitoring, to health monitoring, to human-robot interactions, and more. To work seamlessly in real-world contexts, such systems must meet certain requirements, including for speed, customizability, ease of use, robustness, and live integration of both acoustic and lexical cues. This demo introduces SenSay AnalyticsTM, a platform that performs real-time speaker-state classification from spoken audio. SenSay is easily configured and is customizable to new domains, while its underlying architecture offers extensibility and scalability. Andreas Tsiartas, C. Albright, Nikoletta Bassiou, Michael W. Frandsen, I. Miller, Elizabeth Shriberg, Jennifer Smith 0006, L. Lynn Voss, Valerie Wagner |
ICASSP | 6 |
| 2017 | Inferring Stance from Prosody
Nigel G. Ward, Jason C. Carlson, Olac Fuentes, Diego Castán, Elizabeth Shriberg, Andreas Tsiartas |
INTERSPEECH | 5 |
| 2016 | Noise and reverberation effects on depression detection from speechabstractSpeech-based depression detection has gained importance in recent years, but most research has used relatively quiet conditions or examined a single corpus per study. Little is thus known about the robustness of speech cues in the wild. This study compares the effect of noise and reverberation on depression prediction using 1 ) standard mel-frequency cepstral coefficients (MFCCs), and 2) features designed for noise robustness, damped oscillator cepstral coefficients (DOCCs). Data come from the 2014 Audio-Visual Emotion Recognition Challenge (AVEC). Results using additive noise and reverberation reveal a consistent pattern of findings for multiple evaluation metrics under both matched and mismatched conditions. First and most notably: standard MFCC features suffer dramatically under test/train mismatch for both noise and reverberation; DOCC features are far more robust. Second, including higher-order cepstral coefficients is generally beneficial. Third, artificial neural networks tend to outperform support vector regression. Fourth, spontaneous speech appears to offer better robustness than read speech. Finally, a cross-corpus (and cross-language) experiment reveals better noise and reverberation robustness for DOCCs than for MFCCs. Implications and future directions for real-world robust depression detection are discussed. Vikramjit Mitra, Andreas Tsiartas, Elizabeth Shriberg |
ICASSP | 3 |
| 2016 | Privacy-Preserving Speech Analytics for Automatic Assessment of Student Collaboration
Nikoletta Bassiou, Andreas Tsiartas, Jennifer Smith 0006, Harry Bratt, Colleen Richey, Elizabeth Shriberg, Cynthia M. D'Angelo, Nonye Alozie |
INTERSPEECH | 6 |
| 2016 | The SRI CLEO Speaker-State Corpus
Andreas Kathol, Elizabeth Shriberg, Massimiliano de Zambotti |
INTERSPEECH | 2 |
| 2016 | The SRI Speech-Based Collaborative Learning CorpusabstractSRI Speech-Based Collaborative Learning Corpus was developed by SRI International and is comprised of approximately 120 hours of English speech from 134 US middle school students working collaboratively. The data set also contains orthographic transcriptions, manual annotation of collaboration, log files, and supporting documentation. This collection was part of a project investigating the utility of a speech-based learning analytics approach to collaborative learning. The goal was to determine whether detectable patterns exist in student speech that correlate with collaborative learning indicators and to provide a means of assessing collaboration quality. The participants were students in middle schools (grades six, seven and eight) located in California. Students worked in groups of three on sets of short mathematics problems based on the “cloze” task in which each student was assigned one blank and each problem required the students to work together and talk to each other to coordinate their three answers. The problems were presented on iPads with a custom software application. Data The audio data was captured by both head-mounted and table-top microphones and is released as 16 kHz, 16-bit flac compressed pcm wav. Recording sessions were manually annotated with codes that mark indicators of collaboration (I codes) and that assess the overall collaboration quality of the interaction (Q codes). Annotations are presented as UTF-8 csv files. Also included in this corpus are orthorgraphic transcripts for a subset of the audio recordings and log files for iPad usage; both are released as UTF-8 encoded plain text. Colleen Richey, Cynthia M. D'Angelo, Nonye Alozie, Harry Bratt, Elizabeth Shriberg |
INTERSPEECH | 5 |
| 2015 | Effects of feature type, learning algorithm and speaking style for depression detection from speechabstractComputational methods for speech-based detection of depression are still relatively new, and have focused on either a standard set of features or on specific additional approaches. We systematically study the effects of feature type, machine learning approach, and speaking style (read versus spontaneous) on depression prediction in the AVEC-2014 evaluation corpus, using features related to speech production, perception, acoustic phonetics, and prosody. Using a multilayer ANN we find that one feature type, MMEDuSA [2], results in a 25% relative error reduction over the AVEC-2014 baseline system [1] for both mean absolute error (MAE) and root mean squared error (RMSE). Other individual feature types perform comparably to the baseline, but have much lower dimensionality and simpler to interpret. Further improvements were achieved from fusing diverse features and systems. Finally, results suggest that the relative contribution of different feature types depends on whether the speech is spontaneous or read. Overall, spontaneous speech led to lower error rates than read speech, an important consideration for the collection of future clinical data. Vikramjit Mitra, Elizabeth Shriberg |
ICASSP | 2 |
| 2015 | Cross-corpus depression prediction from speechabstractResearch on detecting depression from speech has advanced in recent years, but most work has focused on the analysis of one corpus at a time. Given that clinical corpora are typically small, it is important to explore approaches that generalize across corpora and that could ultimately be adapted to new data. We study a new corpus of patient-clinician interactions recorded when patients are admitted to a hospital for suicide risk and again when they are released. To train prediction models, we use the 2014 AVEC challenge German speech dataset, which differs from our data in many factors (including language, context, speakers, and recording conditions). Results reveal that some of the AVEC-trained models predict scores for the clinical data that correlate with both HAM-D depression scores and with the pre-/post-admission ordering. A KL-divergence analysis within the clinical data confirms that the same feature set captures changes correlated with the HAM-D scores. Finally, read versus spontaneous speech samples in both corpora behave differently with respect to the best features and modeling approaches. Implications for the cross-corpus prediction of depression are discussed. Vikramjit Mitra, Elizabeth Shriberg, Dimitra Vergyri, Bruce Knoth, Ronald M. Salomon |
ICASSP | 2 |
| 2015 | Audio-based affect detection in web videosabstractWe present a new technique for detecting audio concepts in web content as well outline the technique's applications to video sequence parsing. Our focus is primarily on affective concepts and in order to study them we have collected a new dataset, consisting of videos where a speaker is persuading a crowd, called “Rallying a Crowd”. We develop new classifiers for graded levels of arousal in speech as well as crowd noise and music and demonstrate their effectiveness on web content. These techniques achieve high detection accuracy (58.2%) for affective concepts on this new dataset and outperform (36.8%) state-of-the-art techniques (33.1%) for semantic concepts on a previously collected dataset. We also develop a new audio sequence segmentation technique which enables us to rapidly classify subsections of test sequence audio into the aforementioned audio classes. We are thus able to robustly address the detection of affective concepts in highly variable web content as well as the computational challenge of quick classification so as to enable web scale processing. Dave Chisholm, Behjat Siddiquie, Ajay Divakaran, Elizabeth Shriberg |
ICME | 4 |
| 2015 | Prediction of heart rate changes from speech features during interaction with a misbehaving dialog systemabstractMost research on detecting a speaker’s cognitive state when interacting with a dialog system has been based on selfreports, or on hand-coded subjective judgments based on audio or audio-visual observations. This study examines two questions: (1) how do undesirable system responses affect people physiologically, and (2) to what extent can we predict physiological changes from the speech signal alone? To address these questions, we use a new corpus of simultaneous speech and high-quality physiological recordings in the product returns domain (the SRI BioFrustration Corpus). “Triggers” were used to frustrate users at specific times during the interaction to produce emotional responses at similar times during the experiment across participants. For each of eight return tasks per participant, we compared speaker-normalized pre-trigger (cooperative system behavior) regions to posttrigger (uncooperative system behavior) regions. Results using random forest classifiers show that changes in spectral and temporal features of speech can predict heart rate changes with an accuracy of ~70%. Implications for future research and applications are discussed. Andreas Tsiartas, Andreas Kathol, Elizabeth Shriberg, Massimiliano de Zambotti, Adrian Willoughby |
INTERSPEECH | 3 |
| 2015 | Speech-based assessment of PTSD in a military population using diverse feature classesabstractThere is a critical need for detection and monitoring of PostTraumatic Stress Disorder (PTSD) in both military and civilian populations. Current diagnosis is based on clinical interviews, but clinicians cannot keep up with the growing need. We examined the feasibility of using speech for assessment in a military population. We analyzed recordings of the Clinician-Administered PTSD Scale (CAPS) interview from military personnel diagnosed as PTSD positive versus negative. Three feature types were explored: frame-level spectral features, longer-range prosodic features, and lexical features. Results using gaussian backend, decision tree and neural network classifiers (for spectral and prosodic features) and boosting (for lexical features) showed an accuracy of 77% correct in split-half cross validation experiments, a figure significantly above chance (which was 61.5% for our dataset). Spectral and prosodic features outperformed lexical features, and feature combination yielded further gains. An important finding was that sparser prosodic features offered more robustness than acoustic features to channel-based variation in the interview recordings. Implications and future work are discussed. Index Terms: PTSD assessment, mental health assessment. Dimitra Vergyri, Bruce Knoth, Elizabeth Shriberg, Vikramjit Mitra, Mitchell McLaren, Luciana Ferrer, Charles Marmar |
INTERSPEECH | 3 |
| 2014 | Computationally-efficient endpointing features for natural spoken interaction with personal-assistant systemsabstractCurrent speech-input systems typically use a nonspeech threshold for end-of-utterance detection. While usually sufficient for short utterances, the approach can cut speakers off during pauses in more complex utterances. We elicit personal-assistant speech (reminders, calendar entries, messaging, search) using a recognizer with a dramatically increased endpoint threshold, and find frequent nonfinal pauses. A standard endpointer with a 500 ms threshold (latency) results in a 36% cutoff rate for this corpus. Based on the new data, we develop low-cost acoustic features to discriminate nonfinal from final pauses. Features capture periodicity, speaking rate, spectral constancy, duration/intensity, and pitch of prepausal speech - using no speech recognition, speaker or session information. Classification experiments yield 20% EER at a 100 ms latency, thereby reducing both cutoffs and latency compared with the threshold-only baseline. Additional results on computational cost, feature importance, and speaker differences are discussed. Harish Arsikere, Elizabeth Shriberg, Umut Ozertem |
ICASSP | 2 |
| 2014 | Automatic characterization of speaking styles in educational videosabstractRecent studies have shown the importance of using online videos along with textual material in educational instruction, especially for better content retention and improved concept understanding. A key question is how to select videos to maximize student engagement, particularly when there are multiple possible videos on the same topic. While there are many aspects that drive student engagement, in this paper we focus on presenter speaking styles in the video. We use crowd-sourcing to explore speaking style dimensions in online educational videos, and identify six broad dimensions: liveliness, speaking rate, pleasantness, clarity, formality and confidence. We then propose techniques based solely on acoustic features for automatically identifying a subset of the dimensions. Finally, we perform video re-ranking experiments to learn how users apply their speaking style preferences to augment textbook material. Our findings also indicate how certain dimensions are correlated with perceptions of general pleasantness of the voice. Soroosh Mariooryad, Anitha Kannan, Dilek Hakkani-Tür, Elizabeth Shriberg |
ICASSP | 4 |
| 2014 | Detecting Inappropriate Clarification Requests in Spoken Dialogue SystemsabstractAlex Liu, Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg |
SIGDIAL Conference | 6 |
| 2013 | Addressee detection for dialog systems using temporal and spectral dimensions of speaking styleabstractAs dialog systems evolve to handle unconstrained input and for use in open environments, addressee detection (detecting speech to the system versus to other people) becomes an increasingly important challenge. We study a corpus in which speakers talk both to a system and to each other, and model two dimensions of speaking style that talkers modify when changing addressee: speech rhythm and vocal effort. For each dimension we design features that do not require speech recognition output, session normalization, speaker normalization, or dialog context. Detection experiments show that rhythm and effort features are complementary, outperform lexical models based on recognized words, and reduce error rates even if word recognition is error-free. Simulated online processing experiments show that all features need only the first couple seconds of speech. Finally, we find that temporal and spectral stylistic models can be trained on outside corpora, such as ATIS and ICSI meetings, with reasonable generalization to the target task, thus showing promise for domain-independent computerversus-human addressee detectors. Elizabeth Shriberg, Andreas Stolcke, Suman V. Ravuri |
INTERSPEECH | 1 |
| 2013 | Pitch-gesture modeling using subband autocorrelation change detectionabstractCalculating speaker pitch (or f0) is typically the first computational step in modeling tone and intonation for spoken language understanding. Usually pitch is treated as a fixed, single-valued quantity. The inherent ambiguity judging the octave of pitch, as well as spurious values, leads to errors in modeling pitch gestures that propagate in a computational pipeline. We present an alternative that instead measures changes in the harmonic structure using a subband autocorrelation change detector (SACD). This approach builds upon new machine-learning ideas for how to integrate autocorrelation information across subbands. Importantly however, for modeling gestures, we preserve multiple hypotheses and integrate information from all harmonics over time. The benefits of SACD over standard pitch approaches include robustness to noise and amount of voicing. This is important for real-world data in terms of both acoustic conditions and speaking style. We discuss applications in tone and intonation modeling, and demonstrate the efficacy of the approach in a Mandarin Chinese tone-classification experiment. Results suggest that SACD could replace conventional pitch-based methods for modeling gestures in selected spoken-language processing tasks. Malcolm Slaney, Elizabeth Shriberg, Jui-Ting Huang |
INTERSPEECH | 2 |
| 2013 | Using Out-of-Domain Data for Lexical Addressee Detection in Human-Human-Computer Dialog
Andreas Stolcke, Elizabeth Shriberg |
HLT-NAACL | 3 |
| 2012 | Corpus-independent history compression for stochastic turn-taking modelsabstractStochastic turn-taking models use a truncated representation of past speech activity to specify how likely a speaker is to talk at the next instant. An unanswered question in such modeling is how far back to extend the conditioning context. We study this question using Switchboard (English, telephone) and Spontal (Swedish, face-to-face) conversations. We also explore whether to trade off precision with range when moving backward in the history. We find that (1) a nearly logarithmic compression of history is optimal, for both speaker and interlocutor; (2) the absolute duration of the conditioning context is at least 7 seconds; and (3) the compression scheme generalizes remarkably well across the two different corpora. Kornel Laskowski, Elizabeth Shriberg |
ICASSP | 2 |
| 2012 | Speaker recognition with region-constrained MLLR transformsabstractIt has been shown that standard cepstral speaker recognition models can be enhanced by region-constrained models, where features are extracted only from certain speech regions defined by linguistic or prosodic criteria. Such region-constrained models can capture features that are more stable, highly idiosyncratic, or simply complementary to the baseline system. In this paper we ask if another major class of speaker recognition models, those based on MLLR speaker adaptation transforms, can also benefit from region-constrained feature extraction. In our approach, we define regions based on phonetic and prosodic criteria, based on automatic speech recognition output, and perform MLLR estimation using only frames selected by these criteria. The resulting transform features are appended to those of a state-of-the-art MLLR speaker recognition system and jointly modeled by SVMs. Multiple regions can be added in this fashion. We find consistent gains over the baseline system in the SRE2010 speaker verification task. Andreas Stolcke, Arindam Mandal, Elizabeth Shriberg |
ICASSP | 3 |
| 2012 | ProTK: An Improved Prosody Toolkit
Jacob Okamoto, Serguei V. S. Pakhomov, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 3 |
| 2012 | Learning When to Listen: Detecting System-Addressed Speech in Human-Human-Computer DialogabstractNew challenges arise for addressee detection when multiple people interact jointly with a spoken dialog system using unconstrained natural language. We study the problem of discriminating computer-directed from human-directed speech in a new corpus of human-human-computer (H-H-C) dialog, using lexical and prosodic features. The prosodic features use no word, context, or speaker information. Results with 19% WER speech recognition show improvements from lexical features (EER=23.1%) to prosodic features (EER=12.6%) to a combined model (EER=11.1%). Prosodic features also provide a 35% error reduction over a lexical model using true words (EER from 10.2% to 6.7%). Modeling energy contours with GMMs provides a particularly good prosodic model. While lexical models perform well for commands, they confuse free-form system-directed speech with human-human speech. Prosodic models dramatically reduce these confusions, implying that users change speaking style as they shift addressees (computer versus human) within a session. Overall results provide strong support for combining simple acoustic-prosodic models with lexical models to detect speaking style differences for this task. Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 1 |
| 2011 | Bird species recognition combining acoustic and sequence modelingabstractThe goal of this work was to explore modeling techniques to improve bird species classification from audio samples. We first developed an unsupervised approach to obtain approximate note models from acoustic features. From these note models we created a bird species recognition system by leveraging a phone n-gram statistical model developed for speaker recognition applications. We found competitive performance from the note n-gram system compared to a Gaussian mixture model baseline using the same acoustic features. We found an important gain by doing score-level combination relative to the best individual system results. We verified that on most of the bird species under study there was a gain from system combination. Martin Graciarena, Michelle Delplanche, Elizabeth Shriberg, Andreas Stolcke |
ICASSP | 3 |
| 2011 | Recent progress in prosodic speaker verificationabstractWe describe recent progress in the field of prosodic modeling for speaker verification. In a previous paper, we proposed a technique for modeling syllable-based prosodic features that uses a multinomial subspace model for feature extraction and within-class covariance normalization or linear discriminant analysis for session variability compensation. In this paper, we show that performance can be significantly improved with the use of probabilistic linear discriminant analysis (PLDA) for session variability compensation. This system does not require score normalization. We report an equal error rate below 7% on a NIST 2008 task. To our knowledge, this is the best reported result to date for a prosodic system for speaker recognition. Fusion of this system with a state-of-the-art acoustic baseline system yields 10% relative improvement in the new detection cost function (DCF) as defined by NIST. Marcel Kockmann, Luciana Ferrer, Lukás Burget, Elizabeth Shriberg, Jan Cernocký |
ICASSP | 4 |
| 2011 | The SRI NIST 2010 speaker recognition evaluation systemabstractThe SRI speaker recognition system for the 2010 NIST speaker recognition evaluation (SRE) incorporates multiple subsystems with a variety of features and modeling techniques. We describe our strategy for this year's evaluation, from the use of speech recognition and speech segmentation to the individual system descriptions as well as the final combination. Our results show that under most conditions, the cepstral systems tend to perform the best, but that other, non-cepstral systems have the most complementarity. The combination of several subsystems with the use of adequate side information gives a 35% improvement on the standard telephone condition. We also show that a constrained cepstral system based on nasal syllables tends to be more robust to vocal effort variabilities. Nicolas Scheffer, Luciana Ferrer, Martin Graciarena, Sachin S. Kajarekar, Elizabeth Shriberg, Andreas Stolcke |
ICASSP | 5 |
| 2011 | Language-independent constrained cepstral features for speaker recognitionabstractConstrained cepstral systems, which select frames to match various linguistic "constraints" in enrollment and test, have shown significant improvements for speaker verification performance. Past work, however, relied on word recognition, making the approach language dependent (LD). We develop language-independent (LI) versions of constraints and compare results to parallel LD versions for English data on the NIST 2008 interview task. Results indicate that (1) LI versions show surprisingly little degradation from associated LD versions, (2) some LI constraints outperform their LD counterparts, (3) useful constraint types include phonetic, syllable position, prosodic, and speaking-rate regions, (4) benefits generally hold for different train/test lengths, and (5) constraints provide particular benefit in reducing false alarms. Overall, we conclude that constrained cepstral modeling can benefit speaker recognition without the need for language-dependent automatic speech recognition. Elizabeth Shriberg, Andreas Stolcke |
ICASSP | 1 |
| 2011 | Bootstrapping Domain Detection Using Query Click Logs for New DomainsabstractDomain detection in spoken dialog systems is usually treated as a multi-class, multi-label classification problem, and training of domain classifiers requires collection and manual annotation of example utterances. In order to extend a dialog system to new domains in a way that is seamless for users, domain detection should be able to handle utterances from the new domain as soon as it is introduced. In this work, we propose using web search query logs, which include queries entered by users and the links they subsequently click on, to bootstrap domain detection for new domains. While sampling user queries from the query click logs to train new domain classifiers, we introduce two types of measures based on the behavior of the users who entered a query and the form of the query. We show that both types of measures result in reductions in the error rate as compared to randomly sampling training queries. In controlled experiments over five domains, we achieve the best gain from the combination of the two types of sampling criteria. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Elizabeth Shriberg |
INTERSPEECH | 4 |
| 2011 | Constrained Cepstral Speaker Recognition Using Matched UBM and JFA TrainingabstractWe study constrained speaker recognition systems, or systems that model standard cepstral features that fall within particular types of speech regions. A question in modeling such systems is whether to constrain universal background model (UBM) training, joint factor analysis (JFA), or both. We explore this question, as well as how to optimize UBM model size, using a corpus of Arabic male speakers. Over a large set of phonetic and prosodic constraints, we find that the performance of a system using constrained JFA and UBM is on average 5.24 % better than when using constraint-independent (all frames) JFA and UBM. We find further improvement from optimizing UBM size based on the percentage of frames covered by the constraint. Michelle Hewlett Sanchez, Luciana Ferrer, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 3 |
| 2010 | A comparison of approaches for modeling prosodic features in speaker recognitionabstractProsodic information has been successfully used for speaker recognition for more than a decade. The best-performing prosodic system to date has been one based on features extracted over syllables obtained automatically from speech recognition output. The features are then transformed using a Fisher kernel, and speaker models are trained using support vector machines (SVMs). Recently, a simpler version of these features, based on pseudo-syllables was shown to perform well when modeled using joint factor analysis (JFA). In this work, we study the two modeling techniques for the simpler set of features. We show that, for these features, a combination of JFA systems for different sequence lengths greatly outperforms both original modeling methods. Furthermore, we show that the combination of both methods gives significant improvements over the best single system. Overall, a performance improvement of 30% in the detection cost function (DCF) with respect to the two previously published methods is achieved using very simple strategies. Luciana Ferrer, Nicolas Scheffer, Elizabeth Shriberg |
ICASSP | 3 |
| 2010 | Acoustic front-end optimization for bird species recognitionabstractThe goal of this work was to explore the optimization of the feature extraction module (front-end) parameters to improve bird species recognition. We explored optimizing the spectral and temporal parameters of a Mel cepstrum feature-based front-end, starting from common parameter values used in speech processing experiments. These features were modeled using a Gaussian mixture model (GMM) system. We found an important improvement when increasing the spectral bandwidth and increasing the number of filter banks. We found no improvement when switching the filter bank distribution from the perceptually based Mel frequency scale to a linear frequency scale. In addition, no improvement was found when we either reduced or increased the time resolution. On the other hand, we found that the best time resolution is species dependent. We did find great improvements from a species-specific combination of different front-ends with different time resolutions relative to using the same front-end time resolution for all species. Martin Graciarena, Michelle Delplanche, Elizabeth Shriberg, Andreas Stolcke, Luciana Ferrer |
ICASSP | 3 |
| 2010 | Comparing the contributions of context and prosody in text-independent dialog act recognitionabstractAutomatic segmentation and classification of dialog acts (DAs; e.g., statements versus questions) is important for spoken language understanding (SLU). While most systems have relied on word and word boundary information, interest in privacy-sensitive applications and non-ASR-based processing requires an approach that is text-independent. We propose a framework for employing both speech/non-speech-based (“contextual”) features and prosodic features, and apply it to DA segmentation and classification in multiparty meetings. We find that: (1) contextual features are better for recognizing turn edge DA types and DA boundary types, while prosodic features are better for finding floor mechanisms and backchannels; (2) the two knowledge sources are complementary for most of the DA types studied; and (3) the performance of the resulting system approaches that achieved using oracle lexical information for several DA types. These results suggest that there is significant promise in text-independent features for DA recognition, and possibly for other SLU tasks, particularly when words are not available. Kornel Laskowski, Elizabeth Shriberg |
ICASSP | 2 |
| 2010 | Speaker adaptation of language and prosodic models for automatic dialog act segmentation of speech
Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
Speech Commun. | 3 |
| 2010 | The CALO Meeting Assistant SystemabstractThe CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
IEEE Trans. Speech Audio Process. | 19 |
| 2009 | Speaker recognition using syllable-based constraints for cepstral frame selectionabstractWe describe a new GMM-UBM speaker recognition system that uses standard cepstral features, but selects different frames of speech for different subsystems. Subsystems, or ldquoconstraintsrdquo, are based on syllable-level information and combined at the score level. Results on both the NIST 2006 and 2008 test data sets for the English telephone train and test condition reveal that a set of eight constraints performs extremely well, resulting in better performance than other commonly-used cepstral models. Given the still largely-unexplored world of possible constraints and combinations, it is likely that the approach can be even further improved. Tobias Bocklet, Elizabeth Shriberg |
ICASSP | 2 |
| 2009 | Syntactically-informed models for comma predictionabstractProviding punctuation in speech transcripts not only improves readability, but it also helps downstream text processing such as information extraction or machine translation. In this paper, we improve by 7% the accuracy of comma prediction in English broadcast news by introducing syntactic features inspired by the role of commas as described in linguistics studies. We conduct an analysis of the impact of those features on other subsets of features (prosody, words...) when combined through CRFs. The syntactic cues can help characterizing large syntactic patterns such as appositions and lists which are not necessarily marked by prosody. Benoît Favre, Dilek Hakkani-Tür, Elizabeth Shriberg |
ICASSP | 3 |
| 2009 | THE SRI NIST 2008 speaker recognition evaluation systemabstractThe SRI speaker recognition system for the 2008 NIST speaker recognition evaluation (SRE) incorporates a variety of models and features, both cepstral and stylistic. We highlight the improvements made to specific subsystems and analyze the performance of various subsystem combinations in different data conditions. We show the importance of language and nativeness conditioning, as well as the role of ASR for speaker verification. Sachin S. Kajarekar, Nicolas Scheffer, Martin Graciarena, Elizabeth Shriberg, Andreas Stolcke, Luciana Ferrer, Tobias Bocklet |
ICASSP | 4 |
| 2009 | Genre effects on automatic sentence segmentation of speech: A comparison of broadcast news and broadcast conversationsabstractWe investigate genre effects on the task of automatic sentence segmentation, focusing on two important domains - broadcast news (BN) and broadcast conversation (BC). We employ an HMM model based on textual and prosodic information and analyze differences in segmentation accuracy and feature usage between the two genres using both manual and automatic speech transcripts. Experiments are evaluated using Czech broadcast corpora annotated for sentence-like units (SUs). Prosodic features capture information about pause, duration, pitch, and energy patterns. Textual knowledge sources include words, part-of-speech, and automatically induced classes. We also analyze effects of using additional textual data that is not annotated for SUs. Feature analysis reveals significant differences in both textual and prosodic feature usage patterns between the two genres. The analysis is important for building automatic understanding systems when limited matched-genre data are available, or for designing eventual genre-independent systems. Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
ICASSP | 3 |
| 2009 | Feature-based and channel-based analyses of intrinsic variability in speaker verificationabstractWe explore how intrinsic variations (those associated with the speaker rather than the recording environment) affect textindependent speaker verification performance. In a previous paper we introduced the SRI-FRTIV corpus and provided speaker verification results using a Gaussian mixture model (GMM) system on telephone-channel speech. In this paper we explore the use of other speaker verification systems on the telephone channel data and compare against the GMM baseline. We found the GMM system to be one of the more robust across all conditions. Systems relying on recognition hypotheses had a significant degradation in low vocal effort conditions. We also explore the use of the GMM system on several other channels. We found improved performance on table-top microphones compared to the telephone channel in furtive conditions and gradual degradations as a function of the distance from the microphone to the speaker. Therefore distant microphones further degrade the speaker verification performance due to intrinsic variability. Index Terms: speaker recognition, vocal effort, speaking style, intrinsic variation, furtive speech, interview speech, read Martin Graciarena, Tobias Bocklet, Elizabeth Shriberg, Andreas Stolcke, Sachin S. Kajarekar |
INTERSPEECH | 3 |
| 2009 | Modeling other talkers for improved dialog act recognition in meetingsabstractAutomatic dialog act (DA) modeling has been shown to benefit meeting understanding, but current approaches to DA recognition tend to suffer from a common problem: they underrepresent behaviors found at turn edges, during which the “floor” is negotiated among meeting participants. We propose a new approach that takes into account speech from other talkers, relying only on speech/non-speech information from all participants. We find (1) that modeling other participants improves DA detection, even in the absence of other information, (2) that only the single locally most talkative other participant matters, and (3) that 10 seconds provides a sufficiently large local context. Results further show significant performance improvements over a lexical-only system — particularly for the DAs of interest. We conclude that interaction-based modeling at turn edges can be achieved by relatively simple features and should be incorporated for improved meeting understanding. Kornel Laskowski, Elizabeth Shriberg |
INTERSPEECH | 2 |
| 2009 | Does session variability compensation in speaker recognition model intrinsic variation under mismatched conditions?abstractIntersession variability (ISV) compensation in speaker recognition is well studied with respect to extrinsic variation, but little is known about its ability to model intrinsic variation. We find that ISV compensation is remarkably successful on a corpus of intrinsic variation that is highly controlled for channel (a dominant component of ISV). The results are particularly surprising because the ISV training data come from a different corpus than do speaker train and test data. We further find that relative improvements are (1) inversely related to uncompensated performance, (2) reduced more by vocal effort train/test mismatch than by speaking style mismatch, and (3) reduced additionally for mismatches in both style and level. Results demonstrate that intersession variability compensation does model intrinsic variation, and suggest that mismatched data may be more useful than previously expected for modeling certain types of withinspeaker variability in speech. Index terms: speaker recognition, channel compensation, intersession variability compensation, intrinsic variation, speaking style, vocal effort. 1. Elizabeth Shriberg, Sachin S. Kajarekar, Nicolas Scheffer |
INTERSPEECH | 1 |
| 2009 | An Anticorrelation Kernel for Subsystem Training in Multiple Classifier Systems
Luciana Ferrer, M. Kemal Sönmez, Elizabeth Shriberg |
J. Mach. Learn. Res. | 3 |
| 2008 | System combination using auxiliary information for speaker verificationabstractRecent studies in speaker recognition have shown that score- level combination of subsystems can yield significant performance gains over individual subsystems. We explore the use of auxiliary information to aid the combination procedure. We propose a modified linear logistic regression procedure that conditions combination weights on the auxiliary information. A regularization procedure is used to control the complexity of the extended model. Several auxiliary features are explored. Results are presented for data from the 2006 NIST speaker recognition evaluation (SRE). When an estimated degree of nonnativeness for the speaker is used as auxiliary information, the proposed combination results in a 15% relative reduction in equal error rate over methods based on standard linear logistic regression, support vector machines, and neural networks. Luciana Ferrer, Martin Graciarena, Argyris Zymnis, Elizabeth Shriberg |
ICASSP | 4 |
| 2008 | Exploiting dialogue act tagging and prosodic information for action item identificationabstractAn important task for multiparty meeting understanding is extracting action items. Action items are a set of tasks that are agreed on by the participants for execution after the meeting, with specific due dates and owners. Dialogue acts, the pragmatic function of an utterance, such as question or backchannel, have been reported to be useful for various dialogue understanding tasks. On the other hand, prosodic information, such as pitch, volume, and speech rate, has been reported to be useful for segmenting a dialogue into utterances or detecting questions. In this paper we investigate the use of dialogue act tagging to improve the identification of action item descriptions and prosodic information to improve action item agreements. Our results indicate that dialogue act tagging improves the identification of action item descriptions by 5% over lexical information, and prosodic information helps discriminating backchannels from agreements with 25% absolute improvement over a baseline. Gökhan Tür, Elizabeth Shriberg |
ICASSP | 3 |
| 2008 | Effects of vocal effort and speaking style on text-independent speaker verificationabstractWe study the question of how intrinsic variations (associated with the speaker rather than the recording environment) affect text-independent speaker verification performance. Experiments using the SRI-FRTIV corpus, which systematically varies both vocal effort and speaking style, reveal that (1) “furtive” speech poses a significant challenge; (2) conversations and interviews, despite stylistic differences, are well matched; (3) high-effort oration, in contrast to high-effort read speech, shares characteristics with conversational and interview styles; and (4) train/test pairings are generally symmetrical. Implications for further work in the area are discussed. Index Terms: speaker recognition, vocal effort, speaking style, intrinsic variation, furtive speech, interview speech, read Elizabeth Shriberg, Martin Graciarena, Harry Bratt, Andreas Kathol, Sachin S. Kajarekar, Huda Jameel, Colleen Richey, Fred Goodman |
INTERSPEECH | 1 |
| 2008 | The case for automatic higher-level features in forensic speaker recognitionabstractApproaches from standard automatic speaker recognition, which rely on cepstral features, suffer the problem of lack of interpretability for forensic applications. But the growing practice of using “higher-level ” features in automatic systems offers promise in this regard. We provide an overview of automatic higher-level systems and discuss potential advantages, as well as issues, for their use in the forensic context. Index Terms: speaker recognition, higher-level features, forensics 1. Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 1 |
| 2008 | The CALO meeting speech recognition and understanding systemabstractThe CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
SLT | 19 |
| 2007 | Parameterization of Prosodic Feature Distributions for SVM Modeling in Speaker RecognitionabstractMultiple recent studies have shown that speaker recognition performance using frame-based cepstral features is improved by adding higher-level information, including prosodic and lexical features. This paper explores the important question of finding a good kernel for a system that models syllable-based prosodic features using support vector machines (SVMs). The system has been the best performing of our high-level systems in the last two NIST evaluations, and gives significant improvements when combined with cepstral-based systems. We introduce two new methods for transforming the syllable-level features into a single high-dimensional vector that can be well modeled by SVMs, resulting in significant gains in speaker recognition performance. Luciana Ferrer, Elizabeth Shriberg, Sachin S. Kajarekar, M. Kemal Sönmez |
ICASSP (4) | 2 |
| 2007 | Noise Robust Speaker Identification for Spontaneous Arabic SpeechabstractTwo important challenges for speaker recognition applications are noise robustness and portability to new languages. We present an approach that integrates multiple components and models for improved speaker identification in spontaneous Arabic speech in adverse acoustic conditions. We used two different acoustic speaker models: cepstral Gaussian mixture models (GMM) and maximum likelihood linear regression support vector machine (MLLR-SVM) models and a neural network combiner. The noise-robust components are Wiener filtering, speech-nonspeech segmentation, and frame selection. We present baselines and results on the Arabic portion of the NIST mixer data, in clean conditions and with added noise at different signal-to-noise ratios. We used two realistic noises: babble and city traffic. In both noisy scenarios, we found significant equal error rate (EER) reductions over the no-compensation condition. The various noise robustness methods gave complementary gains for both acoustic models. Finally, the combiner provides a reduction in EER over the individual systems in noisy conditions. Martin Graciarena, Sachin S. Kajarekar, Andreas Stolcke, Elizabeth Shriberg |
ICASSP (4) | 4 |
| 2007 | Comparing Evaluation Metrics for Sentence Boundary DetectionabstractIn recent NIST evaluations on sentence boundary detection, a single error metric was used to describe performance. Additional metrics, however, are available for such tasks, in which a word stream is partitioned into subunits. This paper compares alternative evaluation metrics - including the NIST error rate, classification error rate per word boundary, precision and recall, ROC curves, DET curves, precision-recall curves, and area under the curves - and discusses advantages and disadvantages of each. Unlike many studies in machine learning, we use real data for a real task. We find benefit from using curves in addition to a single metric. Furthermore, we find that data skew has an impact on metrics, and that differences among different system outputs are more visible in precision-recall curves. Results are expected to help us better understand evaluation metrics that should be generalizable to similar language processing tasks. Yang Liu 0004, Elizabeth Shriberg |
ICASSP (4) | 2 |
| 2007 | Entropy Based Classifier Combination for Sentence SegmentationabstractWe describe recent extensions to our previous work, where we explored the use of individual classifiers, namely, boosting and maximum entropy models for sentence segmentation. In this paper we extend the set of classification methods with support vector machine (SVM). We propose a new dynamic entropy-based classifier combination approach to combine these classifiers, and compare it with the traditional classifier combination techniques, namely, voting, linear regression and logistic regression. Furthermore, we also investigate the combination of hidden event language models with the output of the proposed classifier combination, and the output of individual classifiers. Experimental studies conducted on the Mandarin TDT4 broadcast news database shows that the SVM classifier as an individual classifier improves over our previous best system. However, the proposed entropy-based classifier combination approach shows the best improvement in F-measure of 1% absolute, and the voting approach shows the best reduction in NIST error rate of 2.7% absolute when compared to the previous best system. Mathew Magimai-Doss, Dilek Hakkani-Tür, Özgür Çetin, Elizabeth Shriberg, James G. Fung, Nikki Mirghafori |
ICASSP (4) | 4 |
| 2007 | Detecting deception using critical segmentsabstractWe present an investigation of segments that map to GLOBAL LIES, that is, the intent to deceive with respect to salient topics of the discourse. We propose that identifying the truth or falsity of these CRITICAL SEGMENTS may be important in determining a speaker’s veracity over the larger topic of discourse. Further, answers to key questions, which can be identified a priori, may represent emotional and cognitive HOT SPOTS, analogous to those observed by psychologists who study gestural and facial cues to deception. We present results of experiments that use two different definitions of CRITICAL SEGMENTS and employ machine learning techniques that compensate for imbalances in the dataset. Using this approach, we achieve a performance gain of 23.8% relative to chance, in contrast with human performance on a similar task, which averages substantially below chance. We discuss the features used by the models, and consider how these findings can influence future research. Frank Enos, Elizabeth Shriberg, Martin Graciarena, Julia Hirschberg, Andreas Stolcke |
INTERSPEECH | 2 |
| 2007 | A smoothing kernel for spatially related features and its application to speaker verificationabstractMost commonly used kernels are invariant to permutations of the feature vector components. This characteristic may make machine learning methods that use such kernels suboptimal in cases where the feature vector has an underlying structure. In this paper we will consider one such case, where the features are spatially related. We show a way to modify the objective function of the support vector machine (SVM) optimization problem to account for this structure. The new optimization problem can be implemented as a standard SVM using a particular smoothing kernel. Results are shown on a speaker verification task using prosodic features that are transformed using a particular implementation of the Fisher score. The proposed method leads to improvements of as much as 15 % in equal error rate (EER). Luciana Ferrer, M. Kemal Sönmez, Elizabeth Shriberg |
INTERSPEECH | 3 |
| 2007 | Cross-linguistic analysis of prosodic features for sentence segmentationabstractIn this paper, we perform a cross-linguistic study of prosodic features in sentence segmentation by using two different feature selection approaches: a forward search wrapper and feature filtering. Experiments in Arabic, English, and Mandarin show that prosodic features make significant contributions in all three languages. Feature selection results indicate that feature relevancy can vary greatly depending on the target language, and therefore the optimal feature subset varies considerably between languages. We observe patterns in the feature selection and the affinity of the different languages toward certain feature types, which gives us insight into future feature selection and feature design. Index Terms: prosodic features, cross lingual, feature selection, sentence segmentation James G. Fung, Dilek Hakkani-Tür, Mathew Magimai-Doss, Elizabeth Shriberg, Sébastien Cuendet, Nikki Mirghafori |
INTERSPEECH | 4 |
| 2007 | Speaker adaptation of language models for automatic dialog act segmentation of meetingsabstractAbstract Dialog act (DA) segmentation in meeting speech is important for meeting understanding. In this paper, we explore speaker adaptation of hidden event language models (LMs) for DA segmentation using the ICSI Meeting Corpus. Speaker adaptation is performed using a linear combination of the generic speaker independent LM and an LM trained on only the data from individual speakers. We test the method on 20 frequent speakers, on both reference word transcripts and the output of automatic speech recognition. Results indicate improvements for 17 speakers on reference transcripts, and for 15 speakers on automatic transcripts. Overall, the speaker-adapted LM yields statistically significant improvement over the baseline LM for both test conditions. Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
INTERSPEECH | 3 |
| 2007 | A text-constrained prosodic system for speaker verificationabstractWe describe four improvements to a prosody SVM system, including a new method based on textand part-of-speechconstrained prosodic features. The improved system shows remarkably good performance on NIST SRE06 data, reducing the error rate of an MLLR system by as much as 23% after combination. In addition, an N -best system analysis using eight systems reveals that the prosody SVM is the third and second most important system for 1and 8-side training conditions, respectively—providing more complementary information than other state-of-the-art cepstral systems. We conclude that as cepstral systems continue to improve, it should become only more important to develop systems based on higher-level features. Elizabeth Shriberg, Luciana Ferrer |
INTERSPEECH | 1 |
| 2007 | Duration and pronunciation conditioned lexical modeling for speaker verificationabstractWe propose a method to improve speaker recognition lexical model performance using acoustic-prosodic information. More specifically, the lexical model is trained using durationand pronunciation-conditioned word N-grams, simultaneously modeling lexical information along with their acoustic and prosodic characteristics. Support vector machines are used for modeling and scoring, with N-gram frequency vectors serving as features. Experimental results using NIST Speaker Recognition Evaluation data sets show that this method outperforms the regular word N-gram-based lexical models. Furthermore, our approach gives additional information when combined with a high-accuracy acoustic speaker model. We believe that this is a promising step toward integrated speaker recognition models that combine multiple types of high-level features. Gökhan Tür, Elizabeth Shriberg, Andreas Stolcke, Sachin S. Kajarekar |
INTERSPEECH | 2 |
| 2006 | Speaker Overlaps and ASR Errors in Meetings: Effects Before, During, and After the OverlapabstractWe analyze automatic speech recognition (ASR) errors made by a state-of-the-art meeting recognizer, with respect to locations of overlapping speech. Our analysis focuses on recognition errors made both during an overlap and in the regions immediately preceding and following the location of overlapped speech. We devise an experimental paradigm to allow examination of the same foreground speech both with and without naturally occurring cross-talk. We then analyze ASR errors with respect to a number of factors, including the severity of the cross-talk and distance from the overlap region. In addition to reporting effects on ASR errors, we discover a number of interesting phenomena. First, we find that overlaps tend to occur at high-perplexity regions in the foreground talker's speech. Second, word sequences within overlaps have higher perplexity than those in nonoverlaps, if using trigrams or 4-grams, but the unigram perplexity within overlaps is considerably lower than that of nonoverlaps. An explanation for this behavior is proposed, based on the preponderance of multiple short dialog acts found in overlap regions. Third, we discover that the word error rate (WER) after overlaps is consistently lower than that before the overlap. This finding cannot be explained by the recognition process itself; rather, the foreground speaker appears to reduce perplexity shortly after being overlapped. Taken together, these observations suggest that the automatic modeling of meetings could benefit from a broader view of the relationship between speaker overlap and ASR in natural conversation Özgür Çetin, Elizabeth Shriberg |
ICASSP (1) | 2 |
| 2006 | The Contribution of Cepstral and Stylistic Features to SRI's 2005 NIST Speaker Recognition Evaluation SystemabstractRecent work in speaker recognition has demonstrated the advantage of modeling stylistic features in addition to traditional cepstral features, but to date there has been little study of the relative contributions of these different feature types to a state-of-the-art system. In this paper we provide such an analysis, based on SRI's submission to the NIST 2005 speaker recognition evaluation. The system consists of 7 subsystems (3 cepstral 4 stylistic). By running independent N-way subsystem combinations for increasing values of N, we fines that (1) a monotonic pattern in the choice of the best N systems allows for the inference of subsystem importance; (2) the ordering of subsystems alternates between cepstral and stylistic; (3) syllable-based prosodic features are the strongest stylistic features, and (4) overall subsystem ordering depends crucially on the amount of training data (1 versus 8 conversation sides). Improvements over the baseline cepstral system, when all systems are combined, range from 47% to 67%, with larger improvements for the 8-side condition. These results provide direct evidence of the complementary contributions of cepstral and stylistic features to speaker discrimination Luciana Ferrer, Elizabeth Shriberg, Sachin S. Kajarekar, Andreas Stolcke, M. Kemal Sönmez, Anand Venkataraman, Harry Bratt |
ICASSP (1) | 2 |
| 2006 | Combining Prosodic Lexical and Cepstral Systems for Deceptive Speech DetectionabstractWe report on machine learning experiments to distinguish deceptive from nondeceptive speech in the Columbia-SRI-Colorado (CSC) corpus. Specifically, we propose a system combination approach using different models and features for deception detection. Scores from an SVM system based on prosodic/lexical features are combined with scores from a Gaussian mixture model system based on acoustic features, resulting in improved accuracy over the individual systems. Finally, we compare results from the prosodic-only SVM system using features derived either from recognized words or from human transcriptions. Martin Graciarena, Elizabeth Shriberg, Andreas Stolcke, Frank Enos, Julia Hirschberg, Sachin S. Kajarekar |
ICASSP (1) | 2 |
| 2006 | Joint Segmentation and Classification of Dialog Acts in Multiparty MeetingsabstractThis paper investigates a scheme for joint segmentation and classification of dialog acts (DAs) of the ICSI Meeting Corpus based on hidden-event language models and a maximum entropy classifier for the modeling of word boundary types. Specifically, the modeling of the boundary types takes into account dependencies between the duration of a pause and its surrounding words. Results for the proposed method compare favorably with our previous work on the same task Matthias Zimmermann, Andreas Stolcke, Elizabeth Shriberg |
ICASSP (1) | 3 |
| 2006 | Analysis of overlaps in meetings by dialog factors, hot spots, speakers, and collection site: insights for automatic speech recognitionabstractIn previous work we found that automatic speech recognition (ASR) results on meetings show interesting patterns with respect to speaker overlaps, including a robust asymmetry in word error rates (WERs) before and after overlaps. The paradigm used allowed us to infer that these correlations are not due to crosstalk itself but to changes in how a person speaks around overlap regions. To better understand these ASR and perplexity results, we analyze speaker overlaps with respect to various factors, including collection site, speakers, dialog acts, and hot spots. We examine a total of 101 meetings from the ICSI meeting corpus and the NIST meeting transcription evaluations of the last four years. We find that overlaps tend to occur at high-perplexity regions in the foreground talker’s speech. We also find that overlap regions tend to have higher perplexity than those in nonoverlaps, if trigrams or 4-grams are used, but unigram perplexity within overlaps is considerably lower than that of nonoverlaps. These appear to be robust findings, because they hold in general across meetings from different collection sites, even though meeting style and absolute rates of overlap vary by site. Further analyses of overlap with respect to speakers and meeting content reveal interesting relationships between overlap and dialog acts, as well as between overlap and “hot spots ” (points of increased participant involvement). Finally, results from the ICSI meeting corpus show that individual speakers have widely varying rates of being overlapped. Index Terms: automatic speech recognition, meeting recognition, crosstalk, speaker overlap, and dialog acts Özgür Çetin, Elizabeth Shriberg |
INTERSPEECH | 2 |
| 2006 | Personality factors in human deception detection: comparing human to machine performanceabstractPrevious studies of human performance in deception detection have found that humans generally are quite poor at this task, comparing unfavorably even to the performance of automated procedures.However, different scenarios and speakers may be harder or easier to judge.In this paper we compare human to machine performance detecting deception on a single corpus, the Columbia-SRI-Colorado Corpus of deceptive speech.On average, our human judges scored worse than chance -and worse than current best machine learning performance on this corpus.However, not all judges scored poorly.Based on personality tests given before the task, we find that several personality factors appear to correlate with the ability of a judge to detect deception in speech. Frank Enos, Stefan Benus, Robin L. Cautin, Martin Graciarena, Julia Hirschberg, Elizabeth Shriberg |
INTERSPEECH | 6 |
| 2006 | On speaker-specific prosodic models for automatic dialog act segmentation of multi-party meetingsabstractTento článek zkoumá prozodické modely specifické pro jednotlivé řečníky, které jsou používány pro automatickou segmentaci řeči z ICSI meetings korpusu na dialogové akty. Zkoumáme, zda-li je výhodné trénovat tyto modely pouze na řeči konkrétního řečníka. Jáchym Kolár, Elizabeth Shriberg, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2006 | CHAT: a conversational helper for automotive tasksabstractSpoken dialogue interfaces, mostly command-and-control, become more visible in applications where attention needs to be shared with other tasks, such as driving a car. The deployment of the simple dialog systems, instead of more sophisticated ones, is partly because the computing platforms used for such tasks have been less powerful and partly because certain issues from these cognitively challenging tasks have not been well addressed even in the most advanced dialog systems. This paper reports the progress of our research effort in developing a robust, wide-coverage, and cognitive load-sensitive spoken dialog interface called CHAT: Conversational Helper for Automotive Tasks. Our research in the past few years has led to promising results, including high task completion rate, dialog efficiency, and improved user experience. Index Terms: dialog systems, cognitive load, robustness Fuliang Weng, Sebastian Varges, Badri Raghunathan, Florin Ratiu, Heather Pon-Barry, Brian Lathrop, Harry Bratt, Tobias Scheideck, Matthew Purver, Annie Lien, Madhuri Raya, Stanley Peters, J. Russell, Lawrence Cavedon, Elizabeth Shriberg, Hauke Schmidt, R. Prieto |
INTERSPEECH | 19 |
| 2006 | The ICSI+ multilingual sentence segmentation systemabstractThe ICSI+ multilingual sentence segmentation with results for English and Mandarin broadcast news automatic speech recognizer transcriptions represents a joint effort involving ICSI, SRI, and UT Dallas. Our approach is based on using hidden event language models for exploiting lexical information, and maximum entropy and boosting classifiers for exploiting lexical, as well as prosodic, speaker change and syntactic information. We demonstrate that the proposed methodology including pitch- and energy-related prosodic features performs significantly better than a baseline system that uses words and simple pause features only. Furthermore, the obtained improvements are consistent across both languages, and no language-specific adaptation of the methodology is necessary. The best results were achieved by combining hidden event language models with a boosting-based classifier that to our knowledge has not previously been applied for this task. M. Zimmerman, Dilek Hakkani-Tür, James G. Fung, Nikki Mirghafori, Luke R. Gottlieb, Elizabeth Shriberg, Yang Liu 0004 |
INTERSPEECH | 6 |
| 2006 | A study in machine learning from imbalanced data for sentence boundary detection in speech
Yang Liu 0004, Nitesh V. Chawla, Mary P. Harper, Elizabeth Shriberg, Andreas Stolcke |
Comput. Speech Lang. | 4 |
| 2006 | Enriching speech recognition with automatic detection of sentence boundaries and disfluenciesabstractEffective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Mary P. Harper |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Using Conditional Random Fields for Sentence Boundary Detection in SpeechabstractSentence boundary detection in speech is important for enriching speech recognition output, making it easier for humans to read and downstream modules to process. In previous work, we have developed hidden Markov model (HMM) and maximum entropy (Maxent) classifiers that integrate textual and prosodic knowledge sources for detecting sentence boundaries. In this paper, we evaluate the use of a conditional random field (CRF) for this task and relate results with this model to our prior work. We evaluate across two corpora (conversational telephone speech and broadcast news speech) on both human transcriptions and speech recognition output. In general, our CRF model yields a lower error rate than the HMM and Maxent models on the NIST sentence boundary detection task in speech, although it is interesting to note that the best results are achieved by three-way voting among the classifiers. This probably occurs because each model has different strengths and weaknesses for modeling the knowledge sources. Yang Liu 0004, Andreas Stolcke, Elizabeth Shriberg, Mary P. Harper |
ACL | 3 |
| 2005 | Automatic Dialog Act Segmentation and Classification in Multiparty MeetingsabstractWe explore the two related tasks of dialog act (DA) segmentation and DA classification for speech from the ICSI Meeting Corpus. We employ simple lexical and prosodic knowledge sources, and compare results for human-transcribed versus automatically recognized words. Since there is little previous work on DA segmentation and classification in the meeting domain, our study provides baseline performance rates for both tasks. We introduce a range of metrics for use in evaluation, each of which measures different aspects of interest. Results show that both tasks are difficult, particularly for a fully automatic system. We find that a very simple prosodic model aids performance over lexical information alone, especially for segmentation. Both tasks, but particularly word-based segmentation, are degraded by word recognition errors. Finally, while classification results for meeting data show some similarities to previous results for telephone conversations, findings also suggest a potential difference with respect to the effect of modeling DA context. Jeremy Ang, Yang Liu 0004, Elizabeth Shriberg |
ICASSP (1) | 3 |
| 2005 | SRI's 2004 NIST Speaker Recognition Evaluation SystemabstractThe paper describes our recent efforts in exploring longer-range features and their statistical modeling techniques for speaker recognition. In particular, we describe a system that uses discriminant features from cepstral coefficients, and systems that use discriminant models from word n-grams and syllable-based NERF n-grams. These systems together with a cepstral baseline system are evaluated on the 2004 NIST speaker recognition evaluation dataset. The effect of the development set is measured using two different datasets, one from Switchboard databases and another from the FISHER database. Results show that the difference between the development and evaluation sets affects the performance of the systems only when more training data is available. Results also show that systems using longer-range features combined with the baseline result in about a 31% improvement with 1-side training over the baseline system and about a 61% improvement with 8-side training over the baseline system. Sachin S. Kajarekar, Luciana Ferrer, Elizabeth Shriberg, M. Kemal Sönmez, Andreas Stolcke, Anand Venkataraman, Jing Zheng 0001 |
ICASSP (1) | 3 |
| 2005 | Structural metadata research in the EARS programabstractBoth human and automatic processing of speech require recognition of more than just words. In this paper we provide a brief overview of research on structural metadata extraction in the DARPA EARS rich transcription program. Tasks include detection of sentence boundaries, filler words, and disfluencies. Modeling approaches combine lexical, prosodic, and syntactic information, using various modeling techniques for knowledge source integration. The performance of these methods is evaluated by task, by data source (broadcast news versus spontaneous telephone conversations) and by whether transcriptions come from humans or from an (errorful) automatic speech recognizer. A representative sample of results shows that combining multiple knowledge sources (words, prosody, syntactic information) is helpful, that prosody is more helpful for news speech than for conversational speech, that word errors significantly impact performance, and that discriminative models generally provide benefit over maximum likelihood models. Important remaining issues, both technical and programmatic, are also discussed. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Barbara Peskin, Jeremy Ang, Dustin Hillard, Mari Ostendorf, Marcus Tomalin, Philip C. Woodland, Mary P. Harper |
ICASSP (5) | 2 |
| 2005 | Human language technology: opportunities and challengesabstractIn recent years, there has been dramatic progress in both speech and language processing, in many cases leveraging some of the same underlying methods. This progress and the growing technical ties motivate efforts to combine speech and language technologies in spoken document processing applications. This paper outlines some of the issues involved, as well as the opportunities, presenting an overview of the special double session on this topic. Mari Ostendorf, Elizabeth Shriberg, Andreas Stolcke |
ICASSP (5) | 2 |
| 2005 | Distinguishing deceptive from non-deceptive speechabstractTo date, studies of deceptive speech have largely been confined to descriptive studies and observations from subjects, researchers, or practitioners, with few empirical studies of the specific lexical or acoustic/prosodic features which may characterize deceptive speech.We present results from a study seeking to distinguish deceptive from non-deceptive speech using machine learning techniques on features extracted from a large corpus of deceptive and non-deceptive speech.This corpus employs an interview paradigm that includes subject reports of truth vs. lie at multiple temporal scales.We present current results comparing the performance of acoustic/prosodic, lexical, and speaker-dependent features and discuss future research directions. Julia Hirschberg, Stefan Benus, Jason M. Brenier, Frank Enos, Sarah Friedman, Sarah Gilman, Cynthia Girand, Martin Graciarena, Andreas Kathol, Laura A. Michaelis, Bryan L. Pellom, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 12 |
| 2005 | Two experiments comparing reading with listening for human processing of conversational telephone speechabstractWe report on results of two experiments designed to compare subjects ’ ability to extract information from audio recordings of conversational telephone speech (CTS) with their ability to extract information from text transcripts of these conversations, with and without the ability to hear the audio recordings. Although progress in machine processing of CTS speech is well documented, human processing of these materials has not been as well studied. These experiments compare subject’s processing time and comprehension of widely-available CTS data in audio and written formats – one experiment involves careful reading and one involves visual scanning for information. We observed a very modest improvement using transcripts compared with the audio-only condition for the careful reading task (speed-up by a factor of 1.2) and a much more dramatic improvement using transcripts in the visual scanning task (speed-up by a factor of 2.9). The implications of the experiments are twofold: (1) we expect to see similar gains in human productivity for comparable applications outside the laboratory environment and (2) the gains can vary widely, depending on the specific tasks involved. 1. Douglas A. Jones, Wade Shen, Elizabeth Shriberg, Andreas Stolcke, Teresa M. Kamm, Douglas A. Reynolds |
INTERSPEECH | 3 |
| 2005 | Comparing HMM, maximum entropy, and conditional random fields for disfluency detectionabstractAutomatic detection of disfluencies in spoken language is important for making speech recognition output more readable, and for aiding downstream language processing modules. We compare a generative hidden Markov model (HMM)-based approach and two conditional models — a maximum entropy (Maxent) model and a conditional random field (CRF) — for detecting disfluencies in speech. The conditional modeling approaches provide a more principled way to model correlated features. In particular, the CRF approach directly detects the reparandum regions, and thus avoids the use of ad-hoc heuristic rules. We evaluate performance of these three models across two different corpora (conversational speech and broadcast news) and for two types of transcriptions (human transcriptions and recognition output). Overall we find that that the conditional modeling approaches (Maxent and CRF) provide benefit over the HMM approach. Effects of speaking style, word recognition errors, and future directions are also discussed. 1. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Mary P. Harper |
INTERSPEECH | 2 |
| 2005 | Spontaneous speech: how people really talk and why engineers should careabstractSpontaneous conversation is optimized for human-human communication, but differs in some important ways from the types of speech for which human language technology is often developed. This overview describes four fundamental properties of spontaneousspeech that present challenges for spoken language applications because they violate assumptions often applied in automatic processing technology. 1. Elizabeth Shriberg |
INTERSPEECH | 1 |
| 2005 | MLLR transforms as features in speaker recognitionabstractWe explore the use of adaptation transforms employed in speech recognition systems as features for speaker recognition. This approach is attractive because, unlike standard framebased cepstral speaker recognition models, it normalizes for the choice of spoken words in text-independent speaker verification. Affine transforms are computed for the Gaussian means of the acoustic models used in a recognizer, using maximum likelihood linear regression (MLLR). The high-dimensional vectors formed by the transform coefficients are then modeled as speaker features using support vector machines (SVMs). The resulting speaker verification system is competitive, and in some cases significantly more accurate, than state-of-the-art cepstral gaussian mixture and SVM systems. Further improvements are obtained by combining baseline and MLLR-based systems. 1. Andreas Stolcke, Luciana Ferrer, Sachin S. Kajarekar, Elizabeth Shriberg, Anand Venkataraman |
INTERSPEECH | 4 |
| 2005 | Does active learning help automatic dialog act tagging in meeting data?abstractKnowledge of Dialog Acts (DAs) is important for the automatic understanding and summarization of meetings. Current approaches rely on a lot of hand labeled data to train automatic taggers. One approach that has been successful in reducing the amount of training data in other areas of NLP is active learning. We ask if active learning with lexical cues can help for this task and this domain. To better address this question, we explore active learning for two different types of DA models – hidden Markov models (HMMs) and maximum entropy (maxent). Anand Venkataraman, Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 3 |
| 2005 | Modeling prosodic feature sequences for speaker recognition
Elizabeth Shriberg, Luciana Ferrer, Sachin S. Kajarekar, Anand Venkataraman, Andreas Stolcke |
Speech Commun. | 1 |
| 2004 | Identifying Agreement and Disagreement in Conversational Speech: Use of Bayesian Networks to Model Pragmatic DependenciesabstractWe describe a statistical approach for modeling agreements and disagreements in conversational interaction. Our approach first identifies adjacency pairs using maximum entropy ranking based on a set of lexical, durational, and structural features that look both forward and backward in the discourse. We then classify utterances as agreement or disagreement using these adjacency pairs and features that represent various pragmatic influences of previous agreement or disagreement on the current utterance. Our approach achieves 86.9% accuracy, a 4.9% increase over previous work. Michel Galley, Kathy McKeown, Julia Hirschberg, Elizabeth Shriberg |
ACL | 4 |
| 2004 | Comparing and Combining Generative and Posterior Probability Models: Some Advances in Sentence Boundary Detection in Speech
Yang Liu 0004, Andreas Stolcke, Elizabeth Shriberg, Mary P. Harper |
EMNLP | 3 |
| 2004 | Multimodal model integration for sentence unit detectionabstractIn this paper, we adopt a direct modeling approach to utilize conversational gesture cues in detecting sentence boundaries, called SUs, in video taped conversations. We treat the detection of SUs as a classification task such that for each inter-word boundary, the classifier decides whether there is an SU boundary or not. In addition to gesture cues, we also utilize prosody and lexical knowledge sources. In this first investigation, we find that gesture features complement the prosodic and lexical knowledge sources for this task. By using all of the knowledge sources, the model is able to achieve the lowest overall SU detection error rate. Mary P. Harper, Elizabeth Shriberg |
ICMI | 2 |
| 2004 | Using machine learning to cope with imbalanced classes in natural speech: evidence from sentence boundary and disfluency detectionabstractWe investigate machine learning techniques for coping with highly skewed class distributions in two spontaneous speech processing tasks. Both tasks, sentence boundary and disfluency detection, provide important structural information for downstream language processing modules. We examine the effect of data set size, task, sampling method (no sampling, downsampling, oversampling, and ensemble sampling), and learning method (bagging, ensemble bagging, and boosting) for a decision tree prosody model. Results show that (1) bagging benefits both tasks, but to different degrees, (2) the benefit from ensemble bagging decreases as data size increases, and (3) boosting can outperform bagging under certain conditions. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Mary P. Harper |
INTERSPEECH | 2 |
| 2004 | A wizard of oz framework for collecting spoken human-computer dialogsabstractAbstract This paper describes a data collection process aimed atgathering human-computer dialogs in high-stress or “busy”do-mains where the user is concentrating on tasks other than theconversation, for example, when driving a car. Designing spo-ken dialog interfaces for suchdomains is extremely challengingand the data collected will help us improve the dialog systemin-terfaceand performance,understandhowhumansperformthesetasks with respect to stressful situations, and obtain speech ut-terances for extracting prosodic features. This paper describesthe experimental design for collecting speech data in a simu-lated driving environment. 1. Background Research in human-computer interfaces has been carried outinapplications where the useris focusedon taskssuch as driving acar [4] or operating other machinery,with the goal of designinginterfaces that will help reduce the user’s overall cognitive load.In such applications, the user normally controls several devicessimultaneously. Existing applications maintain little or no dia-log context, and require the userto learn and remember compli-cated sets of device-specific commands. To overcome some ofthe shortcomings of such systems, researchers have been inves-tigating designing spokeninterface systems which can conversewith the user more naturally, allowing more flexibility in th euser’s speech and keeping track of the dialog context, similarto how a human speech partner would [6, 7]. However, human-humanspeechin suchscenariosis highly context-and situation-dependent,full of disfluencies(e.g., false starts and paus es)andsentence fragments (abandoned or repaired utterances), and ishighly interactive and collaborative. We believe that the easiestinterfaces to use will be those that mimic human-human inter-action in some, though perhaps not all, respects. Therefore,our data collection focuses on collecting the kind of speech thatwould occur between a human and a system that is as flexibleand capable as that user would desire.Our goal is a system should mimic human-human interac-tions by understanding the user’s requests and producing re-sponsesbased on the user’s knowledge, the conversationalcon-text, and the external situation. We use the car-driving domainas a testbed ofsucha dialog interface for operating in-carequip-ment, such as obtaining navigation information (e.g., turn-by-turn instructions) and information about local points of interest.Figure 1 illustrates the systemcomponents,which include a lan-guage understanding component,a response generator, a dialogmanager and a prosody classifier. We use off-the-shelf tech-nologies and tools for speechrecognition, speechsynthesis,andknowledge management.1.1. Purposes for Data CollectionThe ultimate goal of our dialog system is to enable natural in-teractions between the driver and the system to be like thosebetween humans. Therefore collecting human-human dialogsfor the above tasks helps us to develop and tune the system tosimulate such interactions. As the first step of our system de -velopment, data collection has the following specific purpo ses.Improve the system interface and performance: Languagecoverage has been a bottleneck for existing dialog systems. Arobust dialog system should allow the user to speak freely andbe ableto understandthe user’sintention expressedthroughvar-ious utterances. The robustness of a system can only be en-hanced using a large amount of data that are expected to covermost language phenomena in the target application. Thereforewe aim to collect dialogs from many subjects and to use thesedata to train the language understanding component. The datawill also provide evidence as to what features users would de-sire in an in-car conversationalsystem.Understand how humans give navigation instructions in adriving situation: Although human navigation data has beencollected for developing systems that automatically generatenavigation instructions, e.g. [1], the data is often written de-scriptions based on the subject’s mental recap of the route. Ina driving environment,humans might chooseto give navigationinformation differently with respect to the current position ofthe vehicle (e.g., close to a turn) and external situations (e.g.,emergency stop). There is a need, therefore, to collect new datato discover what kinds of strategies humans would use to con-vey navigation information in a real-time setting.Obtain speech utterances for extracting prosody features:Drivers are likely to produce disfluent and distracted speec hwith potentially complex syntax when focusing on tasks otherthan talking. Such data contain rich prosodic information thatcaptures variations in timing (e.g., lengthened sounds, pauses),intonation (e.g., pitch rise/fall at the end of utterance), and loud-ness. These features convey information beyond that carried bythe words themselves. They can help a dialog system detectutterance boundaries, driver intention and stress level, and sub-sequently generate appropriate responses which take into ac-count the driver’s emotional state. They can also help augmentthe information available to a natural language parser, to help Elizabeth Shriberg, Sandra Upson, Joyce Chen, Fuliang Weng, Stanley Peters, Lawrence Cavedon, John Niekrasz, Harry Bratt |
INTERSPEECH | 2 |
| 2004 | SVM modeling of "SNERF-grams" for speaker recognitionabstractWe describe a new approach to modeling idiosyncratic prosodic behavior for automatic speaker recognition. The approach computes prosodic features by syllable (syllablebased nonuniform extraction region features, or “SNERFs”), and models the syllable-feature sequences (“SNERF-grams”) using support vector machines (SVMs). We evaluate performance on development data for a system submitted to the NIST 2004 Speaker Recognition Evaluation. Results show that SNERF-grams provide significant performance gains when combined with a state-of-the-art baseline system, as well as with both prosodic and word-based noncepstral systems. 1. Elizabeth Shriberg, Luciana Ferrer, Anand Venkataraman, Sachin S. Kajarekar |
INTERSPEECH | 1 |
| 2004 | The ICSI-SRI-UW metadata extraction systemabstractBoth human and automatic processing of speech require recognizing more than just the words. We describe a state-of-the-art system for automatic detection of “metadata” (information beyond the words) in both broadcast news and spontaneous telephone conversations, developed as part of the DARPA EARS Rich Transcription program. System tasks include sentence boundary detection, filler word detection, and detection/correction of disfluencies. To achieve best performance, we combine information from different types of language models (based on words, part-of-speech classes, and automatically induced classes) with information from a prosodic classifier. The prosodic classifier employs bagging and ensemble approaches to better estimate posterior probabilities. We use confusion networks to improve robustness to speech recognition errors. Most recently, we have investigated a maximum entropy approach for the sentence boundary detection task, yielding a gain over our standard HMM approach. We report results for these techniques on the official NIST Rich Transcription metadata tasks. Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Barbara Peskin, Mary P. Harper, Yang Liu 0004 |
INTERSPEECH | 1 |
| 2004 | A conversational dialogue system for cognitively overloaded usersabstractSpoken dialogue interfaces are gaining increased acceptance in a wide range of applications. Most current examples of such systems, however, rely on using restricted language and scripted dialogue interactions. We argue that speech interfaces in highly stressed or cognitively overloaded domains, i.e. those involving a user concentrating on other tasks, call for more flexible dialogue with robust, wide-coverage language understanding. We describe an initial effort at addressing flexible and rich dialogue in a system with a number of features, such as full spoken language understanding, a multithreaded dialogue manager, dynamic update of information, and recognition of partial proper names. Fuliang Weng, Lawrence Cavedon, Badri Raghunathan, Danilo Mirkovic, Hauke Schmidt, Harry Bratt, Stanley Peters, Sandra Upson, Elizabeth Shriberg, Carsten Bergmann |
INTERSPEECH | 11 |
| 2003 | A prosody-based approach to end-of-utterance detection that does not require speech recognitionabstractIn previous work we showed that state-of-the-art end-of-utterance detection (as used, for example, in dialog systems) can be improved significantly by making use of prosodic and/or language models that predict utterance endpoints, based on word and alignment output from a speech recognizer. However, using a recognizer in endpointing might not be practical in certain applications. We demonstrate that the improvements due to the prosodic knowledge can be realized largely without alignment information, i.e., without requiring a speech recognizer. A prosodic end-of-utterance detector using only speech/nonspeech detection output is still considerably more accurate and has lower latency than a baseline system based on pause-length thresholding. Luciana Ferrer, Elizabeth Shriberg, Andreas Stolcke |
ICASSP (1) | 2 |
| 2003 | The ICSI Meeting CorpusabstractWe have collected a corpus of data from natural meetings that occurred at the International Computer Science Institute (ICSI) in Berkeley, California over the last three years. The corpus contains audio recorded simultaneously from head-worn and table-top microphones, word-level transcripts of meetings, and various metadata on participants, meetings, and hardware. Such a corpus supports work in automatic speech recognition, noise robustness, dialog modeling, prosody, rich transcription, information retrieval, and more. We present details on the contents of the corpus, as well as rationales for the decisions that led to its configuration. The corpus were delivered to the Linguistic Data Consortium (LDC). Adam Janin, Don Baron, Jane Edwards, Daniel P. W. Ellis, David Gelbart, Nelson Morgan, Barbara Peskin, Thilo Pfau, Elizabeth Shriberg, Andreas Stolcke, Chuck Wooters |
ICASSP (1) | 9 |
| 2003 | Meetings about meetings: research at ICSI on speech in multiparty conversationsabstractIn early 2001, we reported (at the Human Language Technology meeting) the early stages of an ICSI (International Computer Science Institute) project on processing speech from meetings (in collaboration with other sites, principally SRI, Columbia, and UW). We report our progress from the first few years of this effort, including: the collection and subsequent release of a 75-meeting corpus (over 70 meeting-hours and up to 16 channels for each meeting); the development of a prosodic database for a large subset of these meetings, and its subsequent use for punctuation and disfluency detection; the development of a dialog annotation scheme and its implementation for a large subset of the meetings; and the improvement of both near-mic and far-mic speech recognition results for meeting speech test sets. Nelson Morgan, Don Baron, Sonali Bhagat, Hannah Carvey, Rajdip Dhillon, Jane Edwards, David Gelbart, Adam Janin, Ashley Krupski, Barbara Peskin, Thilo Pfau, Elizabeth Shriberg, Andreas Stolcke, Chuck Wooters |
ICASSP (4) | 12 |
| 2003 | Training a prosody-based dialog act tagger from unlabeled dataabstractDialog act tagging is an important step toward speech understanding, yet training such taggers usually requires large amounts of data labeled by linguistic experts. Here we investigate the use of unlabeled data for training HMM-based dialog act taggers. Three techniques are shown to be effective for bootstrapping a tagger from very small amounts of labeled data: iterative relabeling and retraining on unlabeled data; a dialog grammar to model dialog act context, and a model of the prosodic correlates of dialog acts. On the SPINE dialog corpus, the combined use of prosodic information and unlabeled data reduces the tagging error between 12% and 16%, compared to baseline systems using word information and various amounts of labeled data only. Anand Venkataraman, Luciana Ferrer, Andreas Stolcke, Elizabeth Shriberg |
ICASSP (1) | 4 |
| 2003 | Prosodic knowledge sources for automatic speech recognitionabstractIn this work, different prosodic knowledge sources are integrated into a state-of-the-art large vocabulary speech recognition system. Prosody manifests itself on different levels in the speech signal: within the words as a change in phone durations and pitch, in between the words as a variation in the pause length, and beyond the words, correlating with higher linguistic structures and nonlexical phenomena. We investigate three models, each exploiting a different level of prosodic information, in rescoring N-best hypotheses according to how well recognized words correspond to prosodic features of the utterance. Experiments on the Switchboard corpus show word accuracy improvements with each prosodic knowledge source. A further improvement is observed with the combination of all models, demonstrating that they each capture somewhat different prosodic characteristics of the speech signal. Dimitra Vergyri, Andreas Stolcke, Venkata Ramana Rao Gadde, Luciana Ferrer, Elizabeth Shriberg |
ICASSP (1) | 5 |
| 2003 | Modeling duration patterns for speaker recognitionabstractWe present a method for speaker recognition that uses the duration patterns of speech units to aid speaker classification. The approach represents each word and/or phone by a feature vector comprised of either the durations of the individual phones making up the word, or the HMM states making up the phone. We model the vectors using mixtures of Gaussians. The speaker specific models are obtained through adaptation of a “background” model that is trained on a large pool of speakers. Speaker models are then used to score the test data; they are normalized by subtracting the scores obtained with the background model. We find that this approach yields significant perfomance improvement when combined with a state-of-the-art speaker recognition system based on standard cepstral features. Furthermore, the improvement persists even after combination with lexical features. Finally, the improvement continues to increase with longer test sample durations, beyond the test duration at which standard system accuracy level off. Luciana Ferrer, Harry Bratt, Venkata Ramana Rao Gadde, Sachin S. Kajarekar, Elizabeth Shriberg, M. Kemal Sönmez, Andreas Stolcke, Anand Venkataraman |
INTERSPEECH | 5 |
| 2003 | Automatic disfluency identification in conversational speech using multiple knowledge sourcesabstractDisfluencies occur frequently in spontaneous speech. Detection and correction of disfluencies can make automatic speech recognition transcripts more readable for human readers, and can aid downstream processing by machine. This work investigates a number of knowledge sources for disfluency detection, including acoustic-prosodic features, a language model (LM) to account for repetition patterns, a part-of-speech (POS) based LM, and rule-based knowledge. Different components are designed for different purposes in the system. Results show that detection of disfluency interruption points is best achieved by a combination of prosodic cues, word-based cues, and POS-based cues. The onset of a disfluency to be removed, in contrast, is best found using knowledge-based rules. Finally, specific disfluency types can be aided by the modeling of word patterns. 1. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 2 |
| 2003 | Spotting "hot spots" in meetings: human judgments and prosodic cuesabstractRecent interest in the automatic processing of meetings is motivated by a desire to summarize, browse, and retrieve important information from lengthy archives of spoken data. One of the most useful capabilities such a technology could provide is a way for users to locate “hot spots” or regions in which participants are highly involved in the discussion (e.g. heated arguments, points of excitement, etc.). We ask two questions about hot spots in meetings in the ICSI Meeting Recorder corpus. First, we ask whether involvement can be judged reliably by human listeners. Results show that despite the subjective nature of the task, raters show significant agreement in distinguishing involved from non-involved utterances. Second, we ask whether there is a relationship between human judgments of involvement and automatically extracted prosodic features of the associated regions. Results show that there are significant differences in both F0 and energy between involved and non-involved utterances. These findings suggest that humans do agree to some extent on the judgment of hot spots, and that acoustic-only cues could be used for automatic detection of hot spots in natural meetings. Britta Wrede, Elizabeth Shriberg |
INTERSPEECH | 2 |
| 2003 | "TalkPrinting": Improving Speaker Recognition by Modeling Stylistic Features
Sachin S. Kajarekar, M. Kemal Sönmez, Luciana Ferrer, Venkata Ramana Rao Gadde, Anand Venkataraman, Elizabeth Shriberg, Andreas Stolcke, Harry Bratt |
ISI | 6 |
| 2003 | Detection Of Agreement vs. Disagreement In Meetings: Training With Unlabeled Data
Dustin Hillard, Mari Ostendorf, Elizabeth Shriberg |
HLT-NAACL | 3 |
| 2002 | Using prosodic and lexical information for speaker identificationabstractWe investigate the incorporation of larger time-scale information, such as prosody, into standard speaker ID systems. Our study is based on the Extended Data Task of the NIST 2001 Speaker ID evaluation, which provides much more test and training data than has traditionally been available to similar speaker ID investigations. In addition, we have had access to a detailed prosodic feature database of Switchboard-I conversations, including data not previously applied to speaker ID. We describe two baseline acoustic systems, an approach using Gaussian Mixture Models, and an LVCSR-based speaker ID system. These results are compared to and combined with two larger time-scale systems: a system based on an “idiolect” language model. and a system making use of the contents of the prosody database. We find that, with sufficient test and training data, suprasegmental information can significantly enhance the performance of traditional speaker ID systems. Frederick Weber, Linda Manganaro, Barbara Peskin, Elizabeth Shriberg |
ICASSP | 4 |
| 2002 | Prosody-based automatic detection of annoyance and frustration in human-computer dialogabstractWe investigate the use of prosody for the detection of frustration and annoyance in natural human-computer dialog. In addition to prosodic features, we examine the contribution of language model information and speaking "style". Results show that a prosodic model can predict whether an utterance is neutral versus "annoyed or frustrated" with an accuracy on par with that of human interlabeler agreement. Accuracy increases when discriminating only "frustrated" from other utterances, and when using only those utterances on which labelers originally agreed. Furthermore, prosodic model accuracy degrades only slightly when using recognized versus true words. Language model features, even if based on true words, are relatively poor predictors of frustration. Finally, we find that hyperarticulation is not a good predictor of emotion; the two phenomena often occur independently. Jeremy Ang, Rajdip Dhillon, Ashley Krupski, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 4 |
| 2002 | Automatic punctuation and disfluency detection in multi-party meetings using prosodic and lexical cuesabstractWe investigate automatic approaches to finding "hidden" spontaneous speech events, such as sentence boundaries and disfluencies, in multi-party meetings. Hidden events are characterized prosodically by a large array of automatically extracted energy, duration, and pitch features, and are modeled by decision tree classifiers; lexical cues are modeled by N-gram language models. Both sources of information are combined in a hidden Markov model framework. Results show that combined classifiers achieve higher accuracy than either single knowledge source alone. We also study classifiers that use only the preceding context for predicting events, simulating online processing. We find that prosodic features are more robust than are language model features to this constraint. Finally, we examine the effect of automatic word recognition errors, in both training and testing, on classification accuracy. We find that lexical models degrade much more severely than do prosodic models in this case, again showing the relative robustness of prosodic information for hidden-event detection in natural conversation. Don Baron, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 2 |
| 2002 | Is the speaker done yet? faster and more accurate end-of-utterance detection using prosodyabstractWe examine the problem of end-of-utterance (EOU) detection for real-time speech recognition, particularly in the context of a human-computer dialog system. Current EOU detection algorithms use only a simple pause threshold for making this decision, leading to two problems. First, especially as speech-driven interfaces become more natural, users often pause inside utterances, resulting in a premature cut off by the system. Second, when users really are done, the minimum system wait is always the threshold value, needlessly adding time to the interaction. We have developed a new approach to EOU detection that uses prosodic features to address both of these problems. Prosodic features are modeled by decision trees and combined with an event N-gram language model to obtain a score that measures the likelihood that any nonspeech region is an EOU. We find that this approach dramatically improves both the accuracy and speed of online EOU detection. 1. Luciana Ferrer, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 2 |
| 2001 | Observations on overlap: findings and implications for automatic processing of multi-party conversationabstractWe examine the distribution of overlapping speech in different corpora of natural multi-party conversations, including two types of meetings, and two corpora of telephone conversations. Analyses are based on forced alignment and speech recognition using an identical recognizer across tasks. Three results are discussed. First, all corpora show high overall rates of overlap, with similar rates for meetings and telephone conversations. Second, speech recognition performance in non-overlapped regions of meetings is no worse than that in single-channel telephone conversations, while recognition in overlap regions degrades considerably. Finally, interrupt locations are associated with endpoints of word-level events in a speaker's turn, including backchannels, discourse markers, and disfluencies. Results suggest that overlap is an important inherent characteristic of conversational speech that should not be ignored; on the contrary, it should be jointly modeled with acoustic and language model information in machine processing of conversation. 1. Elizabeth Shriberg, Andreas Stolcke, Don Baron |
INTERSPEECH | 1 |
| 2001 | Integrating Prosodic and Lexical Cues for Automatic Topic SegmentationabstractWe present a probabilistic model that uses both prosodic and lexical cues for the automatic segmentation of speech into topically coherent units. We propose two methods for combining lexical and prosodic information using hidden Markov models and decision trees. Lexical information is obtained from a speech recognizer, and prosodic features are extracted automatically from speech waveforms. We evaluate our approach on the Broadcast News corpus, using the DARPA-TDT evaluation metrics. Results show that the prosodic model alone is competitive with word-based segmentation methods. Furthermore, we achieve a significant reduction in error by combining the prosodic and word-based knowledge sources. Gökhan Tür, Dilek Hakkani-Tür, Andreas Stolcke, Elizabeth Shriberg |
Comput. Linguistics | 4 |
| 2000 | Consonant discrimination in elicited and spontaneous speech: a case for signal-adaptive front ends in ASR
M. Kemal Sönmez, Madelaine Plauché, Elizabeth Shriberg, Horacio Franco |
INTERSPEECH | 3 |
| 2000 | Prosodic features for automatic text-independent evaluation of degree of nativeness for language learnersabstractPredicting the degree of nativeness of a student's utterance is an important issue in computer-aided language learning. This task has been addressed by many studies focusing on the segmental assessment of the speech signal. Toachieve improved correlations between human and automatic nativeness scores, other aspects of speechshould also be considered, such as prosody. The goal of this study is to evaluate the use of prosodic information to help predict the degree of nativeness of pronunciation, independent of the text. A supervised strategy based on human grades is used in an attempt to select promising features for this task. Preliminary results show improvements in the correlation between human and automatic scores. Horacio Franco, Elizabeth Shriberg, Kristin Precoda, M. Kemal Sönmez |
INTERSPEECH | 3 |
| 2000 | Dialog Act Modeling for Automatic Tagging and Recognition of Conversational SpeechabstractWe describe a statistical approach for modeling dialogue acts in conversational speech, i.e., speech-act-like units such as STATEMENT, Question, BACKCHANNEL, Agreement, Disagreement, and Apology. Our model detects and predicts dialogue acts based on lexical, collocational, and prosodic cues, as well as on the discourse coherence of the dialogue act sequence. The dialogue model is based on treating the discourse structure of a conversation as a hidden Markov model and the individual dialogue acts as observations emanating from the model states. Constraints on the likely sequence of dialogue acts are modeled via a dialogue act n-gram. The statistical dialogue grammar is combined with word n-grams, decision trees, and neural networks modeling the idiosyncratic lexical and prosodic manifestations of each dialogue act. We develop a probabilistic integration of speech recognition with dialogue modeling, to improve both speech recognition and dialogue act classification accuracy. Models are trained and evaluated using a large hand-labeled database of 1,155 conversations from the Switchboard corpus of spontaneous human-to-human telephone speech. We achieved good dialogue act labeling accuracy (65% based on errorful, automatically recognized words and prosody, and 71% based on word transcripts, compared to a chance baseline accuracy of 35% and human accuracy of 84%) and a small reduction in word recognition error. Andreas Stolcke, Klaus Ries 0001, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates 0001, Daniel Jurafsky, Paul Taylor 0001, Rachel Martin, Carol Van Ess-Dykema, Marie Meteer |
Comput. Linguistics | 4 |
| 2000 | Prosody-based automatic segmentation of speech into sentences and topics
Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, Gökhan Tür |
Speech Commun. | 1 |
| 1999 | Combining words and prosody for information extraction from speechabstractThe design principles and collection procedures behind a speech synthesis corpus directly impact the performance of the resulting text-to-speech system. This paper describes the design and collection of the Victoria corpus, created to support speech synthesis research and development at Apple Computer. This corpus is composed of ve constituent parts, each designed to cover a speci c aspect of speech synthesis: polyphones, prosodic contexts, reiterant speech, function word sequences, and continuous speech. It was spoken in general U.S. English by one linguisticallytrained adult female. Portions of the corpus are being used in the statistical estimation of duration and pitch models for Apple's next-generation textto-speech system, MacinTalk 4. Dilek Hakkani-Tür, Gökhan Tür, Andreas Stolcke, Elizabeth Shriberg |
EUROSPEECH | 4 |
| 1999 | Modeling the prosody of hidden events for improved word recognitionabstractWe investigate a new approach for using speech prosody as a knowledge source for speech recognition. The idea is to penalize word hypotheses that are inconsistent with prosodic features such as duration and pitch. To model the interaction between words and prosody we modify the language model to represent hidden events such as sentence boundaries and various forms of disfluency, and combine with it decision trees that predict such events from prosodic features. N-best rescoring experiments on the Switchboard corpus show a small but consistent reduction of word error as a result of this modeling. We conclude with a preliminary analysis of the types of errors that are corrected by the prosodically informed model. 1. Andreas Stolcke, Elizabeth Shriberg, Dilek Hakkani-Tür, Gökhan Tür |
EUROSPEECH | 2 |
| 1998 | Collection and detailed transcription of a speech database for development of language learning technologiesabstractWe describe the methodologies for collecting and annotating a Latin-American Spanish speech database. The database includes recordings by native and nonnative speakers. The nonnative recordings are annotated with ratings of pronunciation quality and detailed phonetic transcriptions. We use the annotated database to investigate rater reliability, the effect of each phone on overall perceived nonnativeness, and the frequency of specific pronunciation errors. 1. INTRODUCTION In this paper we describe the methodologies for collecting and annotating a Latin-American Spanish speech database. The database includes recordings by native and nonnative speakers. A panel of listeners rated the pronunciation quality of the nonnative data, and a group of expert phoneticians phonetically transcribed a subset of the nonnative data. The database was intended for use in the development of hidden-Markov model (HMM) based speech technologies for language learning [4], including robust speech recognition... Harry Bratt, Leonardo Neumeyer, Elizabeth Shriberg, Horacio Franco |
ICSLP | 3 |
| 1998 | Crosslinguistic disfluency modelling: a comparative analysis of Swedish and american English human-human and human-machine dialoguesabstractWe report results from a cross-language study of disfluencies (DFs) in Swedish and American English human--machine and human--human dialogs. The focus is on comparisons not directly affected by differences in overall rates since these could be associates with task details. Rather, we focus on differences of how speakers utilize DFs in the different languages, including: relative rates of the use of hesitation forms, the location of hesitations, and surface characteristics of DFs. Results suggest that although the languages differ in some respects (such as the ability to insert filled pauses within 'words'), in many analyses the languages show similar behavior. Such results provide suggestions for cross-linguistic DF modeling in both theoretical and applied fields. Robert Eklund, Elizabeth Shriberg |
ICSLP | 2 |
| 1998 | How far do speakers back up in repairs? a quantitatve modelabstractSpeakers frequently retrace one or more words when continuing after a break in fluency. Syntactic principles constrain the points from which speakers retrace; however syntactic principles do not provide predictions about the relative usage of different allowable retrace points. Such predictions are useful for automatic processing of repairs in speech technology, particularly if they use information readily available to a speech recognizer. We propose a quantitative model that predicts the overall distribution of retrace lengths in a large corpus of spontaneous speech, based only on word position. The model has two components: (1) a constant, position-independent probability for extending a retrace by one more word; and (2) a position-dependent probability to “skip” to the beginning of the sentence. Results have implications for modeling repairs in speech applications and constrain explanatory models in psycholinguistics. Elizabeth Shriberg, Andreas Stolcke |
ICSLP | 1 |
| 1998 | Modeling dynamic prosodic variation for speaker verificationabstractStatistics of frame-level pitch have recently been used in speaker recognition systems with good results [1, 2, 3]. Although they convey useful long-term information about a speaker's distribution of f 0 values, such statistics fail to capture information about local dynamics in intonation that characterize an individual's speaking style. In this work, we take a first step toward capturing such suprasegmental patterns for automatic speaker verification. Specifically, we model the speaker's f 0 movements by fitting a piecewise linear model to the f 0 track to obtain a stylized f 0 contour. Parameters of the model are then used as statistical features for speaker verification. We report results on 1998 NIST speaker verification evaluation. Prosody modeling improves the verification performance of a cepstrum-based Gaussian mixture model system (as measured by a task-specific Bayes risk) by 10%. 1. INTRODUCTION Statistics of frame-level pitch have recently been shown to improve the perf... M. Kemal Sönmez, Elizabeth Shriberg, Larry Heck, Mitch Weintraub |
ICSLP | 2 |
| 1998 | Automatic detection of sentence boundaries and disfluencies based on recognized wordsabstractWe study the problem of detecting linguistic events at interword boundaries, such as sentence boundaries and disfluency locations, in speech transcribed by an automatic recognizer. Recovering such events is crucial to facilitate speech understanding and other natural language processing tasks. Our approach is based on a combination of prosodic cues modeled by decision trees, and word-based event N-gram language models. Several model combination approaches are investigated. The techniques are evaluated on conversational speech from the Switchboard corpus. Model combination is shown to give a significant win over individual knowledge sources. 1. INTRODUCTION Current automatic speech recognition systems output a string of words. Most natural language understanding systems, however, require structural information such as punctuation, which is present in text but not overtly indicated in spoken language. Similarly, for speech understanding and information extraction, it is important to fi... Andreas Stolcke, Elizabeth Shriberg, Rebecca Bates 0001, Mari Ostendorf, Dilek Hakkani-Tür, Madelaine Plauché, Gökhan Tür |
ICSLP | 2 |
| 1997 | A prosody only decision-tree model for disfluency detectionabstractSpeech disfluencies (filled pauses, repetitions, repairs, and false starts) are pervasive in spontaneous speech. The ability to detect and correct disfluencies automatically is important for effective natural language understanding, as well as to improve speech models in general. Previous approaches to disfluency detection have relied heavily on lexical information, which makes them less applicable when word recognition is unreliable. We have developed a disfluency detection method using decision tree classifiers that use only local and automatically extracted prosodic features. Because the model doesn’t rely on lexical information, it is widely applicable even when word recognition is unreliable. The model performed significantly better than chance at detecting four disfluency types. It also outperformed a language model in the detection of false starts, given the correct transcription. Combining the prosody model with a specialized language model improved accuracy over either model alone for the detection of false starts. Results suggest that a prosody-only model can aid the automatic detection of disfluencies in spontaneous speech. 1. Elizabeth Shriberg, Rebecca Bates 0001, Andreas Stolcke |
EUROSPEECH | 1 |
| 1997 | A lognormal tied mixture model of pitch for prosody based speaker recognitionabstractStatistics of pitch have recently been used in speaker recognition systems with good results. The success of such systems depends on robust and accurate computation of pitch statistics in the presence of pitch tracking errors. In this work, we develop a statistical model of pitch that allows unbiased estimation of pitch statistics from pitch tracks which are subject to doubling and/or halving. We first argue by a simple correlation model and empirically demonstrate by QQ plots that "clean" pitch is distributed with a lognormal distribution rather than the often assumed normal distribution. Second, we present a probabilistic model for estimated pitch via a pitch tracker in the presence of doubling/halving, which leads to a mixture of three lognormal distributions with tied means and variances for a total of four free parameters. We use the obtained pitch statistics as features in speaker verification on the March 1996 NIST Speaker Recognition Evaluation data (subset of Switchboard) and ... M. Kemal Sönmez, Larry Heck, Mitch Weintraub, Elizabeth Shriberg |
EUROSPEECH | 4 |
| 1996 | Statistical language modeling for speech disfluenciesabstractSpeech disfluencies (such as filled pauses, repetitions, restarts) are among the characteristics distinguishing spontaneous speech from planned or read speech. We introduce a language model that predicts disfluencies probabilistically and uses an edited, fluent context to predict following words. The model is based on a generalization of the standard N-gram language model. It uses dynamic programming to compute the probability of a word sequence, taking into account possible hidden disfluency events. We analyze the model's performance for various disfluency types on the Switchboard corpus. We find that the model reduces the word perplexity in the neighborhood of disfluency events; however, overall differences are small and have no significant impact on the recognition accuracy. We also note that for modeling of the most frequent type of disfluency, filled pauses, a segmentation of utterances into linguistic (rather than acoustic) units is required. Our analysis illustrates a generally useful technique for language model evaluation based on local perplexity comparisons. Andreas Stolcke, Elizabeth Shriberg |
ICASSP | 2 |
| 1996 | Modeling intra-speaker pitch range variation: predicting F0 targets when "speaking up"
Elizabeth Shriberg, D. Robert Ladd, Jacques M. B. Terken |
ICSLP | 1 |
| 1996 | Word predictability after hesitations: a corpus-based study
Elizabeth Shriberg, Andreas Stolcke |
ICSLP | 1 |
| 1996 | Automatic linguistic segmentation of conversational speech
Andreas Stolcke, Elizabeth Shriberg |
ICSLP | 2 |
| 1992 | Integrating Multiple Knowledge Sources for Detection and Correction of Repairs in Human-Computer DialogabstractWe have analyzed 607 sentences of spontaneous human-computer speech data containing repairs, drawn from a total corpus of 10,718 sentences. We present here criteria and techniques for automatically detecting the presence of a repair, its location, and making the appropriate correction. The criteria involve integration of knowledge from several sources: pattern matching, syntactic and semantic analysis, and acoustics. John Bear, John Dowding, Elizabeth Shriberg |
ACL | 3 |
| 1992 | Intonation of clause-internal filled pausesabstractClause-internal filled pauses and preceding peak fundamental frequency (F0) values were analyzed to determine whether the intonation of filled pauses is relative to, or independent of, prior prosodic context. Higher peaks were found to be systematically associated with higher filled-pause values, supporting the 'relative' hypothesis. A linear model, in which filled-pause F0 was expressed as an invariant (over speakers) proportion of the distance between preceding peak F0 and a speaker-dependent baseline F0, produced results nearly identical to those of a two-parameter model in which the coefficients of peak and baseline were allowed to vary freely. The model was less appropriate for filled pauses after sentence-initial peaks, but unaffected by temporal variables. Elizabeth Shriberg, Robin J. Lickley |
ICSLP | 1 |
| 1992 | User behaviors affecting speech recognition
Elizabeth Wade, Elizabeth Shriberg, Patti Price |
ICSLP | 2 |
| 1990 | Hypercorrection in speech perception
John J. Ohala, Elizabeth Shriberg |
ICSLP | 2 |