VLDB 2026 Research / reviewers in the wild / expert
Carol Y. Espy-Wilson
dblp:12/1594 · also Carol Y. Espy
· DBLP profile ↗
106ranked-venue papers
8as first author
32since 2021 · last 2025
0000-0002-1012-183XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 96 · 8 first-author · 27 since 2021Artificial intelligence and machine learning · 75 · 6 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal InsufficiencyabstractTraditional clinical approaches for assessing nasality, such as nasopharyngoscopy and nasometry, involve unpleasant experiences and are problematic for children. Speech Inversion (SI), a noninvasive technique, offers a promising alternative for estimating articulatory movement without the need for physical instrumentation. In this study, an SI system trained on nasalance data from healthy adults is augmented with source information from electroglottography and acoustically derived F0, periodic and aperiodic energy estimates as proxies for glottal control. This model achieves $16.92 \%$ relative improvement in Pearson Product-Moment Correlation (PPMC) compared to a previous SI system for nasalance estimation. To adapt the SI system for nasalance estimation in children with Velopharyngeal Insufficiency (VPI), the model initially trained on adult speech was fine-tuned using children with VPI data, yielding an $7.90 \%$ relative improvement in PPMC compared to its performance before fine-tuning. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
ASRU | 5 |
| 2025 | Speaking with Robots in Noisy EnvironmentsabstractA fundamental limitation for speech-enabled human-robot interaction (HRI) is automatic speech recognition (ASR), or the process of converting a raw speech signal into text. If this fails then all downstream tasks will fail. While commercial ASR systems have improved substantially in recent years, they are still well known to struggle with noisy speech, which is commonplace in a variety of social, medical, and military environments. In this paper, we address this challenge by introducing a dataset for training and evaluating ASR systems to be robust to background noise. The dataset is comprised of 100 hours from the Librispeech ASR Corpus, which we supplemented with a variety of noises, including office, traffic, crowd, and others. We evaluate the state-of-the-art Whisper ASR model on the dataset and demonstrate improved ASR accuracy on noisy speech. We also investigate how several methods of speech enhancement can be trained on the data to further improve performance. Finally, we demonstrate this approach in a real-world HRI task involving noisy speech. The dataset and models are available for researchers to freely use to improve the performance of their ASR systems for HRI in noisy environments. Shuubham Ojha, Felix Gervits, Carol Y. Espy-Wilson |
HRI | 3 |
| 2025 | CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom EnvironmentsabstractCreating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model’s robustness to different noises, microphones and classroom conditions. Ahmed Adel Attia, Dorottya Demszky, Tolúlopé Ògúnrèmí, Jing Liu 0064, Carol Y. Espy-Wilson |
ICASSP | 5 |
| 2025 | Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia SymptomsabstractMultimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variational Auto-Encoder (VQ-VAE) based Multimodal Representation Learning (MRL) model to produce task-agnostic speech representations from vocal Tract Variables (TVs) and Facial Action Units (FAUs). These representations are then used in a Multi-Task Learning (MTL) based downstream prediction model to obtain class labels and an overall severity score. The proposed framework outperforms the previous works on the multi-class classification task across all evaluation metrics (Weighted F1 score, AUC-ROC score, and Weighted Accuracy). Additionally, it estimates the schizophrenia severity score, a task not addressed by earlier approaches. Gowtham Premananth, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2025 | From Weak Labels to Strong Results: Utilizing 5, 000 Hours of Noisy Classroom Transcripts with Minimal Accurate DataabstractRecent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios. Ahmed Adel Attia, Dorottya Demszky, Jing Liu 0064, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2025 | Subtyping Speech Errors in Childhood Speech Sound Disorders with Acoustic-to-Articulatory Speech Inversion
Nina Benway, Saba Tabatabaee, Benjamin Munson, Jonathan Preston, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2025 | Speech Kinematic Analysis from Acoustics: Scientific, Clinical and Practical Applications
Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2025 | Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory PerspectivesabstractAdvancements in AI-driven speech-based applications have transformed diverse industries ranging from healthcare to customer service. However, the increasing prevalence of non-native accented speech in global interactions poses significant challenges for speech-processing systems, which are often trained on datasets dominated by native speech. This study investigates accented English speech through articulatory and acoustic analysis, identifying simpler coordination patterns and higher average pitch than native speech. Using eigenspectra and Vocal Tract Variable-based coordination features, we establish an efficient method for quantifying accent strength without relying on resource-intensive phonetic transcriptions. Our findings provide a new avenue for research on the impacts of accents on speech intelligibility and offer insights for developing inclusive, robust speech processing systems that accommodate diverse linguistic communities. Gowtham Premananth, Vinith Kugathasan, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2025 | Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity EstimationabstractStudies on schizophrenia assessments using deep learning typically treat it as a classification task to detect the presence or absence of the disorder, oversimplifying the condition and reducing its clinical applicability. This traditional approach overlooks the complexity of schizophrenia, limiting its practical value in healthcare settings. This study shifts the focus to individual symptom severity estimation using a multimodal approach that integrates speech, video, and text inputs. We develop unimodal models for each modality and a multimodal framework to improve accuracy and robustness. By capturing a more detailed symptom profile, this approach can help in enhancing diagnostic precision and support personalized treatment, offering a scalable and objective tool for mental health assessment. Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2025 | Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating NasalityabstractSpeech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue constrictions, called oral tract variables (TVs), which were later enhanced by including source information (periodic and aperiodic energies, and F0 frequency) as proxies for glottal control. Comparison of the nasometric measures with high-speed nasopharyngoscopy showed that nasalance can serve as ground truth, and that an SI system trained with it reliably recovers velum movement patterns for American English speakers. Here, two SI training approaches are compared: baseline models that estimate oral TVs and nasalance independently, and a synergistic model that combines oral TVs and source features with nasalance. The synergistic model shows relative improvements of 5% in oral TVs estimation and 9% in nasalance estimation compared to the baseline models. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2025 | FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom EnvironmentsabstractCreating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy of finetuning with augmented children datasets to adapt the x-vector and ECAPA-TDNN to classroom environments. We demonstrate that finetuning with augmented children's datasets is powerful in that regard and reduces the Equal Error Rate (EER) of x-vector and ECAPA-TDNN models for both classroom datasets and children speech datasets. Notably, this method reduces EER of the ECAPA-TDNN model on average by half (a 5 % improvement) for classrooms in the MPT dataset compared to the ECAPA-TDNN baseline model. The x-vector model shows an 8 % average improvement for classrooms in the NCTE dataset compared to its baseline. Saba Tabatabaee, Jing Liu 0064, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2024 | Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. AdultsabstractRecent advancements in Automatic Speech Recognition (ASR) systems, exemplified by Whisper, have demonstrated the potential of these systems to approach human-level performance given sufficient data. However, this progress doesn’t readily extend to ASR for children due to the lim- ited availability of suitable child-specific databases and the distinct characteristics of children’s speech. A recent study investigated leveraging the My Science Tutor (MyST) chil- dren’s speech corpus to enhance Whisper’s performance in recognizing children’s speech. They were able to demon- strate some improvement on a limited testset. This paper builds on these findings by enhancing the utility of the MyST dataset through more efficient data preprocessing. We reduce the Word Error Rate (WER) on the MyST testset 13.93% to 9.11% with Whisper-Small and from 13.23% to 8.61% with Whisper-Medium and show that this improvement can be generalized to unseen datasets. We also highlight important challenges towards improving children’s ASR performance and the effect of fine-tuning in improving the transcription of disfluent speech. Ahmed Adel Attia, Jing Liu 0064, Wei Ai 0002, Dorottya Demszky, Carol Y. Espy-Wilson |
AIES (1) | 5 |
| 2024 | Examining Vocal Tract Coordination in Childhood Apraxia of Speech with Acoustic-to-Articulatory Speech Inversion Feature Sets
Nina Benway, Jonathan L. Preston, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2024 | A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
Gowtham Premananth, Yashish M. Siriwardena, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Y. Espy-Wilson |
INTERSPEECH | 6 |
| 2024 | Accent Conversion with Articulatory Representations
Yashish M. Siriwardena, Nathan Swedlow, Audrey Howard, Evan Gitterman, Dan Darcy, Carol Y. Espy-Wilson, Andrea Fanelli |
INTERSPEECH | 6 |
| 2023 | Masked Autoencoders are Articulatory LearnersabstractArticulatory recordings track the positions and motion of different articulators along the vocal tract and are widely used to study speech production and to develop speech technologies such as articulatory based speech synthesizers and speech inversion systems. The University of Wisconsin X-Ray Mi-crobeam (XRMB) dataset is one of various datasets that provide articulatory recordings synced with audio recordings. The XRMB articulatory recordings employ pellets placed on a number of articulators which can be tracked by the mi-crobeam. However, a significant portion of the articulatory recordings are mistracked, and have been so far unusable. In this work, we present a deep learning based approach using Masked Autoencoders to accurately reconstruct the mistracked articulatory recordings for 41 out of 47 speakers of the XRMB dataset. Our model is able to reconstruct articulatory trajectories that closely match ground truth, even when three out of eight articulators are mistracked, and retrieve 3.28 out of 3.4 hours of previously unusable recordings. Ahmed Adel Attia, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2023 | The Secret Source : Incorporating Source Features to Improve Acoustic-To-Articulatory Speech InversionabstractIn this work, we incorporated acoustically derived source features, aperiodicity, periodicity and pitch as additional targets to an acoustic-to-articulatory speech inversion (SI) system. We also propose a Temporal Convolution based SI system, which uses auditory spectrograms as the input speech representation, to learn long-range dependencies and complex interactions between the source and vocal tract, to improve the SI task. The experiments are conducted with both the Wisconsin X-ray microbeam (XRMB) and Haskins Production Rate Comparison (HPRC) datasets, with comparisons done with respect to three baseline SI model architectures. The proposed SI system with the HPRC dataset gains an improvement of close to 28% when the source features are used as additional targets. The same SI system outperforms the current best performing SI models by around 9% on the XRMB dataset. Yashish M. Siriwardena, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2023 | Enhancing Speech Articulation Analysis Using A Geometric Transformation of the X-ray Microbeam DatasetabstractAccurate analysis of speech articulation is crucial for speech analysis.However, X-Y coordinates of articulators strongly depend on the anatomy of the speakers and variability of pellet placements, and existing methods for mapping anatomical landmarks in the X-ray Microbeam Dataset (XRMB) fail to capture the entire anatomy of the vocal tract.In this paper, we propose a new geometric transformation that improves the accuracy of these measurements.Our transformation maps anatomical landmarks' X-Y coordinates along the midsagittal plane onto six relative measures: Lip Aperture (LA), Lip Protrusion (LP), Tongue Body Constriction Location (TBCL), Degree (TBCD), Tongue Tip Constriction Location (TTCL), and Degree (TTCD).Our novel contribution is the extension of the palate trace towards the inferred anterior pharyngeal line, which improves measurements of tongue body constriction. Ahmed Adel Attia, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2023 | Acoustic-to-Articulatory Speech Inversion Features for Mispronunciation Detection of /ɹ/ in Child Speech Sound DisordersabstractAcoustic-to-articulatory speech inversion could enhance automated clinical mispronunciation detection to provide detailed articulatory feedback unattainable by formant-based mispronunciation detection algorithms; however, it is unclear the extent to which a speech inversion system trained on adult speech performs in the context of (1) child and (2) clinical speech.In the absence of an articulatory dataset in children with rhotic speech sound disorders, we show that classifiers trained on tract variables from acoustic-to-articulatory speech inversion meet or exceed the performance of state-of-the-art features when predicting clinician judgment of rhoticity. Nina Benway, Yashish M. Siriwardena, Jonathan L. Preston, Elaine Hitchcock, Tara McAllister Byun, Carol Y. Espy-Wilson |
INTERSPEECH | 6 |
| 2023 | Speaker-independent Speech Inversion for Estimation of Nasalance
Yashish M. Siriwardena, Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede, Liran Oren |
INTERSPEECH | 2 |
| 2023 | Learning to Compute the Articulatory Representations of Speech with the MIRRORNET
Yashish M. Siriwardena, Carol Y. Espy-Wilson, Shihab A. Shamma |
INTERSPEECH | 2 |
| 2022 | Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation SystemsabstractRecent advancements in deep learning have led to drastic improvements in speech segregation models. Despite their success and growing applicability, few efforts have been made to analyze the underlying principles that these networks learn to perform segregation. Here we analyze the role of harmonicity on two state-of-the-art Deep Neural Networks (DNN)-based models- Conv-TasNet and DPT-Net [1],[2]. We evaluate their performance with mixtures of natural speech versus slightly manipulated inharmonic speech, where harmonics are slightly frequency jittered. We find that performance deteriorates significantly if one source is even slightly harmonically jittered, e.g., an imperceptible 3% harmonic jitter degrades performance of Conv-TasNet from 15.4 dB to 0.70 dB. Training the model on inharmonic speech does not remedy this sensitivity, instead resulting in worse performance on natural speech mixtures, making inharmonicity a powerful adversarial factor in DNN models. Furthermore, additional analyses reveal that DNN algorithms deviate markedly from biologically inspired algorithms [3] that rely primarily on timing cues and not harmonicity to segregate speech. Rahil Parikh, Ilya Kavalerov, Carol Y. Espy-Wilson, Shihab A. Shamma |
ICASSP | 3 |
| 2022 | Multimodal Depression Classification using Articulatory Coordination Features and Hierarchical Attention Based text EmbeddingsabstractMultimodal depression classification has gained immense popularity over the recent years. We develop a multimodal depression classification system using articulatory coordination features extracted from vocal tract variables and text transcriptions obtained from an automatic speech recognition tool that yields improvements of area under the receiver operating characteristics curve compared to unimodal classifiers (7.5% and 13.7% for audio and text respectively). We show that in the case of limited training data, a segment-level classifier can first be trained to then obtain a session-wise prediction without hindering the performance, using a multi-stage convolutional recurrent neural network. A text model is trained using a Hierarchical Attention Network (HAN). The multimodal system is developed by combining embeddings from the session-level audio model and the HAN text model. Nadee Seneviratne, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2022 | An Empirical Analysis on the Vulnerabilities of End-to-End Speech Segregation ModelsabstractInternational audience Rahil Parikh, Gaspar Rochette, Carol Y. Espy-Wilson, Shihab A. Shamma |
INTERSPEECH | 3 |
| 2022 | Acoustic To Articulatory Speech Inversion Using Multi-Resolution Spectro-Temporal Representations Of Speech SignalsabstractMulti-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of the speech signals. The purpose of this paper is to evaluate how well the auditory cortex representation of speech signals contribute to estimate articulatory features of those corresponding signals. Since obtaining articulatory features from acoustic features of speech signals has been a challenging topic of interest for different speech communities, we investigate the possibility of using this multi-resolution representation of speech signals as acoustic features. We used U. of Wisconsin X-ray Microbeam (XRMB) database of clean speech signals to train a feed-forward deep neural network (DNN) to estimate articulatory trajectories of six tract variables. The optimal set of multi-resolution spectro-temporal features to train the model were chosen using appropriate scale and rate vector parameters to obtain the best performing model. Experiments achieved a correlation of 0.675 with ground-truth tract variables. We compared the performance of this speech inversion system with prior experiments conducted using Mel Frequency Cepstral Coefficients (MFCCs). Rahil Parikh, Nadee Seneviratne, Ganesh Sivaraman, Shihab A. Shamma, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2022 | Multimodal Depression Severity Score Prediction Using Articulatory Coordination Features and Hierarchical Attention Based Text Embeddings
Nadee Seneviratne, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2022 | Acoustic-to-articulatory Speech Inversion with Multi-task LearningabstractMulti-task learning (MTL) frameworks have proven to be effective in diverse speech related tasks like automatic speech recognition (ASR) and speech emotion recognition. This paper proposes a MTL framework to perform acoustic-to-articulatory speech inversion by simultaneously learning an acoustic to phoneme mapping as a shared task. We use the Haskins Production Rate Comparison (HPRC) database which has both the electromagnetic articulography (EMA) data and the corresponding phonetic transcriptions. Performance of the system was measured by computing the correlation between estimated and actual tract variables (TVs) from the acoustic to articulatory speech inversion task. The proposed MTL based Bidirectional Gated Recurrent Neural Network (RNN) model learns to map the input acoustic features to nine TVs while outperforming the baseline model trained to perform only acoustic to articulatory inversion. Yashish M. Siriwardena, Ganesh Sivaraman, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 2 |
| 2022 | Modeling Feature Representations for Affective Speech Using Generative Adversarial NetworksabstractEmotion recognition is a classic field of research with a typical setup extracting features and feeding them through a classifier for prediction. On the other hand, generative models jointly capture the distributional relationship between emotions and the feature profiles. Recently, Generative Adversarial Networks (GANs) have surfaced as a new class of generative models and have shown considerable success in modeling distributions in the fields of computer vision and natural language understanding. In this article, we experiment with variants of GAN architectures to generate feature vectors corresponding to an emotion in two ways: (i) A generator is trained with samples from a mixture prior. Each mixture component corresponds to an emotional class and can be sampled to generate features from the corresponding emotion. (ii) A one-hot vector corresponding to an emotion can be explicitly used to generate the features. We perform analysis on such models and also propose different metrics used to measure the performance of the GAN models in their ability to generate realistic synthetic samples. Apart from evaluation on a given dataset of interest, we perform a cross-corpus study where we study the utility of the synthetic samples as additional training data in low resource conditions. Saurabh Sahu, Rahul Gupta 0001, Carol Y. Espy-Wilson |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Multimodal Approach for Assessing Neuromotor Coordination in Schizophrenia Using Convolutional Neural NetworksabstractThis study investigates the speech articulatory coordination in schizophrenia subjects exhibiting strong positive symptoms (e.g. hallucinations and delusions), using two distinct channel-delay correlation methods. We show that the schizophrenic subjects with strong positive symptoms and who are markedly ill pose complex articulatory coordination pattern in facial and speech gestures than what is observed in healthy subjects. This distinction in speech coordination pattern is used to train a multimodal convolutional neural network (CNN) which uses video and audio data during speech to distinguish schizophrenic patients with strong positive symptoms from healthy subjects. We also show that the vocal tract variables (TVs) which correspond to place of articulation and glottal source outperform the Mel-frequency Cepstral Coefficients (MFCCs) when fused with Facial Action Units (FAUs) in the proposed multimodal network. For the clinical dataset we collected, our best performing multimodal network improves the mean F1 score for detecting schizophrenia by around 18% with respect to the full vocal tract coordination (FVTC) baseline method implemented with fusing FAUs and MFCCs. Yashish M. Siriwardena, Carol Y. Espy-Wilson, Christopher Kitchen, Deanna L. Kelly |
ICMI | 2 |
| 2021 | Speech Based Depression Severity Level Classification Using a Multi-Stage Dilated CNN-LSTM ModelabstractSpeech based depression classification has gained immense popularity over the recent years. However, most of the classification studies have focused on binary classification to distinguish depressed subjects from non-depressed subjects. In this paper, we formulate the depression classification task as a severity level classification problem to provide more granularity to the classification outcomes. We use articulatory coordination features (ACFs) developed to capture the changes of neuromotor coordination that happens as a result of psychomotor slowing, a necessary feature of Major Depressive Disorder. The ACFs derived from the vocal tract variables (TVs) are used to train a dilated Convolutional Neural Network based depression classification model to obtain segment-level predictions. Then, we propose a Recurrent Neural Network based approach to obtain session-level predictions from segment-level predictions. We show that strengths of the segment-wise classifier are amplified when a session-wise classifier is trained on embeddings obtained from it. The model trained on ACFs derived from TVs show relative improvement of 27.47% in Unweighted Average Recall (UAR) at the session-level classification task, compared to the ACFs derived from Mel Frequency Cepstral Coefficients (MFCCs). Nadee Seneviratne, Carol Y. Espy-Wilson |
Interspeech | 2 |
| 2021 | Generalized Dilated CNN Models for Depression Detection Using Inverted Vocal Tract VariablesabstractDepression detection using vocal biomarkers is a highly researched area. Articulatory coordination features (ACFs) are developed based on the changes in neuromotor coordination due to psychomotor slowing, a key feature of Major Depressive Disorder. However findings of existing studies are mostly validated on a single database which limits the generalizability of results. Variability across different depression databases adversely affects the results in cross corpus evaluations (CCEs). We propose to develop a generalized classifier for depression detection using a dilated Convolutional Neural Network which is trained on ACFs extracted from two depression databases. We show that ACFs derived from Vocal Tract Variables (TVs) show promise as a robust set of features for depression detection. Our model achieves relative accuracy improvements of ~10% compared to CCEs performed on models trained on a single database. We extend the study to show that fusing TVs and Mel-Frequency Cepstral Coefficients can further improve the performance of this classifier. Nadee Seneviratne, Carol Y. Espy-Wilson |
Interspeech | 2 |
| 2020 | Extended Study on the Use of Vocal Tract Variables to Quantify Neuromotor Coordination in Depression
Nadee Seneviratne, James R. Williamson, Adam C. Lammert, Thomas F. Quatieri, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2019 | Assessing Neuromotor Coordination in Depression Using Inverted Vocal Tract Variables
Carol Y. Espy-Wilson, Adam C. Lammert, Nadee Seneviratne, Thomas F. Quatieri |
INTERSPEECH | 1 |
| 2019 | Multi-Modal Learning for Speech Emotion Recognition: An Analysis and Comparison of ASR Outputs with Ground Truth Transcription
Saurabh Sahu, Vikramjit Mitra, Nadee Seneviratne, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2019 | Multi-Corpus Acoustic-to-Articulatory Speech Inversion
Nadee Seneviratne, Ganesh Sivaraman, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2018 | Semi-Supervised and Transfer Learning Approaches for Low Resource Sentiment ClassificationabstractSentiment classification involves quantifying the affective reaction of a human to a document, media item or an event. Although researchers have investigated several methods to reliably infer sentiment from lexical, speech and body language cues, training a model with a small set of labeled datasets is still a challenge. For instance, in expanding sentiment analysis to new languages and cultures, it may not always be possible to obtain comprehensive labeled datasets. In this paper, we investigate the application of semi- supervised and transfer learning methods to improve performances on low resource sentiment classification tasks. We experiment with extracting dense feature representations, pre-training and manifold regularization in enhancing the performance of sentiment classification systems. Our goal is a coherent implementation of these methods and we evaluate the gains achieved by these methods in matched setting involving training and testing on a single corpus setting as well as two cross corpora settings. In both the cases, our experiments demonstrate that the proposed methods can significantly enhance the model performance against a purely supervised approach, particularly in cases involving a handful of training data. Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan |
ICASSP | 3 |
| 2018 | Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion RecognitionabstractTraining discriminative classifiers involves learning a conditional distribution p(yi|xi), given a set of feature vectors xiand the corresponding labels yi, i=1...N. For a classifier to be generalizable and not overfit to training data, the resulting conditional distribution p(yi|xi) is desired to be smoothly varying over the inputs xi. Adversarial training procedures enforce this smoothness using manifold regularization techniques. Manifold regularization makes the model's output distribution more robust to local perturbation added to a datapoint xi. In this paper, we experiment with the application of adversarial training procedures to increase the accuracy of a deep neural network based emotion recognition system using speech cues. Specifically, we investigate two training procedures: (i) adversarial training where we determine the adversarial direction based on the given labels for the training data and, (ii) virtual adversarial training where we determine the adversarial direction based only on the output distribution of the training data. We demonstrate the efficacy of adversarial training procedures by performing a k-fold cross validation experiment on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and a cross-corpus performance analysis on three separate corpora. Results show improvement over a purely supervised approach, as well as better generalization capability to cross-corpus settings. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Carol Y. Espy-Wilson |
ICASSP | 4 |
| 2018 | On Enhancing Speech Emotion Recognition Using Generative Adversarial NetworksabstractGenerative Adversarial Networks (GANs) have gained a lot of attention from machine learning community due to their ability to learn and mimic an input data distribution. GANs consist of a discriminator and a generator working in tandem playing a min-max game to learn a target underlying data distribution; when fed with data-points sampled from a simpler distribution (like uniform or Gaussian distribution). Once trained, they allow synthetic generation of examples sampled from the target distribution. We investigate the application of GANs to generate synthetic feature vectors used for speech emotion recognition. Specifically, we investigate two set ups: (i) a vanilla GAN that learns the distribution of a lower dimensional representation of the actual higher dimensional feature vector and, (ii) a conditional GAN that learns the distribution of the higher dimensional feature vectors conditioned on the labels or the emotional class to which it belongs. As a potential practical application of these synthetically generated samples, we measure any improvement in a classifier's performance when the synthetic data is used along with real data for training. We perform cross-validation analyses followed by a cross-corpus study. Saurabh Sahu, Rahul Gupta 0001, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2018 | Noise Robust Acoustic to Articulatory Speech Inversion
Nadee Seneviratne, Ganesh Sivaraman, Vikramjit Mitra, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2017 | Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasksabstractArticulatory information can effectively model variability in speech and can improve speech recognition performance under varying acoustic conditions. Learning speaker-independent articulatory models has always been challenging, as speaker-specific information in the articulatory and acoustic spaces increases the complexity of the speech-to-articulatory space inverse modeling, which is already an ill-posed problem due to its inherent nonlinearity and non-uniqueness. This paper investigates using deep neural networks (DNN) and convolutional neural networks (CNNs) for mapping speech data into its corresponding articulatory space. Our results indicate that the CNN models perform better than their DNN counterparts for speech inversion. In addition, we used the inverse models to generate articulatory trajectories from speech for three different standard speech recognition tasks. To effectively model the articulatory features' temporal modulations while retaining the acoustic features' spatiotemporal signatures, we explored a joint modeling strategy to simultaneously learn both the acoustic and articulatory spaces. The results from multiple speech recognition tasks indicate that articulatory features can improve recognition performance when the acoustic and articulatory spaces are jointly learned with one common objective function. Vikramjit Mitra, Ganesh Sivaraman, Chris Bartels, Hosung Nam, Wen Wang 0001, Carol Y. Espy-Wilson, Dimitra Vergyri, Horacio Franco |
ICASSP | 6 |
| 2017 | An Affect Prediction Approach Through Depression Severity Parameter Incorporation in Neural Networks
Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan |
INTERSPEECH | 3 |
| 2017 | Adversarial Auto-Encoders for Speech Based Emotion RecognitionabstractRecently, generative adversarial networks and adversarial autoencoders have gained a lot of attention in machine learning community due to their exceptional performance in tasks such as digit classification and face recognition. They map the autoencoder's bottleneck layer output (termed as code vectors) to different noise Probability Distribution Functions (PDFs), that can be further regularized to cluster based on class information. In addition, they also allow a generation of synthetic samples by sampling the code vectors from the mapped PDFs. Inspired by these properties, we investigate the application of adversarial autoencoders to the domain of emotion recognition. Specifically, we conduct experiments on the following two aspects: (i) their ability to encode high dimensional feature vector representations for emotional utterances into a compressed space (with a minimal loss of emotion class discriminability in the compressed space), and (ii) their ability to regenerate synthetic samples in the original feature space, to be later used for purposes such as training emotion recognition classifiers. We demonstrate the promise of adversarial autoencoders with regards to these aspects on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpus and present our analysis. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Wael Abd-Almageed, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2017 | Analysis of Acoustic-to-Articulatory Speech Inversion Across Different Accents and Languages
Ganesh Sivaraman, Carol Y. Espy-Wilson, Martijn Wieling 0001 |
INTERSPEECH | 2 |
| 2017 | Hybrid convolutional neural networks for articulatory and acoustic information based speech recognition
Vikramjit Mitra, Ganesh Sivaraman, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Mark K. Tiede |
Speech Commun. | 4 |
| 2016 | Speech Features for Depression Detection
Saurabh Sahu, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2016 | Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
Ganesh Sivaraman, Vikramjit Mitra, Hosung Nam, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2015 | Analysis of coarticulated speech using estimated articulatory trajectories
Ganesh Sivaraman, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson |
INTERSPEECH | 6 |
| 2014 | Articulatory features from deep neural networks and their role in speech recognitionabstractThis paper presents a deep neural network (DNN) to extract articulatory information from the speech signal and explores different ways to use such information in a continuous speech recognition task. The DNN was trained to estimate articulatory trajectories from input speech, where the training data is a corpus of synthetic English words generated by the Haskins Laboratories' task-dynamic model of speech production. Speech parameterized as cepstral features were used to train the DNN, where we explored different cepstral features to observe their role in the accuracy of articulatory trajectory estimation. The best feature was used to train the final DNN system, where the system was used to predict articulatory trajectories for the training and testing set of Aurora-4, the noisy Wall Street Journal (WSJ0) corpus. This study also explored the use of hidden variables in the DNN pipeline as a potential acoustic feature candidate for speech recognition and the results were encouraging. Word recognition results from Aurora-4 indicate that the articulatory features from the DNN provide improvement in speech recognition performance when fused with other standard cepstral features; however when tried by themselves, they failed to match the baseline performance. Vikramjit Mitra, Ganesh Sivaraman, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman |
ICASSP | 4 |
| 2013 | A cine MRI-based study of sibilant fricatives production in post-glossectomy speakersabstractGlossectomy changes properties of the tongue and negatively affects patients' speech production. Among the most difficult consonants to produce in the post-glossectomy speakers, the sibilant fricatives /s/ and /sh/ are often problematic. To better understand these problems in production, this study analyzed acoustic and articulatory data of /s/ and /sh/ from three subjects: one normal speaker and two post-glossectomy speakers with abnormal /s/ or /sh. Based on cine magnetic resonance images, three dimensional vocal tract reconstructions, tongue surface shapes behind constrictions, and area functions were analyzed. Our results show that in each patient, contrary to normal, /s/ and /sh/ were quite similar in acoustic spectra, tongue surface shapes, and constriction locations. In the abnormal /s/, the missing unilateral tongue tissue created an air flow bypass which made the constriction further backward. The abnormal /sh/ may be explained by the lack of precise tongue control after surgery. In addition, the tongue surfaces in the patients were more asymmetric in the back and were not grooved for /s/ anterior to the constriction. Xinhui Zhou, Jonghye Woo, Maureen Stone 0001, Carol Y. Espy-Wilson |
ICASSP | 4 |
| 2012 | Multicondition training of Gaussian PLDA models in i-vector space for noise and reverberation robust speaker recognitionabstractWe present a multicondition training strategy for Gaussian Probabilistic Linear Discriminant Analysis (PLDA) modeling of i-vector representations of speech utterances. The proposed approach uses a multicondition set to train a collection of individual subsystems that are tuned to specific conditions. A final verification score is obtained by combining the individual scores according to the posterior probability of each condition given the trial at hand. The performance of our approach is demonstrated on a subset of the interview data of NIST SRE 2010. Significant robustness to the adverse noise and reverberation conditions included in the multicondition training set are obtained. The system is also shown to generalize to unseen conditions. Daniel Garcia-Romero, Xinhui Zhou, Carol Y. Espy-Wilson |
ICASSP | 3 |
| 2012 | Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulationsabstractOral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Depending on the tumor size, location and staging, patients are treated by radical surgery, radiology, chemotherapy or a combination of those treatments. As a result, their anatomical structures for speech are impaired and this leads to some negative impact on their speech intelligibility. As a part of the INTERSPEECH 2012 speaker trait Pathology sub-challenge, this study explored the use of auditory-inspired spectro-temporal modulation features for automatic speech intelligibility assessment of those pathologic speech. The averaged spectro-temporal modulations of speech considered as either intelligible or non-intelligible in the challenge database were analyzed and it was found that the non-intelligible speech tends to have its modulation amplitude peaks shift towards a smaller rate and scale. Based on SVM and GMM, variants of spectro-temporal modulation features were tested on the speaker trait challenge problem and the resulting performances on both the development and the test datasets are comparable to the baseline performance. Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen Stone 0001, Carol Y. Espy-Wilson, Shihab A. Shamma |
INTERSPEECH | 5 |
| 2011 | Robust speech recognition using articulatory gestures in a Dynamic Bayesian Network frameworkabstractArticulatory Phonology models speech as spatio-temporal constellation of constricting events (e.g. raising tongue tip, narrowing lips etc.), known as articulatory gestures. These gestures are associated with distinct organs (lips, tongue tip, tongue body, velum and glottis) along the vocal tract. In this paper we present a Dynamic Bayesian Network based speech recognition architecture that models the articulatory gestures as hidden variables and uses them for speech recognition. Using the proposed architecture we performed: (a) word recognition experiments on the noisy data of Aurora-2 and (b) phone recognition experiments on the University of Wisconsin X-ray microbeam database. Our results indicate that the use of gestural information helps to improve the performance of the recognition system compared to the system using acoustic information only. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson |
ASRU | 3 |
| 2011 | Linear versus mel frequency cepstral coefficients for speaker recognitionabstractMel-frequency cepstral coefficients (MFCC) have been dominantly used in speaker recognition as well as in speech recognition. However, based on theories in speech production, some speaker characteristics associated with the structure of the vocal tract, particularly the vocal tract length, are reflected more in the high frequency range of speech. This insight suggests that a linear scale in frequency may provide some advantages in speaker recognition over the mel scale. Based on two state-of-the-art speaker recognition back-end systems (one Joint Factor Analysis system and one Probabilistic Linear Discriminant Analysis system), this study compares the performances between MFCC and LFCC (Linear frequency cepstral coefficients) in the NIST SRE (Speaker Recognition Evaluation) 2010 extended-core task. Our results in SRE10 show that, while they are complementary to each other, LFCC consistently outperforms MFCC, mainly due to its better performance in the female trials. This can be explained by the relatively shorter vocal tract in females and the resulting higher formant frequencies in speech. LFCC benefits more in female speech by better capturing the spectral characteristics in the high frequency region. In addition, our results show some advantage of LFCC over MFCC in reverberant speech. LFCC is as robust as MFCC in the babble noise, but not in the white noise. It is concluded that LFCC should be more widely used, at least for the female trials, by the mainstream of the speaker recognition community. Xinhui Zhou, Daniel Garcia-Romero, Ramani Duraiswami, Carol Y. Espy-Wilson, Shihab A. Shamma |
ASRU | 4 |
| 2011 | Gesture-based Dynamic Bayesian Network for noise robust speech recognitionabstractPreviously we have proposed different models for estimating articulatory gestures and vocal tract variable (TV) trajectories from synthetic speech. We have shown that when deployed on natural speech, such models can help to improve the noise robustness of a hidden Markov model (HMM) based speech recognition system. In this paper we propose a model for estimating TVs trained on natural speech and present a Dynamic Bayesian Network (DBN) based speech recognition architecture that treats vocal tract constriction gestures as hidden variables, eliminating the necessity for explicit gesture recognition. Using the proposed architecture we performed a word recognition task for the noisy data of Aurora 2. Significant improvement was observed in using the gestural information as hidden variables in a DBN architecture over using only the mel-frequency cepstral coefficient based HMM or DBN backend. We also compare our results with other noise-robust front ends. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
ICASSP | 3 |
| 2011 | Speech inversion: Benefits of tract variables over pellet trajectoriesabstractSpeech inversion is a way of estimating articulatory trajectories or vocal tract configurations from the acoustic speech signal. Traditionally, articulator flesh-point or pellet trajectories have been used in speech-inversion research; however such information introduces additional variability into the inverse problem given they are head-centered, task-neutral measures. This paper proposes the use of vocal tract constriction variables (TVs) that are less variable for speech-inversion since they are constriction-based, task-specific measures. TVs considered in this study consist of five constriction degree variables, lip aperture (LA), tongue body constriction degree (TBCD), tongue tip constriction degree (TTCD), velum (VEL), and glottis (GLO); and three constriction location variables, lip protrusion (LP), tongue tip constriction location (TTCL) and tongue body constriction location (TBCL). Six different flesh-point trajectories were considered that were measured with transducers placed on the upper lip (UL), lower lip (LL) and four positions on the tongue (T1, T2, T3 and T4) between the tongue tip and the tongue dorsum. Speech inversion using a simple neural network architecture shows that the TVs can be estimated relatively more accurately than the pellet trajectories. Further statistical investigation reveals that the non-uniqueness is reduced in the TVs compared to the pellet trajectories for phones which are known to appreciably suffer from non-uniqueness. Finally we perform word recognition experiments using the estimated TVs as opposed to the pellet trajectories and show that the former offers greater word recognition accuracy both in clean and noisy speech, indicating that the TVs are a better choice for speech recognition systems. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
ICASSP | 3 |
| 2011 | Analysis of i-vector Length Normalization in Speaker Recognition SystemsabstractWe present a method to boost the performance of probabilistic generative models that work with i-vector representations. The proposed approach deals with the nonGaussian behavior of i-vectors by performing a simple length normalization. This non-linear transformation allows the use of probabilistic models with Gaussian assumptions that yield equivalent performance to that of more complicated systems based on Heavy-Tailed assumptions. Significant performance improvements are demonstrated on the telephone portion of NIST SRE 2010. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2011 | Automatic Speech Codec Identification with Applications to Tampering Detection of Speech RecordingsabstractIn this work many versions of CELP codecs are explored, and an observation is made that different codebooks are used to encode noisy part of residual. Taking advantage of noise patterns they generated, an algorithm was proposed to detect GSM-AMR,EFR,HR and SILK codecs. Another partly knowledge-based and partly data driven algorithm is also proposed to improve the performance for SILK. Then it's extended to identify subframe offset to do tampering detection of cellphone speech recordings. Jingting Zhou, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2011 | A Comparative Acoustic Study on Speech of Glossectomy Patients and Normal SubjectsabstractOral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Tongue cancer patients are treated by glossectomy, a surgical procedure to remove the cancerous tumor. As a result, the tongue properties such as volume, shape, muscle structure, and motility are affected. As a result, the vocal tract acoustics are affected too. This study compares the speech acoustics between normal subjects and partial glossectomy patients with T1 or T2 tumors. The acoustic signal of four vowels (/iy/, /uw/, /eh/, and /ah/) and two fricatives (/s/ and /sh/) were analyzed. Our results show that, while the average formants (F1-F3) for the four vowels between the normal subjects and the glossectomy patients are very similar, the average centers of gravity for the two fricatives differ significantly. These differences in fricatives can be explained by the more posterior constriction in patients due to the glossectomy (or the cancer tumor) and its resulting longer front cavity. Xinhui Zhou, Maureen Stone 0001, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2011 | Articulatory Information for Noise Robust Speech RecognitionabstractPrior research has shown that articulatory information, if extracted properly from the speech signal, can improve the performance of automatic speech recognition systems. However, such information is not readily available in the signal. The challenge posed by the estimation of articulatory information from speech acoustics has led to a new line of research known as “acoustic-to-articulatory inversion” or “speech-inversion.” While most of the research in this area has focused on estimating articulatory information more accurately, few have explored ways to apply this information in speech recognition tasks. In this paper, we first estimated articulatory information in the form of vocal tract constriction variables (abbreviated as TVs) from the Aurora-2 speech corpus using a neural network based speech-inversion model. Word recognition tasks were then performed for both noisy and clean speech using articulatory information in conjunction with traditional acoustic features. Our results indicate that incorporating TVs can significantly improve word recognition rates when used in conjunction with traditional acoustic features. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Automatic acquisition device identification from speech recordingsabstractIn this paper we present a study on the automatic identification of acquisition devices when only access to the output speech recordings is possible. A statistical characterization of the frequency response of the device contextualized by the speech content is proposed. In particular, the intrinsic characteristics of the device are captured by a template, constructed by appending together the means of a Gaussian mixture trained on the device speech recordings. This study focuses on two classes of acquisition devices, namely, landline telephone handsets and microphones. Three publicly available databases are used to assess the performance of linear- and mel-scaled cepstral coefficients. A Support Vector Machine classifier was used to perform closed-set identification experiments. The results show classification accuracies higher than 90 percent among the eight telephone handsets and eight microphones tested. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2010 | An MRI-based articulatory and acoustic study of lateral sound in American EnglishabstractThe production of the lateral sounds generally involves a linguo-alveolar contact and one or two lateral channels along the parasagittal sides of the tongue. The acoustic effect of these articulatory features is not clearly understood. In this study, we compare two productions of /l/ in American English by one subject, one for a dark /l/ and the other for a light /l/. Three-dimensional vocal tract models derived from the magnetic resonance images were analyzed. It was shown that zeros in the vocal tract acoustic response are produced in the F3-F5 region in both /l/ productions, but the number of zeros and their frequencies are affected by the length of the linguo-alveolar contact and by the presence or absence of lateral linguopalatal contacts. The dark /l/ has one zero below 5 kHz, produced by the cross mode posterior to the linguo-alveolar contact, while the light /l/ has three zeros below 5 kHz, produced by the asymmetrical lateral channels, the supralingual cavity and the cross mode posterior to linguo-alveolar contact. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
ICASSP | 2 |
| 2010 | Robust word recognition using articulatory trajectories and gesturesabstractArticulatory Phonology views speech as an ensemble of constricting events (e.g. narrowing lips, raising tongue tip), gestures, at distinct organs (lips, tongue tip, tongue body, velum, and glottis) along the vocal tract. This study shows that articulatory information in the form of gestures and their output trajectories (tract variable time functions or TVs) can help to improve the performance of automatic speech recognition systems. The lack of any natural speech database containing such articulatory information prompted us to use a synthetic speech dataset (obtained from Haskins Laboratories TAsk Dynamic model of speech production) that contains acoustic waveform for a given utterance and its corresponding gestures and TVs. First, we propose neural network based models to recognize the gestures and estimate the TVs from acoustic information. Second, the “synthetic-data trained” articulatory models were applied to the natural speech utterances in Aurora-2 corpus to estimate their gestures and TVs. Finally, we show that the estimated articulatory information helps to improve the noise robustness of a word recognition system when used along with the cepstral Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
INTERSPEECH | 3 |
| 2010 | A procedure for estimating gestural scores from natural speechabstractAbstract * Speech can be represented as a constellation of constricting events, gestures, , which are defined at distinct vocal tract sites, in the form of a gestural score.. Gestures and their output trajectories, tract variables, , which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given an utterance, the Haskins Laboratories Task Dynamics and Application (TADA) model was used to generate its prototype gestural score and the corresponding synthetic acoustic output. An optimal gestural score was estimated through iterative time-warping processes such that the distance between original and TADA-synthesized synthesized speech is minimized. We compared the performance of our approach to that of a conventional dynamic time warping procedure using Log-Spectral and Itakura Distance measures. We also performed a word recognition experiment using the gestural annotations to show that the gestural scores are suitable for word recognition. Hosung Nam, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson |
INTERSPEECH | 6 |
| 2009 | From acoustics to Vocal Tract time functionsabstractIn this paper we present a technique for obtaining Vocal Tract (VT) time functions from the acoustic speech signal. Knowledge-based Acoustic Parameters (APs) are extracted from the speech signal and a pertinent subset is used to obtain the mapping between them and the VT time functions. Eight different vocal tract constriction variables consisting of five constriction degree variables, lip aperture (LA), tongue body (TBCD), tongue tip (TTCD), velum (VEL), and glottis (GLO); and three constriction location variables, lip protrusion (LP), tongue tip (TTCL), tongue body (TBCL) were considered in this study. The TAsk Dynamics Application model (TADA [1]) is used to create a synthetic speech dataset along with its corresponding VT time functions. We explore Support Vector Regression (SVR) followed by Kalman smoothing to achieve mapping between the APs and the VT time functions. Vikramjit Mitra, I. Yücel Özbek, Hosung Nam, Xinhui Zhou, Carol Y. Espy-Wilson |
ICASSP | 5 |
| 2009 | An algorithm for speech segregation of co-channel speechabstractThis paper introduces an algorithm to separate speech streams from a single-channel speech mixture. Most current speech segregation algorithms allocate speech regions to participating speakers depending on which speaker dominates in which spectro-temporal region. The proposed method is a different approach to speech segregation, in that it separates the participating speaker streams rather than decide in the favor of the dominating speaker. The algorithm depends on a lease-squares fitting approach to model the speech mixture as a sum of complex exponentials. The algorithm gives results that are better than an existent algorithm when tested on the same task. The performance on a different database yielded good segregation results, even for Target-to-Masker ratios as low as -15 dB. The algorithm has immense promise for improvement and practical implementation. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2009 | A noise-type and level-dependent MPO-based speech enhancement architecture with variable frame analysis for noise-robust speech recognitionabstractIn previous work, a speech enhancement algorithm based on phase opponency and a periodicity measure (MPO-APP) was developed for speech recognition. Axiomatic thresholds were used in the MPO-APP regardless of the signal-to-noise ratio (SNR) of the corrupted speech or any characterization of the noise. The current work developed an algorithm for adjusting the threshold in the MPO-APP based on the SNR and whether the speech signal is clean, corrupted by aperiodic noise or corrupted with noise with periodic components. In addition, variable frame rate (VFR) analysis has been incorporated so that dynamic regions in the speech signal are more heavily sampled than steady-state regions. The result is a 2-stage algorithm that gives superior performance to the previous MPO-APP, and to several other state-of-the-art speech enhancement algorithms. Index Terms: Speech enhancement, robust speech recognition, SNR estimation, variable frame rate analysis, phase opponency. Vikramjit Mitra, Bengt J. Borgstrom, Carol Y. Espy-Wilson, Abeer Alwan |
INTERSPEECH | 3 |
| 2009 | Noise robustness of tract variables and their application to speech recognition
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
INTERSPEECH | 3 |
| 2008 | Language detection in audio content analysisabstractExperiments have shown that Language Identification systems for telephonic speech using shifted delta cepstra as the feature set and Gaussian mixture models as the backend, offers superior performance than other competing techniques. This paper aims to address the task of Language Identification for audio signals. The abundance of digital music from the Internet calls for a reliable real-time system for analyzing and properly categorizing them. Previous research has mainly focused on categorizing audio files into appropriate genres; however genre types vary with language. This paper proposes a systematic audio content analysis strategy by initially detecting whether an audio file has any vocals present in it and, if present, then detecting the language of the song. Given the language of the song, genre detection becomes a closed set classification problem. Vikramjit Mitra, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
ICASSP | 3 |
| 2008 | Intersession variability in speaker recognition: a behind the scene analysisabstractThe representation of a speaker’s identity by means of Gaussian supervectors (GSV) is at the heart of most of the state-of-the-art recognition systems. In this paper we present a novel procedure for the visualization of GSV by which qualitative insight about the information being captured can be obtained. Based on this visualization approach, the Switchboard-I database (SWB-I) is used to study the relationship between a data-driven partition of the acoustic space and a knowledge based partition (i.e., broad phonetic classes). Moreover, the structure of an intersession variability subspace (IVS), computed from the SWB-I database, is analyzed by displaying the projection of a speaker’s GSV into the set of eigenvectors with highest eigenvalues. This analysis reveals a strong presence of linguistic information in the IVS components with highest energy. Finally, after projecting away the information contained in the IVS from the speaker’s GSV, a visualization of the resulting GSV provides information about the characteristic patterns of spectral allocation of energy of a speaker. Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2008 | Language and genre detection in audio content analysisabstractThis paper presents an audio genre detection framework that can be used for a multi-language audio corpus. Cepstral coefficients are considered and analyzed as the feature set for both a language dependent and language independent genre identification (GID) task. Language information is found to increase the overall detection accuracy on an average by at least 2.6% from its language independent counterpart. Melfrequency cepstral coefficients have been widely used for Music Information Retrieval (MIR), however, the present study shows that Linear-frequency cepstral coefficients (LFCC) with a higher number of frequency bands can improve the detection accuracy. Two other GID architectures have also been considered, but the results show that the logenergy amplitudes from triangular linearly spaced filter banks and their deltas can offer average detection accuracy as high as 98.2%, when language information is taken into account. Vikramjit Mitra, Daniel Garcia-Romero, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2008 | An algorithm for multi-pitch tracking in co-channel speechabstractMost multi-pitch algorithms are tested for performance only in voiced regions of speech, and are prone to yield pitch estimates even when the participating speakers are unvoiced. This paper presents a multi-pitch algorithm that detects the voiced and unvoiced regions in a mixture of two speakers, identifies the number of speakers in voiced regions, and yields the pitch estimates of each speaker in those regions. The algorithm relies on the 2-Dimensional AMDF for estimating the periodicity of the signal, and uses the temporal evolution of the 2-D AMDF to estimate the number of speakers present in periodic regions. Evaluation of this algorithm on a frame-wise basis demonstrates accurate voiced / unvoiced decisions and also gives pitch estimation results comparable to the state of the art. The pitch estimation errors are quantitatively analyzed and shown to be resulting partly from speaker domination & pitch matching between speakers. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2007 | Landmark-based approach to speech recognition: an alternative to HMMsabstractIn this paper, we compare a Probabilistic Landmark-Based speech recognition System (LBS) which uses Knowledge-based Acoustic Parameters (APs) as the front-end with an HMMbased recognition system that uses the Mel-Frequency Cepstral Coefficients as its front end. The advantages of LBS based on APs are (1) the APs are normalized for extra-linguistic information, (2) acoustic analysis at different landmarks may be performed with different resolutions and with different APs, (3) LBS outputs multiple acoustic landmark sequences that signal perceptually significant regions in the speech signal, (4) it may be easier to port this system to another language since the phonetic features captured by the APs are universal, and (5) LBS can be used as a tool for uncovering and subsequently understanding variability. LBS also has a probabilistic framework that can be combined with pronunciation and language models in order to make it more scalable to large vocabulary recognition tasks. Index Terms: landmark, speech recognition, acoustic parameters, phonetic features. Carol Y. Espy-Wilson, Tarun Pruthi, Amit Juneja, Om Deshmukh |
INTERSPEECH | 1 |
| 2007 | A semi-automatic approach for speaker mining of tapped telephone conversations
Sandeep Manocha, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2007 | Acoustic parameters for the automatic detection of vowel nasalizationabstractThe aim of this work was to propose Acoustic Parameters (APs) for the automatic detection of vowel nasalization based on prior knowledge of the acoustics of nasalized vowels. Nine automatically extractable APs were proposed to capture the most important acoustic correlates of vowel nasalization (extra pole-zero pairs, F1 amplitude reduction, F1 bandwidth increase and spectral flattening). The performance of these APs was tested on several databases with different sampling rates and recording conditions. Accuracies of 96.28%, 77.90% and 69.58% were obtained by using these APs on StoryDB, TIMIT and WS96/97 databases, respectively, in a Support Vector Machine classifier framework. To our knowledge these results are the best anyone has achieved on this task. Index Terms: nasal, nasalization, acoustic parameters, landmark, speech recognition. Tarun Pruthi, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2007 | An articulatory and acoustic study of "retroflex" and "bunched" american English rhotic sound based on MRIabstractThe North American rhotic liquid has two maximally distinct articulatory variants, the classic ”retroflex” and the classic ”bunched” tongue postures. The evidence for acoustic differences between these two variants is reexamined using magnetic resonance images of the vocal tract in this study. Two subjects with similar vocal tract dimensions but different tongue postures for sustained /r/ are used. It is shown that these two variants have similar patterns of F1-F3 and zero frequencies. However, the ”retroflex” variant has a larger difference between F4 and F5 than the ”bunched” one (around 1400 Hz vs. around 700 Hz). This difference can be explained by the geometry differences between these two variants, in particular, the shorter and more forward palatal constriction of the ”retroflex” /r/ and the sharper transition between palatal constriction and its anterior and posterior cavities. This formant pattern difference is confirmed by measurement from acoustic data of several additional subjects. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
INTERSPEECH | 2 |
| 2007 | Report on the NSF-sponsored Human Language Technology Workshop on Industrial Centers
Mary P. Harper, Alex Acero, Srinivas Bangalore, Jordan Cohen, Barbara Cuthill, Carol Y. Espy-Wilson, Christiane Fellbaum, John Garofolo, Chin-Hui Lee 0001, Jim Lester, Andrew McCallum, Nelson Morgan, Michael Picheney, Joseph Picone, Lance Ramshaw, Jeffrey C. Reynar, Hadar Shemtov, Clare Voss |
MTSummit | 7 |
| 2006 | Modified phase opponency based solution to the speech separation challengeabstractIn this work, we present a single-channel speech enhancement technique called the Modified Phase Opponency (MPO) model as a solution to the Speech Separation Challenge. The MPO model is based on a neural model for detection of tones-in-noise called the Phase Opponency (PO) model. Replacing the noisy speech signals by the corresponding MPO-processed signals increases the accuracy by 31%when the speech signals are corrupted by speech-shaped noise at 0 dB Signal-to-Noise Ratio (SNR). It is worth men-tioning that the MPO enhancement scheme was developed using the noisy connected-digit Aurora database and was not tailored in any way to fit the Grid database used in this challenge. One of the salient features of the MPO-based speech enhancement scheme is that it does not need to estimate the noise characteristics, nor does it assume that the noise satisfies any statistical model. Index Terms: speech separation, robust speech recognition. 1. Om Deshmukh, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2006 | Speech enhancement using modified phase opponency model
Om Deshmukh, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2006 | A new set of features for text-independent speaker identificationabstractThe success of a speaker identification system depends largely on the set of features used to characterize speaker-specific information. In this paper, we discuss a small set of low-level acoustic parameters that capture information about the speaker’s source, vocal tract size and vocal tract shape. We demonstrate that the set of eight acoustic parameters has comparable performance to the standard sets of 26 or 39 MFCCs for the speaker identification task. Gaussian Mixture Models were used for constructing speaker models. Index Terms: speaker identification, acoustic parameters, Carol Y. Espy-Wilson, Sandeep Manocha, Srikanth Vishnubhotla |
INTERSPEECH | 1 |
| 2006 | An MRI based study of the acoustic effects of sinus cavities and its application to speaker recognitionabstractThe goal of this paper is to explore the effects of changes in velar coupling area and oral cavity configuration on the poles and zeros introduced in the nasalized vowel and nasal consonant spectra due to the sphenoidal and maxillary sinuses. MRI data for the vocal tract and nasal tract of one speaker was used to simulate the spectra of the nasalized vowels, and nasal consonants with different coupling areas. It is shown that during nasalized vowels, the frequencies of both poles and zeros due to the sinuses change with a change in the velar coupling area or the vowel. It is also shown that during nasal consonants, the zero frequencies are constant, and the pole frequencies are more stable as compared to nasalized vowels. This study, therefore, corroborates the use of nasal consonant spectra for speaker recognition and raises doubts on the potential benefits of using nasalization during vowels for that purpose. Index Terms: speaker recognition, nasal, sinus, MRI. 1. Tarun Pruthi, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2006 | Automatic detection of irregular phonation in continuous speechabstractVoice quality is one of the most important source characteristics of a speaker’s speech production process, and creakiness is one of the variations of voice quality. This paper describes the development of an algorithm to automatically detect irregular phonation, including creakiness and other variations, in continuous running speech. The algorithm is an extension of the Aperiodicity, Periodicity and Pitch (APP) Detector. The algorithm has been run on 485 files of the TIMIT database, which contained 677 instances of irregular phonation. The test set comprised of 97 speakers, of which 57 were male and 40 were female. The algorithm has been found to give an accuracy of 86.7 % on average, with performance being almost the same for both male and female speakers. Automatic detection of irregular phonation should help characterize speakers for speaker identification applications. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2005 | Modeling of the Front Cavity and Sublingual Space in American English Rhotic SoundsabstractThe production of American English (AE) /r/ sounds is variable but generally involves a large-volume front cavity. In some cases, there is also a sublingual cavity. Previous work has shown that the large front-cavity volume and the sublingual cavity are directly or indirectly responsible for the characteristically low frequency of the third formant (F3) of /r/. The entire front cavity is normally modeled as a single tube. The sublingual cavity, if present, is modeled as a side branch to the front cavity. However, given the dimensions of the front cavity, it is possible that high order acoustic modes are excited which may produce zeros and affect formant locations. Detailed information of the flow field involved can help to understand better and model accurately the front cavity acoustics. A finite element study of the flow field in the front cavity and sublingual cavity is described, using dimensions measured from MRI studies of subjects producing AE /r/. The results show that the large-volume front cavity is better modeled as a single tube with a side branch rather than as a single tube alone. The effective length of this side branch is further increased by the presence of a sublingual cavity, giving a zero in the range of F5 in the resulting spectrum. Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede |
ICASSP (1) | 2 |
| 2005 | Speech enhancement using auditory phase opponency modelabstractIn this work we address the problem of single-channel speech enhancement when the speech is corrupted by additive noise. The model presented here, called the Modified Phase Opponency (MPO) model, is based on the auditory PO model, proposed by Carney et. al., for detection of tones in noise. The PO model includes a physiologically realistic mechanism for processing the information in neural discharge times and exploits the frequency-dependent phase properties of the tuned filters in the auditory periphery by using a cross-auditory-nerve-fiber coincidence detection for extracting temporal cues. Initial evaluation of the MPO model on speech corrupted by white noise at different SNRs shows that the MPO model is able to enhance the spectral peaks while suppressing the noise-only regions. Om Deshmukh, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2005 | Use of Temporal Information: Detection of Periodicity, Aperiodicity, and Pitch in SpeechabstractIn this paper, we present a time domain aperiodicity, periodicity, and pitch (APP) detector that estimates 1) the proportion of periodic and aperiodic energy in a speech signal and 2) the pitch period of the periodic component. The APP system is particularly useful in situations where the speech signal contains simultaneous periodic and aperiodic energy, as in the case of breathy vowels and some voiced obstruents. The performance of the APP system was evaluated on synthetic speech-like signals corrupted with noise at various levels of signal-to-noise ratio (SNR) and on three different natural speech databases that consist of simultaneously recorded electroglottograph (EGG) and acoustic data. When compared on a frame basis (at a frame rate of 2.5 ms) the results show excellent agreement between the periodic/aperiodic decisions made by the APP system and the estimates obtained from the EGG data (94.43% for periodicity and 96.32% for aperiodicity). The results also support previous studies that show that voiced obstruents are frequently manifested with either little or no aperiodic energy, or with strong periodic and aperiodic components. The EGG data were used as a reference for evaluating the pitch detection algorithm. The ground truth was not manually checked to rectify or exclude incorrect estimates. The overall gross error rate in pitch prediction across the three speech databases was 5.67%. In the case of synthetic speech-like data, the estimated SNR was found to be in close proportion to the actual SNR, and the pitch was always accurately found regardless of the presence of any shimmer or jitter. Om Deshmukh, Carol Y. Espy-Wilson, Ariel Salomon, Jawahar Singh |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | A novel method for computation of periodicity, aperiodicity and pitch of speech signalsabstractThe paper presents improvements to our previously proposed algorithm to compute the proportion of periodic and aperiodic energies in speech signals and to estimate the pitch period. Although previously the periodic and aperiodic energies were estimated independently of each other at each frame, a binary decision was made at each of the non-silent channels. We present an extension that replaces the binary decision with a measure of the degree of periodicity and aperiodicity in each channel. Evaluation on synthetic speech-like data shows a better agreement in the estimated SNR and the actual SNR by using this improvement. Moreover, in the task of estimating the SNRs, this method significantly outperforms a method based on cepstral coefficients. When the method is evaluated on a speech database, the periodicity and aperiodicity accuracy increase significantly. The previous pitch detector was prone to committing pitch doubling and pitch halving errors and was unable to detect pitch reliably in weakly periodic regions. Significant changes have reduced the error rate by 28.7%. The pitch detector is also able to detect accurately the pitch of the synthetic speech-like signals and to capture the jitter present in the signals. Om Deshmukh, Jawahar Singh, Carol Y. Espy-Wilson |
ICASSP (1) | 3 |
| 2004 | Acoustic parameters for automatic detection of nasal manner
Tarun Pruthi, Carol Y. Espy-Wilson |
Speech Commun. | 2 |
| 2003 | A measure of aperiodicity and periodicity in speechabstractIn this paper, we discuss a direct measure for aperiodic energy and periodic energy in speech signals. Most measures for aperiodicity have been indirect, such as zero crossing rate, high-frequency energy and the ratio of high-frequency energy to low-frequency energy. Such indirect measurements will usually fail in situations where there is both strong periodic and aperiodic energy in the speech signal, as in the case of some voiced fricatives or when there is a need to distinguish between high frequency periodic versus high frequency aperiodic energy. We propose an AMDF based temporal method to estimate directly the amount of periodic and aperiodic energy in the speech signal. The algorithm also gives an estimate of the pitch period in periodic regions. Om Deshmukh, Carol Y. Espy-Wilson |
ICASSP (1) | 2 |
| 2003 | A measure of aperiodicity and periodicity in speechabstractIn this paper, we discuss a direct measure for aperiodic energy and periodic energy in speech signals. Most measures for aperiodicity have been indirect, such as zero crossing rate, high- frequency energy and the ratio of high-frequency energy to low- frequency energy. Such indirect measurements will usually fail in situations where there is both strong periodic and aperiodic energy in the speech signal, as in the case of some voiced fricatives or when there is a need to distinguish between high frequency periodic versus high frequency aperiodic energy. We propose an AMDF based temporal method to estimate directly the amount of periodic and aperiodic energy in the speech signal. The algorithm also gives an estimate of the pitch period in periodic regions. Om Deshmukh, Carol Y. Espy-Wilson |
ICME | 2 |
| 2003 | Speech segmentation using probabilistic phonetic feature hierarchy and support vector machinesabstractWe propose a method that combines a probabilistic phonetic feature hierarchy with support vector machines for segmentation of continuous speech into five classes - vowel, sonorant consonant, fricative, stop and silence. We show that by using the hierarchy, only four binary classifiers are required to recognize the five classes. Due to the probabilistic nature of the hierarchy, the method overcomes the disadvantage of the traditional acoustic-phonetic methods where the error is carried down the hierarchy. In addition, the hierarchical approach allows the use of comparable amount of training data of two classes that each binary classifier is designed to discriminate. The segmentation method with 13 knowledge based parameters performs considerably better than a context-dependent hidden Markov model (HMM) based approach that uses 39 mel-cepstrum based parameters. Amit Juneja, Carol Y. Espy-Wilson |
IJCNN | 2 |
| 2003 | Acoustic modeling of american English lateral approximantsabstractA vocal tract model for an American English /l / production with lateral channels and a supralingual side branch has been developed. Acoustic modeling of an /l / production using MRI-derived vocal tract dimensions shows that both the lateral channels and the supralingual side branch contribute to the production of zeros in the F3 to F5 frequency range, thereby resulting in pole-zero clusters around 2-5 kHz in the spectrum of the /l / sound. 1. Carol Y. Espy-Wilson, Mark K. Tiede |
INTERSPEECH | 2 |
| 2002 | Acoustic-phonetic speech parameters for speaker-independent speech recognitionabstractCoping with inter-speaker variability (i.e., differences in the vocal tract characteristics of speakers) is still a major challenge for Automatic Speech Recognizers. In this paper, we discuss a method that compensates for differences in speaker characteristics. In particular, we demonstrate that when continuous density hidden Markov model based system is used as the back-end, a Knowledge-Based Front End (KBFE) can outperform the traditional Mel-Frequency Cepstral Coefficients (MFCCs), particularly when there is a mismatch in the gender and ages of the subjects used to train and test the recognizer. This work was supported by NSF grant # SBR-9729688 and NIH grant # IK02DCOOI49. Om Deshmukh, Carol Y. Espy-Wilson, Amit Juneja |
ICASSP | 2 |
| 2002 | An event-based acoustic-phonetic approach for speech segmentation and E-set recognitionabstractIn this paper, we discuss an event-based recognition system (EBS) which is based on phonetic feature theory and acoustic phonetics. First, acoustic events related to the manner phonetic features are extracted from the speech signal. Second, based on the manner acoustic events, information related to the place phonetic features and voicing are extracted. Most recently, we focused on place and voicing information needed to distinguish among the stop consonants /t,d,p,b/. Using the E-set utterances from the TI46 database, EBS achieved 75.7% overall word accuracy. Further, the knowledge-based acoustic parameters (APs) optimized within the EBS framework were compared to the mel-frequency cepstral coefficients in an HMM-based recognition system. The results on the E-set task showed that the APs achieve a higher recognition accuracy. Amit Juneja, Om Deshmukh, Carol Y. Espy-Wilson |
ICASSP | 3 |
| 2000 | Detection of speech landmarks using temporal cues
Ariel Salomon, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2000 | A new strategy of formant tracking based on dynamic programming
Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 1999 | Improvement of electrolaryngeal speech by introducing normal excitation information
Pelin Demirel, Carol Y. Espy-Wilson, Joel MacAuslan |
EUROSPEECH | 3 |
| 1999 | Automatic detection of manner events based on temporal parameters
Ariel Salomon, Carol Y. Espy-Wilson |
EUROSPEECH | 2 |
| 1997 | The design of acoustic parameters for speaker-independent speech recognition
Nabil N. Bitar, Carol Y. Espy-Wilson |
EUROSPEECH | 2 |
| 1997 | Acoustic modelling of American English /r/
Carol Y. Espy-Wilson, Shri Narayanan, Suzanne Boyce, Abeer Alwan |
EUROSPEECH | 1 |
| 1996 | Knowledge-based parameters for HMM speech recognitionabstractThis paper presents acoustic parameters (APs) that were motivated by phonetic feature theory and employed as a signal representation of speech in a hidden Markov model (HMM) recognition framework. Presently, the phonetic features considered are the manner features: sonorant, syllabic, nonsyllabic, noncontinuant and fricated. The objective of the parameters is to directly target the linguistic information in the signal and to reduce the speaker-dependent information that may yield large speech variability. To achieve these goals, the APs were defined in a relational manner across time or frequency. For evaluation, broad-class recognition experiments were conducted comparing the APs to cepstral-based parameters. The results of the experiments indicate that the APs are able to capture the phonetically relevant information in the speech signal and that, in comparison to the cepstral-based parameters, they are more able to reduce the interspeaker variability. Nabil N. Bitar, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 1996 | Coarticulatory stability in american English /r/
Suzanne Boyce, Carol Y. Espy-Wilson |
ICSLP | 2 |
| 1996 | Enhancement of alaryngeal speech by adaptive filteringabstractArticial larynxes enable adequate communication for people who are unable to use their larynxes.However, the resulting speech has an unnatural quality and is signicantly less intelligible than normal speech.One of the major problems with the widely-used Transcutaneous Articial Larynx (TAL) is the presence of a steady background noise due to the leakage of acoustic energy.In the present study, a novel adaptive ltering architecture was designed and implemented for the purpose of removing the background noise.Perceptual tests were conducted to assess speech from 2 laryngectomees and 2 normal speakers using the Servox T AL, before and after processing by the adaptive lter.Results from the perceptual tests indicate a clear preference for the processed speech and spectral analysis of the reveals a signicant reduction in the background source radiation. Carol Y. Espy-Wilson, Venkatesh R. Chari, Caroline B. Huang |
ICSLP | 1 |
| 1995 | Speech parameterization based on phonetic features: application to speech recognition
Nabil N. Bitar, Carol Y. Espy-Wilson |
EUROSPEECH | 2 |
| 1995 | Adaptive enhancement of Fourier spectraabstractAn adaptive enhancement procedure is presented which emphasizes continuant spectral features such as formant frequencies, by imposing frequency and amplitude continuity constraints on a short-time Fourier representation of the speech signal. At each point in the time-frequency field, the direction of maximum energy correlation is determined by the angle of a linear window at which the energy density within it is closest in magnitude to the point under consideration. Weighted smoothing is then performed in that direction to enhance continuant features.> Venkatesh R. Chari, Carol Y. Espy-Wilson |
IEEE Trans. Speech Audio Process. | 2 |
| 1986 | A phonetically based semivowel recognition systemabstractA phonetically based approach to speech recognition uses speech specific knowledge obtained from phonotactics, phonology and acoustic phonetics to capture relevant phonetic information. Thus, a recognition system based on this approach can make broad classifications as well as detailed phonetic distinctions. This paper discusses a framework for developing a phonetically based recognition system. The recognition task is the class of sounds known as the semivowels. The recognition results reported, though incomplete, are encouraging. Carol Y. Espy-Wilson |
ICASSP | 1 |
| 1982 | Effects of noise on signal reconstruction from Fourier transform phaseabstractThe effects of noise in the given phase on signal reconstruction from the Fourier transform phase are studied. Specifically, the effects of different methods of sampling the degraded phase, of the number of non-zero points in the sequence, and of the noise level on the sequence reconstruction are examined. A sampling method is developed to significantly reduce the error in the reconstructed sequence, and the error is found to increase as the number of non-zero points in the sequence increases and as the noise level increases. In addition, an averaging technique is developed which reduces the effects of noise when the continuous phase function is known. Finally, as an illustration of how the results in this paper may be applied in practice, Fourier transform signal coding is considered. Coding only the Fourier transform phase and reconstructing the signal from the coded phase is found to be considerably less efficient (i.e. a higher bit rate is required for the same mean square error) than reconstructing from both the coded phase and magnitude. Carol Y. Espy-Wilson, Jae S. Lim |
ICASSP | 1 |