EDBT 2026 Demo / reviewers in the wild / expert
Jordan R. Green
dblp:68/10650
· DBLP profile ↗
36ranked-venue papers
2as first author
13since 2021 · last 2024
0000-0002-1464-1373ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Large Language Models As A Proxy For Human Evaluation In Assessing The Comprehensibility Of Disordered Speech TranscriptionabstractAutomatic Speech Recognition (ASR) systems, despite significant advances in recent years, still have much room for improvement particularly in the recognition of disordered speech. Even so, erroneous transcripts from ASR models can help people with disordered speech be better understood, especially if the transcription doesn’t significantly change the intended meaning. Evaluating the efficacy of ASR for this use case requires a methodology for measuring the impact of transcription errors on the intended meaning and comprehensibility. Human evaluation is the gold standard for this, but it can be laborious, slow, and expensive. In this work, we tune and evaluate large language models for this task and find them to be a much better proxy for human evaluators than other metrics commonly used. We further present a case-study using the presented approach to assess the quality of personalized ASR models to make model deployment decisions and correctly set user expectations for model quality as part of our trusted tester program. Katrin Tomanek, Jimmy Tobin, Subhashini Venugopalan, Richard Cave, Katie Seaver, Jordan R. Green, Rus Heywood |
ICASSP | 6 |
| 2024 | Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speechabstractProject Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report describes the project's latest advancements in data collection and annotation methodologies, such as expanding speaker diversity in the database, adding human-reviewed transcript corrections and audio quality tags to 350K (of the 1.2M total) audio recordings, and amassing a comprehensive set of metadata (including more than 40 speech characteristic labels) for over 75% of the speakers in the database. We report on the impact of transcript corrections on our machine-learning (ML) research, inter-rater variability of assessments of disordered speech patterns, and our rationale for gathering speech metadata. We also consider the limitations of using automated off-the-shelf annotation methods for assessing disordered speech. Pan-Pan Jiang, Jimmy Tobin, Katrin Tomanek, Robert L. MacDonald, Katie Seaver, Richard Cave, Marilyn A. Ladewig, Rus Heywood, Jordan R. Green |
INTERSPEECH | 9 |
| 2023 | An Analysis of Degenerating Speech Due to Progressive Dysarthria on ASR PerformanceabstractAlthough personalized automatic speech recognition (ASR) models have recently been improved to recognize even severely impaired speech, model performance may degrade over time for persons with degenerating speech. The aims of this study were to (1) analyze the change of performance of ASR over time in individuals with degrading speech, and (2) explore mitigation strategies to optimize recognition throughout disease progression. Speech was recorded by four individuals with degrading speech due to amyotrophic lateral sclerosis (ALS). Word error rates (WER) across recording sessions were computed for three ASR models: Unadapted Speaker Independent (U-SI), Adapted Speaker Independent (A-SI), and Adapted Speaker Dependent (A-SD or personalized). The performance of all models degraded significantly over time as speech became more impaired, but the A-SD model improved markedly when updated with recordings from the severe stages of speech progression. Recording additional utterances early in the disease before significant speech degradation did not improve the performance of A-SD models. This emphasizes the importance of continuous recording (and model retraining) when providing personalized models for individuals with progressive speech impairments. Katrin Tomanek, Katie Seaver, Pan-Pan Jiang, Richard Cave, Lauren Harrell, Jordan R. Green |
ICASSP | 6 |
| 2023 | Speech Intelligibility Classifiers from 550k Disordered Speech SamplesabstractWe developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following different deep learning approaches and evaluated them on ~ 94K utterances from 100 speakers. We further found the models to generalize well (without further training) on the TORGO database[1] (100% accuracy), UASpeech[2] (0.93 correlation), ALS-TDI PMP[3] (0.81 AUC) datasets as well as on a dataset of realistic unprompted speech we gathered (106 dysarthric and 76 control speakers, ~ 2300 samples). We share our model1to advance research in this domain. Subhashini Venugopalan, Jimmy Tobin, Samuel J. Yang, Katie Seaver, Richard Cave, Pan-Pan Jiang, Neil Zeghidour, Rus Heywood, Jordan R. Green, Michael P. Brenner |
ICASSP | 9 |
| 2023 | Responsiveness, Sensitivity and Clinical Utility of Timing-Related Speech Biomarkers for Remote Monitoring of ALS Disease Progressionabstract= 94). We further evaluated the sensitivity of speech metrics in tracking disease progression in pALS while their ALSFRS-R speech score remained unchanged at 3 out of a total possible score of 4. We observed that timing-related speech metrics showed significant longitudinal changes even after accounting for learning effects. The findings of this study have the potential to inform disease prognosis and functional outcomes of clinical trials. Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Jordan R. Green, Vikram Ramanarayanan |
INTERSPEECH | 4 |
| 2023 | Validation of a Task-Independent Cepstral Peak Prominence Measure with Voice Activity Detection
Olivia M. Murton, Abigail E. Haenssler, Marc F. Maffei, Kathryn P. Connaghan, Jordan R. Green |
INTERSPEECH | 5 |
| 2023 | Remote Assessment for ALS using Multimodal Dialog Agents: Data Quality, Feasibility and Task ComplianceabstractWe investigate the feasibility, task compliance and audiovisual data quality of a multimodal dialog-based solution for remote assessment of Amyotrophic Lateral Sclerosis (ALS). 53 people with ALS and 52 healthy controls interacted with Tina, a cloud-based conversational agent, in performing speech tasks designed to probe various aspects of motor speech function while their audio and video was recorded. We rated a total of 250 recordings for audio/video quality and participant task compliance, along with the relative frequency of different issues observed. We observed excellent compliance (98%) and audio (95.2%) and visual quality rates (84.8%), resulting in an overall yield of 80.8% recordings that were both compliant and of high quality. Furthermore, recording quality and compliance were not affected by level of speech severity and did not differ significantly across end devices. These findings support the utility of dialog systems for remote monitoring of speech in ALS. Vanessa Richter, Michael Neumann 0001, Jordan R. Green, Brian Richburg, Oliver Roesler, Hardik Kothare, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2021 | A Voice-Activated Switch for Persons with Motor and Speech Impairments: Isolated-Vowel Spotting Using Neural Networks
Shanqing Cai, Lisie Lillianfeld, Katie Seaver, Jordan R. Green, Michael P. Brenner, Philip C. Nelson, D. Sculley |
Interspeech | 4 |
| 2021 | Automatic Speech Recognition of Disordered Speech: Personalized Models Outperforming Human Listeners on Short Phrases
Jordan R. Green, Robert L. MacDonald, Pan-Pan Jiang, Julie Cattiau, Rus Heywood, Richard Cave, Katie Seaver, Marilyn A. Ladewig, Jimmy Tobin, Michael P. Brenner, Philip C. Nelson, Katrin Tomanek |
Interspeech | 1 |
| 2021 | Speaking with a KN95 Face Mask: ASR Performance and Speaker CompensationabstractThe increasing prevalence of face masks in the United States due to the COVID-19 pandemic necessitates serious consideration of the functional impact of wearing a mask on speech. This study considers how the presence of a KN95 mask affects the performance of a commercial ASR system, Google Cloud Speech. We present evidence that wearing a mask does not impact ASR performance at the sentence level. Moreover, speakers may be naturally adapting to the mask by increasing their vowel space area. However, when speakers intentionally altered their speech by speaking clearly or loudly (though not slowly), ASR performance improved. These findings suggest that ASR users can employ speech strategies to achieve better ASR results when wearing a mask. Beyond healthy speakers, our study has implications for mask-wearing ASR users with otherwise reduced speech intelligibility. Copyright © 2021 ISCA. Sarah E. Gutz, Hannah P. Rowe, Jordan R. Green |
Interspeech | 3 |
| 2021 | Disordered Speech Data Collection: Lessons Learned at 1 Million Utterances from Project Euphonia
Robert L. MacDonald, Pan-Pan Jiang, Julie Cattiau, Rus Heywood, Richard Cave, Katie Seaver, Marilyn A. Ladewig, Jimmy Tobin, Michael P. Brenner, Philip C. Nelson, Jordan R. Green, Katrin Tomanek |
Interspeech | 11 |
| 2021 | Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at ScaleabstractWe propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS. Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan |
Interspeech | 16 |
| 2021 | Comparing Supervised Models and Learned Speech Representations for Classifying Intelligibility of Disordered Speech on Selected PhrasesabstractAutomatic classification of disordered speech can provide an objective tool for identifying the presence and severity of speech impairment. Classification approaches can also help identify hard-to-recognize speech samples to teach ASR systems about the variable manifestations of impaired speech. Here, we develop and compare different deep learning techniques to classify the intelligibility of disordered speech on selected phrases. We collected samples from a diverse set of 661 speakers with a variety of self-reported disorders speaking 29 words or phrases, which were rated by speech-language pathologists for their overall intelligibility using a five-point Likert scale. We then evaluated classifiers developed using 3 approaches: (1) a convolutional neural network (CNN) trained for the task, (2) classifiers trained on non-semantic speech representations from CNNs that used an unsupervised objective [1], and (3) classifiers trained on the acoustic (encoder) embeddings from an ASR system trained on typical speech [2]. We found that the ASR encoder's embeddings considerably outperform the other two on detecting and classifying disordered speech. Further analysis shows that the ASR embeddings cluster speech by the spoken phrase, while the non-semantic embeddings cluster speech by speaker. Also, longer phrases are more indicative of intelligibility deficits than single words. Subhashini Venugopalan, Joel Shor, Manoj Plakal, Jimmy Tobin, Katrin Tomanek, Jordan R. Green, Michael P. Brenner |
Interspeech | 6 |
| 2020 | Acoustic-Based Articulatory Phenotypes of Amyotrophic Lateral Sclerosis and Parkinson's Disease: Towards an Interpretable, Hypothesis-Driven Framework of Motor Control
Hannah P. Rowe, Sarah E. Gutz, Marc F. Maffei, Jordan R. Green |
INTERSPEECH | 4 |
| 2019 | Use of Beiwe Smartphone App to Identify and Track Speech Decline in Amyotrophic Lateral Sclerosis (ALS)
Kathryn P. Connaghan, Jordan R. Green, Sabrina Paganoni, James Chan, Harli Weber, Ella Collins, Brian Richburg, Marziye Eshghi, Jukka-Pekka Onnela, James D. Berry |
INTERSPEECH | 2 |
| 2019 | Reduced Task Adaptation in Alternating Motion Rate Tasks as an Early Marker of Bulbar Involvement in Amyotrophic Lateral Sclerosis
Marziye Eshghi, Panying Rong, Antje S. Mefferd, Kaila L. Stipancic, Yana Yunusova, Jordan R. Green |
INTERSPEECH | 6 |
| 2019 | Early Identification of Speech Changes Due to Amyotrophic Lateral Sclerosis Using Machine Classification
Sarah E. Gutz, Jun Wang 0037, Yana Yunusova, Jordan R. Green |
INTERSPEECH | 4 |
| 2019 | Vocal Biomarker Assessment Following Pediatric Traumatic Brain Injury: A Retrospective Cohort Study
Camille Noufi, Adam C. Lammert, Daryush D. Mehta, James R. Williamson, Gregory A. Ciccarelli, Douglas E. Sturim, Jordan R. Green, Thomas F. Campbell, Thomas F. Quatieri |
INTERSPEECH | 7 |
| 2019 | Profiling Speech Motor Impairments in Persons with Amyotrophic Lateral Sclerosis: An Acoustic-Based Approach
Hannah P. Rowe, Jordan R. Green |
INTERSPEECH | 2 |
| 2019 | A Sparse Non-Negative Matrix Factorization Framework for Identifying Functional Units of Tongue Behavior From MRIabstractMuscle coordination patterns of lingual behaviors are synergies generated by deforming local muscle groups in a variety of ways. Functional units are functional muscle groups of local structural elements within the tongue that compress, expand, and move in a cohesive and consistent manner. Identifying the functional units using tagged-magnetic resonance imaging (MRI) sheds light on the mechanisms of normal and pathological muscle coordination patterns, yielding improvement in surgical planning, treatment, or rehabilitation procedures. In this paper, to mine this information, we propose a matrix factorization and probabilistic graphical model framework to produce building blocks and their associated weighting map using motion quantities extracted from tagged-MRI. Our tagged-MRI imaging and accurate voxel-level tracking provide previously unavailable internal tongue motion patterns, thus revealing the inner workings of the tongue during speech or other lingual behaviors. We then employ spectral clustering on the weighting map to identify the cohesive regions defined by the tongue motion that may involve multiple or undocumented regions. To evaluate our method, we perform a series of experiments. We first use two-dimensional images and synthetic data to demonstrate the accuracy of our method. We then use three-dimensional synthetic and in vivo tongue motion data using protrusion and simple speech tasks to identify subject-specific and data-driven functional units of the tongue in localized regions. Jonghye Woo, Jerry L. Prince, Maureen Stone 0001, Fangxu Xing, Arnold D. Gomez, Jordan R. Green, Christopher J. Hartnick, Thomas J. Brady, Timothy G. Reese, Van J. Wedeen, Georges El Fakhri |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Automatic Detection of Amyotrophic Lateral Sclerosis (ALS) from Video-Based Analysis of Facial Movements: Speech and Non-Speech TasksabstractThe analysis of facial movements in patients with amyotrophic lateral sclerosis (ALS) can provide important information about early diagnosis and tracking disease progression. However, the use of expensive motion tracking systems has limited the clinical utility of the assessment. In this study, we propose a marker-less video-based approach to discriminate patients with ALS from neurotypical subjects. Facial movements were recorded using a depth sensor (Intel® RealSense" SR300) during speech and nonspeech tasks. A small set of kinematic features of lips was extracted in order to mirror the perceptual evaluation performed by clinicians, considering the following aspects: (1) range of motion, (2) speed of motion, (3) symmetry, and (4) shape. Our results demonstrate that it is possible to distinguish patients with ALS from neurotypical subjects with high overall accuracy (up to 88.9%) during repetitions of sentences, syllables, and labial non-speech movements (e.g., lip spreading). This paper provides strong rationale for the development of automated systems to detect neurological diseases from facial movements. This work has a high social impact, as it opens new possibilities to develop intelligent systems to support clinicians in their diagnosis, introducing novel standards for assessing the oro-facial impairment in ALS, and tracking disease progression remotely from home. Andrea Bandini, Jordan R. Green, Babak Taati, Silvia Orlandi, Lorne Zinman, Yana Yunusova |
FG | 2 |
| 2018 | Automatic Early Detection of Amyotrophic Lateral Sclerosis from Intelligible Speech Using Convolutional Neural Networks
Kwanghoon An, Myung Jong Kim, Kristin Teplansky, Jordan R. Green, Thomas F. Campbell, Yana Yunusova, Daragh Heitzman, Jun Wang 0037 |
INTERSPEECH | 4 |
| 2018 | Automatic Detection of Orofacial Impairment in Stroke
Andrea Bandini, Jordan R. Green, Brian Richburg, Yana Yunusova |
INTERSPEECH | 2 |
| 2017 | Classification of Bulbar ALS from Kinematic Features of the Jaw and Lips: Towards Computer-Mediated Assessment
Andrea Bandini, Jordan R. Green, Lorne Zinman, Yana Yunusova |
INTERSPEECH | 2 |
| 2016 | Towards an Automated Screening Tool for Developmental Speech and Language Impairments
Jen J. Gong, Maryann Gong, Dina Levy-Lambert, Jordan R. Green, Tiffany P. Hogan, John V. Guttag |
INTERSPEECH | 4 |
| 2016 | Relation of Automatically Extracted Formant Trajectories with Intelligibility Loss and Speaking Rate Decline in Amyotrophic Lateral Sclerosis
Rachelle L. Horwitz-Martin, Thomas F. Quatieri, Adam C. Lammert, James R. Williamson, Yana Yunusova, Elizabeth Godoy, Daryush D. Mehta, Jordan R. Green |
INTERSPEECH | 8 |
| 2016 | Differential Effects of Velopharyngeal Dysfunction on Speech Intelligibility During Early and Late Stages of Amyotrophic Lateral Sclerosis
Panying Rong, Yana Yunusova, Jordan R. Green |
INTERSPEECH | 3 |
| 2015 | Speech intelligibility decline in individuals with fast and slow rates of ALS progression
Panying Rong, Yana Yunusova, Jordan R. Green |
INTERSPEECH | 3 |
| 2014 | Parameterization of articulatory pattern in speakers with ALSabstractA combination of parallel factor analysis (PARAFAC) and principal component analysis (PCA) was used to parameterize the articulatory pattern of tongue, jaw and lip movements in 8 English vowels produced by 7 subjects with amyotrophic lateral sclerosis (ALS). A two-factor PARAFAC model derived an overall articulatory pattern represented by two basic modes dominated by tongue raising and advancement, respectively. The relation between the two articulatory modes and the acoustic formants (F1, F2) followed a simple one-to-one linear mapping. The PCA on the residuals of the PARAFAC model showed various individualized articulatory features superimposed on the overall pattern. These articulatory features contributed in a systematic way to the acoustic deviation across different subjects. The parameterization approach (1) provided a simple and generalizable way to explore the underlying articulatory mechanism of speech decline in ALS and (2) accounted for the articulatory features across affected individuals. With further development of the approach and a comparison with the articulatory pattern for healthy subjects, it is possible to derive a set of quantitative articulatory indicators of speech impairment in ALS. Index Terms: parameterization, articulatory-acoustic mapping, ALS Panying Rong, Yana Yunusova, James D. Berry, Lorne Zinman, Jordan R. Green |
INTERSPEECH | 5 |
| 2014 | Across-speaker articulatory normalization for speaker-independent silent speech recognitionabstractSilent speech interfaces (SSIs), which recognize speech from articulatory information (i.e., without using audio information), have the potential to enable persons with laryngectomy or a neurological disease to produce synthesized speech with a natural sounding voice using their tongue and lips. Current approaches to SSIs have largely relied on speaker-dependent recognition models to minimize the negative effects of talker variation on recognition accuracy. Speaker-independent approaches are needed to reduce the large amount of training data required from each user; only limited articulatory samples are often available for persons with moderate to severe speech impairments, due to the logistic difficulty of data collection. This paper reported an across-speaker articulatory normalization approach based on Procrustes matching, a bidimensional regression technique for removing translational, scaling, and rotational effects of spatial data. A dataset of short functional sentences was collected from seven English talkers. A support vector machine was then trained to classify sentences based on normalized tongue and lip movements. Speaker-independent classification accuracy (tested using leave-one-subject-out cross validation) improved significantly, from 68.63 % to 95.90%, following normalization. These results support the feasibility of a speaker-independent SSI using Procrustes matching as the basis for articulatory normalization across speakers. Index Terms: silent speech recognition, speech kinematics, Procrustes analysis, support vector machine Jun Wang 0037, Ashok Samal, Jordan R. Green |
INTERSPEECH | 3 |
| 2013 | Individual articulator's contribution to phoneme productionabstractSpeech sounds are the result of coordinated movements of individual articulators. Understanding each articulator's role in speech is fundamental not only for understanding how speech is produced, but also for optimizing speech assessments and treatments. In this paper, we studied the individual contributions of six articulators, tongue tip, tongue blade, tongue body front, tongue body back, upper lip, and lower lip to phoneme classification. A total of 3,838 vowel and consonant production samples were collected from eleven native English speakers. The results of speech movement classification using a support vector machine indicated that the tongue encoded significantly more information than lips, and that the tongue tip may be the most important single articulator among all of the six for phoneme production. Furthermore, our results suggested that the tracking of four articulators (i.e., tongue tip, tongue body back, upper lip, and lower lip) may be sufficient for distinguishing major English phonemes based on articulatory movements. Jun Wang 0037, Jordan R. Green, Ashok Samal |
ICASSP | 2 |
| 2013 | SMASH: a tool for articulatory data processing and analysisabstractRecent innovations in 3D motion capture technology such as electromagnetic articulography (EMA) are providing unprecedented access to the intricate movements of the articulators during speech production. Although these technological advances afford exciting opportunities for advancing the assessment and treatment of speech, they have presented new challenges associated with data collection, processing, and analysis. To address these challenges, we have standardized our EMA data collection protocols and developed a Matlab-based software tool, SMASH, for processing, visualizing, and analyzing speech movement data. The goal of the software is to advance research on speech production by improving the efficiency and reliability of speech movement analyses. Jordan R. Green, Jun Wang 0037, David L. Wilson |
INTERSPEECH | 1 |
| 2012 | Sentence recognition from articulatory movements for silent speech interfacesabstractRecent research has demonstrated the potential of using an articulation-based silent speech interface for command-and-control systems. Such an interface converts articulation to words that can then drive a text-to-speech synthesizer. In this paper, we have proposed a novel near-time algorithm to recognize whole-sentences from continuous tongue and lip movements. Our goal is to assist persons who are aphonic or have a severe motor speech impairment to produce functional speech using their tongue and lips. Our algorithm was tested using a functional sentence data set collected from ten speakers (3012 utterances). The average accuracy was 94.89% with an average latency of 3.11 seconds for each sentence prediction. The results indicate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for clinical applications. Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz |
ICASSP | 3 |
| 2012 | Whole-Word Recognition from Articulatory Movements for Silent Speech InterfacesabstractArticulation-based silent speech interfaces convert silently produced speech movements into audible words. These systems are still in their experimental stages, but have significant potential for facilitating oral communication in persons with laryngectomy or speech impairments. In this paper, we report the result of a novel, real-time algorithm that recognizes whole-words based on articulatory movements. This approach differs from prior work that has focused primarily on phoneme-level recognition based on articulatory features. On average, our algorithm missed 1.93 words in a sequence of twenty-five words with an average latency of 0.79 seconds for each word prediction using a data set of 5,500 isolated word samples collected from ten speakers. The results demonstrate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for health applications. Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz |
INTERSPEECH | 3 |
| 2011 | Quantifying Articulatory Distinctiveness of VowelsabstractThe articulatory distinctiveness among vowels has been frequently characterized descriptively based on tongue height and front-back position; however, very few empirical methods have been proposed to characterize vowels based on time-varying articulatory characteristics. Such information is not only needed to improve knowledge about the articulation of vowels but also to determine the contribution of articulatory imprecision to poor speech intelligibility. In this paper, a novel statistical shape analysis was used to derive a vowel space that depicted the quantified articulatory distinctiveness among vowels based on tongue and lip movements. The effectiveness of the approach was supported by vowel classification accuracy of up to 91.7%. The theoretical relevance and clinical implication of the derived vowel space were discussed. Index Terms: speech production, articulatory vowel space, Procrustes analysis, multi-dimensional scaling Jun Wang 0037, Jordan R. Green, Ashok Samal, David Marx |
INTERSPEECH | 2 |
| 1994 | Relationship between acoustic measures of vocal perturbation and perceptual judgments of breathiness, harshness, and hoarseness
Fred D. Minifie, Daniel Z. Huang, Jordan R. Green |
ICSLP | 3 |