VLDB 2026 Research / reviewers in the wild / expert
Louis Goldstein
dblp:87/5550 · also Louis M. Goldstein
· DBLP profile ↗
54ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 12 since 2021Artificial intelligence and machine learning · 47 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model
Xuan Shi, Tiantian Feng, Jay Park, Christina Hagedorn, Louis Goldstein, Shri Narayanan |
Comput. Speech Lang. | 5 |
| 2025 | Instantaneous changes in acoustic signals reflect syllable progression and cross-linguistic syllable variation
Haley Hsu, Dani Byrd, Khalil Iskarous, Louis Goldstein |
INTERSPEECH | 4 |
| 2025 | Articulatory Feature Prediction from Surface EMG during Speech Production
Kleanthis Avramidis, Simon Pistrosch, Monica González Machorro, Yoonjeong Lee, Björn W. Schuller, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 8 |
| 2025 | Towards a dynamical model of transitions between fluent and stuttered speech
Yijing Lu, Khalil Iskarous, Louis Goldstein |
INTERSPEECH | 3 |
| 2025 | 75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment
Xuan Shi, Yubin Zhang, Yijing Lu, Marcus Ma, Tiantian Feng, Asterios Toutios, Haley Hsu, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 8 |
| 2025 | Co-registration of real-time MRI and respiration for speech research
Yubin Zhang, Prakash Kumar, Xuan Shi, Haley Hsu, Shri Narayanan, Krishna S. Nayak, Louis Goldstein |
INTERSPEECH | 11 |
| 2024 | Analysis of articulatory setting for L1 and L2 English speakers using MRI dataabstractThis paper investigates the extent to which the geographical region (country) where a speaker acquired their English language affects the articulatory setting in their speech.To obtain accurate measurements for evaluating articulatory setting, we utilized a large real-time MRI corpus of vocal tract articulation.The corpus was obtained from speakers from a variety of linguistic backgrounds producing continuous English speech.We use an automated pipeline to process and extract articulatory positional information from the MRI video data.This data is used to draw comparisons between English language speakers from the United States and speakers who acquired their English in India, Korea, and China.Analysis of the speaker groups reveals statistically significant articulatory setting posture differences in multiple places of articulation. Jack Goldberg, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2023 | Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix FactorizationabstractArticulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with human speech production mechanism, and corresponding gestural scores. We continue with this line of work by raising two concerns: (1) The articulators are entangled together in the original algorithm such that some of the articulators do not leverage effective moving patterns, which limits the interpretability of both gestures and gestural scores; (2) The EMA data is sparsely sampled from articulators, which limits the intelligibility of learned representations. In this work, we propose a novel articulatory representation decomposition algorithm that takes the advantage of guided factor analysis to derive the articulatory-specific factors and factor scores. A neural convolutive matrix factorization algorithm is then employed on the factor scores to derive the new gestures and gestural scores. We experiment with the rtMRI corpus that captures the fine-grained vocal tract contours. Both subjective and objective evaluation results suggest that the newly proposed system delivers the articulatory representations that are intelligible, generalizable, efficient and interpretable. Jiachen Lian, Alan W. Black, Yijing Lu, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli |
ICASSP | 4 |
| 2023 | Speaker-Independent Acoustic-to-Articulatory Speech InversionabstractTo build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech production. To this end, we build an acoustic-to-articulatory inversion (AAI) model that leverages autoregression, adversarial training, and self supervision to generalize to unseen speakers. Our approach obtains 0.784 correlation on an electromagnetic articulography (EMA) dataset, improving the state-of-the-art by 12.5%. Additionally, we show the interpretability of these representations through directly com-paring the behavior of estimated representations with speech production behavior. Finally, we propose a resynthesis-based AAI evaluation metric that does not rely on articulatory labels, demonstrating its efficacy with an 18-speaker dataset. Peter Wu, Cheol Jun Cho, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli |
ICASSP | 5 |
| 2023 | Deep Speech Synthesis from MRI-Based Articulatory Representations
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W. Black, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli |
INTERSPEECH | 7 |
| 2022 | Deep Neural Convolutive Matrix Factorization for Articulatory Representation DecompositionabstractMost of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure.This work, investigating the speech representations derived from articulatory kinematics signals, uses a neural implementation of convolutive sparse matrix factorization to decompose the articulatory data into interpretable gestures and gestural scores.By applying sparse constraints, the gestural scores leverage the discrete combinatorial properties of phonological gestures.Phoneme recognition experiments were additionally performed to show that gestural scores indeed code phonological information successfully.The proposed work thus makes a bridge between articulatory phonology and deep neural networks to leverage informative, intelligible, interpretable,and efficient speech representations. Jiachen Lian, Alan W. Black, Louis Goldstein, Gopala Krishna Anumanchipalli |
INTERSPEECH | 3 |
| 2022 | Deep Speech Synthesis from Articulatory Representations
Peter Wu, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli |
INTERSPEECH | 3 |
| 2021 | Who converges? Variation reveals individual speaker adaptabilityabstractLittle is known about the cognitive capacities underlying real-time accommodation in spoken language and how they may allow conversing speakers to adapt their speech production behaviors. This study first presents a simple attunement model that incorporates hypothesized capacities, with a focus on individual variability as one of those capacities. The model makes explicit predictions about observable convergence behaviors in interacting speakers, including that: i) the intrinsically more variable speaker of the two will be the one who converges to their partner, ii) this flexible speaker with higher baseline variability will exhibit a substantial decrease in variability and iii) a greater change in the variability between speaking solo and interacting with their partner. These predictions are supported by the results of the modeling simulations. To further test the model's predictions, we analyzed a behavioral dataset including acoustic and articulatory data from three pairs of interacting speakers participating in a maze navigation task as well as a like solo speech task. The amount of variability in the speech parameters of each dyad member was quantified using coefficient of variation. The experimental results parallel the simulation results, and taken together, this work indicates that structured variability is an illuminating index of individual speaker adaptability and convergence behavior. Yoon-Jeong Lee, Louis Goldstein, Benjamin Parrell, Dani Byrd |
Speech Commun. | 2 |
| 2018 | Analysis of speech production real-time MRI
Vikram Ramanarayanan, Sam Tilsen, Michael I. Proctor, Johannes Töger, Louis Goldstein, Krishna S. Nayak, Shri Narayanan |
Comput. Speech Lang. | 5 |
| 2017 | Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 9 |
| 2017 | Test-Retest Repeatability of Articulatory Strategies Using Real-Time Magnetic Resonance Imaging
Tanner Sorensen, Asterios Toutios, Johannes Töger, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 4 |
| 2016 | Velum Control for Oral Sounds
Reed Blaylock, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 2 |
| 2016 | L2 Acquisition and Production of the English Rhotic Pharyngeal Gesture
Sarah Harper, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 2 |
| 2016 | Perceptual Lateralization of Coda Rhotic Production in Puerto Rican Spanish
Mairym Lloréns Monteserín, Shri Narayanan, Louis Goldstein |
INTERSPEECH | 3 |
| 2016 | A New Model of Speech Motor Control Based on Task Dynamics and State Feedback
Vikram Ramanarayanan, Benjamin Parrell, Louis Goldstein, Srikantan S. Nagarajan, John F. Houde |
INTERSPEECH | 3 |
| 2016 | Characterizing Vocal Tract Dynamics Across Speakers Using Real-Time MRI
Tanner Sorensen, Asterios Toutios, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2016 | Illustrating the Production of the International Phonetic Alphabet Sounds Using Fast Real-Time Magnetic Resonance Imaging
Asterios Toutios, Sajan Goud Lingala, Colin Vaz, Jangwon Kim, John H. Esling, Patricia A. Keating, Matthew Gordon, Dani Byrd, Louis Goldstein, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 9 |
| 2015 | Analysis of coarticulated speech using estimated articulatory trajectories
Ganesh Sivaraman, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson |
INTERSPEECH | 5 |
| 2015 | Experimental assessment of the tongue incompressibility hypothesis during speech productionabstractThe human tongue is an important organ for speech production. Its deformation and motion control the shape of the vocal tract significantly and thereby the acoustic properties of the speech signal produced. Thus, much effort in the speech research com-munity has been directed towards its biomechanical modeling. A common assumption incorporated into many models of the human tongue is the tissue incompressibility hypothesis: the tongue is considered a muscular hydrostat and therefore its vol-ume should remain constant regardless of its posture. To the best of our knowledge, experimental assessment of the constant volume hypothesis during actual speech production is limited. In this work, the aim is to experimentally assess the incom-pressibility hypothesis during actual speech production using a dataset of volumetric Magnetic Resonance (MR) images of 17 subjects sustaining contextualized continuants (27 continu-ants per subject). A seeded region growing based algorithm is used to segment the tongue and calculate its volume. Then the intra-subject variability of the tongue volume along the differ-ent tongue postures is examined. Within the accuracy of our tongue volume measurements, our empirical results seem con-sistent with the incompressibility hypothesis. Index Terms: speech production, tongue volume, muscular hy-drostat, tissue incompressibility, volumetric MRI Z.-I. Skordilis, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2014 | A real-time MRI study of articulatory setting in second language speechabstractPrevious work has shown that languages differ in their articulatory setting, the postural configuration that the vocal tract articulators tend to adopt when they are not engaged in any active speech gesture, and that this posture might be specified as part of the phonological knowledge speakers have of the language. This study tests whether the articulatory setting of a language can be acquired by non-native speakers. Three native speakers of German who had learned English as a second language were imaged using real-time MRI of the vocal tract while reading passages in German and English, and features that capture vocal tract posture were extracted from the inter-speech pauses in their native and non-native languages. Results show that the speakers exhibit distinct inter-speech postures in each language, with a lower and more retracted tongue in English, consistent with classic descriptions of the differences between the German and the English articulatory settings. This supports the view that non-native speakers may acquire relevant features of the articulatory setting of a second language, and also lends further support to the idea that articulatory setting is part of a speaker’s phonological competence in a language. Index Terms: articulatory setting, speech production, second language speech acquisition, real-time MRI. Andrés Benítez, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2014 | Motor control primitives arising from a learned dynamical systems model of speech articulationabstractWe present a method to derive a small number of speech motor control “primitives” that can produce linguisticallyinterpretable articulatory movements. We envision that such a dictionary of primitives can be useful for speech motor control, particularly in finding a low-dimensional subspace for such control. First, we use the iterative Linear Quadratic Gaussian with Learned Dynamics (iLQG-LD) algorithm to derive (for a set of utterances) a set of stochastically optimal control inputs to a learned dynamical systems model of the vocal tract that produces desired movement sequences. Second, we use a convolutive Nonnegative Matrix Factorization with sparseness constraints (cNMFsc) algorithm to find a small dictionary of control input primitives that can be used to reproduce the aforementioned optimal control inputs that produce the observed articulatory movements. The method performs favorably on both qualitative and quantitative evaluations conducted on synthetic data produced by an articulatory synthesizer. Such a primitivesbased framework could help inform theories of speech motor control and coordination. Index Terms: speech motor control, motor primitives, synergies, dynamical systems, iLQG, NMF. Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Truncation of pharyngeal gesture in English diphthong [aɪ]abstractIt is well acknowledged that [a] in English diphthongs (e.g. [a] in “pie’d”) has a different formant structure from its closest corresponding monophthong (e.g. [a] in “pod”). The current study proposes that these two sounds share the same cognitive unit, i.e. the pharyngeal constriction gesture that produces [a], and the surface difference can be modeled as a consequence of truncating the same articulatory movement in time by the following palatal glide in the diphthongal environment. Formation of pharyngeal constriction gesture during the production of [a] in a diphthong and in its corresponding monophthong was observed in various timing contexts using Realtime MRI; and the collected production data were quantitatively analyzed using the direct image analysis (DIA) technique, which infers tissue movement by tracking pixel intensity change over time in regions of interest. Results support our truncation account in that: (1) formation time of pharyngeal constriction is significantly longer in monophthongs than in diphthongs; (2) this duration correlates with the resulting constriction degree; and (3) the resulting constriction degree predicts the acoustic difference in the F2 dimension as predicted by our hypothesis. Index Terms: English diphthongs, speech production, rtMRI Fang-Ying Hsieh, Louis Goldstein, Dani Byrd, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Velic coordination in French nasals: a real-time magnetic resonance imaging studyabstractProduction of nasal vowels in French, and nasal consonants in French and English, was examined using real-time magnetic resonance imaging (rtMRI). The coordination of velic and lin-gual gestures was found to be tightly controlled across differ-ent prosodic contexts in French nasals. Velum lowering in En-glish nasal consonants did not show the same control, although the timing of the corresponding lingual gestures varied with prosodic context in the same way as for French nasals, suggest-ing a coordinative relationship in which oral and velic articula-tors are consistently phased in French nasal production. These findings illustrate the utility of real-time MRI as a method for studying velic activity and articulatory coordination in vocalic and nasal phonology. Index Terms: speech production, velum, nasals, nasal vowels, French, articulation, real-time MRI Michael I. Proctor, Louis Goldstein, Adam C. Lammert, Dani Byrd, Asterios Toutios, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Articulatory settings facilitate mechanically advantageous motor control of vocal tract articulatorsabstractIt was recently shown that vocal tract postures assumed during pauses in read speech are significantly different from those assumed at absolute rest. This paper examines whether the former category of “articulatory settings” are more mechanically advantageous than absolute rest postures with respect to speech articulation. Appropriate task and articulator variables are extracted from real-time Magnetic Resonance Imaging (rtMRI) data of five speakers reading aloud. Locally-weighted regression is then used to calculate Jacobian matrices representing the transformation between articulatory task velocities and postural velocities. A measure of mechanical advantage is proposed based on the obtained Jacobian. Speech-ready postures and postures during inter-speech pauses are observed to be significantly more mechanically advantageous as compared to rest postures. Furthermore, other postures, such as those that occur during the production of different vowels and consonants, are shown to have mechanical advantages that lie in between this continuum. These results could provide insights into understanding postural motor control and other linguistic phenomena, such as sonority hierarchies, in speech production. Index Terms: speech production, real-time MRI, articulatory setting, postural motor control, task dynamics, forward kinematics, vocal tract shaping. Vikram Ramanarayanan, Adam C. Lammert, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Stable articulatory tasks and their variable formation: tamil retroflex consonantsabstractA real-time MRI examination of retroflex stops and rhotics in Tamil reveals that in some contexts these consonants may in fact be achieved with little or no retroflexion of the tongue tip. Rather, maneuvering and shaping of the tongue in order to achieve post-alveolar contact varies across vowel contexts. Between back vowels /a / and /u/, post-alveolar constriction involves curling back of the tongue tip, but in the context of high front vowel /i/, the same constriction is achieved by bunching of the tongue. It appears that though there is a stable constriction target in the post-alveolar region, its achievement is not fixed but is instead a consequence of the variable state of the vocal tract in different vowel contexts. Articulatory configurations of the tongue across these vowel contexts were examined by comparing measures of Gaussian curvature at evenly spaced points along the vocal tract. The results support the notion that so-called retroflex consonants have a specified target constriction in the post-alveolar region, but that the specific articulations employed to achieve this constriction are not fixed, in keeping with the task dynamic model of speech production. Index Terms: retroflex, Tamil, real-time MRI 1. Caitlin Smith, Michael I. Proctor, Khalil Iskarous, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 4 |
| 2013 | Statistical methods for estimation of direct and differential kinematics of the vocal tract
Adam C. Lammert, Louis Goldstein, Shri Narayanan, Khalil Iskarous |
Speech Commun. | 2 |
| 2012 | Characterizing Covert Articulation in Apraxic Speech Using real-time MRIabstractWe explore the use of real-time magnetic resonance imaging (rtMRI) as a tool to investigate apraxic speech, in particular, by examining articulatory behavior. Our pilot data reveal that covert (silent) gestural intrusion errors (employing an intrinsically simple 1:1 mode of coupling) are made more frequently by an apraxic subject than by fluent speakers. Covert intrusion errors are also found to be pervasive in non-repetitious apraxic speech. We demonstrate that acoustically silent periods observed before the initiation of apraxic speech oftentimes contain completely covert gestures that occur frequently with multigestural segments. Covert gestures corresponding to entire words are also observed. These data demonstrate that rtMRI can provide important new insights into apraxic speech that are not available using traditional methods of transcription based on acoustic data alone. Index Terms: Apraxia, speech production, covert articulation, Christina Hagedorn, Michael I. Proctor, Louis Goldstein, Maria Luisa Gorno-Tempini, Shri Narayanan |
INTERSPEECH | 3 |
| 2012 | Emphatic segments and emphasis spread in Lebanese Arabic: a Real-time Magnetic Resonance Imaging StudyabstractProduction of emphatic consonants by a speaker of Lebanese Arabic was examined using real-time magnetic resonance imag-ing (rtMRI). Emphatic consonants were found to be articulated with a lowered, more retracted tongue body than their non-empatic counterparts, with the narrowest emphatic constriction observed in the upper pharynx. Both progressive and regressive emphasis spread was observed; spreading was not blocked by an intervening palatal approximant [j]. Emphaticized segments exhibit similar retraction and depression, with magnitudes that vary depending on the direction of spreading. These data suggest that emphasis spread may operate in a phonetically-complex way, not currently accounted for by phonological the-ory, and in addition, illustrate the advantage of real-time MRI as a method for studying emphasis in Semitic phonology. Assaf Israel, Michael I. Proctor, Louis Goldstein, Khalil Iskarous, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | Gesture-based Dynamic Bayesian Network for noise robust speech recognitionabstractPreviously we have proposed different models for estimating articulatory gestures and vocal tract variable (TV) trajectories from synthetic speech. We have shown that when deployed on natural speech, such models can help to improve the noise robustness of a hidden Markov model (HMM) based speech recognition system. In this paper we propose a model for estimating TVs trained on natural speech and present a Dynamic Bayesian Network (DBN) based speech recognition architecture that treats vocal tract constriction gestures as hidden variables, eliminating the necessity for explicit gesture recognition. Using the proposed architecture we performed a word recognition task for the noisy data of Aurora 2. Significant improvement was observed in using the gestural information as hidden variables in a DBN architecture over using only the mel-frequency cepstral coefficient based HMM or DBN backend. We also compare our results with other noise-robust front ends. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
ICASSP | 5 |
| 2011 | Speech inversion: Benefits of tract variables over pellet trajectoriesabstractSpeech inversion is a way of estimating articulatory trajectories or vocal tract configurations from the acoustic speech signal. Traditionally, articulator flesh-point or pellet trajectories have been used in speech-inversion research; however such information introduces additional variability into the inverse problem given they are head-centered, task-neutral measures. This paper proposes the use of vocal tract constriction variables (TVs) that are less variable for speech-inversion since they are constriction-based, task-specific measures. TVs considered in this study consist of five constriction degree variables, lip aperture (LA), tongue body constriction degree (TBCD), tongue tip constriction degree (TTCD), velum (VEL), and glottis (GLO); and three constriction location variables, lip protrusion (LP), tongue tip constriction location (TTCL) and tongue body constriction location (TBCL). Six different flesh-point trajectories were considered that were measured with transducers placed on the upper lip (UL), lower lip (LL) and four positions on the tongue (T1, T2, T3 and T4) between the tongue tip and the tongue dorsum. Speech inversion using a simple neural network architecture shows that the TVs can be estimated relatively more accurately than the pellet trajectories. Further statistical investigation reveals that the non-uniqueness is reduced in the TVs compared to the pellet trajectories for phones which are known to appreciably suffer from non-uniqueness. Finally we perform word recognition experiments using the estimated TVs as opposed to the pellet trajectories and show that the former offers greater word recognition accuracy both in clean and noisy speech, indicating that the TVs are a better choice for speech recognition systems. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
ICASSP | 5 |
| 2011 | Automatic Analysis of Singleton and Geminate Consonant Articulation Using Real-Time Magnetic Resonance ImagingabstractWe explore robust methods of automatically quantifying constriction location, constriction degree and gestural kinematics of Italian short and long consonants using direct image analysis techniques applied to rtMRI data. Articulatory kinematics are estimated from correlated regional changes in pixel intensity. We demonstrate that these methods are capable of quantifying differences in constriction duration exhibited by short and long Italian consonants for labial, coronal and dorsal segments, and differences in constriction degree for labial and coronal consonants. No difference in constriction location is observed for geminates and singletons, while systematic differences in constriction location are observed between (i) coronal oral stops and coronal sonorants and (ii) dorsal stops flanked by vowels differing in backness. Index Terms: speech production, real-time MRI, consonant articulation, Italian, geminates, articulatory phonology. Christina Hagedorn, Michael I. Proctor, Louis Goldstein |
INTERSPEECH | 3 |
| 2011 | A Multimodal Real-Time MRI Articulatory Corpus for Speech ResearchabstractWe present MRI-TIMIT: a large-scale database of synchronized audio and real-time magnetic resonance imaging (rtMRI) data for speech research. The database currently consists of speech data acquired from two male and two female speakers of Amer-ican English. Subjects ’ upper airways were imaged in the mid-sagittal plane while reading the same 460 sentence corpus used in the MOCHA-TIMIT corpus [1]. Accompanying acoustic recordings were phonemically transcribed using forced align-ment. Vocal tract tissue boundaries were automatically identi-fied in each video frame, allowing for dynamic quantification of each speaker’s midsagittal articulation. The database and com-panion toolset provide a unique resource with which to examine articulatory-acoustic relationships in speech production. Index Terms: speech production, speech corpora, real-time MRI, multi-modal database, large-scale phonetic tools Shri Narayanan, Erik Bresch, Prasanta Kumar Ghosh, Louis Goldstein, Athanasios Katsamanis, Adam C. Lammert, Michael I. Proctor, Vikram Ramanarayanan, Yinghua Zhu |
INTERSPEECH | 4 |
| 2011 | Direct Estimation of Articulatory Kinematics from Real-Time Magnetic Resonance Image SequencesabstractA method of rapid, automatic extraction of consonantal artic-ulatory trajectories from real-time magnetic resonance image sequences is described. Constriction location targets are esti-mated by identifying regions of maximally-dynamic correlated pixel activity along the palate, the alveolar ridge, and at the lips. Tissue movement into and out of the constriction location is es-timated by calculating the change in mean pixel intensity in a circle located at the center of the region of interest. Closure and release gesture timings are estimated from landmarks in the ve-locity profile derived from the smoothed intensity function. We demonstrate the utility of the technique in the analysis of Italian intervocalic consonant production. Index Terms: speech production, real-time MRI, consonant ar-ticulation, tongue shaping, articulatory phonology Michael I. Proctor, Adam C. Lammert, Athanasios Katsamanis, Louis Goldstein, Christina Hagedorn, Shri Narayanan |
INTERSPEECH | 4 |
| 2011 | Articulatory Information for Noise Robust Speech RecognitionabstractPrior research has shown that articulatory information, if extracted properly from the speech signal, can improve the performance of automatic speech recognition systems. However, such information is not readily available in the signal. The challenge posed by the estimation of articulatory information from speech acoustics has led to a new line of research known as “acoustic-to-articulatory inversion” or “speech-inversion.” While most of the research in this area has focused on estimating articulatory information more accurately, few have explored ways to apply this information in speech recognition tasks. In this paper, we first estimated articulatory information in the form of vocal tract constriction variables (abbreviated as TVs) from the Aurora-2 speech corpus using a neural network based speech-inversion model. Word recognition tasks were then performed for both noisy and clean speech using articulatory information in conjunction with traditional acoustic features. Our results indicate that incorporating TVs can significantly improve word recognition rates when used in conjunction with traditional acoustic features. Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
IEEE Trans. Speech Audio Process. | 5 |
| 2010 | Statistical multi-stream modeling of real-time MRI articulatory speech dataabstractThis paper investigates different statistical modeling frameworks for articulatory speech data obtained using real-time (RT) magnetic resonance imaging (MRI). To quantitatively capture the spatio-temporal shaping process of the human vocal tract during speech production a multi-dimensional stream of direct image features is extracted automatically from the MRI recordings. The features are closely related, though not identical, to the tract variables commonly defined in the articulatory phonology theory. The modeling of the shaping process aims at decomposing the articulatory data streams into primitives by segmentation. A variety of approaches are investigated for carrying out the segmentation task including vector quantizers, Gaussian Mixture Models, Hidden Markov Models, and a coupled Hidden Markov Model. We evaluate the performance of the different segmentation schemes qualitatively with the help of a well understood data set which was used in an earlier study of inter-articulatory timing phenomena of American English nasal sounds. Index Terms: speech production, articulatory modeling, realtime magnetic resonance imaging Erik Bresch, Athanasios Katsamanis, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | Locally-weighted regression for estimating the forward kinematics of a geometric vocal tract modelabstractTask-space control is well studied in modeling speech production [1, 2, 3, 4]. Implementing control of this kind requires an accurate kinematic forward model. Despite debate about how to define the tasks for speech (i.e., acoustical vs. articulatory), a faithful forward model will be complex and infeasible to express analytically. Thus, it is necessary to learn the forward model from data. Artificial Neural Networks (ANNs) have previously been suggested for this [3, 4, 6]. We argue for the use of locally-linear methods, such as Locally-Weighted Regression (LWR). While ANNs are capable of learning complex forward maps, LWR is more appropriate. Common formulations of control assume locally-linearity, whereas ANNs fit a nonlinear model to the entire map. Likewise, training LWR is simple compared to the complex optimization for ANNs. We provide an empirical comparison of these methods for learning a vocal tract forward model, discussing theoretical and practical aspects of each. Adam C. Lammert, Louis Goldstein, Khalil Iskarous |
INTERSPEECH | 2 |
| 2010 | Robust word recognition using articulatory trajectories and gesturesabstractArticulatory Phonology views speech as an ensemble of constricting events (e.g. narrowing lips, raising tongue tip), gestures, at distinct organs (lips, tongue tip, tongue body, velum, and glottis) along the vocal tract. This study shows that articulatory information in the form of gestures and their output trajectories (tract variable time functions or TVs) can help to improve the performance of automatic speech recognition systems. The lack of any natural speech database containing such articulatory information prompted us to use a synthetic speech dataset (obtained from Haskins Laboratories TAsk Dynamic model of speech production) that contains acoustic waveform for a given utterance and its corresponding gestures and TVs. First, we propose neural network based models to recognize the gestures and estimate the TVs from acoustic information. Second, the “synthetic-data trained” articulatory models were applied to the natural speech utterances in Aurora-2 corpus to estimate their gestures and TVs. Finally, we show that the estimated articulatory information helps to improve the noise robustness of a word recognition system when used along with the cepstral Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
INTERSPEECH | 5 |
| 2010 | A procedure for estimating gestural scores from natural speechabstractAbstract * Speech can be represented as a constellation of constricting events, gestures, , which are defined at distinct vocal tract sites, in the form of a gestural score.. Gestures and their output trajectories, tract variables, , which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given an utterance, the Haskins Laboratories Task Dynamics and Application (TADA) model was used to generate its prototype gestural score and the corresponding synthetic acoustic output. An optimal gestural score was estimated through iterative time-warping processes such that the distance between original and TADA-synthesized synthesized speech is minimized. We compared the performance of our approach to that of a conventional dynamic time warping procedure using Log-Spectral and Itakura Distance measures. We also performed a word recognition experiment using the gestural annotations to show that the gestural scores are suitable for word recognition. Hosung Nam, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson |
INTERSPEECH | 5 |
| 2010 | Investigating articulatory setting - pauses, ready position, and rest - using real-time MRIabstractWe present a novel automatic procedure to analyze ―articulatory setting (AS) ‖ or ―basis of articulation ‖ using realtime magnetic resonance images (rt-MRI) of the human vocal tract recorded for read and spontaneously spoken speech. We extract relevant frames of inter-speech pauses (ISPs) and rest positions from MRI sequences of read and spontaneous speech and use automatically-extracted features to quantify areas of different regions of the vocal tract as well as the angle of the jaw. Significant differences were found between the ASs adopted for ISPs in read and spontaneous speech, as well as those between ISPs and absolute rest positions. We further contrast differences between ASs adopted when the person is ready to speak as opposed to an absolute rest position. Index Terms — speech production, real-time MRI, basis of articulation, articulatory setting, pause articulation, read speech, spontaneous speech. 1. Vikram Ramanarayanan, Dani Byrd, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 3 |
| 2009 | Estimation of articulatory gesture patterns from speech acousticsabstractWe investigated dynamic programming (DP) and statemodel (SM) approaches for estimating gestural scores from speech acoustics. We performed a word-identification task using the gestural pattern vector sequences estimated by each approach. For a set of 75 randomly chosen words, we obtained the best word-identification accuracy (66.67%) using the DP approach. This result implies that considerable support for lexical access during speech perception might be provided by such a method of recovering gestural information from acoustics. Index Terms: gestural patterns, acoustic to gesture inversion Prasanta Kumar Ghosh, Shri Narayanan, Pierre L. Divenyi, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 4 |
| 2009 | Noise robustness of tract variables and their application to speech recognition
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein |
INTERSPEECH | 5 |
| 2009 | Connecting rhythm and prominence in automatic ESL pronunciation scoringabstractPast studies have shown that a native Spanish speaker’s use of phrasal prominence is a good indicator of her level of English prosody acquisition. Because of the cross-linguistic differences in the organization of phrasal prominence and durational contrasts, we hypothesize that those speakers with English-like prominence in their L2 speech are also expected to have acquired English-like rhythm. Statistics from a corpus of native and nonnative English confirm that speakers with an Englishlike phrasal prominence are also the ones who use English-like rhythm. Additionally, two methods of automatic score generation based on vowel duration times demonstrate a correlation of at least 0.6 between these automatic scores and subjective scores for phrasal prominence. These findings suggest that simple vowel duration measures obtained from standard automatic speech recognition methods can be salient cues for estimating subjective scores of prosodic acquisition, and of pronunciation in general. Emily Nava, Joseph Tepperman, Louis Goldstein, Maria Luisa Zubizarreta, Shri Narayanan |
INTERSPEECH | 3 |
| 2009 | An articulatory analysis of phonological transfer using real-time MRIabstractPhonological transfer is the influence of a first language on phonological variations made when speaking a second language. With automatic pronunciation assessment applications in mind, this study intends to uncover evidence of phonological transfer in terms of articulation. Real-time MRI videos from three German speakers of English and three native English speakers are compared to uncover the influence of German consonants on close English consonants not found in German. Results show that nonnative speakers demonstrate the effects of L1 transfer through the absence of articulatory contrasts seen in native speakers, while still maintaining minimal articulatory contrasts that are necessary for automatic detection of pronunciation errors, encouraging the further use of articulatory models for speech error characterization and detection. Index Terms: real-time MRI, nonnative speech, articulation, phonological transfer Joseph Tepperman, Erik Bresch, Yoon-Chul Kim, Sungbok Lee, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 5 |
| 2009 | Automatically rating pronunciation through articulatory phonologyabstractArticulatory Phonology’s link between cognitive speech planning and the physical realizations of vocal tract constrictions has implications for speech acoustic and duration modeling that should be useful in assigning subjective ratings of pronunciation quality to nonnative speech. In this work, we compare traditional phoneme models used in automatic speech recognition to similar models for articulatory gestural pattern vectors, each with associated duration models. What we find is that, on the CDT corpus, gestural models outperform the phonemelevel baseline in terms of correlation with listener ratings, and in combination phoneme and gestural models outperform either one alone. This also validates previous findings with a similar (but not gesture-based) pseudo-articulatory representation. Index Terms: pronunciation modeling, nonnative speech, articulatory phonology Joseph Tepperman, Louis Goldstein, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2009 | Articulatory phonological code for word classificationabstractWe propose a framework that leverages articulatory phonology for speech recognition. “Gestural pattern vectors ” (GPV) encode the instantaneous gestural activations that exist across all tract variables at each time. Given a speech observation, recognizing the sequence of GPV recovers the ensemble of gestural activations, i.e., the gestural score. For each word in the vocabulary, we use a task dynamic model of inter-articulator speech coordination to generate the “canonical ” gestural score. Speech recognition is achieved by matching the ensemble of gestural activations. In particular, we estimate the likelihood of the recognized GPV sequence on word-dependent GPV sequence models trained using the “canonical” gestural scores. These likelihoods, weighted by confidence score of the recognized GPVs, are used in a Bayesian speech recognizer. Pilot gestural score recovery and word classification experiments are carried out using synthesized data from one speaker. The observation distribution of each GPV is modeled by an artificial neural network and Gaussian mixture tandem model. Bigram GPV sequence models are used to distinguish gestural scores of different words. Given the tract variable time functions, about 80 % of the instantaneous gestural activation is correctly recovered. Word recognition accuracy is over 85 % for a vocabulary of 139 words with no training observations. These results suggest that the proposed framework might be a viable alternative to the classic sequence-of-phones model. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 4 |
| 2008 | An analysis of vocal tract shaping in English sibilant fricatives using real-time magnetic resonance imagingabstractThis study uses real-time MRI to investigate shaping aspects of two English sibilant fricatives. The purpose of this article is to 1) develop linguistically meaningful quantitative measurements based on vocal tract features that robustly capture the shaping aspects of the two fricatives, and 2) provide qualitative analyses of fricative shaping. Data was recorded in both midsagittal and coronal planes. The proposed three quantitative measures of this study provide robust results in categorizing shape. The qualitative analyses describe tongue shape in terms of grooving and doming and they support previous research. Erik Bresch, Daylen Riggs, Louis Goldstein, Dani Byrd, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 3 |
| 2008 | Six- and twelve-month-olds' discrimination of native versus non-native between- and within-organ fricative place contrasts
Michael D. Tyler, Catherine T. Best, Louis Goldstein, Mark Antoniou, Lidija Krebs-Lazendic |
INTERSPEECH | 3 |
| 2008 | The entropy of the articulatory phonological code: recognizing gestures from tract variablesabstractWe propose an instantaneous “gestural pattern vector ” to encode the instantaneous pattern of gesture activations across tract variables in the gestural score. The design of these gestural pattern vectors is the first step towards an automatic speech recognizer motivated by articulatory phonology, which is expected to be more invariant to speech coarticulation and reduction than conventional speech recognizers built with the sequenceof-phones assumption. We use a tandem model to recover the instantaneous gestural pattern vectors from tract variable time functions in local time windows, and achieve classification accuracy up to 84.5% for synthesized data from one speaker. Recognizing all gestural pattern vectors is equivalent to recognizing the ensemble of gestures. This result suggests that the proposed gestural pattern vector might be a viable unit in statistical models for speech recognition. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 4 |
| 2007 | Inverting mappings from smooth paths through Rn to paths through Rm: A technique applied to recovering articulation from acoustics
John Hogden, Philip Rubin, Erik McDermott, Shigeru Katagiri, Louis Goldstein |
Speech Commun. | 5 |