Louis Goldstein

dblp:87/5550 · also Louis M. Goldstein · DBLP profile ↗
← Back
54ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 12 since 2021Artificial intelligence and machine learning · 47 · 10 since 2021
YearPublicationVenuePosition
2026 Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model
Xuan Shi, Tiantian Feng, Jay Park, Christina Hagedorn, Louis Goldstein, Shri Narayanan
Comput. Speech Lang.5
2025 Instantaneous changes in acoustic signals reflect syllable progression and cross-linguistic syllable variation
Haley Hsu, Dani Byrd, Khalil Iskarous, Louis Goldstein
INTERSPEECH4
2025 Articulatory Feature Prediction from Surface EMG during Speech Production
Kleanthis Avramidis, Simon Pistrosch, Monica González Machorro, Yoonjeong Lee, Björn W. Schuller, Louis Goldstein, Shri Narayanan
INTERSPEECH8
2025 Towards a dynamical model of transitions between fluent and stuttered speech
Yijing Lu, Khalil Iskarous, Louis Goldstein
INTERSPEECH3
2025 75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment
Xuan Shi, Yubin Zhang, Yijing Lu, Marcus Ma, Tiantian Feng, Asterios Toutios, Haley Hsu, Louis Goldstein, Shri Narayanan
INTERSPEECH8
2025 Co-registration of real-time MRI and respiration for speech research
Yubin Zhang, Prakash Kumar, Xuan Shi, Haley Hsu, Shri Narayanan, Krishna S. Nayak, Louis Goldstein
INTERSPEECH11
2024 Analysis of articulatory setting for L1 and L2 English speakers using MRI data
abstract
This paper investigates the extent to which the geographical region (country) where a speaker acquired their English language affects the articulatory setting in their speech.To obtain accurate measurements for evaluating articulatory setting, we utilized a large real-time MRI corpus of vocal tract articulation.The corpus was obtained from speakers from a variety of linguistic backgrounds producing continuous English speech.We use an automated pipeline to process and extract articulatory positional information from the MRI video data.This data is used to draw comparisons between English language speakers from the United States and speakers who acquired their English in India, Korea, and China.Analysis of the speaker groups reveals statistically significant articulatory setting posture differences in multiple places of articulation.
Jack Goldberg, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2023 Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization
abstract
Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with human speech production mechanism, and corresponding gestural scores. We continue with this line of work by raising two concerns: (1) The articulators are entangled together in the original algorithm such that some of the articulators do not leverage effective moving patterns, which limits the interpretability of both gestures and gestural scores; (2) The EMA data is sparsely sampled from articulators, which limits the intelligibility of learned representations. In this work, we propose a novel articulatory representation decomposition algorithm that takes the advantage of guided factor analysis to derive the articulatory-specific factors and factor scores. A neural convolutive matrix factorization algorithm is then employed on the factor scores to derive the new gestures and gestural scores. We experiment with the rtMRI corpus that captures the fine-grained vocal tract contours. Both subjective and objective evaluation results suggest that the newly proposed system delivers the articulatory representations that are intelligible, generalizable, efficient and interpretable.
Jiachen Lian, Alan W. Black, Yijing Lu, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli
ICASSP4
2023 Speaker-Independent Acoustic-to-Articulatory Speech Inversion
abstract
To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech production. To this end, we build an acoustic-to-articulatory inversion (AAI) model that leverages autoregression, adversarial training, and self supervision to generalize to unseen speakers. Our approach obtains 0.784 correlation on an electromagnetic articulography (EMA) dataset, improving the state-of-the-art by 12.5%. Additionally, we show the interpretability of these representations through directly com-paring the behavior of estimated representations with speech production behavior. Finally, we propose a resynthesis-based AAI evaluation metric that does not rely on articulatory labels, demonstrating its efficacy with an 18-speaker dataset.
Peter Wu, Cheol Jun Cho, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli
ICASSP5
2023 Deep Speech Synthesis from MRI-Based Articulatory Representations
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W. Black, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli
INTERSPEECH7
2022 Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition
abstract
Most of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure.This work, investigating the speech representations derived from articulatory kinematics signals, uses a neural implementation of convolutive sparse matrix factorization to decompose the articulatory data into interpretable gestures and gestural scores.By applying sparse constraints, the gestural scores leverage the discrete combinatorial properties of phonological gestures.Phoneme recognition experiments were additionally performed to show that gestural scores indeed code phonological information successfully.The proposed work thus makes a bridge between articulatory phonology and deep neural networks to leverage informative, intelligible, interpretable,and efficient speech representations.
Jiachen Lian, Alan W. Black, Louis Goldstein, Gopala Krishna Anumanchipalli
INTERSPEECH3
2022 Deep Speech Synthesis from Articulatory Representations
Peter Wu, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli
INTERSPEECH3
2021 Who converges? Variation reveals individual speaker adaptability
abstract
Little is known about the cognitive capacities underlying real-time accommodation in spoken language and how they may allow conversing speakers to adapt their speech production behaviors. This study first presents a simple attunement model that incorporates hypothesized capacities, with a focus on individual variability as one of those capacities. The model makes explicit predictions about observable convergence behaviors in interacting speakers, including that: i) the intrinsically more variable speaker of the two will be the one who converges to their partner, ii) this flexible speaker with higher baseline variability will exhibit a substantial decrease in variability and iii) a greater change in the variability between speaking solo and interacting with their partner. These predictions are supported by the results of the modeling simulations. To further test the model's predictions, we analyzed a behavioral dataset including acoustic and articulatory data from three pairs of interacting speakers participating in a maze navigation task as well as a like solo speech task. The amount of variability in the speech parameters of each dyad member was quantified using coefficient of variation. The experimental results parallel the simulation results, and taken together, this work indicates that structured variability is an illuminating index of individual speaker adaptability and convergence behavior.
Yoon-Jeong Lee, Louis Goldstein, Benjamin Parrell, Dani Byrd
Speech Commun.2
2018 Analysis of speech production real-time MRI
Vikram Ramanarayanan, Sam Tilsen, Michael I. Proctor, Johannes Töger, Louis Goldstein, Krishna S. Nayak, Shri Narayanan
Comput. Speech Lang.5
2017 Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan
INTERSPEECH9
2017 Test-Retest Repeatability of Articulatory Strategies Using Real-Time Magnetic Resonance Imaging
Tanner Sorensen, Asterios Toutios, Johannes Töger, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2016 Velum Control for Oral Sounds
Reed Blaylock, Louis Goldstein, Shri Narayanan
INTERSPEECH2
2016 L2 Acquisition and Production of the English Rhotic Pharyngeal Gesture
Sarah Harper, Louis Goldstein, Shri Narayanan
INTERSPEECH2
2016 Perceptual Lateralization of Coda Rhotic Production in Puerto Rican Spanish
Mairym Lloréns Monteserín, Shri Narayanan, Louis Goldstein
INTERSPEECH3
2016 A New Model of Speech Motor Control Based on Task Dynamics and State Feedback
Vikram Ramanarayanan, Benjamin Parrell, Louis Goldstein, Srikantan S. Nagarajan, John F. Houde
INTERSPEECH3
2016 Characterizing Vocal Tract Dynamics Across Speakers Using Real-Time MRI
Tanner Sorensen, Asterios Toutios, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2016 Illustrating the Production of the International Phonetic Alphabet Sounds Using Fast Real-Time Magnetic Resonance Imaging
Asterios Toutios, Sajan Goud Lingala, Colin Vaz, Jangwon Kim, John H. Esling, Patricia A. Keating, Matthew Gordon, Dani Byrd, Louis Goldstein, Krishna S. Nayak, Shri Narayanan
INTERSPEECH9
2015 Analysis of coarticulated speech using estimated articulatory trajectories
Ganesh Sivaraman, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson
INTERSPEECH5
2015 Experimental assessment of the tongue incompressibility hypothesis during speech production
abstract
The human tongue is an important organ for speech production. Its deformation and motion control the shape of the vocal tract significantly and thereby the acoustic properties of the speech signal produced. Thus, much effort in the speech research com-munity has been directed towards its biomechanical modeling. A common assumption incorporated into many models of the human tongue is the tissue incompressibility hypothesis: the tongue is considered a muscular hydrostat and therefore its vol-ume should remain constant regardless of its posture. To the best of our knowledge, experimental assessment of the constant volume hypothesis during actual speech production is limited. In this work, the aim is to experimentally assess the incom-pressibility hypothesis during actual speech production using a dataset of volumetric Magnetic Resonance (MR) images of 17 subjects sustaining contextualized continuants (27 continu-ants per subject). A seeded region growing based algorithm is used to segment the tongue and calculate its volume. Then the intra-subject variability of the tongue volume along the differ-ent tongue postures is examined. Within the accuracy of our tongue volume measurements, our empirical results seem con-sistent with the incompressibility hypothesis. Index Terms: speech production, tongue volume, muscular hy-drostat, tissue incompressibility, volumetric MRI
Z.-I. Skordilis, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2014 A real-time MRI study of articulatory setting in second language speech
abstract
Previous work has shown that languages differ in their articulatory setting, the postural configuration that the vocal tract articulators tend to adopt when they are not engaged in any active speech gesture, and that this posture might be specified as part of the phonological knowledge speakers have of the language. This study tests whether the articulatory setting of a language can be acquired by non-native speakers. Three native speakers of German who had learned English as a second language were imaged using real-time MRI of the vocal tract while reading passages in German and English, and features that capture vocal tract posture were extracted from the inter-speech pauses in their native and non-native languages. Results show that the speakers exhibit distinct inter-speech postures in each language, with a lower and more retracted tongue in English, consistent with classic descriptions of the differences between the German and the English articulatory settings. This supports the view that non-native speakers may acquire relevant features of the articulatory setting of a second language, and also lends further support to the idea that articulatory setting is part of a speaker’s phonological competence in a language. Index Terms: articulatory setting, speech production, second language speech acquisition, real-time MRI.
Andrés Benítez, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2014 Motor control primitives arising from a learned dynamical systems model of speech articulation
abstract
We present a method to derive a small number of speech motor control “primitives” that can produce linguisticallyinterpretable articulatory movements. We envision that such a dictionary of primitives can be useful for speech motor control, particularly in finding a low-dimensional subspace for such control. First, we use the iterative Linear Quadratic Gaussian with Learned Dynamics (iLQG-LD) algorithm to derive (for a set of utterances) a set of stochastically optimal control inputs to a learned dynamical systems model of the vocal tract that produces desired movement sequences. Second, we use a convolutive Nonnegative Matrix Factorization with sparseness constraints (cNMFsc) algorithm to find a small dictionary of control input primitives that can be used to reproduce the aforementioned optimal control inputs that produce the observed articulatory movements. The method performs favorably on both qualitative and quantitative evaluations conducted on synthetic data produced by an articulatory synthesizer. Such a primitivesbased framework could help inform theories of speech motor control and coordination. Index Terms: speech motor control, motor primitives, synergies, dynamical systems, iLQG, NMF.
Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH2
2013 Truncation of pharyngeal gesture in English diphthong [aɪ]
abstract
It is well acknowledged that [a] in English diphthongs (e.g. [a] in “pie’d”) has a different formant structure from its closest corresponding monophthong (e.g. [a] in “pod”). The current study proposes that these two sounds share the same cognitive unit, i.e. the pharyngeal constriction gesture that produces [a], and the surface difference can be modeled as a consequence of truncating the same articulatory movement in time by the following palatal glide in the diphthongal environment. Formation of pharyngeal constriction gesture during the production of [a] in a diphthong and in its corresponding monophthong was observed in various timing contexts using Realtime MRI; and the collected production data were quantitatively analyzed using the direct image analysis (DIA) technique, which infers tissue movement by tracking pixel intensity change over time in regions of interest. Results support our truncation account in that: (1) formation time of pharyngeal constriction is significantly longer in monophthongs than in diphthongs; (2) this duration correlates with the resulting constriction degree; and (3) the resulting constriction degree predicts the acoustic difference in the F2 dimension as predicted by our hypothesis. Index Terms: English diphthongs, speech production, rtMRI
Fang-Ying Hsieh, Louis Goldstein, Dani Byrd, Shri Narayanan
INTERSPEECH2
2013 Velic coordination in French nasals: a real-time magnetic resonance imaging study
abstract
Production of nasal vowels in French, and nasal consonants in French and English, was examined using real-time magnetic resonance imaging (rtMRI). The coordination of velic and lin-gual gestures was found to be tightly controlled across differ-ent prosodic contexts in French nasals. Velum lowering in En-glish nasal consonants did not show the same control, although the timing of the corresponding lingual gestures varied with prosodic context in the same way as for French nasals, suggest-ing a coordinative relationship in which oral and velic articula-tors are consistently phased in French nasal production. These findings illustrate the utility of real-time MRI as a method for studying velic activity and articulatory coordination in vocalic and nasal phonology. Index Terms: speech production, velum, nasals, nasal vowels, French, articulation, real-time MRI
Michael I. Proctor, Louis Goldstein, Adam C. Lammert, Dani Byrd, Asterios Toutios, Shri Narayanan
INTERSPEECH2
2013 Articulatory settings facilitate mechanically advantageous motor control of vocal tract articulators
abstract
It was recently shown that vocal tract postures assumed during pauses in read speech are significantly different from those assumed at absolute rest. This paper examines whether the former category of “articulatory settings” are more mechanically advantageous than absolute rest postures with respect to speech articulation. Appropriate task and articulator variables are extracted from real-time Magnetic Resonance Imaging (rtMRI) data of five speakers reading aloud. Locally-weighted regression is then used to calculate Jacobian matrices representing the transformation between articulatory task velocities and postural velocities. A measure of mechanical advantage is proposed based on the obtained Jacobian. Speech-ready postures and postures during inter-speech pauses are observed to be significantly more mechanically advantageous as compared to rest postures. Furthermore, other postures, such as those that occur during the production of different vowels and consonants, are shown to have mechanical advantages that lie in between this continuum. These results could provide insights into understanding postural motor control and other linguistic phenomena, such as sonority hierarchies, in speech production. Index Terms: speech production, real-time MRI, articulatory setting, postural motor control, task dynamics, forward kinematics, vocal tract shaping.
Vikram Ramanarayanan, Adam C. Lammert, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2013 Stable articulatory tasks and their variable formation: tamil retroflex consonants
abstract
A real-time MRI examination of retroflex stops and rhotics in Tamil reveals that in some contexts these consonants may in fact be achieved with little or no retroflexion of the tongue tip. Rather, maneuvering and shaping of the tongue in order to achieve post-alveolar contact varies across vowel contexts. Between back vowels /a / and /u/, post-alveolar constriction involves curling back of the tongue tip, but in the context of high front vowel /i/, the same constriction is achieved by bunching of the tongue. It appears that though there is a stable constriction target in the post-alveolar region, its achievement is not fixed but is instead a consequence of the variable state of the vocal tract in different vowel contexts. Articulatory configurations of the tongue across these vowel contexts were examined by comparing measures of Gaussian curvature at evenly spaced points along the vocal tract. The results support the notion that so-called retroflex consonants have a specified target constriction in the post-alveolar region, but that the specific articulations employed to achieve this constriction are not fixed, in keeping with the task dynamic model of speech production. Index Terms: retroflex, Tamil, real-time MRI 1.
Caitlin Smith, Michael I. Proctor, Khalil Iskarous, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2013 Statistical methods for estimation of direct and differential kinematics of the vocal tract
Adam C. Lammert, Louis Goldstein, Shri Narayanan, Khalil Iskarous
Speech Commun.2
2012 Characterizing Covert Articulation in Apraxic Speech Using real-time MRI
abstract
We explore the use of real-time magnetic resonance imaging (rtMRI) as a tool to investigate apraxic speech, in particular, by examining articulatory behavior. Our pilot data reveal that covert (silent) gestural intrusion errors (employing an intrinsically simple 1:1 mode of coupling) are made more frequently by an apraxic subject than by fluent speakers. Covert intrusion errors are also found to be pervasive in non-repetitious apraxic speech. We demonstrate that acoustically silent periods observed before the initiation of apraxic speech oftentimes contain completely covert gestures that occur frequently with multigestural segments. Covert gestures corresponding to entire words are also observed. These data demonstrate that rtMRI can provide important new insights into apraxic speech that are not available using traditional methods of transcription based on acoustic data alone. Index Terms: Apraxia, speech production, covert articulation,
Christina Hagedorn, Michael I. Proctor, Louis Goldstein, Maria Luisa Gorno-Tempini, Shri Narayanan
INTERSPEECH3
2012 Emphatic segments and emphasis spread in Lebanese Arabic: a Real-time Magnetic Resonance Imaging Study
abstract
Production of emphatic consonants by a speaker of Lebanese Arabic was examined using real-time magnetic resonance imag-ing (rtMRI). Emphatic consonants were found to be articulated with a lowered, more retracted tongue body than their non-empatic counterparts, with the narrowest emphatic constriction observed in the upper pharynx. Both progressive and regressive emphasis spread was observed; spreading was not blocked by an intervening palatal approximant [j]. Emphaticized segments exhibit similar retraction and depression, with magnitudes that vary depending on the direction of spreading. These data suggest that emphasis spread may operate in a phonetically-complex way, not currently accounted for by phonological the-ory, and in addition, illustrate the advantage of real-time MRI as a method for studying emphasis in Semitic phonology.
Assaf Israel, Michael I. Proctor, Louis Goldstein, Khalil Iskarous, Shri Narayanan
INTERSPEECH3
2011 Gesture-based Dynamic Bayesian Network for noise robust speech recognition
abstract
Previously we have proposed different models for estimating articulatory gestures and vocal tract variable (TV) trajectories from synthetic speech. We have shown that when deployed on natural speech, such models can help to improve the noise robustness of a hidden Markov model (HMM) based speech recognition system. In this paper we propose a model for estimating TVs trained on natural speech and present a Dynamic Bayesian Network (DBN) based speech recognition architecture that treats vocal tract constriction gestures as hidden variables, eliminating the necessity for explicit gesture recognition. Using the proposed architecture we performed a word recognition task for the noisy data of Aurora 2. Significant improvement was observed in using the gestural information as hidden variables in a DBN architecture over using only the mel-frequency cepstral coefficient based HMM or DBN backend. We also compare our results with other noise-robust front ends.
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein
ICASSP5
2011 Speech inversion: Benefits of tract variables over pellet trajectories
abstract
Speech inversion is a way of estimating articulatory trajectories or vocal tract configurations from the acoustic speech signal. Traditionally, articulator flesh-point or pellet trajectories have been used in speech-inversion research; however such information introduces additional variability into the inverse problem given they are head-centered, task-neutral measures. This paper proposes the use of vocal tract constriction variables (TVs) that are less variable for speech-inversion since they are constriction-based, task-specific measures. TVs considered in this study consist of five constriction degree variables, lip aperture (LA), tongue body constriction degree (TBCD), tongue tip constriction degree (TTCD), velum (VEL), and glottis (GLO); and three constriction location variables, lip protrusion (LP), tongue tip constriction location (TTCL) and tongue body constriction location (TBCL). Six different flesh-point trajectories were considered that were measured with transducers placed on the upper lip (UL), lower lip (LL) and four positions on the tongue (T1, T2, T3 and T4) between the tongue tip and the tongue dorsum. Speech inversion using a simple neural network architecture shows that the TVs can be estimated relatively more accurately than the pellet trajectories. Further statistical investigation reveals that the non-uniqueness is reduced in the TVs compared to the pellet trajectories for phones which are known to appreciably suffer from non-uniqueness. Finally we perform word recognition experiments using the estimated TVs as opposed to the pellet trajectories and show that the former offers greater word recognition accuracy both in clean and noisy speech, indicating that the TVs are a better choice for speech recognition systems.
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein
ICASSP5
2011 Automatic Analysis of Singleton and Geminate Consonant Articulation Using Real-Time Magnetic Resonance Imaging
abstract
We explore robust methods of automatically quantifying constriction location, constriction degree and gestural kinematics of Italian short and long consonants using direct image analysis techniques applied to rtMRI data. Articulatory kinematics are estimated from correlated regional changes in pixel intensity. We demonstrate that these methods are capable of quantifying differences in constriction duration exhibited by short and long Italian consonants for labial, coronal and dorsal segments, and differences in constriction degree for labial and coronal consonants. No difference in constriction location is observed for geminates and singletons, while systematic differences in constriction location are observed between (i) coronal oral stops and coronal sonorants and (ii) dorsal stops flanked by vowels differing in backness. Index Terms: speech production, real-time MRI, consonant articulation, Italian, geminates, articulatory phonology.
Christina Hagedorn, Michael I. Proctor, Louis Goldstein
INTERSPEECH3
2011 A Multimodal Real-Time MRI Articulatory Corpus for Speech Research
abstract
We present MRI-TIMIT: a large-scale database of synchronized audio and real-time magnetic resonance imaging (rtMRI) data for speech research. The database currently consists of speech data acquired from two male and two female speakers of Amer-ican English. Subjects ’ upper airways were imaged in the mid-sagittal plane while reading the same 460 sentence corpus used in the MOCHA-TIMIT corpus [1]. Accompanying acoustic recordings were phonemically transcribed using forced align-ment. Vocal tract tissue boundaries were automatically identi-fied in each video frame, allowing for dynamic quantification of each speaker’s midsagittal articulation. The database and com-panion toolset provide a unique resource with which to examine articulatory-acoustic relationships in speech production. Index Terms: speech production, speech corpora, real-time MRI, multi-modal database, large-scale phonetic tools
Shri Narayanan, Erik Bresch, Prasanta Kumar Ghosh, Louis Goldstein, Athanasios Katsamanis, Adam C. Lammert, Michael I. Proctor, Vikram Ramanarayanan, Yinghua Zhu
INTERSPEECH4
2011 Direct Estimation of Articulatory Kinematics from Real-Time Magnetic Resonance Image Sequences
abstract
A method of rapid, automatic extraction of consonantal artic-ulatory trajectories from real-time magnetic resonance image sequences is described. Constriction location targets are esti-mated by identifying regions of maximally-dynamic correlated pixel activity along the palate, the alveolar ridge, and at the lips. Tissue movement into and out of the constriction location is es-timated by calculating the change in mean pixel intensity in a circle located at the center of the region of interest. Closure and release gesture timings are estimated from landmarks in the ve-locity profile derived from the smoothed intensity function. We demonstrate the utility of the technique in the analysis of Italian intervocalic consonant production. Index Terms: speech production, real-time MRI, consonant ar-ticulation, tongue shaping, articulatory phonology
Michael I. Proctor, Adam C. Lammert, Athanasios Katsamanis, Louis Goldstein, Christina Hagedorn, Shri Narayanan
INTERSPEECH4
2011 Articulatory Information for Noise Robust Speech Recognition
abstract
Prior research has shown that articulatory information, if extracted properly from the speech signal, can improve the performance of automatic speech recognition systems. However, such information is not readily available in the signal. The challenge posed by the estimation of articulatory information from speech acoustics has led to a new line of research known as “acoustic-to-articulatory inversion” or “speech-inversion.” While most of the research in this area has focused on estimating articulatory information more accurately, few have explored ways to apply this information in speech recognition tasks. In this paper, we first estimated articulatory information in the form of vocal tract constriction variables (abbreviated as TVs) from the Aurora-2 speech corpus using a neural network based speech-inversion model. Word recognition tasks were then performed for both noisy and clean speech using articulatory information in conjunction with traditional acoustic features. Our results indicate that incorporating TVs can significantly improve word recognition rates when used in conjunction with traditional acoustic features.
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein
IEEE Trans. Speech Audio Process.5
2010 Statistical multi-stream modeling of real-time MRI articulatory speech data
abstract
This paper investigates different statistical modeling frameworks for articulatory speech data obtained using real-time (RT) magnetic resonance imaging (MRI). To quantitatively capture the spatio-temporal shaping process of the human vocal tract during speech production a multi-dimensional stream of direct image features is extracted automatically from the MRI recordings. The features are closely related, though not identical, to the tract variables commonly defined in the articulatory phonology theory. The modeling of the shaping process aims at decomposing the articulatory data streams into primitives by segmentation. A variety of approaches are investigated for carrying out the segmentation task including vector quantizers, Gaussian Mixture Models, Hidden Markov Models, and a coupled Hidden Markov Model. We evaluate the performance of the different segmentation schemes qualitatively with the help of a well understood data set which was used in an earlier study of inter-articulatory timing phenomena of American English nasal sounds. Index Terms: speech production, articulatory modeling, realtime magnetic resonance imaging
Erik Bresch, Athanasios Katsamanis, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2010 Locally-weighted regression for estimating the forward kinematics of a geometric vocal tract model
abstract
Task-space control is well studied in modeling speech production [1, 2, 3, 4]. Implementing control of this kind requires an accurate kinematic forward model. Despite debate about how to define the tasks for speech (i.e., acoustical vs. articulatory), a faithful forward model will be complex and infeasible to express analytically. Thus, it is necessary to learn the forward model from data. Artificial Neural Networks (ANNs) have previously been suggested for this [3, 4, 6]. We argue for the use of locally-linear methods, such as Locally-Weighted Regression (LWR). While ANNs are capable of learning complex forward maps, LWR is more appropriate. Common formulations of control assume locally-linearity, whereas ANNs fit a nonlinear model to the entire map. Likewise, training LWR is simple compared to the complex optimization for ANNs. We provide an empirical comparison of these methods for learning a vocal tract forward model, discussing theoretical and practical aspects of each.
Adam C. Lammert, Louis Goldstein, Khalil Iskarous
INTERSPEECH2
2010 Robust word recognition using articulatory trajectories and gestures
abstract
Articulatory Phonology views speech as an ensemble of constricting events (e.g. narrowing lips, raising tongue tip), gestures, at distinct organs (lips, tongue tip, tongue body, velum, and glottis) along the vocal tract. This study shows that articulatory information in the form of gestures and their output trajectories (tract variable time functions or TVs) can help to improve the performance of automatic speech recognition systems. The lack of any natural speech database containing such articulatory information prompted us to use a synthetic speech dataset (obtained from Haskins Laboratories TAsk Dynamic model of speech production) that contains acoustic waveform for a given utterance and its corresponding gestures and TVs. First, we propose neural network based models to recognize the gestures and estimate the TVs from acoustic information. Second, the “synthetic-data trained” articulatory models were applied to the natural speech utterances in Aurora-2 corpus to estimate their gestures and TVs. Finally, we show that the estimated articulatory information helps to improve the noise robustness of a word recognition system when used along with the cepstral
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein
INTERSPEECH5
2010 A procedure for estimating gestural scores from natural speech
abstract
Abstract * Speech can be represented as a constellation of constricting events, gestures, , which are defined at distinct vocal tract sites, in the form of a gestural score.. Gestures and their output trajectories, tract variables, , which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given an utterance, the Haskins Laboratories Task Dynamics and Application (TADA) model was used to generate its prototype gestural score and the corresponding synthetic acoustic output. An optimal gestural score was estimated through iterative time-warping processes such that the distance between original and TADA-synthesized synthesized speech is minimized. We compared the performance of our approach to that of a conventional dynamic time warping procedure using Log-Spectral and Itakura Distance measures. We also performed a word recognition experiment using the gestural annotations to show that the gestural scores are suitable for word recognition.
Hosung Nam, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson
INTERSPEECH5
2010 Investigating articulatory setting - pauses, ready position, and rest - using real-time MRI
abstract
We present a novel automatic procedure to analyze ―articulatory setting (AS) ‖ or ―basis of articulation ‖ using realtime magnetic resonance images (rt-MRI) of the human vocal tract recorded for read and spontaneously spoken speech. We extract relevant frames of inter-speech pauses (ISPs) and rest positions from MRI sequences of read and spontaneous speech and use automatically-extracted features to quantify areas of different regions of the vocal tract as well as the angle of the jaw. Significant differences were found between the ASs adopted for ISPs in read and spontaneous speech, as well as those between ISPs and absolute rest positions. We further contrast differences between ASs adopted when the person is ready to speak as opposed to an absolute rest position. Index Terms — speech production, real-time MRI, basis of articulation, articulatory setting, pause articulation, read speech, spontaneous speech. 1.
Vikram Ramanarayanan, Dani Byrd, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2009 Estimation of articulatory gesture patterns from speech acoustics
abstract
We investigated dynamic programming (DP) and statemodel (SM) approaches for estimating gestural scores from speech acoustics. We performed a word-identification task using the gestural pattern vector sequences estimated by each approach. For a set of 75 randomly chosen words, we obtained the best word-identification accuracy (66.67%) using the DP approach. This result implies that considerable support for lexical access during speech perception might be provided by such a method of recovering gestural information from acoustics. Index Terms: gestural patterns, acoustic to gesture inversion
Prasanta Kumar Ghosh, Shri Narayanan, Pierre L. Divenyi, Louis Goldstein, Elliot Saltzman
INTERSPEECH4
2009 Noise robustness of tract variables and their application to speech recognition
Vikramjit Mitra, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Louis Goldstein
INTERSPEECH5
2009 Connecting rhythm and prominence in automatic ESL pronunciation scoring
abstract
Past studies have shown that a native Spanish speaker’s use of phrasal prominence is a good indicator of her level of English prosody acquisition. Because of the cross-linguistic differences in the organization of phrasal prominence and durational contrasts, we hypothesize that those speakers with English-like prominence in their L2 speech are also expected to have acquired English-like rhythm. Statistics from a corpus of native and nonnative English confirm that speakers with an Englishlike phrasal prominence are also the ones who use English-like rhythm. Additionally, two methods of automatic score generation based on vowel duration times demonstrate a correlation of at least 0.6 between these automatic scores and subjective scores for phrasal prominence. These findings suggest that simple vowel duration measures obtained from standard automatic speech recognition methods can be salient cues for estimating subjective scores of prosodic acquisition, and of pronunciation in general.
Emily Nava, Joseph Tepperman, Louis Goldstein, Maria Luisa Zubizarreta, Shri Narayanan
INTERSPEECH3
2009 An articulatory analysis of phonological transfer using real-time MRI
abstract
Phonological transfer is the influence of a first language on phonological variations made when speaking a second language. With automatic pronunciation assessment applications in mind, this study intends to uncover evidence of phonological transfer in terms of articulation. Real-time MRI videos from three German speakers of English and three native English speakers are compared to uncover the influence of German consonants on close English consonants not found in German. Results show that nonnative speakers demonstrate the effects of L1 transfer through the absence of articulatory contrasts seen in native speakers, while still maintaining minimal articulatory contrasts that are necessary for automatic detection of pronunciation errors, encouraging the further use of articulatory models for speech error characterization and detection. Index Terms: real-time MRI, nonnative speech, articulation, phonological transfer
Joseph Tepperman, Erik Bresch, Yoon-Chul Kim, Sungbok Lee, Louis Goldstein, Shri Narayanan
INTERSPEECH5
2009 Automatically rating pronunciation through articulatory phonology
abstract
Articulatory Phonology’s link between cognitive speech planning and the physical realizations of vocal tract constrictions has implications for speech acoustic and duration modeling that should be useful in assigning subjective ratings of pronunciation quality to nonnative speech. In this work, we compare traditional phoneme models used in automatic speech recognition to similar models for articulatory gestural pattern vectors, each with associated duration models. What we find is that, on the CDT corpus, gestural models outperform the phonemelevel baseline in terms of correlation with listener ratings, and in combination phoneme and gestural models outperform either one alone. This also validates previous findings with a similar (but not gesture-based) pseudo-articulatory representation. Index Terms: pronunciation modeling, nonnative speech, articulatory phonology
Joseph Tepperman, Louis Goldstein, Sungbok Lee, Shri Narayanan
INTERSPEECH2
2009 Articulatory phonological code for word classification
abstract
We propose a framework that leverages articulatory phonology for speech recognition. “Gestural pattern vectors ” (GPV) encode the instantaneous gestural activations that exist across all tract variables at each time. Given a speech observation, recognizing the sequence of GPV recovers the ensemble of gestural activations, i.e., the gestural score. For each word in the vocabulary, we use a task dynamic model of inter-articulator speech coordination to generate the “canonical ” gestural score. Speech recognition is achieved by matching the ensemble of gestural activations. In particular, we estimate the likelihood of the recognized GPV sequence on word-dependent GPV sequence models trained using the “canonical” gestural scores. These likelihoods, weighted by confidence score of the recognized GPVs, are used in a Bayesian speech recognizer. Pilot gestural score recovery and word classification experiments are carried out using synthesized data from one speaker. The observation distribution of each GPV is modeled by an artificial neural network and Gaussian mixture tandem model. Bigram GPV sequence models are used to distinguish gestural scores of different words. Given the tract variable time functions, about 80 % of the instantaneous gestural activation is correctly recovered. Word recognition accuracy is over 85 % for a vocabulary of 139 words with no training observations. These results suggest that the proposed framework might be a viable alternative to the classic sequence-of-phones model. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model
Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman
INTERSPEECH4
2008 An analysis of vocal tract shaping in English sibilant fricatives using real-time magnetic resonance imaging
abstract
This study uses real-time MRI to investigate shaping aspects of two English sibilant fricatives. The purpose of this article is to 1) develop linguistically meaningful quantitative measurements based on vocal tract features that robustly capture the shaping aspects of the two fricatives, and 2) provide qualitative analyses of fricative shaping. Data was recorded in both midsagittal and coronal planes. The proposed three quantitative measures of this study provide robust results in categorizing shape. The qualitative analyses describe tongue shape in terms of grooving and doming and they support previous research.
Erik Bresch, Daylen Riggs, Louis Goldstein, Dani Byrd, Sungbok Lee, Shri Narayanan
INTERSPEECH3
2008 Six- and twelve-month-olds' discrimination of native versus non-native between- and within-organ fricative place contrasts
Michael D. Tyler, Catherine T. Best, Louis Goldstein, Mark Antoniou, Lidija Krebs-Lazendic
INTERSPEECH3
2008 The entropy of the articulatory phonological code: recognizing gestures from tract variables
abstract
We propose an instantaneous “gestural pattern vector ” to encode the instantaneous pattern of gesture activations across tract variables in the gestural score. The design of these gestural pattern vectors is the first step towards an automatic speech recognizer motivated by articulatory phonology, which is expected to be more invariant to speech coarticulation and reduction than conventional speech recognizers built with the sequenceof-phones assumption. We use a tandem model to recover the instantaneous gestural pattern vectors from tract variable time functions in local time windows, and achieve classification accuracy up to 84.5% for synthesized data from one speaker. Recognizing all gestural pattern vectors is equivalent to recognizing the ensemble of gestures. This result suggests that the proposed gestural pattern vector might be a viable unit in statistical models for speech recognition. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model
Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman
INTERSPEECH4
2007 Inverting mappings from smooth paths through Rn to paths through Rm: A technique applied to recovering articulation from acoustics
John Hogden, Philip Rubin, Erik McDermott, Shigeru Katagiri, Louis Goldstein
Speech Commun.5