Maureen Stone 0001

dblp:16/1888 · also Maureen L. Stone · DBLP profile ↗
← Back
48ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0004-3222-9771ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 since 2021
YearPublicationVenuePosition
2026 A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Fangxu Xing, Xiaofeng Liu 0001, Maureen Stone 0001, Jiachen Zhuo, Juan Rafael Orozco-Arroyave, Elmar Nöth, Jana Hutter, Jerry L. Prince, Andreas K. Maier, Jonghye Woo
Medical Image Anal.5
2024 Contrastive Learning Approach for Assessment of Phonological Precision in Patients with Tongue Cancer Using MRI Data
abstract
Magnetic Resonance Imaging (MRI) allows analyzing speech production by capturing high-resolution images of the dynamic processes in the vocal tract. In clinical applications, combining MRI with synchronized speech recordings leads to improved patient outcomes, especially if a phonological-based approach is used for assessment. However, when audio signals are unavailable, the recognition accuracy of sounds is decreased when using only MRI data. We propose a contrastive learning approach to improve the detection of phonological classes from MRI data when acoustic signals are not available at inference time. We demonstrate that frame-wise recognition of phonological classes improves from an f1 of 0.74 to 0.85 when the contrastive loss approach is implemented. Furthermore, we show the utility of our approach in the clinical application of using such phonological classes to assess speech disorders in patients with tongue cancer, yielding promising results in the recognition task.
Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Maria Schuster, Elmar Nöth, Jonghye Woo, Andreas K. Maier
INTERSPEECH5
2024 Tagged-to-Cine MRI Sequence Synthesis via Light Spatial-Temporal Transformer
Xiaofeng Liu 0001, Fangxu Xing, Zhangxing Bian, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Andreas K. Maier, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Jonghye Woo
MICCAI (7)7
2023 Motor Control Similarity Between Speakers Saying "A Souk" Using Inverse Atlas Tongue Modeling
abstract
Finite element models (FEM) of the tongue have facilitated speech studies through analysis of internal muscle forces indirectly derived from imaging data. In this work, we build a uniform hexahedral FEM of a tongue atlas constructed from magnetic resonance imaging data of a healthy population. The FEM is driven by inverse internal tongue tissue kinematics of speakers temporally aligned and deformed into the same atlas space, while performing the speech task "a souk" allowing muscle activation predictions. This work aims to investigate the commonalities in tongue motor strategies in the articulation of "a souk" predicted by the inverse tongue atlas model. Our findings report variability among five speakers for estimated muscle activations with a similarity index using a dynamic time warp function. Two speakers show similarity index > 0.9 and two others < 0.7 with respect to a reference speaker for most tongue muscles. The relative motion tracking error of the model is less than 2% which is promising for speech study applications.
Ursa Maity, Fangxu Xing, Jerry L. Prince, Maureen Stone 0001, Georges El Fakhri, Jonghye Woo, Sidney S. Fels
INTERSPEECH4
2023 Speech Audio Synthesis from Tagged MRI and Non-negative Matrix Factorization via Plastic Transformer
Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Sidney S. Fels, Jerry L. Prince, Georges El Fakhri, Jonghye Woo
MICCAI (7)3
2023 Attentive continuous generative self-training for unsupervised domain adaptive medical image translation
Xiaofeng Liu 0001, Jerry L. Prince, Fangxu Xing, Jiachen Zhuo, Timothy G. Reese, Maureen Stone 0001, Georges El Fakhri, Jonghye Woo
Medical Image Anal.6
2022 Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous Translator
abstract
Multimodal representation learning using visual movements from cine magnetic resonance imaging (MRI) and their acoustics has shown great potential to learn shared representation and to predict one modality from another. Here, we propose a new synthesis framework to translate from cine MRI sequences to spectrograms with a limited dataset size. Our framework hinges on a novel fully convolutional heterogeneous translator, with a 3D CNN encoder for efficient sequence encoding and a 2D transpose convolution decoder. In addition, a pairwise correlation of the samples with the same speech word is utilized with a latent space representation disentanglement scheme. Furthermore, an adversarial training approach with generative adversarial networks is incorporated to provide enhanced realism on our generated spectrograms. Our experimental results, carried out with a total of 63 cine MRI sequences alongside speech acoustics, show that our framework improves synthesis accuracy, compared with competing methods. Our framework thereby has shown the potential to aid in better understanding the relationship between the two modalities.
Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jerry L. Prince, Jangwon Kim, Georges El Fakhri, Jonghye Woo
ICASSP3
2022 Tagged-MRI Sequence to Audio Synthesis via Self Residual Attention Guided Heterogeneous Translator
Xiaofeng Liu 0001, Fangxu Xing, Jerry L. Prince, Jiachen Zhuo, Maureen Stone 0001, Georges El Fakhri, Jonghye Woo
MICCAI (6)5
2021 Generative Self-training for Cross-Domain Unsupervised Tagged-to-Cine MRI Synthesis
Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Timothy G. Reese, Jerry L. Prince, Georges El Fakhri, Jonghye Woo
MICCAI (3)3
2021 A deep joint sparse non-negative matrix factorization framework for identifying the common and subject-specific functional units of tongue motion during speech
Jonghye Woo, Fangxu Xing, Jerry L. Prince, Maureen Stone 0001, Arnold D. Gomez, Timothy G. Reese, Van J. Wedeen, Georges El Fakhri
Medical Image Anal.4
2019 A Sparse Non-Negative Matrix Factorization Framework for Identifying Functional Units of Tongue Behavior From MRI
abstract
Muscle coordination patterns of lingual behaviors are synergies generated by deforming local muscle groups in a variety of ways. Functional units are functional muscle groups of local structural elements within the tongue that compress, expand, and move in a cohesive and consistent manner. Identifying the functional units using tagged-magnetic resonance imaging (MRI) sheds light on the mechanisms of normal and pathological muscle coordination patterns, yielding improvement in surgical planning, treatment, or rehabilitation procedures. In this paper, to mine this information, we propose a matrix factorization and probabilistic graphical model framework to produce building blocks and their associated weighting map using motion quantities extracted from tagged-MRI. Our tagged-MRI imaging and accurate voxel-level tracking provide previously unavailable internal tongue motion patterns, thus revealing the inner workings of the tongue during speech or other lingual behaviors. We then employ spectral clustering on the weighting map to identify the cohesive regions defined by the tongue motion that may involve multiple or undocumented regions. To evaluate our method, we perform a series of experiments. We first use two-dimensional images and synthetic data to demonstrate the accuracy of our method. We then use three-dimensional synthetic and in vivo tongue motion data using protrusion and simple speech tasks to identify subject-specific and data-driven functional units of the tongue in localized regions.
Jonghye Woo, Jerry L. Prince, Maureen Stone 0001, Fangxu Xing, Arnold D. Gomez, Jordan R. Green, Christopher J. Hartnick, Thomas J. Brady, Timothy G. Reese, Van J. Wedeen, Georges El Fakhri
IEEE Trans. Medical Imaging3
2018 Quantifying Tensor Field Similarity with Global Distributions and Optimal Transport
Arnold D. Gomez, Maureen Stone 0001, Philip V. Bayly, Jerry L. Prince
MICCAI (2)2
2017 Speaker-Specific Biomechanical Model-Based Investigation of a Simple Speech Task Based on Tagged-MRI
Keyi Tang, Negar M. Harandi, Jonghye Woo, Georges El Fakhri, Maureen Stone 0001, Sidney S. Fels
INTERSPEECH5
2017 Phase Vector Incompressible Registration Algorithm for Motion Estimation From Tagged Magnetic Resonance Images
abstract
Tagged magnetic resonance imaging has been used for decades to observe and quantify motion and strain of deforming tissue. It is challenging to obtain 3-D motion estimates due to a tradeoff between image slice density and acquisition time. Typically, interpolation methods are used either to combine 2-D motion extracted from sparse slice acquisitions into 3-D motion or to construct a dense volume from sparse acquisitions before image registration methods are applied. This paper proposes a new phase-based 3-D motion estimation technique that first computes harmonic phase volumes from interpolated tagged slices and then matches them using an image registration framework. The approach uses several concepts from diffeomorphic image registration with a key novelty that defines a symmetric similarity metric on harmonic phase volumes from multiple orientations. The material property of harmonic phase solves the aperture problem of optical flow and intensity-based methods and is robust to tag fading. A harmonic magnitude volume is used in enforcing incompressibility in the tissue regions. The estimated motion fields are dense, incompressible, diffeomorphic, and inverse-consistent at a 3-D voxel level. The method was evaluated using simulated phantoms, human brain data in mild head accelerations, human tongue data during speech, and an open cardiac data set. The method shows comparable accuracy to three existing methods while demonstrating low computation time and robustness to tag fading and noise.
Fangxu Xing, Jonghye Woo, Arnold D. Gomez, Dzung L. Pham, Philip V. Bayly, Maureen Stone 0001, Jerry L. Prince
IEEE Trans. Medical Imaging6
2015 Segmentation of tongue muscles from super-resolution magnetic resonance images
Bulat Ibragimov, Jerry L. Prince, Emi Z. Murano, Jonghye Woo, Maureen Stone 0001, Bostjan Likar, Franjo Pernus, Tomaz Vrtovec
Medical Image Anal.5
2015 Multimodal Registration via Mutual Information Incorporating Geometric and Spatial Context
abstract
Multimodal image registration is a class of algorithms to find correspondence from different modalities. Since different modalities do not exhibit the same characteristics, finding accurate correspondence still remains a challenge. To deal with this, mutual information (MI)-based registration has been a preferred choice as MI is based on the statistical relationship between both volumes to be registered. However, MI has some limitations. First, MI-based registration often fails when there are local intensity variations in the volumes. Second, MI only considers the statistical intensity relationships between both volumes and ignores the spatial and geometric information about the voxel. In this work, we propose to address these limitations by incorporating spatial and geometric information via a 3D Harris operator. In particular, we focus on the registration between a high-resolution image and a low-resolution image. The MI cost function is computed in the regions where there are large spatial variations such as corner or edge. In addition, the MI cost function is augmented with geometric information derived from the 3D Harris operator applied to the high-resolution image. The robustness and accuracy of the proposed method were demonstrated using experiments on synthetic and clinical data including the brain and the tongue. The proposed method provided accurate registration and yielded better performance over standard registration methods.
Jonghye Woo, Maureen Stone 0001, Jerry L. Prince
IEEE Trans. Image Process.2
2014 An educational platform to capture, visualize and analyze rare singing
Patrick Chawah, Samer Al Kork, Thibaut Fux, Martine Adda-Decker, Angélique Amelot, Nicolas Audibert, Bruce Denby, Gérard Dreyfus, Aurore Jaumard-Hakoun, Claire Pillot-Loiseau, Pierre Roussel-Ragot, Maureen Stone 0001, Kele Xu, Lise Crevier-Buchman
INTERSPEECH12
2014 3d tongue motion visualization based on ultrasound image sequences
Kele Xu, Yin Yang 0002, Aurore Jaumard-Hakoun, Martine Adda-Decker, Angélique Amelot, Samer Al Kork, Lise Crevier-Buchman, Patrick Chawah, Gérard Dreyfus, Thibaut Fux, Claire Pillot-Loiseau, Pierre Roussel-Ragot, Maureen Stone 0001, Bruce Denby
INTERSPEECH13
2014 Determining Functional Units of Tongue Motion via Graph-Regularized Sparse Non-negative Matrix Factorization
Jonghye Woo, Fangxu Xing, Maureen Stone 0001, Jerry L. Prince
MICCAI (2)4
2013 A cine MRI-based study of sibilant fricatives production in post-glossectomy speakers
abstract
Glossectomy changes properties of the tongue and negatively affects patients' speech production. Among the most difficult consonants to produce in the post-glossectomy speakers, the sibilant fricatives /s/ and /sh/ are often problematic. To better understand these problems in production, this study analyzed acoustic and articulatory data of /s/ and /sh/ from three subjects: one normal speaker and two post-glossectomy speakers with abnormal /s/ or /sh. Based on cine magnetic resonance images, three dimensional vocal tract reconstructions, tongue surface shapes behind constrictions, and area functions were analyzed. Our results show that in each patient, contrary to normal, /s/ and /sh/ were quite similar in acoustic spectra, tongue surface shapes, and constriction locations. In the abnormal /s/, the missing unilateral tongue tissue created an air flow bypass which made the constriction further backward. The abnormal /sh/ may be explained by the lack of precise tongue control after surgery. In addition, the tongue surfaces in the patients were more asymmetric in the back and were not grooved for /s/ anterior to the constriction.
Xinhui Zhou, Jonghye Woo, Maureen Stone 0001, Carol Y. Espy-Wilson
ICASSP3
2013 3D Tongue Motion from Tagged and Cine MR Images
Fangxu Xing, Jonghye Woo, Emi Z. Murano, Maureen Stone 0001, Jerry L. Prince
MICCAI (3)5
2012 Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulations
abstract
Oral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Depending on the tumor size, location and staging, patients are treated by radical surgery, radiology, chemotherapy or a combination of those treatments. As a result, their anatomical structures for speech are impaired and this leads to some negative impact on their speech intelligibility. As a part of the INTERSPEECH 2012 speaker trait Pathology sub-challenge, this study explored the use of auditory-inspired spectro-temporal modulation features for automatic speech intelligibility assessment of those pathologic speech. The averaged spectro-temporal modulations of speech considered as either intelligible or non-intelligible in the challenge database were analyzed and it was found that the non-intelligible speech tends to have its modulation amplitude peaks shift towards a smaller rate and scale. Based on SVM and GMM, variants of spectro-temporal modulation features were tested on the speaker trait challenge problem and the resulting performances on both the development and the test datasets are comparable to the baseline performance.
Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen Stone 0001, Carol Y. Espy-Wilson, Shihab A. Shamma
INTERSPEECH4
2012 Incompressible Deformation Estimation Algorithm (IDEA) From Tagged MR Images
abstract
Measuring the 3D motion of muscular tissues, e.g., the heart or the tongue, using magnetic resonance (MR) tagging is typically carried out by interpolating the 2D motion information measured on orthogonal stacks of images. The incompressibility of muscle tissue is an important constraint on the reconstructed motion field and can significantly help to counter the sparsity and incompleteness of the available motion information. Previous methods utilizing this fact produced incompressible motions with limited accuracy. In this paper, we present an incompressible deformation estimation algorithm (IDEA) that reconstructs a dense representation of the 3D displacement field from tagged MR images and the estimated motion field is incompressible to high precision. At each imaged time frame, the tagged images are first processed to determine components of the displacement vector at each pixel relative to the reference time. IDEA then applies a smoothing, divergence-free, vector spline to interpolate velocity fields at intermediate discrete times such that the collection of velocity fields integrate over time to match the observed displacement components. Through this process, IDEA yields a dense estimate of a 3D displacement field that matches our observations and also corresponds to an incompressible motion. The method was validated with both numerical simulation and in vivo human experiments on the heart and the tongue.
Xiaofeng Liu 0001, Khaled Z. Abd-Elmoniem, Maureen Stone 0001, Emi Z. Murano, Jiachen Zhuo, Rao P. Gullapalli, Jerry L. Prince
IEEE Trans. Medical Imaging3
2011 A Comparative Acoustic Study on Speech of Glossectomy Patients and Normal Subjects
abstract
Oral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Tongue cancer patients are treated by glossectomy, a surgical procedure to remove the cancerous tumor. As a result, the tongue properties such as volume, shape, muscle structure, and motility are affected. As a result, the vocal tract acoustics are affected too. This study compares the speech acoustics between normal subjects and partial glossectomy patients with T1 or T2 tumors. The acoustic signal of four vowels (/iy/, /uw/, /eh/, and /ah/) and two fricatives (/s/ and /sh/) were analyzed. Our results show that, while the average formants (F1-F3) for the four vowels between the normal subjects and the glossectomy patients are very similar, the average centers of gravity for the two fricatives differ significantly. These differences in fricatives can be explained by the more posterior constriction in patients due to the glossectomy (or the cancer tumor) and its resulting longer front cavity.
Xinhui Zhou, Maureen Stone 0001, Carol Y. Espy-Wilson
INTERSPEECH2
2011 Deformable Registration of High-Resolution and Cine MR Tongue Images
Jonghye Woo, Maureen Stone 0001, Jerry L. Prince
MICCAI (1)2
2010 Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips
Thomas Hueber, Elie-Laurent Benaroya, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001
Speech Commun.6
2009 Visuo-phonetic decoding using multi-stream and context-dependent models for an ultrasound-based silent speech interface
abstract
Recent improvements are presented for phonetic decoding of continuous-speech from ultrasound and optical observations of the tongue and lips in a silent speech interface application. In a new approach to this critical step, the visual streams are modeled by context-dependent multi-stream Hidden Markov Models (CD-MSHMM). Results are compared to a baseline system using context-independent modeling and a visual feature fusion strategy, with both systems evaluated on a one-hour, phonetically balanced English speech database. Tongue and lip images are coded using PCA-based feature extraction techniques. The uttered speech signal, also recorded, is used to initialize the training of the visual HMMs. Visual phonetic decoding performance is evaluated successively with and without the help of linguistic constraints introduced via a 2.5k-word decoding dictionary. Index Terms: silent speech interface, visual speech recognition, multi-stream modeling 1.
Thomas Hueber, Elie-Laurent Benaroya, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001
INTERSPEECH6
2008 Towards a segmental vocoder driven by ultrasound and optical images of the tongue and lips
abstract
This article presents a framework for a phonetic vocoder driven by ultrasound and optical images of the tongue and lips for a “silent speech interface” application. The system is built around an HMM-based visual phone recognition step which provides target phonetic sequences from a continuous visual observation stream. The phonetic target constrains the search for the optimal sequence of diphones that maximizes similarity to the input test data in visual space subject to a unit concatenation cost in the acoustic domain. The final speech waveform is generated using “Harmonic plus Noise Model” synthesis techniques. Experimental results are based on a onehour continuous speech audiovisual database comprising ultrasound images of the tongue and both frontal and lateral view of the speaker’s lips.
Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001
INTERSPEECH5
2008 Phone recognition from ultrasound and optical video sequences for a silent speech interface
abstract
Latest results on continuous speech phone recognition from video observations of the tongue and lips are described in the context of an ultrasound-based silent speech interface. The study is based on a new 61-minute audiovisual database containing ultrasound sequences of the tongue as well as both frontal and lateral view of the speaker’s lips. Phonetically balanced and exhibiting good diphone coverage, this database is designed both for recognition and corpus-based synthesis purposes. Acoustic waveforms are phonetically labeled, and visual sequences coded using PCA-based robust feature extraction techniques. Visual and acoustic observations of each phonetic class are modeled by continuous HMMs, allowing the performance of the visual phone recognizer to be compared to a traditional acoustic-based phone recognition experiment. The phone recognition confusion matrix is also discussed in detail.
Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001
INTERSPEECH5
2007 Eigentongue Feature Extraction for an Ultrasound-Based Silent Speech Interface
abstract
The article compares two approaches to the description of ultrasound vocal tract images for application in a "silent speech interface," one based on tongue contour modeling, and a second, global coding approach in which images are projected onto a feature space of Eigentongues. A curvature-based lip profile feature extraction method is also presented. Extracted visual features are input to a neural network which learns the relation between the vocal tract configuration and line spectrum frequencies (LSF) contained in a one-hour speech corpus. An examination of the quality of LSFs derived from the two approaches demonstrates that the Eigemongues approach has a more efficient implementation and provides superior results based on a normalized mean squared error criterion.
Thomas Hueber, Guido Aversano, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Yacine Oussar, Pierre Roussel-Ragot, Maureen Stone 0001
ICASSP (1)8
2007 Continuous-speech phone recognition from ultrasound and optical images of the tongue and lips
abstract
The article describes a video-only speech recognition system for a “silent speech interface” application, using ultrasound and optical images of the voice organ. A one-hour audiovisual speech corpus was phonetically labeled using an automatic speech alignment procedure and robust visual feature extraction techniques. HMM-based stochastic models were estimated separately on the visual and acoustic corpus. The performance of the visual speech recognition system is compared to a traditional acoustic-based recognizer.
Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001
INTERSPEECH5
2007 Nonrigid motion recovery for 3D surfaces
Chandra Kambhamettu, Maureen Stone 0001
Image Vis. Comput.3
2006 A Level Set Approach for Shape Recovery of Open Contours
Chandra Kambhamettu, Maureen Stone 0001
ACCV (1)3
2006 A Hierarchical Method for 3D Rigid Motion Estimation
Thitiwan Srinark, Chandra Kambhamettu, Maureen Stone 0001
ACCV (2)3
2006 Prospects for a Silent Speech Interface using Ultrasound Imaging
abstract
The feasibility of a silent speech interface using ultrasound (US) imaging and lip profile video is investigated by examining the quality of line spectral frequencies (LSF) derived from the image sequences. It is found that the data do not at present allow reliable identification of silences and fricatives, but that LSF's recovered from vocalized passages are compatible with the synthesis of intelligible speech
Bruce Denby, Yacine Oussar, Gérard Dreyfus, Maureen Stone 0001
ICASSP (1)4
2005 A tagged-cine MRI investigation of German vowels
abstract
This project investigates the articulatory properties of German vowels on the basis of tagged Cine-MRI data. German has 15 monophthongs, which are classified into seven tense-lax pairs. Understanding the phonetic correlate of vowel tenseness has proven elusive, partly due to the difficulty of obtaining global shape information for the tongue, especially the pharyngeal portion. Tagged Cine-MRI allows motion trajectories for individual tissue points to be tracked and strains to be calculated. From these data one can infer contraction patterns of the muscles of the tongue. The current project uses this technique to compare four tense and lax vowel pairs in terms of muscular compression and expansion patterns as well as root-blade motion. 1.
Marianne Pouplier, Maureen Stone 0001
INTERSPEECH2
2004 A General Framework for 2D Multiframe and 3D Surface-to-surface Motion Estimation
abstract
A general framework for 2D multiframe and 3D surface-to-surface motion estimation is presented in this paper. By viewing a 2D contour sequence as a pseudo 3D surface, we solve the motion estimation problem for 2D multiframe and 3D surface-to-surface in a general framework, by estimating the motion of a ”surface”. The deformation of a ”surface ” is modeled using spline-based motion. This spline-based motion model does not constrain the motion type in the temporal domain for 2D multiframe motion estimation. For 3D motion estimation, we focus on the relationship between the underlying nonrigid motion and 3D surface properties. The spline motion model provides our method certain advantages over other nonrigid shapebased methods. For example, we do not need approximation of the orthogonal parameterization. The small deformation constraint introduced by the previous surface-to-surface motion estimation methods is also relaxed in our method. Experiments on both synthetic and real motion are presented in this paper. 1
Chandra Kambhamettu, Maureen Stone 0001
BMVC3
2004 Speech synthesis from real time ultrasound images of the tongue
abstract
A machine learning technique is used to match reconstructed tongue contours in 30 frame per second ultrasound images to speaker vocal tract parameters obtained from a synchronized audio track. Speech synthesized using the learned parameters and noise as an activation function displays many of the time and frequency domain characteristics of the original audio, and, for isolated passages, is remarkably clear - although no articulators other than the tongue are included.
Bruce Denby, Maureen Stone 0001
ICASSP (1)2
2002 Dynamic programming method for temporal registration of three-dimensional tongue surface motion from multiple utterances
Changsheng Yang, Maureen Stone 0001
Speech Commun.2
1999 Automatic Extraction and Tracking of the Tongue Contours
abstract
Computerized analysis of the tongue surface movement can provide valuable information to speech and swallowing research. Ultrasound technology is currently the most attractive modality for the tongue imaging mainly because of its high video frame rate. However, problems with ultrasound imaging, such as noise and echo artifacts, refractions, and unrelated reflections pose significant challenges for computer analysis of the tongue images and hence specific methods must be developed. This paper presents a system that is developed for automatic extraction and tracking of the tongue surface movements from ultrasound image sequences. The ultrasound images are supplied by the head and transducer support system (HATS), which was developed in order to fix the head and support the transducer under the chin in a known position without disturbing speech. In this work, we propose a novel scheme for the analysis of the tongue images using deformable contours. We incorporate novel mechanisms to 1) impose speech related constraints on the deformations; 2) perform spatiotemporal smoothing using a contour postprocessing stage; 3) utilize optical flow techniques to speed up the search process; and 4) propagate user supplied information to the analysis of all image frames. We tested the system's performance qualitatively and quantitatively in consultation with speech scientists. Our system produced contours that are within the range of manual measurement variations. The results of our system are extremely encouraging and the system can be used in practical speech and swallowing research in the field of otolaryngology.
Yusuf Sinan Akgül, Chandra Kambhamettu, Maureen Stone 0001
IEEE Trans. Medical Imaging3
1998 Extraction and Tracking of the Tongue Surface from Ultrasound Image Sequences
abstract
This paper presents a system for automatic extraction and tracking of 2D contours of the tongue surfaces from digital ultrasound image sequences. The input to the system is provided by a Head and Transducer Support System (HATS), which is developed for use in ultrasound imaging of the tongue movement. We developed a novel active contour (snakes) model that uses several temporally adjacent images during the extraction of the tongue surface contour for an image frame. The user supplies an initial contour model for a single image frame in the whole sequence. Using optical flow and multi-resolution methods, this initial contour is then used to find the candidate contour points in the temporally immediate adjacent images. Subsequently, the new snake mechanism is applied to estimate optimal contours for each image frame using these candidate points. In turn, the extracted contours are used as models for the extraction process of new adjacent frames. Finally, the system uses a novel postprocessing technique to refine the positions of the contours. We tested the system on 11 different speech sequences, each containing about 25 images. Visual inspection of the detected contours by the speech experts shows that the results are very promising and this system can be effectively employed in speech and swallowing research.
Yusuf Sinan Akgül, Chandra Kambhamettu, Maureen Stone 0001
CVPR3
1998 Reconstructing the tongue surface from six cross-sectional contours: ultrasound data
Andrew J. Lundberg, Maureen Stone 0001
ICSLP2
1998 Analysis of the tongue surface movement using a spatiotemporally coherent deformable model
abstract
We present a system that extracts and tracks 2D mid-sagital tongue surfaces from ultrasound image sequences produced by a head and transducer support system (HATS). Such a system can be a valuable and practical tool for speech and swallowing research. We extend our previous work by introducing a deformable contour model to impose restrictions of spatiotemporal coherency by assuming that the 2D mid-sagital tongue surface contours sweep a coherent 3D structure in spatiotemporal 3D space. Another novel contribution of this paper is a dynamic programming minimization method to optimize the new energy functional, which can be employed in general snake minimization tasks. We tested this system on a number of speech and swallowing sequences each containing 24 to 30 frames. We verified with speech experts and found that our system produces promising results, especially for the swallowing sequences, which are more problematic.
Yusuf Sinan Akgül, Chandra Kambhamettu, Maureen Stone 0001
WACV3
1997 Three-dimensional coarticulatory strategies of tongue movement
abstract
This paper will present three-dimensional tongue “volumes,” reconstructed from three sagittal slices (left, mid, right) made using tagged cine MRI. The volumes will be animated to show CV movement from the consonants /k/ and /s/ to the vowels /i/, /a/, and /u/.
Maureen Stone 0001, Andrew J. Lundberg, Edward P. Davis, Rao P. Gullapalli, Moriel NessAiver
EUROSPEECH1
1997 Principal component analysis of cross sections of tongue shapes in vowel production
abstract
Images of the vocal tract provide the speech researcher with new and unique data. However, quantification of the “biologically important” features present in the image is a significant challenge because these features are typically complex, inherently variable and difficult to describe. The present study quantified cross-sectional tongue shape for vowels in a single plane (post-alveolar) using Principal Component Analysis. A single subject repeated eleven English vowels in two consonant contexts, five times each. Each of the resulting cross-sectional waveshapes (tokens) was represented by 70 samples. The 110 tokens were placed in a 110 × 70 matrix for PCA. The resulting principal components represent a basis set of orthonormal waveshapes with properties particularly suitable for the present project. The first two components accounted for 93% of the variance in the data. The loadings of the eleven vowels on the first two PCs indicated three distinct shape groups. These were examined statistically and found to represent high vowels, front vowels and back vowels. L'imagerie du conduit vocal fournit aux chercheurs en parole des données nouvelles et uniques. Cependant, savoir quantifier l'importance biologique des caractéristiques présentes dans l'image est un enjeu important, car ces caractéristiques sont typiquement complexes, intrinsèquement variables et difficiles à décrire. Ce papier présente l'analyse en composantes principales de la forme de la langue dans le plan coronal (post-alvéolaire) pour les 11 voyelles de l'anglais, répétées 5 fois, par un seul locuteur et dans deux contextes consonantiques. Chacun des contours dans le plan coronal a été représenté par 70 points. Les 110 contours ont été placés dans une matrice 110 × 70 pour l'analyse en composantes principales. Les composantes principales ainsi extraites constituent une base orthonormale dont les propriétés sont particulièrement intéressantes. Les deux premières composantes rendent compte de 93% de la variance mesurée sur les données. Les coefficients de pondération des 11 voyelles selon ces deux premières composantes font apparaître 3 groupes distincts. Ces groupes ont été analysés statistiquement, et se sont révélés représentatifs des voyelles hautes, des voyelles avant et des voyelles arrières.
Maureen Stone 0001, Moise H. Goldstein Jr.
Speech Commun.1
1996 A continuum mechanics representation of tongue deformation
Edward P. Davis, Andrew Douglas, Maureen Stone 0001
ICSLP3
1994 Tongue-palate interactions in consonants vs. vowels
Maureen Stone 0001, Andrew J. Lundberg
ICSLP1
1992 Representing the tongue surface with curve fits
Maureen Stone 0001, Subhash Lele
ICSLP1