EDBT 2026 Demo / reviewers in the wild / expert
Mark K. Tiede
dblp:91/9336 · also Mark Tiede
· DBLP profile ↗
25ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-6118-0776ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal InsufficiencyabstractTraditional clinical approaches for assessing nasality, such as nasopharyngoscopy and nasometry, involve unpleasant experiences and are problematic for children. Speech Inversion (SI), a noninvasive technique, offers a promising alternative for estimating articulatory movement without the need for physical instrumentation. In this study, an SI system trained on nasalance data from healthy adults is augmented with source information from electroglottography and acoustically derived F0, periodic and aperiodic energy estimates as proxies for glottal control. This model achieves $16.92 \%$ relative improvement in Pearson Product-Moment Correlation (PPMC) compared to a previous SI system for nasalance estimation. To adapt the SI system for nasalance estimation in children with Velopharyngeal Insufficiency (VPI), the model initially trained on adult speech was fine-tuned using children with VPI data, yielding an $7.90 \%$ relative improvement in PPMC compared to its performance before fine-tuning. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
ASRU | 4 |
| 2025 | Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating NasalityabstractSpeech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue constrictions, called oral tract variables (TVs), which were later enhanced by including source information (periodic and aperiodic energies, and F0 frequency) as proxies for glottal control. Comparison of the nasometric measures with high-speed nasopharyngoscopy showed that nasalance can serve as ground truth, and that an SI system trained with it reliably recovers velum movement patterns for American English speakers. Here, two SI training approaches are compared: baseline models that estimate oral TVs and nasalance independently, and a synergistic model that combines oral TVs and source features with nasalance. The synergistic model shows relative improvements of 5% in oral TVs estimation and 9% in nasalance estimation compared to the baseline models. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2023 | Enhancing Speech Articulation Analysis Using A Geometric Transformation of the X-ray Microbeam DatasetabstractAccurate analysis of speech articulation is crucial for speech analysis.However, X-Y coordinates of articulators strongly depend on the anatomy of the speakers and variability of pellet placements, and existing methods for mapping anatomical landmarks in the X-ray Microbeam Dataset (XRMB) fail to capture the entire anatomy of the vocal tract.In this paper, we propose a new geometric transformation that improves the accuracy of these measurements.Our transformation maps anatomical landmarks' X-Y coordinates along the midsagittal plane onto six relative measures: Lip Aperture (LA), Lip Protrusion (LP), Tongue Body Constriction Location (TBCL), Degree (TBCD), Tongue Tip Constriction Location (TTCL), and Degree (TTCD).Our novel contribution is the extension of the palate trace towards the inferred anterior pharyngeal line, which improves measurements of tongue body constriction. Ahmed Adel Attia, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2023 | Speaker-independent Speech Inversion for Estimation of Nasalance
Yashish M. Siriwardena, Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede, Liran Oren |
INTERSPEECH | 4 |
| 2023 | Parameters of unit-based measures of speech rate
Sam Tilsen, Mark K. Tiede |
Speech Commun. | 2 |
| 2022 | Speech Driven Tongue AnimationabstractAdvances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not provide realistic animation of the inner mouth. This paper addresses the problem of speech-driven inner mouth animation. Obtaining performance capture data of the tongue and jaw from video alone is difficult because the inner mouth is only partially observable during speech. In this work, we introduce a large-scale speech and mocap dataset that focuses on capturing tongue, jaw, and lip motion. This dataset enables research using data-driven techniques to generate realistic inner mouth animation from speech. We then propose a deep-learning based method for accurate and generalizable speech to tongue and jaw animation, and evaluate several encoder-decoder network architectures and audio feature encoders. We find that recent self-supervised deep learning based audio feature encoders are robust, generalize well to unseen speakers and content, and work best for our task. To demonstrate the practical application of our approach, we show animations on high-quality parametric 3D face models driven by the landmarks generated from our speech-to-tongue animation method. Denis Tomè, Carsten Stoll, Mark K. Tiede, Kevin Munhall, Alex Hauptmann 0001, Iain A. Matthews |
CVPR | 4 |
| 2021 | Importance of Parasagittal Sensor Information in Tongue Motion Capture Through a Diphonic AnalysisabstractOur study examines the information obtained by adding two parasagittal sensors to the standard midsagittal configuration of an Electromagnetic Articulography (EMA) observation of lingual articulation. In this work, we present a large and phonetically balanced corpus obtained from an EMA recording session of a single English native speaker reading 1899 sentences from the Harvard and TIMIT corpora. According to a statistical analysis of the diphones produced during the recording session, the motion captured by the parasagittal sensors has a low correlation to the midsagittal sensors in the mediolateral direction. We perform a geometric analysis of the lateral tongue by the measure of its width and using a proxy of the tongue’s curvature that is computed using the Menger curvature. To provide a better understanding of the tongue sensor motion we present dynamic visualizations of all diphones. Finally, we present a summary of the velocity information computed from the tongue sensor information. Sarah Taylor, Mark K. Tiede, Alex Hauptmann 0001, Iain A. Matthews |
Interspeech | 3 |
| 2017 | Hybrid convolutional neural networks for articulatory and acoustic information based speech recognition
Vikramjit Mitra, Ganesh Sivaraman, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Mark K. Tiede |
Speech Commun. | 6 |
| 2016 | Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
Ganesh Sivaraman, Vikramjit Mitra, Hosung Nam, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2015 | Speech planning in 4-year-old children versus adults: acoustic and articulatory analysesabstractThis study investigates speech motor control in 4-year-old Canadian French children in comparison with adults.It focuses on measures of token-to-token variability in the production of isolated vowels and on anticipatory extrasyllabic coarticulation within V 1 -C-V 2 sequences.Acoustic and ultrasound articulatory data were recorded.Acoustic data from 20 children and 10 adults have been analyzed.Thus far, ultrasound data have been analyzed from a subset of these participants: 6 children and 2 adults.In agreement with former studies, token-to-token variability was greater in children than in adults.Strong anticipation of V 2 in V 1 was found in all adults, but not in children.Most of the children showed no anticipation at all and some of them showed a small amount of anticipation along the antero-posterior dimension only, manifested in the acoustic F2 dimension.These results are interpreted as evidence for the immaturity of children's speech motor control from two perspectives: insufficiently stable motor control patterns for vowel production, and a lack of effectiveness in anticipating forthcoming gestures.In line with theories of optimal motor control, anticipatory coarticulation is assumed to be based on the use of internal models of the speech apparatus and the increasing maturation of these representations as speech develops. Guillaume Barbier, Pascal Perrier, Lucie Ménard, Yohan Payan, Mark K. Tiede, Joseph S. Perkell |
INTERSPEECH | 5 |
| 2015 | Analysis of coarticulated speech using estimated articulatory trajectories
Ganesh Sivaraman, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2013 | Speech planning as an index of speech motor control maturityabstractInternational audience Guillaume Barbier, Pascal Perrier, Lucie Ménard, Yohan Payan, Mark K. Tiede, Joseph S. Perkell |
INTERSPEECH | 5 |
| 2013 | Correlates of contrastive focus in congenitally blind adults and sighted adults
Lucie Ménard, Annie Leclerc, Mark K. Tiede, Amélie Prémont, Christine Turgeon, Paméla Trudeau-Fisette, Dominique Côté |
INTERSPEECH | 3 |
| 2011 | Biomechanical Tongue Models: An Approach to Studying Inter-Speaker VariabilityabstractSpeakers of a given language vary with respect to their acoustics, articulation, and motor commands. This variation is driven by a variety of influences, such as emotional states, communicative interaction, and individual properties of the vocal tract. In this work we focus on the latter. First, we build speaker-specific biomechanical tongue models. Second, we discuss the impact of the relative position of the bending in the vocal tract on the basis of extensive simulations with two different models. We focus on /i,a,u/ by defining target regions in the acoustic space, and discuss the corresponding speaker-specific articulatory and motor command variability observed. Ralf Winkler, Susanne Fuchs, Pascal Perrier, Mark K. Tiede |
INTERSPEECH | 4 |
| 2010 | An MRI-based articulatory and acoustic study of lateral sound in American EnglishabstractThe production of the lateral sounds generally involves a linguo-alveolar contact and one or two lateral channels along the parasagittal sides of the tongue. The acoustic effect of these articulatory features is not clearly understood. In this study, we compare two productions of /l/ in American English by one subject, one for a dark /l/ and the other for a light /l/. Three-dimensional vocal tract models derived from the magnetic resonance images were analyzed. It was shown that zeros in the vocal tract acoustic response are produced in the F3-F5 region in both /l/ productions, but the number of zeros and their frequencies are affected by the length of the linguo-alveolar contact and by the presence or absence of lateral linguopalatal contacts. The dark /l/ has one zero below 5 kHz, produced by the cross mode posterior to the linguo-alveolar contact, while the light /l/ has three zeros below 5 kHz, produced by the asymmetrical lateral channels, the supralingual cavity and the cross mode posterior to linguo-alveolar contact. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
ICASSP | 3 |
| 2010 | A procedure for estimating gestural scores from natural speechabstractAbstract * Speech can be represented as a constellation of constricting events, gestures, , which are defined at distinct vocal tract sites, in the form of a gestural score.. Gestures and their output trajectories, tract variables, , which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given an utterance, the Haskins Laboratories Task Dynamics and Application (TADA) model was used to generate its prototype gestural score and the corresponding synthetic acoustic output. An optimal gestural score was estimated through iterative time-warping processes such that the distance between original and TADA-synthesized synthesized speech is minimized. We compared the performance of our approach to that of a conventional dynamic time warping procedure using Log-Spectral and Itakura Distance measures. We also performed a word recognition experiment using the gestural annotations to show that the gestural scores are suitable for word recognition. Hosung Nam, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2009 | Comparison of vowel structures of Japanese and English in articulatory and auditory spacesabstractIn previous work [1] we investigated the vowel structures of Japanese in both articulatory space and auditory perceptual space using Laplacian eigenmaps, and examined relations between speech production and perception. The results showed that the inherent structures of Japanese vowels were consistent in the two spaces. To verify whether such a property generalizes to other languages, we use the same approach to investigate the more crowded English vowel space. Results show that the vowel structure reflects the articulatory features for both languages. The degree of tongue-palate approximation is the most important feature for vowels, followed by the open ratio of the mouth to oral cavity. The topological relations of the vowel structures are consistent with both the articulatory and auditory perceptual spaces; in particular the lip-protruded vowel /UW / of English was distinct from the unrounded Japanese /�/. The rhotic vowel /ER / was located apart from the surface constructed by the other vowels, where the same phenomena appeared in both spaces. Index Terms: vowels, speech production, speech perception 1. Mark K. Tiede, Jiahong Yuan |
INTERSPEECH | 2 |
| 2007 | An articulatory and acoustic study of "retroflex" and "bunched" american English rhotic sound based on MRIabstractThe North American rhotic liquid has two maximally distinct articulatory variants, the classic ”retroflex” and the classic ”bunched” tongue postures. The evidence for acoustic differences between these two variants is reexamined using magnetic resonance images of the vocal tract in this study. Two subjects with similar vocal tract dimensions but different tongue postures for sustained /r/ are used. It is shown that these two variants have similar patterns of F1-F3 and zero frequencies. However, the ”retroflex” variant has a larger difference between F4 and F5 than the ”bunched” one (around 1400 Hz vs. around 700 Hz). This difference can be explained by the geometry differences between these two variants, in particular, the shorter and more forward palatal constriction of the ”retroflex” /r/ and the sharper transition between palatal constriction and its anterior and posterior cavities. This formant pattern difference is confirmed by measurement from acoustic data of several additional subjects. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
INTERSPEECH | 3 |
| 2005 | Modeling of the Front Cavity and Sublingual Space in American English Rhotic SoundsabstractThe production of American English (AE) /r/ sounds is variable but generally involves a large-volume front cavity. In some cases, there is also a sublingual cavity. Previous work has shown that the large front-cavity volume and the sublingual cavity are directly or indirectly responsible for the characteristically low frequency of the third formant (F3) of /r/. The entire front cavity is normally modeled as a single tube. The sublingual cavity, if present, is modeled as a side branch to the front cavity. However, given the dimensions of the front cavity, it is possible that high order acoustic modes are excited which may produce zeros and affect formant locations. Detailed information of the flow field involved can help to understand better and model accurately the front cavity acoustics. A finite element study of the flow field in the front cavity and sublingual cavity is described, using dimensions measured from MRI studies of subjects producing AE /r/. The results show that the large-volume front cavity is better modeled as a single tube with a side branch rather than as a single tube alone. The effective length of this side branch is further increased by the presence of a sublingual cavity, giving a zero in the range of F5 in the resulting spectrum. Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede |
ICASSP (1) | 4 |
| 2003 | Acoustic modeling of american English lateral approximantsabstractA vocal tract model for an American English /l / production with lateral channels and a supralingual side branch has been developed. Acoustic modeling of an /l / production using MRI-derived vocal tract dimensions shows that both the lateral channels and the supralingual side branch contribute to the production of zeros in the F3 to F5 frequency range, thereby resulting in pole-zero clusters around 2-5 kHz in the spectrum of the /l / sound. 1. Carol Y. Espy-Wilson, Mark K. Tiede |
INTERSPEECH | 3 |
| 1998 | An MRI study on the relationship between oral cavity shape and larynx position
Kiyoshi Honda, Mark K. Tiede |
ICSLP | 2 |
| 1997 | A parametric three-dimensional model of the vocal-tract based on MRI dataabstractTwenty four three-dimensional (3D) vocal-tract (VT) shapes extracted from MRI data are used to derive a parametric model for the vocal-tract. The method is as follows: first, each 3D VT shape is sampled using a semi-cylindrical grid whose position is determined by reference points based on the VT anatomy. After that, the VT projections onto each plane of the grid are represented by their two main components obtained via principal component analysis (PCA). PCA is once again used to parametrize the sequences of coefficients that represent the sections along the tract. It was verified that the first four components can explain about 90% of the total variance of the observed shapes. Following this procedure, 3D VT shapes are approximated by linear combinations of four 3D basis functions. Finally, it is shown that the four parameters of the model can be estimated from the VT midsagittal profiles. Hani Yehia, Mark K. Tiede |
ICASSP | 2 |
| 1996 | An MRI-based analysis of the English /r/ and /l/ articulations
Shinobu Masaki, Reiko Akahane-Yamada, Mark K. Tiede, Yasuhiro Shimada, Ichiro Fujimoto |
ICSLP | 3 |
| 1994 | Extracting articulator movement parameters from a videodisc-based cineradiographic database
Mark K. Tiede, Eric Vatikiotis-Bateson |
ICSLP | 1 |
| 1994 | Phoneme extraction using via point estimation of real speech
Eric Vatikiotis-Bateson, Mark K. Tiede, Yasuhiro Wada, Vincent L. Gracco, Mitsuo Kawato |
ICSLP | 2 |