VLDB 2026 Research / reviewers in the wild / expert
Suzanne Boyce
dblp:78/8760
· DBLP profile ↗
13ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-2105-2486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimized classification of accurate and misarticulated American English rhotic speech sounds for use in gamified real-time ultrasound biofeedback therapyabstractGamified biofeedback has the potential to help individuals improve speech motor patterns. For such systems, feedback provided in response to attempted changes must directly characterize performance improvements. This is particularly the case for ultrasound biofeedback therapy (UBT), which is increasingly used to visualize tongue movement for treating misarticulations of speech sounds (e.g., American English /ɹ/). Previous studies on ultrasound imaging of misarticulated /ɹ/ sounds have established that a single, time-dependent measured parameter, δ (relative difference between normalized tongue dorsum and blade displacements) can classify articulatory accuracy, matching human judgments with about 85% agreement. The δ parameter thus is an ideal basis for real-time gamified UBT. However, during ultrasound imaging of a speech production, there are multiple possible image frames from which δ could be automatically evaluated to represent production accuracy. Accordingly, to identify the optimal frame selection for δ evaluation, ultrasound image sequences depicting 2,944 productions of 15 distinct rhotic phrases were analyzed, while perceptual accuracy was judged by trained listeners. Four different selection criteria were compared, with classification performance and optimal thresholds evaluated using receiver-operating-characteristic curve analysis with 8-fold cross validation. For pre-vocalic and post-vocalic rhotic syllables, highest classification success rates resulted from evaluating δ at productions’ end. For pre-vocalic words preceded by the carrier phrase “a”, best classification resulted from evaluating δ at the estimated temporal midpoint of /ɹ/. Results provide a viable basis for gamified ultrasound biofeedback therapy based on real-time ultrasound measurements of tongue movements. Sarah A. Biehl, Sarah Dugan, Sarah R. Li, Reneé Seward, Michael A. Riley, Suzanne Boyce, T. Douglas Mast |
Speech Commun. | 6 |
| 2025 | Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal InsufficiencyabstractTraditional clinical approaches for assessing nasality, such as nasopharyngoscopy and nasometry, involve unpleasant experiences and are problematic for children. Speech Inversion (SI), a noninvasive technique, offers a promising alternative for estimating articulatory movement without the need for physical instrumentation. In this study, an SI system trained on nasalance data from healthy adults is augmented with source information from electroglottography and acoustically derived F0, periodic and aperiodic energy estimates as proxies for glottal control. This model achieves $16.92 \%$ relative improvement in Pearson Product-Moment Correlation (PPMC) compared to a previous SI system for nasalance estimation. To adapt the SI system for nasalance estimation in children with Velopharyngeal Insufficiency (VPI), the model initially trained on adult speech was fine-tuned using children with VPI data, yielding an $7.90 \%$ relative improvement in PPMC compared to its performance before fine-tuning. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
ASRU | 2 |
| 2025 | Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating NasalityabstractSpeech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue constrictions, called oral tract variables (TVs), which were later enhanced by including source information (periodic and aperiodic energies, and F0 frequency) as proxies for glottal control. Comparison of the nasometric measures with high-speed nasopharyngoscopy showed that nasalance can serve as ground truth, and that an SI system trained with it reliably recovers velum movement patterns for American English speakers. Here, two SI training approaches are compared: baseline models that estimate oral TVs and nasalance independently, and a synergistic model that combines oral TVs and source features with nasalance. The synergistic model shows relative improvements of 5% in oral TVs estimation and 9% in nasalance estimation compared to the baseline models. Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2023 | Speaker-independent Speech Inversion for Estimation of Nasalance
Yashish M. Siriwardena, Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede, Liran Oren |
INTERSPEECH | 3 |
| 2021 | An Automatic, Simple Ultrasound Biofeedback Parameter for Distinguishing Accurate and Misarticulated Rhotic SyllablesabstractCharacterizing accurate vs. misarticulated patterns of tongue movement using ultrasound can be challenging in real time because of the fast, independent movement of tongue regions. The usefulness of ultrasound for biofeedback speech therapy is limited because speakers must mentally track and compare differences between their tongue movement and available models. It is desirable to automate this interpretive task using a single parameter representing deviation from known accurate tongue movements. In this study, displacements recorded automatically by ultrasound image tracking were transformed into a single biofeedback parameter (time-dependent difference between blade and dorsum displacements). Receiver operating characteristic (ROC) curve analysis was used to evaluate this parameter as a predictor of production accuracy over a range of different vowel contexts with initial and final /r/ in American English. Areas under ROC curves were 0.8 or above, indicating that this simple parameter may provide useful real-time biofeedback on /r/ accuracy within a range of rhotic contexts. Sarah R. Li, Colin T. Annand, Sarah Dugan, Sarah M. Schwab, Kathryn J. Eary, Michael Swearengen, Sarah Stack, Suzanne Boyce, Michael A. Riley, T. Douglas Mast |
Interspeech | 8 |
| 2019 | Using Ultrasound Imaging to Create Augmented Visual Biofeedback for Articulatory Practice
Colin T. Annand, Maurice Lamb, Sarah Dugan, Sarah R. Li, Hannah M. Woeste, T. Douglas Mast, Michael A. Riley, Jack A. Masterson, Neeraja Mahalingam, Kathryn J. Eary, Caroline Spencer, Suzanne Boyce, Stephanie Jackson, Anoosha Baxi, Reneé Seward |
INTERSPEECH | 12 |
| 2013 | Speechmark acoustic landmark tool: application to voice pathology
Suzanne Boyce, Marisha Speights, Keiko Ishikawa, Joel MacAuslan |
INTERSPEECH | 1 |
| 2012 | SpeechMark: Landmark Detection Tool for Speech AnalysisabstractLandmark-based software tools are particularly suited to fast, automatic analysis of small, non-lexical differences in production of the same speech material by the same speaker. We are building a suite of independent applications and plugins as toolkits that make our landmark-based software system, SpeechMark, available to the wider scientific community. This will be achieved by extending existing software platforms with “plug-ins” that perform specific measures and report results to the user and by developing a MATLAB toolkit. These tools provide automatic summary statistics for measures of speech acoustics based on Stevens ’ paradigm of landmarks, points in an utterance around which information about articulatory events can be extracted. Index Terms: speech production, articulation, landmark, software. Suzanne Boyce, Harriet J. Fell, Joel MacAuslan |
INTERSPEECH | 1 |
| 2010 | An MRI-based articulatory and acoustic study of lateral sound in American EnglishabstractThe production of the lateral sounds generally involves a linguo-alveolar contact and one or two lateral channels along the parasagittal sides of the tongue. The acoustic effect of these articulatory features is not clearly understood. In this study, we compare two productions of /l/ in American English by one subject, one for a dark /l/ and the other for a light /l/. Three-dimensional vocal tract models derived from the magnetic resonance images were analyzed. It was shown that zeros in the vocal tract acoustic response are produced in the F3-F5 region in both /l/ productions, but the number of zeros and their frequencies are affected by the length of the linguo-alveolar contact and by the presence or absence of lateral linguopalatal contacts. The dark /l/ has one zero below 5 kHz, produced by the cross mode posterior to the linguo-alveolar contact, while the light /l/ has three zeros below 5 kHz, produced by the asymmetrical lateral channels, the supralingual cavity and the cross mode posterior to linguo-alveolar contact. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
ICASSP | 4 |
| 2007 | An articulatory and acoustic study of "retroflex" and "bunched" american English rhotic sound based on MRIabstractThe North American rhotic liquid has two maximally distinct articulatory variants, the classic ”retroflex” and the classic ”bunched” tongue postures. The evidence for acoustic differences between these two variants is reexamined using magnetic resonance images of the vocal tract in this study. Two subjects with similar vocal tract dimensions but different tongue postures for sustained /r/ are used. It is shown that these two variants have similar patterns of F1-F3 and zero frequencies. However, the ”retroflex” variant has a larger difference between F4 and F5 than the ”bunched” one (around 1400 Hz vs. around 700 Hz). This difference can be explained by the geometry differences between these two variants, in particular, the shorter and more forward palatal constriction of the ”retroflex” /r/ and the sharper transition between palatal constriction and its anterior and posterior cavities. This formant pattern difference is confirmed by measurement from acoustic data of several additional subjects. Xinhui Zhou, Carol Y. Espy-Wilson, Mark K. Tiede, Suzanne Boyce |
INTERSPEECH | 4 |
| 2005 | Modeling of the Front Cavity and Sublingual Space in American English Rhotic SoundsabstractThe production of American English (AE) /r/ sounds is variable but generally involves a large-volume front cavity. In some cases, there is also a sublingual cavity. Previous work has shown that the large front-cavity volume and the sublingual cavity are directly or indirectly responsible for the characteristically low frequency of the third formant (F3) of /r/. The entire front cavity is normally modeled as a single tube. The sublingual cavity, if present, is modeled as a side branch to the front cavity. However, given the dimensions of the front cavity, it is possible that high order acoustic modes are excited which may produce zeros and affect formant locations. Detailed information of the flow field involved can help to understand better and model accurately the front cavity acoustics. A finite element study of the flow field in the front cavity and sublingual cavity is described, using dimensions measured from MRI studies of subjects producing AE /r/. The results show that the large-volume front cavity is better modeled as a single tube with a side branch rather than as a single tube alone. The effective length of this side branch is further increased by the presence of a sublingual cavity, giving a zero in the range of F5 in the resulting spectrum. Carol Y. Espy-Wilson, Suzanne Boyce, Mark K. Tiede |
ICASSP (1) | 3 |
| 1997 | Acoustic modelling of American English /r/
Carol Y. Espy-Wilson, Shri Narayanan, Suzanne Boyce, Abeer Alwan |
EUROSPEECH | 3 |
| 1996 | Coarticulatory stability in american English /r/
Suzanne Boyce, Carol Y. Espy-Wilson |
ICSLP | 1 |