EDBT 2026 Demo / reviewers in the wild / expert
Dani Byrd
dblp:38/7167
· DBLP profile ↗
22ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-3319-5871ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the GlobeabstractWe present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the RAIL license at https://github.com/tiantiaf0627/voxlect. Tiantian Feng, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun, Yoonjeong Lee, Dani Byrd, Shri Narayanan |
KDD (1) | 8 |
| 2025 | Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
Tiantian Feng, Thanathai Lertpetchpun, Dani Byrd, Shri Narayanan |
INTERSPEECH | 3 |
| 2025 | Instantaneous changes in acoustic signals reflect syllable progression and cross-linguistic syllable variation
Haley Hsu, Dani Byrd, Khalil Iskarous, Louis Goldstein |
INTERSPEECH | 2 |
| 2025 | On the Relationship between Accent Strength and Articulatory Features
Sean Foley, Yoonjeong Lee, Dani Byrd, Shri Narayanan |
INTERSPEECH | 5 |
| 2025 | Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shri Narayanan |
INTERSPEECH | 3 |
| 2024 | State-of-the-art speech production MRI protocol for new 0.55 Tesla scanners
Prakash Kumar, Yongwan Lim, Sophia X. Cui, Christina Hagedorn, Dani Byrd, Uttam K. Sinha, Shri Narayanan, Krishna S. Nayak |
INTERSPEECH | 6 |
| 2021 | Leveraging Real-Time MRI for Illuminating Linguistic Velum Action
Miran Oh, Dani Byrd, Shri Narayanan |
Interspeech | 2 |
| 2021 | Who converges? Variation reveals individual speaker adaptabilityabstractLittle is known about the cognitive capacities underlying real-time accommodation in spoken language and how they may allow conversing speakers to adapt their speech production behaviors. This study first presents a simple attunement model that incorporates hypothesized capacities, with a focus on individual variability as one of those capacities. The model makes explicit predictions about observable convergence behaviors in interacting speakers, including that: i) the intrinsically more variable speaker of the two will be the one who converges to their partner, ii) this flexible speaker with higher baseline variability will exhibit a substantial decrease in variability and iii) a greater change in the variability between speaking solo and interacting with their partner. These predictions are supported by the results of the modeling simulations. To further test the model's predictions, we analyzed a behavioral dataset including acoustic and articulatory data from three pairs of interacting speakers participating in a maze navigation task as well as a like solo speech task. The amount of variability in the speech parameters of each dyad member was quantified using coefficient of variation. The experimental results parallel the simulation results, and taken together, this work indicates that structured variability is an illuminating index of individual speaker adaptability and convergence behavior. Yoon-Jeong Lee, Louis Goldstein, Benjamin Parrell, Dani Byrd |
Speech Commun. | 4 |
| 2017 | Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 10 |
| 2016 | Illustrating the Production of the International Phonetic Alphabet Sounds Using Fast Real-Time Magnetic Resonance Imaging
Asterios Toutios, Sajan Goud Lingala, Colin Vaz, Jangwon Kim, John H. Esling, Patricia A. Keating, Matthew Gordon, Dani Byrd, Louis Goldstein, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 8 |
| 2013 | Truncation of pharyngeal gesture in English diphthong [aɪ]abstractIt is well acknowledged that [a] in English diphthongs (e.g. [a] in “pie’d”) has a different formant structure from its closest corresponding monophthong (e.g. [a] in “pod”). The current study proposes that these two sounds share the same cognitive unit, i.e. the pharyngeal constriction gesture that produces [a], and the surface difference can be modeled as a consequence of truncating the same articulatory movement in time by the following palatal glide in the diphthongal environment. Formation of pharyngeal constriction gesture during the production of [a] in a diphthong and in its corresponding monophthong was observed in various timing contexts using Realtime MRI; and the collected production data were quantitatively analyzed using the direct image analysis (DIA) technique, which infers tissue movement by tracking pixel intensity change over time in regions of interest. Results support our truncation account in that: (1) formation time of pharyngeal constriction is significantly longer in monophthongs than in diphthongs; (2) this duration correlates with the resulting constriction degree; and (3) the resulting constriction degree predicts the acoustic difference in the F2 dimension as predicted by our hypothesis. Index Terms: English diphthongs, speech production, rtMRI Fang-Ying Hsieh, Louis Goldstein, Dani Byrd, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Velic coordination in French nasals: a real-time magnetic resonance imaging studyabstractProduction of nasal vowels in French, and nasal consonants in French and English, was examined using real-time magnetic resonance imaging (rtMRI). The coordination of velic and lin-gual gestures was found to be tightly controlled across differ-ent prosodic contexts in French nasals. Velum lowering in En-glish nasal consonants did not show the same control, although the timing of the corresponding lingual gestures varied with prosodic context in the same way as for French nasals, suggest-ing a coordinative relationship in which oral and velic articula-tors are consistently phased in French nasal production. These findings illustrate the utility of real-time MRI as a method for studying velic activity and articulatory coordination in vocalic and nasal phonology. Index Terms: speech production, velum, nasals, nasal vowels, French, articulation, real-time MRI Michael I. Proctor, Louis Goldstein, Adam C. Lammert, Dani Byrd, Asterios Toutios, Shri Narayanan |
INTERSPEECH | 4 |
| 2013 | The effect of word frequency and lexical class on articulatory-acoustic couplingabstractWord frequency and lexical class distinction between function and content words have been shown to significantly influence word production. In this paper, we use real-time magnetic resonance imaging to investigate the effect of word frequency and lexical class on articulatory characteristics (the articulator speed) as well as acoustic characteristics (F0 and short-term en-ergy) in word production. Multiple regression analyses showed that word frequency exhibits significantly higher correlation with articulatory and acoustic factors for content words com-pared to function words. A Granger causality analysis uncov-ered a causal relationship from articulatory speed to F0/energy for low-frequency content words. We further observed, us-ing functional canonical correlation analysis, a tight coupling of articulatory and acoustic characteristics for low-frequency content words. These results support the view that word fre-quency distinctly influences the production of function and con-tent words as manifested in their articulation and acoustics, as well as the dynamic coupling of these temporal streams. Index Terms: word frequency, lexical class, articulatory-acoustic coupling, real-time MRI, speech production. Vikram Ramanarayanan, Dani Byrd, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | Investigating articulatory setting - pauses, ready position, and rest - using real-time MRIabstractWe present a novel automatic procedure to analyze ―articulatory setting (AS) ‖ or ―basis of articulation ‖ using realtime magnetic resonance images (rt-MRI) of the human vocal tract recorded for read and spontaneously spoken speech. We extract relevant frames of inter-speech pauses (ISPs) and rest positions from MRI sequences of read and spontaneous speech and use automatically-extracted features to quantify areas of different regions of the vocal tract as well as the angle of the jaw. Significant differences were found between the ASs adopted for ISPs in read and spontaneous speech, as well as those between ISPs and absolute rest positions. We further contrast differences between ASs adopted when the person is ready to speak as opposed to an absolute rest position. Index Terms — speech production, real-time MRI, basis of articulation, articulatory setting, pause articulation, read speech, spontaneous speech. 1. Vikram Ramanarayanan, Dani Byrd, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 2 |
| 2008 | An analysis of vocal tract shaping in English sibilant fricatives using real-time magnetic resonance imagingabstractThis study uses real-time MRI to investigate shaping aspects of two English sibilant fricatives. The purpose of this article is to 1) develop linguistically meaningful quantitative measurements based on vocal tract features that robustly capture the shaping aspects of the two fricatives, and 2) provide qualitative analyses of fricative shaping. Data was recorded in both midsagittal and coronal planes. The proposed three quantitative measures of this study provide robust results in categorizing shape. The qualitative analyses describe tongue shape in terms of grooving and doming and they support previous research. Erik Bresch, Daylen Riggs, Louis Goldstein, Dani Byrd, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 4 |
| 2002 | Analysis of user behavior under error conditions in spoken dialogsabstractWe focus on developing an account of user behavior under error conditions, working with annotated data from real human-machine mixed initiative dialogs. In particular, we examine categories of error perception, user behavior under error, effect of user strategies on error recovery, and the role of user initiative in error situations. A conditional probability model smoothed by weighted ASR error rate is proposed. Results show that users discovering errors through implicit confirmations are less likely to get back on track (or succeed) and take a longer time in doing so than other forms of error discovery such as system reject and reprompts. Further successful user error-recovery strategies included more rephrasing, less contradicting, and a tendency to terminate error episodes (cancel and startover) than to attempt at repairing a chain of errors. JongHo Shin, Shri Narayanan, Laurie Gerber, Abe Kazemzadeh, Dani Byrd |
INTERSPEECH | 5 |
| 2001 | Politeness and frustration language in child-machine interactionsabstractChildren represent a potentially crucial user segment for conversational interfaces. Computer systems interacting with children need to be tailored for these users so that they will understand child intent and so that the child will have a positive and successful experience with the system. This study focuses on discourse analysis of spoken-language childmachine interactions. In particular, politeness and frustration markers were analyzed using a database of child-machine conversations obtained from 160 children using a computer game in a wizard-of-Oz set up. Results indicate that younger children less likely to use overt politeness markers and more polite information requests compared to the older ones, with no apparent gender differences. Younger children, on the other hand, expressed frustration verbally more than the older ones; furthermore, frustration language was more predominant in male children. Sudha Arunachalam, Dylan Gould, Elaine Andersen, Dani Byrd, Shri Narayanan |
INTERSPEECH | 4 |
| 1996 | Liquids in tamil
Shri Narayanan, Abigail Kaun, Dani Byrd, Peter Ladefoged, Abeer Alwan |
ICSLP | 3 |
| 1994 | Relations of sex and dialect to reduction
Dani Byrd |
Speech Communication | 1 |
| 1994 | Phonetic analyses of word and segment variation using the TIMIT corpus of American english
Patricia A. Keating, Dani Byrd, Edward Flemming, Yuichi Todaka |
Speech Commun. | 2 |
| 1992 | Sex, dialects, and reduction
Dani Byrd |
ICSLP | 1 |
| 1992 | Phonetic analyses of the TIMIT corpus of american English
Patricia A. Keating, B. Blankenship, Dani Byrd, Edward Flemming, Yuichi Todaka |
ICSLP | 3 |