EDBT 2026 Demo / reviewers in the wild / expert
Kyle Gorman
dblp:75/9248
· DBLP profile ↗
17ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0002-4233-6595ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 35% Information extraction and text analysis · 33% Probabilistic and Bayesian machine learning · 33% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection |
0.4 | 1 | 2020 | Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation |
0.4 | 1 | 2020 | Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing · EMNLP (1) 2020 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
statistical model comparison |
0.4 | 1 | 2020 | Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
multilingual language modeling |
0.4 | 1 | 2019 | What Kind of Language Is Hard to Language-Model? · ACL (1) 2019 |
Natural language and speech › Language models and text generation
evaluation of language models |
0.1 | 1 | 2019 | What Kind of Language Is Hard to Language-Model? · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging |
0.1 | 1 | 2019 | We Need to Talk about Standard Splits · ACL (1) 2019 |
Methods — techniques the papers use, named apart from their topics
k-fold cross-validation · 0.4bayesian statistics · 0.4surprisal-based difficulty estimation · 0.4statistical testing · 0.4replication study · 0.4mixed-effects modeling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A* shortest string decoding for non-idempotent semiringsabstractThe single shortest path algorithm is undefined for weighted finite-state automata over nonidempotent semirings because such semirings do not guarantee the existence of a shortest path.However, in non-idempotent semirings admitting an order satisfying a monotonicity condition (such as the plus-times or log semirings), the shortest string is well-defined.We describe an algorithm which finds the shortest string for a weighted non-deterministic automaton over such semirings using the backwards shortest distance of an equivalent deterministic automaton (DFA) as a heuristic for A* search performed over a companion idempotent semiring, This algorithm is proven to return the shortest string.There may be exponentially more states in the equivalent DFA, but the proposed algorithm needs to visit only a small fraction of them if determinization is performed "on the fly". Kyle Gorman, Cyril Allauzen |
EACL (1) | 1 |
| 2024 | Quantifying the Hyperparameter Sensitivity of Neural Networks for Character-level Sequence-to-Sequence TasksabstractHyperparameter tuning, the process of searching for suitable hyperparameters, becomes more difficult as the computing resources required to train neural networks continue to grow.This topic continues to receive little attention and discussion-much of it hearsaydespite its obvious importance.We attempt to formalize hyperparameter sensitivity using two metrics: similarity-based sensitivity and performance-based sensitivity.We then use these metrics to quantify two such claims: (1) transformers are more sensitive to hyperparameter choices than LSTMs and (2) transformers are particularly sensitive to batch size.We conduct experiments on two different characterlevel sequence-to-sequence tasks and find that, indeed, the transformer is slightly more sensitive to hyperparameters according to both of our metrics.However, we do not find that it is more sensitive to batch size in particular. Adam Wiemerslage, Kyle Gorman, Katharina von der Wense |
EACL (1) | 2 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 9 |
| 2021 | NeMo Inverse Text Normalization: From Development to ProductionabstractInverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text to improve the readability of the ASR output. Many state-of-the-art ITN systems use hand-written weighted finite-state transducer(WFST) grammars since this task has extremely low tolerance to unrecoverable errors. We introduce an open-source Python WFST-based library for ITN which enables a seamless path from development to production. We describe the specification of ITN grammar rules for English, but the library can be adapted for other languages. It can also be used for written-to-spoken text normalization. We evaluate the NeMo ITN library using a modified version of the Google Text normalization dataset. Yang Zhang 0089, Evelina Bakhturina, Kyle Gorman, Boris Ginsburg |
Interspeech | 3 |
| 2020 | Is the Best Better? Bayesian Statistical Model Comparison for Natural Language ProcessingabstractRecent work raises concerns about the use of standard splits to compare natural language processing models.We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to estimate the likelihood that one model will outperform the other, or that the two will produce practically equivalent results.We use this technique to rank six English part-ofspeech taggers across two data sets and three evaluation metrics. Piotr Szymanski, Kyle Gorman |
EMNLP (1) | 2 |
| 2020 | Massively Multilingual Pronunciation Modeling with WikiPronabstractWe introduce WikiPron, an open-source command-line tool for extracting pronunciation data from Wiktionary, a collaborative multilingual online dictionary. We first describe the design and use of WikiPron. We then discuss the challenges faced scaling this tool to create an automatically-generated database of 1.7 million pronunciations from 165 languages. Finally, we validate the pronunciation database by using it to train and evaluating a collection of generic grapheme-to-phoneme models. The software, pronunciation data, and models are all made available under permissive open-source licenses. Jackson L. Lee, Lucas F. E. Ashby, M. Elizabeth Garza, Yeonju Lee-Sikka, Sean Miller, Arya McCarthy, Kyle Gorman |
LREC | 8 |
| 2020 | UniMorph 3.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological paradigms for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. We have implemented several improvements to the extraction pipeline which creates most of our data, so that it is both more complete and more correct. We have added 66 new languages, as well as new parts of speech for 12 languages. We have also amended the schema in several ways. Finally, we present three new community tools: two to validate data for resource creators, and one to make morphological data available from the command line. UniMorph is based at the Center for Language and Speech Processing (CLSP) at Johns Hopkins University in Baltimore, Maryland. This paper details advances made to the schema, tooling, and dissemination of project resources since the UniMorph 2.0 release described at LREC 2018. Arya McCarthy, Christo Kirov, Matteo Grella, Amrit Nidhi, Patrick Xia 0002, Kyle Gorman, Ekaterina Vylomova, Sabrina J. Mielke, Garrett Nicolai, Miikka Silfverberg, Timofey Arkhangelskiy, Nataly Krizhanovsky, Andrew Krizhanovsky, Elena Klyachko, Alexey Sorokin, John Mansfield, Valts Ernstreits, Yuval Pinter, Cassandra L. Jacobs, Ryan Cotterell, Mans Hulden, David Yarowsky |
LREC | 6 |
| 2019 | We Need to Talk about Standard SplitsabstractIt is standard practice in speech & language technology to rank systems according to performance on a test set held out for evaluation.However, few researchers apply statistical tests to determine whether differences in performance are likely to arise by chance, and few examine the stability of system ranking across multiple training-testing splits.We conduct replication and reproduction experiments with nine part-of-speech taggers published between 2000 and 2018, each of which reports state-of-the-art performance on a widely-used "standard split".We fail to reliably reproduce some rankings using randomly generated splits.We suggest that randomly generated splits should be used in system comparison. Kyle Gorman, Steven Bedrick |
ACL (1) | 1 |
| 2019 | What Kind of Language Is Hard to Language-Model?abstractHow language-agnostic are current state-ofthe-art NLP tools?Are there some types of language that are easier to model with current methods?In prior work (Cotterell et al., 2018) we attempted to address this question for language modeling, and observed that recurrent neural network language models do not perform equally well over all the highresource European languages found in the Europarl corpus.We speculated that inflectional morphology may be the primary culprit for the discrepancy.In this paper, we extend these earlier experiments to cover 69 languages from 13 language families using a multilingual Bible corpus.Methodologically, we introduce a new paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at-least-pairwise parallel corpora.In other words, the model is aware of inter-sentence variation and can handle missing data.Exploiting this model, we show that "translationese" is not any easier to model than natively written language in a fair comparison.Trying to answer the question of what features difficult languages have in common, we try and fail to reproduce our earlier (Cotterell et al., 2018) observation about morphological complexity and instead reveal far simpler statistics of the data that seem to drive complexity in a much larger sample. Difficulty estimation from sentence surprisal Sabrina J. Mielke, Ryan Cotterell, Kyle Gorman, Brian Roark, Jason Eisner |
ACL (1) | 3 |
| 2019 | Weird Inflects but OK: Making Sense of Morphological Generation ErrorsabstractWe conduct a manual error analysis of the CoNLL-SIGMORPHON 2017 Shared Task on Morphological Reinflection.In this task, systems are given a word in citation form (e.g., hug) and asked to produce the corresponding inflected form (e.g., the simple past hugged).This design lets us analyze errors much like we might analyze children's production errors.We propose an error taxonomy and use it to annotate errors made by the top two systems across twelve languages.Many of the observed errors are related to inflectional patterns sensitive to inherent linguistic properties such as animacy or affect; many others are failures to predict truly unpredictable inflectional behaviors.We also find nearly one quarter of the residual "errors" reflect errors in the gold data. Kyle Gorman, Arya McCarthy, Ryan Cotterell, Ekaterina Vylomova, Miikka Silfverberg, Magdalena Markowska |
CoNLL | 1 |
| 2019 | Unified Verbalization for Speech Recognition & Synthesis Across Languages
Sandy Ritchie, Richard Sproat, Kyle Gorman, Daan van Esch, Christian Schallhart, Nikos Bampounis, Benoît Brard, Jonas Fromseier Mortensen, Millie Holt, Eoin Mahon |
INTERSPEECH | 3 |
| 2019 | Neural Models of Text Normalization for Speech ApplicationsabstractMachine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 King Ave. For this task, state-of-the-art industrial systems depend heavily on hand-written language-specific grammars. We propose neural network models that treat text normalization for TTS as a sequence-to-sequence problem, in which the input is a text token in context, and the output is the verbalization of that token. We find that the most effective model, in accuracy and efficiency, is one where the sentential context is computed once and the results of that computation are combined with the computation of each token in sequence to compute the verbalization. This model allows for a great deal of flexibility in terms of representing the context, and also allows us to integrate tagging and segmentation into the process. These models perform very well overall, but occasionally they will predict wildly inappropriate verbalizations, such as reading 3 cm as three kilometers. Although rare, such verbalizations are a major issue for TTS applications. We thus use finite-state covering grammars to guide the neural models, either during training and decoding, or just during decoding, away from such “unrecoverable” errors. Such grammars can largely be learned from data. Hao Zhang 0010, Richard Sproat, Axel H. Ng, Felix Stahlberg, Xiaochang Peng, Kyle Gorman, Brian Roark |
Comput. Linguistics | 6 |
| 2018 | Improving homograph disambiguation with supervised machine learning
Kyle Gorman, Gleb Mazovetskiy, Vitaly Nikolaev |
LREC | 1 |
| 2017 | Minimally supervised written-to-spoken text normalizationabstractText normalization is the task of converting from a written representation into a representation of how the text is to be spoken. For most real-world speech applications, the text normalization engine is developed mostly by hand. For example, a hand-built grammar may be used to enumerate possible ways to say a given token in a given language, and a statistical model used to select the most appropriate verbalizations in context. We examine the tradeoffs associated with using more or less language-specific knowledge for text normalization. In the most data-rich scenario, we have access to a carefully constructed hand-built normalization grammar that for any given token will produce a lattice of all possible verbalizations for that token. We assume a parallel corpus of aligned written-spoken utterances. As a substitute for the hand-built grammar, we consider a language-universal normalization covering grammar, where the developer merely needs to provide a set of lexical items particular to the language. As a substitute for the aligned corpus, we consider a scenario where one only has the spoken side, and the corresponding written side is “hallucinated” by composing the spoken side with the inverted normalization grammar. We report performance of the above scenarios on experiments with English and Russian. Axel H. Ng, Kyle Gorman, Richard Sproat |
ASRU | 2 |
| 2016 | Minimally Supervised Number NormalizationabstractWe propose two models for verbalizing numbers, a key component in speech recognition and synthesis systems. The first model uses an end-to-end recurrent neural network. The second model, drawing inspiration from the linguistics literature, uses finite-state transducers constructed with a minimal amount of training data. While both models achieve near-perfect performance, the latter model can be trained using several orders of magnitude less data than the former, making it particularly useful for low-resource languages. Kyle Gorman, Richard Sproat |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Discriminative pronunciation modeling for dialectal speech recognitionabstractSpeech recognizers are typically trained with data from a stan-dard dialect and do not generalize to non-standard dialects. Mis-match mainly occurs in the acoustic realization of words, which is represented by acoustic models and pronunciation lexicon. Standard techniques for addressing this mismatch are generative in nature and include acoustic model adaptation and expansion of lexicon with pronunciation variants, both of which have lim-ited effectiveness. We present a discriminative pronunciation model whose parameters are learned jointly with parameters from the language models. We tease apart the gains from mod-eling the transitions of canonical phones, the transduction from surface to canonical phones, and the language model. We report experiments on African American Vernacular English (AAVE) using NPR’s StoryCorps corpus. Our models improve the per-formance over the baseline by about 2.1 % on AAVE, of which 0.6 % can be attributed to the pronunciation model. The model learns the most relevant phonetic transformations for AAVE speech. Index Terms: large vocabulary speech recognition, dialec-tal speech recognition, pronunciation modeling, discriminative training 1. Maider Lehr, Kyle Gorman, Izhak Shafran |
INTERSPEECH | 2 |
| 2007 | Perception of disfluency: language differences and listener biasabstractThis paper describes a crosslinguistic disfluency perception experiment. We tested the recognizability of pause fillers and partial words in English, German and Mandarin. Subjects were speakers of English with no knowledge of Mandarin or German. We found that subjects could identify disfluent from fluent utterances at a level above chance. Pause fillers were easier to identify than partial words. Accuracy rates were highest for English, followed by German and then Mandarin. Although German accuracy rates were higher than those for Mandarin, discriminability analysis suggests that this is due to conservative bias towards false negatives rather than non-recognition of the acoustic material. The fact that subjects could identify disfluent speech in languages they did not know shows that there are real phonetic crosslinguistic cues to disfluency. Index Terms: crosslinguistic perception, disfluency, pause filler, partial words. Catherine Lai, Kyle Gorman, Jiahong Yuan, Mark Y. Liberman |
INTERSPEECH | 2 |