VLDB 2026 Research / reviewers in the wild / expert
Arya McCarthy
dblp:219/5712 · also Arya D. McCarthy
· DBLP profile ↗
22ranked-venue papers
8as first author
7since 2021 · last 2023
0000-0001-9440-8792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Information extraction and text analysis · 37% Generative modeling · 18% Motion planning and robot control · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 50% Bioinformatics and computational biology · 50% | |
| Theoretical computer science
2 papers |
Computational complexity · 82% Information theory · 18% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
morphological analysis |
0.9 | 2 | 2020 | Predicting Declension Class from Form and Meaning · ACL 2020 Unsupervised Morphological Paradigm Completion · ACL 2020 |
Machine learning › Generative modeling
energy-based model |
0.6 | 1 | 2022 | On the Uncomputability of Partition Functions in Energy-Based Sequence Models · ICLR 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
partition function |
0.6 | 1 | 2022 | On the Uncomputability of Partition Functions in Energy-Based Sequence Models · ICLR 2022 |
Robotics › Motion planning and robot control › trajectory planning
time allocation |
0.6 | 1 | 2022 | A Major Obstacle for NLP Research: Let's Talk about Time Allocation! · EMNLP 2022 |
Computational complexity › computability theory
uncomputability |
0.6 | 1 | 2022 | On the Uncomputability of Partition Functions in Energy-Based Sequence Models · ICLR 2022 |
Natural language and speech › Machine translation
neural machine translation |
0.4 | 1 | 2020 | Addressing Posterior Collapse with Mutual Information for Improved Variational Neural Machine Translation · ACL 2020 |
Natural language and speech › Information extraction and text analysis › morphological analysis
paradigm completion |
0.4 | 1 | 2020 | Unsupervised Morphological Paradigm Completion · ACL 2020 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2020 | Addressing Posterior Collapse with Mutual Information for Improved Variational Neural Machine Translation · ACL 2020 |
Computational social science and digital humanities
historical linguistics |
0.4 | 1 | 2020 | Measuring the Similarity of Grammatical Gender Systems by Comparing Partitions · EMNLP (1) 2020 |
Bioinformatics and computational biology › phylogenetics
phylogenetic inference |
0.4 | 1 | 2020 | Measuring the Similarity of Grammatical Gender Systems by Comparing Partitions · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › language model interpretability
linguistic representation analysis |
0.4 | 1 | 2019 | Meaning to Form: Measuring Systematicity as Information · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning
mutual information |
0.4 | 1 | 2019 | Meaning to Form: Measuring Systematicity as Information · ACL (1) 2019 |
Natural language and speech › Language models and text generation
low-resource language processing |
0.1 | 1 | 2020 | Unsupervised Morphological Paradigm Completion · ACL 2020 |
Information theory › information measures
mutual information |
0.1 | 1 | 2020 | Predicting Declension Class from Form and Meaning · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
information-theoretic metrics · 0.9information-theoretic analysis · 0.9community detection · 0.9cluster evaluation · 0.9mutual information · 0.8multi-task generalized linear model · 0.5BERT features · 0.5transformer · 0.4inflection generation · 0.4edit tree retrieval · 0.4conditional variational autoencoder · 0.4recurrent neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained ModelsabstractAbteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Abteen Ebrahimi, Arya McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez Lugo, Rolando Coto-Solano, Katharina Kann |
EACL | 2 |
| 2022 | Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich LanguagesabstractThis paper presents a detailed foundational empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian, and a multi-faceted distributional analysis of the underlying word-formation processes that can aid in their compositional translation, tagging, parsing, language modeling, and other NLP tasks. Given that out-of-vocabulary (OOV) words generally present a key open challenge to NLP and machine translation systems, especially toward the lower limit of resource availability, there are useful practical insights, as well as corpus-linguistic insights, from both a detailed manual and automatic taxonomic analysis of the types, multidimensional properties, and processing potential for multiple representative OOV data samples. Georgie Botev, Arya McCarthy, Winston Wu, David Yarowsky |
COLING | 2 |
| 2022 | A Major Obstacle for NLP Research: Let's Talk about Time Allocation!abstractThe field of natural language processing (NLP) has grown over the last few years: conferences have become larger, we have published an incredible amount of papers, and state-of-the-art research has been implemented in a large variety of customer-facing products.However, this paper argues that we have been less successful than we should have been and reflects on where and how the field fails to tap its full potential.Specifically, we demonstrate that, in recent years, subpar time allocation has been a major obstacle for NLP research.We outline multiple concrete problems together with their negative consequences and, importantly, suggest remedies to improve the status quo.We hope that this paper will be a starting point for discussions around which common practices are -or are not -beneficial for NLP research. Katharina Kann, Shiran Dudy, Arya McCarthy |
EMNLP | 3 |
| 2022 | On the Uncomputability of Partition Functions in Energy-Based Sequence Models
Chu-Cheng Lin, Arya McCarthy |
ICLR | 2 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 91 |
| 2022 | Hong Kong: Longitudinal and Synchronic Characterisations of Protest News between 1998 and 2020abstractThis paper showcases the utility and timeliness of the Hong Kong Protest News Dataset, a highly curated collection of news articles from diverse news sources, to investigate longitudinal and synchronic news characterisations of protests in Hong Kong between 1998 and 2020. The properties of the dataset enable us to apply natural language processing to its 4522 articles and thereby study patterns of journalistic practice across newspapers. This paper sheds light on whether depth and/or manner of reporting changed over time, and if so, in what ways, or in response to what. In its focus and methodology, this paper helps bridge the gap between “validity-focused methodological debates” and the use of computational methods of analysis in the social sciences. Arya McCarthy, Giovanna Maria Dora Dore |
LREC | 1 |
| 2021 | Jump-Starting Item Parameters for Adaptive Language TestsabstractA challenge in designing high-stakes language assessments is calibrating the test item difficulties, either a priori or from limited pilot test data.While prior work has addressed 'cold start' estimation of item difficulties without piloting, we devise a multi-task generalized linear model with BERT features to jump-start these estimates, rapidly improving their quality with as few as 500 test-takers and a small sample of item exposures (≈6 each) from a large item bank (≈4,000 items).Our joint model provides a principled way to compare test-taker proficiency, item difficulty, and language proficiency frameworks like the Common European Framework of Reference (CEFR).This also enables new item difficulty estimates without piloting them first, which in turn limits item exposure and thus enhances test security.Finally, using operational data from the Duolingo English Test, a high-stakes English proficiency test, we find that difficulty estimates derived using this method correlate strongly with lexico-grammatical features that correlate with reading complexity. Arya McCarthy, Kevin P. Yancey, Geoffrey T. LaFlair, Jesse Egbert, Manqian Liao, Burr Settles |
EMNLP (1) | 1 |
| 2020 | Unsupervised Morphological Paradigm CompletionabstractWe propose the task of unsupervised morphological paradigm completion.Given only raw text and a lemma list, the task consists of generating the morphological paradigms, i.e., all inflected forms, of the lemmas.From a natural language processing (NLP) perspective, this is a challenging unsupervised task, and high-performing systems have the potential to improve tools for low-resource languages or to assist linguistic annotators.From a cognitive science perspective, this can shed light on how children acquire morphological knowledge.We further introduce a system for the task, which generates morphological paradigms via the following steps: (i) EDIT TREE retrieval, (ii) additional lemma retrieval, (iii) paradigm size discovery, and (iv) inflection generation.We perform an evaluation on 14 typologically diverse languages.Our system outperforms trivial baselines with ease and, for some languages, even obtains a higher accuracy than minimally supervised systems. 1 Sé vigilante y confirma las otras cosas que están para morir , porque no he hallado tus obras bien acabadas delante de Dios .Acuérdate , pues , de lo que has recibido y oído ; guárdalo y arrepiéntete , pues si no velas vendré sobre ti como ladrón y no sabrás a qué hora vendré sobre ti .El vencedor será vestido de vestiduras blancas , y no borraré su nombre del libro de la vida , y confesaré su nombre delante de mi Padre y delante de sus ángeles .El que tiene oído , oiga lo que el Espíritu dice a las iglesias . Huiming Jin, Liwei Cai, Yihui Peng, Chen Xia, Arya McCarthy, Katharina Kann |
ACL | 5 |
| 2020 | Addressing Posterior Collapse with Mutual Information for Improved Variational Neural Machine TranslationabstractThis paper proposes a simple and effective approach to address the problem of posterior collapse in conditional variational autoencoders (CVAEs).It thus improves performance of machine translation models that use noisy or monolingual data, as well as in conventional settings.Extending Transformer and conditional VAEs, our proposed latent variable model measurably prevents posterior collapse by (1) using a modified evidence lower bound (ELBO) objective which promotes mutual information between the latent variable and the target, and (2) guiding the latent variable with an auxiliary bag-of-words prediction task.As a result, the proposed model yields improved translation quality compared to existing variational NMT models on WMT Ro↔En and De↔En.With latent variables being effectively utilized, our model demonstrates improved robustness over non-latent Transformer in handling uncertainty: exploiting noisy source-side monolingual data (up to +3.2 BLEU), and training with weakly aligned web-mined parallel data (up to +4.7 BLEU). Arya McCarthy, Xian Li 0003, Jiatao Gu |
ACL | 1 |
| 2020 | Predicting Declension Class from Form and MeaningabstractThe noun lexica of many natural languages are divided into several declension classes with characteristic morphological properties.Class membership is far from deterministic, but the phonological form of a noun and its meaning can often provide imperfect clues.Here, we investigate the strength of those clues.More specifically, we operationalize "strength" as measuring how much information, in bits, we can glean about declension class from knowing the form and meaning of nouns.We know that form and meaning are often also indicative of grammatical gender-which, as we quantitatively verify, can itself share information with declension class-so we also control for gender.We find for two Indo-European languages (Czech and German) that form and meaning share a significant amount of information with class (and contribute additional information beyond gender).The three-way interaction between class, form, and meaning (given gender) is also significant.Our study is important for two reasons: First, we introduce a new method that provides additional quantitative support for a classic linguistic finding that form and meaning are relevant for the classification of nouns into declensions.Second, we show not only that individual declension classes vary in the strength of their clues within a language, but also that the variations between classes vary across languages.The code is publicly available at https://github.com/ rycolab/declension-mi. Adina Williams, Tiago Pimentel, Hagen Blix, Arya McCarthy, Eleanor Chodroff, Ryan Cotterell |
ACL | 4 |
| 2020 | Neural Transduction for Multilingual Lexical TranslationabstractWe present a method for completing multilingual translation dictionaries.Our probabilistic approach can synthesize new word forms, allowing it to operate in settings where correct translations have not been observed in text (cf.cross-lingual embeddings).In addition, we propose an approximate Maximum Mutual Information (MMI) decoding objective to further improve performance in both many-to-one and one-to-one word level translation tasks where we use either multiple input languages for a single target language or more typical single language pair translation.The model is trained in a many-to-many setting, where it can leverage information from related languages to predict words in each of its many target languages.We focus on 6 languages: French, Spanish, Italian, Portuguese, Romanian, and Turkish.When indirect multilingual information is available, ensembling with mixture-of-experts as well as incorporating related languages leads to a 27% relative improvement in whole-word accuracy of predictions over a single-source baseline.To seed the completion when multilingual data is unavailable, it is better to decode with an MMI objective. Dylan Lewis, Winston Wu, Arya McCarthy, David Yarowsky |
COLING | 3 |
| 2020 | Measuring the Similarity of Grammatical Gender Systems by Comparing PartitionsabstractA grammatical gender system divides a lexicon into a small number of relatively fixed grammatical categories.How similar are these gender systems across languages?To quantify the similarity, we define gender systems extensionally, thereby reducing the problem of comparisons between languages' gender systems to cluster evaluation.We borrow a rich inventory of statistical tools for cluster evaluation from the field of community detection (Driver and Kroeber, 1932;Cattell, 1945), that enable us to craft novel information-theoretic metrics for measuring similarity between gender systems.We first validate our metrics, then use them to measure gender system similarity in 20 languages.Finally, we ask whether our gender system similarities alone are sufficient to reconstruct historical relationships between languages.Towards this end, we make phylogenetic predictions on the popular, but thorny, problem from historical linguistics of inducing a phylogenetic tree over extant Indo-European languages.Languages on the same branch of our phylogenetic tree are notably similar, whereas languages from separate branches are no more similar than chance. Arya McCarthy, Adina Williams, Shijia Liu, David Yarowsky, Ryan Cotterell |
EMNLP (1) | 1 |
| 2020 | SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech TranslationabstractWe propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio thesized to resemble another speaker's voice. Our method compares favorably to SpecAugment on English-French and English-Romanian automatic speech translation (AST) tasks as well as on a low-resource English automatic speech recognition (ASR) task. Further, in ablations, we show the benefits of both quantity and diversity in augmented data. Finally, we show that we can combine our approach with augmentation by machine-translated transcripts to obtain a competitive end-to-end AST model that outperforms a very strong cascade model on an English-French AST task. Our method is sufficiently general that it can be applied to other speech generation and analysis tasks. Arya McCarthy, Liezl Puzon, Juan Pino 0001 |
ICASSP | 1 |
| 2020 | Massively Multilingual Pronunciation Modeling with WikiPronabstractWe introduce WikiPron, an open-source command-line tool for extracting pronunciation data from Wiktionary, a collaborative multilingual online dictionary. We first describe the design and use of WikiPron. We then discuss the challenges faced scaling this tool to create an automatically-generated database of 1.7 million pronunciations from 165 languages. Finally, we validate the pronunciation database by using it to train and evaluating a collection of generic grapheme-to-phoneme models. The software, pronunciation data, and models are all made available under permissive open-source licenses. Jackson L. Lee, Lucas F. E. Ashby, M. Elizabeth Garza, Yeonju Lee-Sikka, Sean Miller, Arya McCarthy, Kyle Gorman |
LREC | 7 |
| 2020 | UniMorph 3.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological paradigms for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. We have implemented several improvements to the extraction pipeline which creates most of our data, so that it is both more complete and more correct. We have added 66 new languages, as well as new parts of speech for 12 languages. We have also amended the schema in several ways. Finally, we present three new community tools: two to validate data for resource creators, and one to make morphological data available from the command line. UniMorph is based at the Center for Language and Speech Processing (CLSP) at Johns Hopkins University in Baltimore, Maryland. This paper details advances made to the schema, tooling, and dissemination of project resources since the UniMorph 2.0 release described at LREC 2018. Arya McCarthy, Christo Kirov, Matteo Grella, Amrit Nidhi, Patrick Xia 0002, Kyle Gorman, Ekaterina Vylomova, Sabrina J. Mielke, Garrett Nicolai, Miikka Silfverberg, Timofey Arkhangelskiy, Nataly Krizhanovsky, Andrew Krizhanovsky, Elena Klyachko, Alexey Sorokin, John Mansfield, Valts Ernstreits, Yuval Pinter, Cassandra L. Jacobs, Ryan Cotterell, Mans Hulden, David Yarowsky |
LREC | 1 |
| 2020 | The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological ExplorationabstractWe present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible and counting. Our data is derived from scraping several online resources and merging them with existing corpora, combining them under a common scheme that is verse-parallel across all translations. We detail our effort to scrape, clean, align, and utilize this ripe multilingual dataset. The corpus captures the great typological variety of the world’s languages. We catalog this by showing highly similar proportions of representation of Ethnologue’s typological features in our corpus. We also give an example application: projecting pronoun features like clusivity across alignments to richly annotate languages which do not mark the distinction. Arya McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky |
LREC | 1 |
| 2020 | An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource LanguagesabstractIn this work, we explore massively multilingual low-resource neural machine translation. Using translations of the Bible (which have parallel structure across languages), we train models with up to 1,107 source languages. We create various multilingual corpora, varying the number and relatedness of source languages. Using these, we investigate the best ways to use this many-way aligned resource for multilingual machine translation. Our experiments employ a grammatically and phylogenetically diverse set of source languages during testing for more representative evaluations. We find that best practices in this domain are highly language-specific: adding more languages to a training set is often better, but too many harms performance—the best number depends on the source language. Furthermore, training on related languages can improve or degrade performance, depending on the language. As there is no one-size-fits-most answer, we find that it is critical to tailor one’s approach to the source language and its typology. Aaron Mueller, Garrett Nicolai, Arya McCarthy, Dylan Lewis, Winston Wu, David Yarowsky |
LREC | 3 |
| 2020 | Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand LanguagesabstractExploiting the broad translation of the Bible into the world’s languages, we train and distribute morphosyntactic tools for approximately one thousand languages, vastly outstripping previous distributions of tools devoted to the processing of inflectional morphology. Evaluation of the tools on a subset of available inflectional dictionaries demonstrates strong initial models, supplemented and improved through ensembling and dictionary-based reranking. Likewise, a novel type-to-token based evaluation metric allows us to confirm that models generalize well across rare and common forms alike Garrett Nicolai, Dylan Lewis, Arya McCarthy, Aaron Mueller, Winston Wu, David Yarowsky |
LREC | 3 |
| 2019 | Meaning to Form: Measuring Systematicity as InformationabstractA longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between a word form and its meaning, or does some systematic phenomenon pervade?For instance, does the character bigram gl have any systematic relationship to the meaning of words like glisten, gleam and glow?In this work, we offer a holistic quantification of the systematicity of the sign using mutual information and recurrent neural networks.We employ these in a data-driven and massively multilingual approach to the question, examining 106 languages.We find a statistically significant reduction in entropy when modeling a word form conditioned on its semantic representation.Encouragingly, we also recover wellattested English examples of systematic affixes.We conclude with the meta-point: Our approximate effect size (measured in bits) is quite small-despite some amount of systematicity between form and meaning, an arbitrary relationship and its resulting benefits dominate human language. Tiago Pimentel, Arya McCarthy, Damián E. Blasi, Brian Roark, Ryan Cotterell |
ACL (1) | 2 |
| 2019 | Weird Inflects but OK: Making Sense of Morphological Generation ErrorsabstractWe conduct a manual error analysis of the CoNLL-SIGMORPHON 2017 Shared Task on Morphological Reinflection.In this task, systems are given a word in citation form (e.g., hug) and asked to produce the corresponding inflected form (e.g., the simple past hugged).This design lets us analyze errors much like we might analyze children's production errors.We propose an error taxonomy and use it to annotate errors made by the top two systems across twelve languages.Many of the observed errors are related to inflectional patterns sensitive to inherent linguistic properties such as animacy or affect; many others are failures to predict truly unpredictable inflectional behaviors.We also find nearly one quarter of the residual "errors" reflect errors in the gold data. Kyle Gorman, Arya McCarthy, Ryan Cotterell, Ekaterina Vylomova, Miikka Silfverberg, Magdalena Markowska |
CoNLL | 2 |
| 2019 | Modeling Color Terminology Across Thousands of LanguagesabstractArya D. McCarthy, Winston Wu, Aaron Mueller, William Watson, David Yarowsky. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Arya McCarthy, Winston Wu, Aaron Mueller, Bill Watson, David Yarowsky |
EMNLP/IJCNLP (1) | 1 |
| 2018 | UniMorph 2.0: Universal Morphology
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia 0002, Manaal Faruqui, Sabrina J. Mielke, Arya McCarthy, Sandra Kübler, David Yarowsky, Jason Eisner, Mans Hulden |
LREC | 9 |