Zoey Liu

dblp:278/8017 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
20since 2021 · last 2026
0009-0008-5313-157XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 11 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Dataset for Oral Reading in Young English Readers
abstract
Madison Rose, Michael Bennie, Valeria Pagliai, Hatice Kubra Karakis, Qian Shen, Xinyi Tai, Walter L. Leite, Zoey Liu. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026.
Madison Rose, Michael Bennie, Valeria Pagliai, Hatice Kubra Karakis, Xinyi Tai, Walter L. Leite, Zoey Liu
CoNLL8
2025 Predictability effects of Spanish-English code-switching: A directionality and part of speech analysis
Joshua Higdon, Valeria Pagliai, Zoey Liu
CogSci3
2025 Studying Cross-linguistic Structural Transfer in Second Language Learning
Zoey Liu, Wenshuo Qin, Haiyin Yang, Joshua K. Hartshorne
CogSci1
2024 Real world event schemas offer modality-independent conceptual bases for verb argument structures
Brennan Gonering, Zoey Liu
CogSci2
2024 Frequency-Dependent Regularization in Mandarin Elastic Word Length
Skyler Jove Reese, Zoey Liu, Masoud Jasbi, Emily Morgan
CogSci2
2024 Enough Is Enough! a Case Study on the Effect of Data Size for Evaluation Using Universal Dependencies
abstract
When creating a new dataset for evaluation, one of the first considerations is the size of the dataset. If our evaluation data is too small, we risk making unsupported claims based on the results on such data. If, on the other hand, the data is too large, we waste valuable annotation time and costs that could have been used to widen the scope of our evaluation (i.e. annotate for more domains/languages). Hence, we investigate the effect of the size and a variety of sampling strategies of evaluation data to optimize annotation efforts, using dependency parsing as a test case. We show that for in-language in-domain datasets, 5,000 tokens is enough to obtain a reliable ranking of different parsers; especially if the data is distant enough from the training split (otherwise, we recommend 10,000). In cross-domain setups, the same amounts are required, but in cross-lingual setups much less (2,000 tokens) is enough.
Rob van der Goot, Zoey Liu, Max Müller-Eberstein
LREC/COLING2
2024 Leveraging Speech Data Diversity to Document Indigenous Heritage and Culture
Allahsera Tapo, Éric Le Ferrand, Zoey Liu, Christopher Homan, Emily Tucker Prud'hommeaux
INTERSPEECH3
2024 The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
abstract
Recent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice.This study addresses these challenges, focusing on morphological segmentation and synthesizing limitations related to language diversity, adoption of multiple datasets and splits, and detailed model comparisons.Our study leverages data from 19 languages, including ten indigenous or endangered languages across 10 language families with diverse morphological systems (polysynthetic, fusional, and agglutinative) and different degrees of data availability.We conduct large-scale experimentation with varying sized combinations of training and evaluation sets as well as new test data.Our results show that, when faced with new test data: (1) models trained from random splits are able to achieve higher numerical scores; (2) model rankings derived from random splits tend to generalize more consistently.
Zoey Liu, Bonnie J. Dorr
NAACL-HLT1
2023 Morphological Inflection: A Reality Check
abstract
Morphological inflection is a popular task in sub-word NLP with both practical and cognitive applications.For years now, state-of-theart systems have reported high, but also highly variable, performance across data sets and languages.We investigate the causes of this high performance and high variability; we find several aspects of data set creation and evaluation which systematically inflate performance and obfuscate differences between languages.To improve generalizability and reliability of results, we propose new data sampling and evaluation strategies that better reflect likely usecases.Using these new strategies, we make new observations on the generalization abilities of current inflection systems.
Jordan Kodner, Sarah R. B. Payne, Salam Khalifa, Zoey Liu
ACL (1)4
2023 Re-Evaluating the Evaluation of Neural Morphological Inflection Models
Jordan Kodner, Salam Khalifa, Sarah R. B. Payne, Zoey Liu
CogSci4
2023 Double PP Constituent Ordering Preferences in English Early Child Language
Zoey Liu, Lauren E. Namdar, Stefanie Wulff, Kenji Sagae
CogSci1
2023 Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation
abstract
Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set.This "holdspeaker(s)-out" data partitioning strategy, however, may not be ideal for data sets in which the number of speakers is very small.This study investigates ten different data split methods for five languages with minimal ASR training resources.We find that (1) model performance varies greatly depending on which speaker is selected for testing; (2) the average word error rate (WER) across all held-out speakers is comparable not only to the average WER over multiple random splits but also to any given individual random split; (3) WER is also generally comparable when the data is split heuristically or adversarially; (4) utterance duration and intensity are comparatively more predictive factors of variability regardless of the data split.These results suggest that the widely used holdspeakers-out approach to ASR data partitioning can yield results that do not reflect model performance on unseen data or speakers.Random splits can yield more reliable and generalizable estimates when facing data sparsity.
Zoey Liu, Justin Spence, Emily Tucker Prud'hommeaux
EACL1
2023 Data-driven Parsing Evaluation for Child-Parent Interactions
abstract
Abstract We present a syntactic dependency treebank for naturalistic child and child-directed spoken English. Our annotations largely follow the guidelines of the Universal Dependencies project (UD [Zeman et al., 2022]), with detailed extensions to lexical and syntactic structures unique to spontaneous spoken language, as opposed to written texts or prepared speech. Compared to existing UD-style spoken treebanks and other dependency corpora of child-parent interactions specifically, our dataset is much larger (44,744 utterances; 233,907 words) and contains data from 10 children covering a wide age range (18–66 months). We conduct thorough dependency parser evaluations using both graph-based and transition-based parsers, trained on three different types of out-of-domain written texts: news, tweets, and learner data. Out-of-domain parsers demonstrate reasonable performance for both child and parent data. In addition, parser performance for child data increases along children’s developmental paths, especially between 18 and 48 months, and gradually approaches the performance for parent data. These results are further validated with in-domain training.
Zoey Liu, Emily Tucker Prud'hommeaux
Trans. Assoc. Comput. Linguistics1
2022 Not always about you: Prioritizing community needs when developing endangered language technology
abstract
Languages are classified as low-resource when they lack the quantity of data necessary for training statistical and machine learning tools and models.Causes of resource scarcity vary but can include poor access to technology for developing these resources, a relatively small population of speakers, or a lack of urgency for collecting such resources in bilingual populations where the second language is highresource.As a result, the languages described as low-resource in the literature are as different as Finnish on the one hand, with millions of speakers using it in every imaginable domain, and Seneca, with only a small-handful of fluent speakers using the language primarily in a restricted domain.While issues stemming from the lack of resources necessary to train models unite this disparate group of languages, many other issues cut across the divide between widely-spoken low-resource languages and endangered languages.In this position paper, we discuss the unique technological, cultural, practical, and ethical challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered language documentation and revitalization.We report the perspectives of language teachers, Master Speakers and elders from indigenous communities, as well as the point of view of academics.We describe an ongoing fruitful collaboration and make recommendations for future partnerships between academic researchers and language community stakeholders.
Zoey Liu, Crystal Richardson, Richard J. Hatcher, Emily Tucker Prud'hommeaux
ACL (1)1
2022 Data-driven Crosslinguistic Syntactic Transfer in Second Language Learning
Zoey Liu, Tiwalayo Eisape, Emily Tucker Prud'hommeaux, Joshua K. Hartshorne
CogSci1
2022 Does One Size Fit all in Crosslinguistic Dependency Length Minimization?
Zoey Liu, Ria Upreti, Mathew A. Kramer, Savithry Namboodiripad
CogSci1
2022 Evaluating the Performance of Transformer-based Language Models for Neuroatypical Language
abstract
Difficulties with social aspects of language are among the hallmarks of autism spectrum disorder (ASD). These communication differences are thought to contribute to the challenges that adults with ASD experience when seeking employment, underscoring the need for interventions that focus on improving areas of weakness in pragmatic and social language. In this paper, we describe a transformer-based framework for identifying linguistic features associated with social aspects of communication using a corpus of conversations between adults with and without ASD and neurotypical conversational partners produced while engaging in collaborative tasks. While our framework yields strong accuracy overall, performance is significantly worse for the language of participants with ASD, suggesting that they use a more diverse set of strategies for some social linguistic functions. These results, while showing promise for the development of automated language analysis tools to support targeted language interventions for ASD, also reveal weaknesses in the ability of large contextualized language models to model neuroatypical language.
Duanchen Liu, Zoey Liu, Qingyun Yang, Yujing Huang, Emily Tucker Prud'hommeaux
COLING2
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC61
2022 Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentation
abstract
Abstract Common designs of model evaluation typically focus on monolingual settings, where different models are compared according to their performance on a single data set that is assumed to be representative of all possible data for the task at hand. While this may be reasonable for a large data set, this assumption is difficult to maintain in low-resource scenarios, where artifacts of the data collection can yield data sets that are outliers, potentially making conclusions about model performance coincidental. To address these concerns, we investigate model generalizability in crosslinguistic low-resource scenarios. Using morphological segmentation as the test case, we compare three broad classes of models with different parameterizations, taking data from 11 languages across 6 language families. In each experimental setting, we evaluate all models on a first data set, then examine their performance consistency when introducing new randomly sampled data sets with the same size and when applying the trained models to unseen test sets of varying sizes. The results demonstrate that the extent of model generalization depends on the characteristics of the data set, and does not necessarily rely heavily on the data set size. Among the characteristics that we studied, the ratio of morpheme overlap and that of the average number of morphemes per word between the training and test sets are the two most prominent factors. Our findings suggest that future work should adopt random sampling to construct data sets with different sizes in order to make more responsible claims about model evaluation.
Zoey Liu, Emily Tucker Prud'hommeaux
Trans. Assoc. Comput. Linguistics1
2021 English Negative Constructions and Communicative Functions in Child Language
Zoey Liu, Masoud Jasbi
CogSci1
2020 Frequency-dependent Regularization in Constituent Ordering Preferences
Zoey Liu, Emily Morgan
CogSci1
2020 A Predicate-Function-Argument Annotation of Natural Language for Open-Domain Information eXpression
abstract
Existing OIE (Open Information Extraction) algorithms are independent of each other such that there exist lots of redundant works; the featured strategies are not reusable and not adaptive to new tasks. This paper proposes a new pipeline to build OIE systems, where an Open-domain Information eXpression (OIX) task is proposed to provide a platform for all OIE strategies. The OIX is an OIE friendly expression of a sentence without information loss. The generation procedure of OIX contains shared works of OIE algorithms so that OIE strategies can be developed on the platform of OIX as inference operations focusing on more critical problems. Based on the same platform of OIX, the OIE strategies are reusable, and people can select a set of strategies to assemble their algorithm for a specific task so that the adaptability may be significantly increased. This paper focuses on the task of OIX and propose a solution – Open Information Annotation (OIA). OIA is a predicate-function-argument annotation for sentences. We label a data set of sentence-OIA pairs and propose a dependency-based rule system to generate OIA annotations from sentences. The evaluation results reveal that learning the OIA from a sentence is a challenge owing to the complexity of natural language sentences, and it is worthy of attracting more attention from the research community.
Mingming Sun 0001, Wenyue Hua, Zoey Liu, Xin Wang 0017, Kangjie Zheng, Ping Li 0001
EMNLP (1)3