Winston Wu

dblp:219/5637 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0002-5888-4836ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 10 first-author · 10 since 2021
YearPublicationVenuePosition
2026 A Modern Online Learning Platform for 'Ōlelo Hawai'i Classrooms
Christian Castro, Keneth Martin, Winston Wu, William H. Wilson
LREC3
2025 Statistical and Neural Methods for Hawaiian Orthography Modernization
abstract
Hawaiian orthography employs two distinct spelling systems, both of which are used by communities of speakers today.These two spelling systems are distinguished by the presence of the 'okina letter and kahakō diacritic, which represent glottal stops and long vowels, respectively.We develop several models ranging in complexity to convert between these two orthographies.Our results demonstrate that simple statistical n-gram models surprisingly outperform neural seq2seq models and LLMs, highlighting the potential for traditional machine learning approaches in a low-resource setting.
Jaden Kapali, Keaton Williamson, Winston Wu
EMNLP3
2024 Analyzing Occupational Distribution Representation in Japanese Language Models
abstract
Recent advances in large language models (LLMs) have enabled users to generate fluent and seemingly convincing text. However, these models have uneven performance in different languages, which is also associated with undesirable societal biases toward marginalized populations. Specifically, there is relatively little work on Japanese models, despite it being the thirteenth most widely spoken language. In this work, we first develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations. We then evaluate a variety of English, multilingual, and Japanese models, correlating the models’ outputs with occupation statistics from the Japanese Census Bureau from the last 100 years. Our findings indicate that models can associate Japanese names with the correct gendered occupations when using constrained decoding. However, with sampling or greedy decoding, Japanese language models have a preference for a small set of stereotypically gendered occupations, and multilingual models, though trained on Japanese, are not always able to understand Japanese prompts.
Katsumi Ibaraki, Winston Wu, Lu Wang 0008, Rada Mihalcea
LREC/COLING2
2024 Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
abstract
Recent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beginning of their careers – worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm.
Oana Ignat, Zhijing Jin 0001, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao 0004, Aylin Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee 0001, Siyang Liu 0003, Do June Min, Shinka Mori, Joan Nwatu, Verónica Pérez-Rosas, Zekun Wang 0002, Winston Wu, Rada Mihalcea
LREC/COLING21
2024 MOKA: Moral Knowledge Augmentation for Moral Event Extraction
abstract
Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang 0008
NAACL-HLT2
2023 Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk Tales
abstract
Folk tales are strong cultural and social influences in children's lives, and they are known to teach morals and values.However, existing studies on folk tales are largely limited to European tales.In our study, we compile a large corpus of over 1,900 tales originating from 27 diverse cultures across six continents.Using a range of lexicons and correlation analyses, we examine how human values, morals, and gender biases are expressed in folk tales across cultures.We discover differences between cultures in prevalent values and morals, as well as cross-cultural trends in problematic gender biases.Furthermore, we find trends of reduced value expression when examining public-domain fiction stories, extrinsically validate our analyses against the multicultural Schwartz Survey of Cultural Values, and find traditional gender biases associated with values, morals, and agency.This largescale cross-cultural study of folk tales paves the way for future studies on how literature influences and reflects cultural norms.
Winston Wu, Lu Wang 0008, Rada Mihalcea
EMNLP1
2022 Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages
abstract
This paper presents a detailed foundational empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian, and a multi-faceted distributional analysis of the underlying word-formation processes that can aid in their compositional translation, tagging, parsing, language modeling, and other NLP tasks. Given that out-of-vocabulary (OOV) words generally present a key open challenge to NLP and machine translation systems, especially toward the lower limit of resource availability, there are useful practical insights, as well as corpus-linguistic insights, from both a detailed manual and automatic taxonomic analysis of the types, multidimensional properties, and processing potential for multiple representative OOV data samples.
Georgie Botev, Arya McCarthy, Winston Wu, David Yarowsky
COLING3
2022 Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis
abstract
Prior work on ideology prediction has largely focused on single modalities, i.e., text or images.In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content.We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter.We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum.Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components.Our bestperforming model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang 0008
EMNLP2
2022 On the Robustness of Cognate Generation Models
abstract
We evaluate two popular neural cognate generation models’ robustness to several types of human-plausible noise (deletion, duplication, swapping, and keyboard errors, as well as a new type of error, phonological errors). We find that duplication and phonological substitution is least harmful, while the other types of errors are harmful. We present an in-depth analysis of the models’ results with respect to each error type to explain how and why these models perform as they do.
Winston Wu, David Yarowsky
LREC1
2021 Evaluating Neural Model Robustness for Machine Comprehension
abstract
We evaluate neural model robustness to adversarial attacks using different types of linguistic unit perturbations -character and word, and propose a new method for strategic sentencelevel perturbations.We experiment with different amounts of perturbations to examine model confidence and misclassification rate, and contrast model performance with different embeddings BERT and ELMo on two benchmark datasets SQuAD and TriviaQA.We demonstrate how to improve model performance during an adversarial attack by using ensembles.Finally, we analyze factors that affect model behavior under adversarial attack, and develop a new model to predict errors during attacks.Our novel findings reveal that (a) unlike BERT, models that use ELMo embeddings are more susceptible to adversarial attacks, (b) unlike word and paraphrase, character perturbations affect the model the most but are most easily compensated for by adversarial training, (c) word perturbations lead to more high-confidence misclassifications compared to sentence-and character-level perturbations, (d) the type of question and model answer length (the longer the answer the more likely it is to be incorrect) is the most predictive of model errors in adversarial setting, and (e) conclusions about model behavior are dataset-specific.
Winston Wu, Dustin Arendt, Svitlana Volkova
EACL1
2020 Neural Transduction for Multilingual Lexical Translation
abstract
We present a method for completing multilingual translation dictionaries.Our probabilistic approach can synthesize new word forms, allowing it to operate in settings where correct translations have not been observed in text (cf.cross-lingual embeddings).In addition, we propose an approximate Maximum Mutual Information (MMI) decoding objective to further improve performance in both many-to-one and one-to-one word level translation tasks where we use either multiple input languages for a single target language or more typical single language pair translation.The model is trained in a many-to-many setting, where it can leverage information from related languages to predict words in each of its many target languages.We focus on 6 languages: French, Spanish, Italian, Portuguese, Romanian, and Turkish.When indirect multilingual information is available, ensembling with mixture-of-experts as well as incorporating related languages leads to a 27% relative improvement in whole-word accuracy of predictions over a single-source baseline.To seed the completion when multilingual data is unavailable, it is better to decode with an MMI objective.
Dylan Lewis, Winston Wu, Arya McCarthy, David Yarowsky
COLING2
2020 Wiktionary Normalization of Translations and Morphological Information
abstract
We extend the Yawipa Wiktionary Parser (Wu and Yarowsky, 2020) to extract and normalize translations from etymology glosses, and morphological form-of relations, resulting in 300K unique translations and over 4 million instances of 168 annotated morphological relations.We propose a method to identify typos in translation annotations.Using the extracted morphological data, we develop multilingual neural models for predicting three types of word formationclipping, contraction, and eye dialect-and improve upon a standard attention baseline by using copy attention.
Winston Wu, David Yarowsky
COLING1
2020 The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration
abstract
We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible and counting. Our data is derived from scraping several online resources and merging them with existing corpora, combining them under a common scheme that is verse-parallel across all translations. We detail our effort to scrape, clean, align, and utilize this ripe multilingual dataset. The corpus captures the great typological variety of the world’s languages. We catalog this by showing highly similar proportions of representation of Ethnologue’s typological features in our corpus. We also give an example application: projecting pronoun features like clusivity across alignments to richly annotate languages which do not mark the distinction.
Arya McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
LREC5
2020 An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages
abstract
In this work, we explore massively multilingual low-resource neural machine translation. Using translations of the Bible (which have parallel structure across languages), we train models with up to 1,107 source languages. We create various multilingual corpora, varying the number and relatedness of source languages. Using these, we investigate the best ways to use this many-way aligned resource for multilingual machine translation. Our experiments employ a grammatically and phylogenetically diverse set of source languages during testing for more representative evaluations. We find that best practices in this domain are highly language-specific: adding more languages to a training set is often better, but too many harms performance—the best number depends on the source language. Furthermore, training on related languages can improve or degrade performance, depending on the language. As there is no one-size-fits-most answer, we find that it is critical to tailor one’s approach to the source language and its typology.
Aaron Mueller, Garrett Nicolai, Arya McCarthy, Dylan Lewis, Winston Wu, David Yarowsky
LREC5
2020 Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages
abstract
Exploiting the broad translation of the Bible into the world’s languages, we train and distribute morphosyntactic tools for approximately one thousand languages, vastly outstripping previous distributions of tools devoted to the processing of inflectional morphology. Evaluation of the tools on a subset of available inflectional dictionaries demonstrates strong initial models, supplemented and improved through ensembling and dictionary-based reranking. Likewise, a novel type-to-token based evaluation metric allows us to confirm that models generalize well across rare and common forms alike
Garrett Nicolai, Dylan Lewis, Arya McCarthy, Aaron Mueller, Winston Wu, David Yarowsky
LREC5
2020 Multilingual Dictionary Based Construction of Core Vocabulary
abstract
We propose a new functional definition and construction method for core vocabulary sets for multiple applications based on the relative coverage of a target concept in thousands of bilingual dictionaries. Our newly developed core concept vocabulary list derived from these dictionary consensus methods achieves high overlap with existing widely utilized core vocabulary lists targeted at applications such as first and second language learning or field linguistics. Our in-depth analysis illustrates multiple desirable properties of our newly proposed core vocabulary set, including their non-compositionality. We employ a cognate prediction method to recover missing coverage of this core vocabulary in massively multilingual dictionary construction, and we argue that this core vocabulary should be prioritized for elicitation when creating new dictionaries for low-resource languages for multiple downstream tasks including machine translation and language learning.
Winston Wu, Garrett Nicolai, David Yarowsky
LREC1
2020 Computational Etymology and Word Emergence
abstract
We developed an extensible, comprehensive Wiktionary parser that improves over several existing parsers. We predict the etymology of a word across the full range of etymology types and languages in Wiktionary, showing improvements over a strong baseline. We also model word emergence and show the application of etymology in modeling this phenomenon. We release our parser to further research in this understudied field.
Winston Wu, David Yarowsky
LREC1
2019 Modeling Color Terminology Across Thousands of Languages
abstract
Arya D. McCarthy, Winston Wu, Aaron Mueller, William Watson, David Yarowsky. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arya McCarthy, Winston Wu, Aaron Mueller, Bill Watson, David Yarowsky
EMNLP/IJCNLP (1)2
2019 An Exploration of Placeholding in Neural Machine Translation
Matt Post, Shuoyang Ding, Marianna J. Martindale, Winston Wu
MTSummit (1)4
2018 Creating a Translation Matrix of the Bible's Names Across 591 Languages
Winston Wu, Nidhi Vyas, David Yarowsky
LREC1
2018 Creating Large-Scale Multilingual Cognate Tables
Winston Wu, David Yarowsky
LREC1
2018 A Comparative Study of Extremely Low-Resource Transliteration of the World's Languages
Winston Wu, David Yarowsky
LREC1
2018 Massively Translingual Compound Analysis and Translation Discovery
Winston Wu, David Yarowsky
LREC1