Aaron Mueller

dblp:248/7949 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
25since 2021 · last 2026
0009-0005-1148-5001ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 9 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CRISP: Persistent Concept Unlearning via Sparse Autoencoders
abstract
As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Feature-level analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features.
Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
ACL (1)3
2026 Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
abstract
Large language models (LLMs) learn nontrivial abstractions during pretraining, such as detecting irregular plural noun subjects.However, because traditional evaluation methods (e.g., benchmarking) fail to reveal how models acquire these concepts and capabilities, it is not well understood when and how these specific linguistic abilities emerge.To bridge this gap and better understand model training at the concept level, we use sparse crosscoders to discover and align features across model checkpoints.Using this approach, we track the evolution of linguistic features during pretraining.We train crosscoders between opensourced checkpoint triplets with significant performance and representation shifts, and introduce a novel metric, Relative Indirect Effects (RELIE), to trace training stages at which individual features become causally important for task performance.We show that crosscoders can detect feature emergence, maintenance, and discontinuation during pretraining.Our approach is architecture-agnostic and scalable, offering a promising path toward more interpretable and fine-grained analysis of representation learning throughout pretraining.1
Deniz Bayazit, Aaron Mueller, Antoine Bosselut
ACL (1)2
2026 From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
abstract
Aaron Mueller, Andrew Lee, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Aaron Mueller, Andrew Lee 0001, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger
ACL (1)1
2026 Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
abstract
Multi-agent debate has been shown to improve reasoning in large language models (LLMs).However, it is compute-intensive, requiring generation of long transcripts before answering questions.To address this inefficiency, we develop a framework that distills multi-agent debate into a single LLM through a two-stage fine-tuning pipeline combining debate structure learning with internalization via dynamic reward scheduling and length clipping.Across multiple models and benchmarks, our internalized models match or exceed explicit multiagent debate performance using up to 93% fewer tokens.We then investigate the mechanistic basis of this capability through activation steering, finding that internalization creates agent-specific subspaces: interpretable directions in activation space corresponding to different agent perspectives.We further demonstrate a practical application: by instilling malicious agents into the LLM through internalized debate, then applying negative steering to suppress them, we show that distillation makes harmful behaviors easier to localize and control with smaller reductions in general performance compared to steering base models.Our findings offer a new perspective for understanding multi-agent capabilities in distilled models and provide practical guidelines for controlling internalized reasoning behaviors.1
John Seon Keun Yi, Aaron Mueller, Dokyun Lee
ACL (1)2
2025 Position-aware Automatic Circuit Discovery
abstract
A widely used strategy to discover and understand language model mechanisms is circuit analysis.A circuit is a minimal subgraph of a model's computation graph that executes a specific task.We identify a gap in existing circuit discovery methods: they assume circuits are position-invariant, treating model components as equally relevant across input positions.This limits their ability to capture cross-positional interactions or mechanisms that vary across positions.To address this gap, we propose two improvements to incorporate positionality into circuits, even on tasks containing variablelength examples.First, we extend edge attribution patching, a gradient-based method for circuit discovery, to differentiate between token positions.Second, we introduce the concept of a dataset schema, which defines token spans with similar semantics across examples, enabling position-aware circuit discovery in datasets with variable length examples.We additionally develop an automated pipeline for schema generation and application using large language models.Our approach enables fully automated discovery of position-sensitive circuits, yielding better trade-offs between circuit size and faithfulness compared to prior work. 1 Belinkov.2021.Causal analysis of syntactic agreement mechanisms in neural language models.
Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov
ACL (1)4
2025 SAEs Are Good for Steering - If You Select the Right Features
abstract
Sparse Autoencoders (SAEs) have been proposed as an unsupervised approach to learn a decomposition of a model's latent space.This enables useful applications such as steeringinfluencing the output of a model towards a desired concept-without requiring labeled data.Current methods identify SAE features to steer by analyzing the input tokens that activate them.However, recent work has highlighted that activations alone do not fully describe the effect of a feature on the model's output.In this work, we draw a distinction between two types of features: input features, which mainly capture patterns in the model's input, and output features, which have a human-understandable effect on the model's output.We propose input and output scores to characterize and locate these types of features, and show that high values for both scores rarely co-occur in the same features.These findings have practical implications: after filtering out features with low output scores, we obtain 2-3x improvements when steering with SAEs, making them competitive with supervised methods. 1
Dana Arad, Aaron Mueller, Yonatan Belinkov
EMNLP2
2025 NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
abstract
We introduce NNsight and NDIF, technologies that work in tandem to enable scientific study of the representations and computations learned by very large neural networks. NNsight is an open-source system that extends PyTorch to introduce deferred remote execution. The National Deep Inference Fabric (NDIF) is a scalable inference service that executes NNsight requests, allowing users to share GPU resources and pretrained models. These technologies are enabled by the Intervention Graph, an architecture developed to decouple experimental design from model runtime. Together, this framework provides transparent and efficient access to the internals of deep neural networks such as very large language models (LLMs) without imposing the cost or complexity of hosting customized models individually. We conduct a quantitative survey of the machine learning literature that reveals a growing gap in the study of the internals of large-scale AI. We demonstrate the design and use of our framework to address this gap by enabling a range of research methods on huge models. Finally, we conduct benchmarks to compare performance with previous approaches. Code, documentation, and tutorials are available at https://nnsight.net/.
Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell 0001, Byron C. Wallace, David Bau
ICLR11
2025 Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
abstract
We introduce methods for discovering and applying **sparse feature circuits**. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors. Circuits identified in prior work consist of polysemantic and difficult-to-interpret units like attention heads or neurons, rendering them unsuitable for many downstream applications. In contrast, sparse feature circuits enable detailed understanding of unanticipated mechanisms in neural networks. Because they are based on fine-grained units, sparse feature circuits are useful for downstream tasks: We introduce SHIFT, where we improve the generalization of a classifier by ablating features that a human judges to be task-irrelevant. Finally, we demonstrate an entirely unsupervised and scalable interpretability pipeline by discovering thousands of sparse feature circuits for automatically discovered model behaviors.
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
ICLR6
2025 Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics
abstract
Do large language models (LLMs) solve reasoning tasks by learning robust generalizable algorithms, or do they memorize training data? To investigate this question, we use arithmetic reasoning as a representative task. Using causal analysis, we identify a subset of the model (a circuit) that explains most of the model's behavior for basic arithmetic logic and examine its functionality. By zooming in on the level of individual circuit neurons, we discover a sparse set of important neurons that implement simple heuristics. Each heuristic identifies a numerical input pattern and outputs corresponding answers. We hypothesize that the combination of these heuristic neurons is the mechanism used to produce correct arithmetic answers. To test this, we categorize each neuron into several heuristic types---such as neurons that activate when an operand falls within a certain range---and find that the unordered combination of these heuristic types is the mechanism that explains most of the model's accuracy on arithmetic prompts. Finally, we demonstrate that this mechanism appears as the main source of arithmetic accuracy early in training. Overall, our experimental results across several LLMs show that LLMs perform arithmetic using neither robust algorithms nor memorization; rather, they rely on a ``bag of heuristics''.
Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov
ICLR3
2025 MIB: A Mechanistic Interpretability Benchmark
abstract
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components---and connections between them---most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field.
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna 0001, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov
ICML1
2025 Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
abstract
Jannik Brinkmann, Chris Wendler, Christian Bartelt, Aaron Mueller. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jannik Brinkmann, Chris Wendler, Christian Bartelt, Aaron Mueller
NAACL (Long Papers)4
2025 Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models
abstract
Michael Hanna, Aaron Mueller. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Michael Hanna 0001, Aaron Mueller
NAACL (Long Papers)2
2025 Characterizing the Role of Similarity in the Property Inferences of Language Models
abstract
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra
NAACL (Long Papers)2
2024 Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen
CogSci2
2024 Function Vectors in Large Language Models
abstract
We report the presence of a simple neural mechanism that represents an input-output function as a vector within autoregressive transformer language models (LMs). Using causal mediation analysis on a diverse range of in-context-learning (ICL) tasks, we find that a small number attention heads transport a compact representation of the demonstrated task, which we call a function vector (FV). FVs are robust to changes in context, i.e., they trigger execution of the task on inputs such as zero-shot and natural text settings that do not resemble the ICL contexts from which they are collected. We test FVs across a range of tasks, models, and layers and find strong causal effects across settings in middle layers. We investigate the internal structure of FVs and find while that they often contain information that encodes the output space of the function, this information alone is not sufficient to reconstruct an FV. Finally, we test semantic vector composition in FVs, and find that to some extent they can be summed to create vectors that trigger new complex tasks. Our findings show that compact, causal internal vector representations of function abstractions can be explicitly extracted from LLMs.
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, David Bau
ICLR4
2024 In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
abstract
Aaron Mueller, Albert Webson, Jackson Petty, Tal Linzen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Aaron Mueller, Albert Webson, Jackson Petty, Tal Linzen
NAACL-HLT1
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)4
2023 How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive Biases
abstract
Accurate syntactic representations are essential for robust generalization in natural language.Recent work has found that pre-training can teach language models to rely on hierarchical syntactic features-as opposed to incorrect linear features-when performing tasks after finetuning.We test what aspects of pre-training are important for endowing encoder-decoder Transformers with an inductive bias that favors hierarchical syntactic generalizations.We focus on architectural features (depth, width, and number of parameters), as well as the genre and size of the pre-training corpus, diagnosing inductive biases using two syntactic transformation tasks: question formation and passivization, both in English.We find that the number of parameters alone does not explain hierarchical generalization: model depth plays greater role than model width.We also find that pre-training on simpler language, such as child-directed speech, induces a hierarchical bias using an order-of-magnitude less data than pre-training on more typical datasets based on web text or Wikipedia; this suggests that in cognitively plausible language acquisition settings, neural language models may be more data-efficient than previously thought.
Aaron Mueller, Tal Linzen
ACL (1)1
2023 Language model acceptability judgements are not always robust to context
abstract
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
ACL (1)3
2022 Label Semantic Aware Pre-training for Few-shot Text Classification
abstract
Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang, Dan Roth. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang 0001, Dan Roth 0001
ACL (1)1
2022 Causal Analysis of Syntactic Agreement Neurons in Multilingual Language Models
abstract
Structural probing work has found evidence for latent syntactic information in pre-trained language models.However, much of this analysis has focused on monolingual models, and analyses of multilingual models have employed correlational methods that are confounded by the choice of probing tasks.In this study, we causally probe multilingual language models (XGLM and multilingual BERT) as well as monolingual BERT-based models across various languages; we do this by performing counterfactual perturbations on neuron activations and observing the effect on models' subjectverb agreement probabilities.We observe where in the model and to what extent syntactic agreement is encoded in each language.We find significant neuron overlap across languages in autoregressive multilingual language models, but not masked language models.We also find two distinct layer-wise effect patterns and two distinct sets of neurons used for syntactic agreement, depending on whether the subject and verb are separated by other tokens.Finally, we find that behavioral analyses of language models are likely underestimating how sensitive masked language models are to syntactic information.
Aaron Mueller, Tal Linzen
CoNLL1
2022 Bernice: A Multilingual Pre-trained Encoder for Twitter
abstract
The language of Twitter differs significantly from that of other domains commonly included in large language model training.While tweets are typically multilingual and contain informal language, including emoji and hashtags, most pre-trained language models for Twitter are either monolingual, adapted from other domains rather than trained exclusively on Twitter, or are trained on a limited amount of in-domain Twitter data.We introduce Bernice, the first multilingual RoBERTa language model trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer.We evaluate on a variety of monolingual and multilingual Twitter benchmarks, finding that our model consistently exceeds or matches the performance of a variety of models adapted to social media data as well as strong multilingual baselines, despite being trained on less data overall.We posit that it is more efficient compute-and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer.
Alexandra DeLucia, Aaron Mueller, Carlos Alejandro Aguirre, Philip Resnik, Mark Dredze
EMNLP3
2021 Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
abstract
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov
ACL/IJCNLP (1)2
2021 Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling
abstract
Neural topic models can augment or replace bag-of-words inputs with the learned representations of deep pre-trained transformer-based word prediction models.One added benefit when using representations from multilingual models is that they facilitate zero-shot polylingual topic modeling.However, while it has been widely observed that pre-trained embeddings should be fine-tuned to a given task, it is not immediately clear what supervision should look like for an unsupervised task such as topic modeling.Thus, we propose several methods for fine-tuning encoders to improve both monolingual and zero-shot polylingual neural topic modeling.We consider fine-tuning on auxiliary tasks, constructing a new topic classification task, integrating the topic classification objective directly into topic model training, and continued pre-training.We find that fine-tuning encoder representations on topic classification and integrating the topic classification task directly into topic modeling improves topic quality, and that fine-tuning encoder representations on any task is the most important factor for facilitating cross-lingual transfer.
Aaron Mueller, Mark Dredze
NAACL-HLT1
2021 Demographic Representation and Collective Storytelling in the Me Too Twitter Hashtag Activism Movement
abstract
The #MeToo movement on Twitter has drawn attention to the pervasive nature of sexual harassment and violence. While #MeToo has been praised for providing support for self-disclosures of harassment or violence and shifting societal response, it has also been criticized for exemplifying how women of color have been discounted for their historical contributions to and excluded from feminist movements. Through an analysis of over 600,000 tweets from over 256,000 unique users, we examine online #MeToo conversations across gender and racial/ethnic identities and the topics that each demographic emphasized. We found that tweets authored by white women were overrepresented in the movement compared to other demographics, aligning with criticism of unequal representation. We found that intersected identities contributed differing narratives to frame the movement, co-opted the movement to raise visibility in parallel ongoing movements, employed the same hashtags both critically and supportively, and revived and created new hashtags in response to pivotal moments. Notably, tweets authored by black women often expressed emotional support and were critical about differential treatment in the justice system and by police. In comparison, tweets authored by white women and men often highlighted sexual harassment and violence by public figures and weaved in more general political discussions. We discuss the implications of this work for digital activism research and design, including suggestions to raise visibility by those who were under-represented in this hashtag activism movement.
Aaron Mueller, Zach Wood-Doughty, Silvio Amir, Mark Dredze, Alicia L. Nobles
Proc. ACM Hum. Comput. Interact.1
2020 Cross-Linguistic Syntactic Evaluation of Word Prediction Models
abstract
A range of studies have concluded that neural word prediction models can distinguish grammatical from ungrammatical sentences with high accuracy.However, these studies are based primarily on monolingual evidence from English.To investigate how these models' ability to learn syntax varies by language, we introduce CLAMS (Cross-Linguistic Assessment of Models on Syntax), a syntactic evaluation suite for monolingual and multilingual models.CLAMS includes subject-verb agreement challenge sets for English, French, German, Hebrew and Russian, generated from grammars we develop.We use CLAMS to evaluate LSTM language models as well as monolingual and multilingual BERT.Across languages, monolingual LSTMs achieved high accuracy on dependencies without attractors, and generally poor accuracy on agreement across object relative clauses.On other constructions, agreement accuracy was generally higher in languages with richer morphology.Multilingual models generally underperformed monolingual models.Multilingual BERT showed high syntactic accuracy on English, but noticeable deficiencies in other languages.
Aaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina, Tal Linzen
ACL1
2020 The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration
abstract
We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible and counting. Our data is derived from scraping several online resources and merging them with existing corpora, combining them under a common scheme that is verse-parallel across all translations. We detail our effort to scrape, clean, align, and utilize this ripe multilingual dataset. The corpus captures the great typological variety of the world’s languages. We catalog this by showing highly similar proportions of representation of Ethnologue’s typological features in our corpus. We also give an example application: projecting pronoun features like clusivity across alignments to richly annotate languages which do not mark the distinction.
Arya McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
LREC4
2020 An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages
abstract
In this work, we explore massively multilingual low-resource neural machine translation. Using translations of the Bible (which have parallel structure across languages), we train models with up to 1,107 source languages. We create various multilingual corpora, varying the number and relatedness of source languages. Using these, we investigate the best ways to use this many-way aligned resource for multilingual machine translation. Our experiments employ a grammatically and phylogenetically diverse set of source languages during testing for more representative evaluations. We find that best practices in this domain are highly language-specific: adding more languages to a training set is often better, but too many harms performance—the best number depends on the source language. Furthermore, training on related languages can improve or degrade performance, depending on the language. As there is no one-size-fits-most answer, we find that it is critical to tailor one’s approach to the source language and its typology.
Aaron Mueller, Garrett Nicolai, Arya McCarthy, Dylan Lewis, Winston Wu, David Yarowsky
LREC1
2020 Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages
abstract
Exploiting the broad translation of the Bible into the world’s languages, we train and distribute morphosyntactic tools for approximately one thousand languages, vastly outstripping previous distributions of tools devoted to the processing of inflectional morphology. Evaluation of the tools on a subset of available inflectional dictionaries demonstrates strong initial models, supplemented and improved through ensembling and dictionary-based reranking. Likewise, a novel type-to-token based evaluation metric allows us to confirm that models generalize well across rare and common forms alike
Garrett Nicolai, Dylan Lewis, Arya McCarthy, Aaron Mueller, Winston Wu, David Yarowsky
LREC4
2019 Modeling Color Terminology Across Thousands of Languages
abstract
Arya D. McCarthy, Winston Wu, Aaron Mueller, William Watson, David Yarowsky. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arya McCarthy, Winston Wu, Aaron Mueller, Bill Watson, David Yarowsky
EMNLP/IJCNLP (1)3
2019 Quantity doesn't buy quality syntax with neural language models
abstract
Marten van Schijndel, Aaron Mueller, Tal Linzen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Marten van Schijndel, Aaron Mueller, Tal Linzen
EMNLP/IJCNLP (1)2