Dan Garrette

dblp:117/4050 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 The Impact of Depth on Compositional Generalization in Transformer Language Models
abstract
Jackson Petty, Sjoerd Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, Tal Linzen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jackson Petty, Sjoerd van Steenkiste, Ishita Dasgupta 0001, Fei Sha, Dan Garrette, Tal Linzen
NAACL-HLT5
2023 Character-Aware Models Improve Visual Text Rendering
abstract
Rosanne Liu, Dan Garrette, Chitwan Saharia, William Chan, Adam Roberts, Sharan Narang, Irina Blok, Rj Mical, Mohammad Norouzi, Noah Constant. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rosanne Liu, Dan Garrette, Chitwan Saharia, Adam Roberts, Sharan Narang, Irina Blok, RJ Mical, Mohammad Norouzi 0002, Noah Constant
ACL (1)2
2023 Dialect-robust Evaluation of Generated Text
abstract
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann
ACL (1)6
2023 How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
abstract
Multilingual language models (MLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languages' data.Impressive performance in zero-shot cross-lingual transfer shows that these models are able to exploit this property.Yet, it remains unclear to what extent, and under which conditions, languages rely on each other's data.To answer this question, we use TracIn (Pruthi et al., 2020), a training data attribution (TDA) method, to retrieve training samples from multilingual data that are most influential for test predictions in a given language.This allows us to analyse cross-lingual sharing mechanisms of MLMs from a new perspective.While previous work studied cross-lingual sharing at the model parameter level, we present the first approach to study it at the data level.We find that MLMs rely on data from multiple languages during fine-tuning and this reliance increases as finetuning progresses.We further find that training samples from other languages can both reinforce and complement the knowledge acquired from data of the test language itself.
Rochelle Choenni, Dan Garrette, Ekaterina Shutova
EMNLP2
2023 Cross-Lingual Transfer with Language-Specific Subnetworks for Low-Resource Dependency Parsing
abstract
Abstract Large multilingual language models typically share their parameters across all languages, which enables cross-lingual task transfer, but learning can also be hindered when training updates from different languages are in conflict. In this article, we propose novel methods for using language-specific subnetworks, which control cross-lingual parameter sharing, to reduce conflicts and increase positive transfer during fine-tuning. We introduce dynamic subnetworks, which are jointly updated with the model, and we combine our methods with meta-learning, an established, but complementary, technique for improving cross-lingual transfer. Finally, we provide extensive analyses of how each of our methods affects the models.
Rochelle Choenni, Dan Garrette, Ekaterina Shutova
Comput. Linguistics2
2023 Scaling Up Models and Data with t5x and seqio
abstract
Scaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio.
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo
J. Mach. Learn. Res.31
2023 FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation
abstract
Abstract We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regional variants each of Portuguese and Mandarin Chinese. Source documents are selected to enable detailed analysis of phenomena of interest, including lexically distinct terms and distractor terms. We explore automatic evaluation metrics for FRMT and validate their correlation with expert human evaluation across both region-matched and mismatched rating scenarios. Finally, we present a number of baseline models for this task, and offer guidelines for how researchers can train, evaluate, and compare their own models. Our dataset and evaluation code are publicly available: https://bit.ly/frmt-task.
Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia, Dan Garrette, Jason Riesa, Orhan Firat, Noah Constant
Trans. Assoc. Comput. Linguistics5
2022 Canine: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
abstract
Abstract Pipelined NLP systems have largely been superseded by end-to-end neural modeling, yet nearly all commonly used models still require an explicit tokenization step. While recent tokenization approaches based on data-derived subword lexicons are less brittle than manually engineered tokenizers, these techniques are not equally suited to all languages, and the use of any fixed vocabulary may limit a model’s ability to adapt. In this paper, we present Canine, a neural encoder that operates directly on character sequences—without explicit tokenization or vocabulary—and a pre-training strategy that operates either directly on characters or optionally uses subwords as a soft inductive bias. To use its finer-grained input effectively and efficiently, Canine combines downsampling, which reduces the input sequence length, with a deep transformer stack, which encodes context. Canine outperforms a comparable mBert model by 5.7 F1 on TyDi QA, a challenging multilingual benchmark, despite having fewer model parameters.
Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting
Trans. Assoc. Comput. Linguistics2
2021 XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation
abstract
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson
EMNLP (1)9
2021 Frequency Effects on Syntactic Rule Learning in Transformers
abstract
Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules.We investigate this question using the case study of BERT's performance on English subject-verb agreement.Unlike prior work, we train multiple instances of BERT from scratch, allowing us to perform a series of controlled interventions at pre-training time.We show that BERT often generalizes well to subject-verb pairs that never occurred in training, suggesting a degree of rule-governed behavior.We also find, however, that performance is heavily influenced by word frequency, with experiments showing that both the absolute frequency of a verb form, as well as the frequency relative to the alternate inflection, are causally implicated in the predictions BERT makes at inference time.Closer analysis of these frequency effects reveals that BERT's behavior is consistent with a system that correctly applies the SVA rule in general but struggles to overcome strong training priors and to estimate agreement features (singular vs. plural) on infrequent lexical items.
Jason Wei, Dan Garrette, Tal Linzen, Ellie Pavlick
EMNLP (1)2
2020 Improving Multilingual Models with Language-Clustered Vocabularies
abstract
State-of-the-art multilingual models depend on vocabularies that cover all of the languages the model will expect to see at inference time, but the standard methods for generating those vocabularies are not ideal for massively multilingual applications.In this work, we introduce a novel procedure for multilingual vocabulary generation that combines the separately trained vocabularies of several automatically derived language clusters, thus balancing the trade-off between cross-lingual subword sharing and language-specific vocabularies.Our experiments show improvements across languages on key multilingual benchmark tasks TYDI QA (+2.9 F1), XNLI (+2.1%), and WikiAnn NER (+2.8 F1) and factor of 8 reduction in out-of-vocabulary rate, all without increasing the size of the model or data.
Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, Jason Riesa
EMNLP (1)2
2020 TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
abstract
Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA—a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology—the set of linguistic features each language expresses—such that we expect models performing well on this set to generalize across a large number of the world’s languages. We present a quantitative analysis of the data quality and example-level qualitative linguistic analyses of observed language phenomena that would not be found in English-only corpora. To provide a realistic information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but don’t know the answer yet, and the data is collected directly in each language without the use of translation.
Jonathan H. Clark, Jennimaria Palomaki, Vitaly Nikolaev, Eunsol Choi, Dan Garrette, Michael Collins 0001, Tom Kwiatkowski
Trans. Assoc. Comput. Linguistics5
2019 How Multilingual is Multilingual BERT?
abstract
In this paper, we show that Multilingual BERT (M-BERT), released by Devlin et al. (2019) as a single language model pre-trained from monolingual corpora in 104 languages, is surprisingly good at zero-shot cross-lingual model transfer, in which task-specific annotations in one language are used to fine-tune the model for evaluation in another language.To understand why, we present a large number of probing experiments, showing that transfer is possible even to languages in different scripts, that transfer works best between typologically similar languages, that monolingual corpora can train models for code-switching, and that the model can find translation pairs.From these results, we can conclude that M-BERT does create multilingual representations, but that these representations exhibit systematic deficiencies affecting certain language pairs.
Telmo Pires, Eva Schlinger, Dan Garrette
ACL (1)3
2018 Part-of-Speech Tagging for Code-Switched, Transliterated Texts without Explicit Language Identification
abstract
Code-switching, the use of more than one language within a single utterance, is ubiquitous in much of the world, but remains a challenge for NLP largely due to the lack of representative data for training models.In this paper, we present a novel model architecture that is trained exclusively on monolingual resources, but can be applied to unseen codeswitched text at inference time.The model accomplishes this by jointly maintaining separate word representations for each of the possible languages-or scripts in the case of transliteration-allowing each to contribute to inferences without forcing the model to commit to a language.Experiments on Hindi-English part-of-speech tagging demonstrate that our approach outperforms standard models when training on monolingual text without transliteration, and testing on code-switched text with alternate scripts.
Kelsey Ball, Dan Garrette
EMNLP2
2016 An Unsupervised Model of Orthographic Variation for Historical Document Transcription
abstract
Historical documents frequently exhibit extensive orthographic variation, including archaic spellings and obsolete shorthand.OCR tools typically seek to produce so-called diplomatic transcriptions that preserve these variants, but many end tasks require transcriptions with normalized orthography.In this paper, we present a novel joint transcription model that learns, unsupervised, a probabilistic mapping between modern orthography and that used in the document.Our system thus produces dual diplomatic and normalized transcriptions simultaneously, and achieves a 35% relative error reduction over a state-of-the-art OCR model on diplomatic transcription, and a 46% reduction on normalized transcription.
Dan Garrette, Hannah Alpert-Abrams
HLT-NAACL1
2015 Weakly-Supervised Grammar-Informed Bayesian CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that, in combination with a small universal set of rules, specify the syntactic configurations in which they may occur. Categories are selected from a large, recursively-defined set; this leads to high word-to-category ambiguity, which is one of the primary factors that make learning CCG parsers difficult, especially in the face of little data. Previous work has shown that learning sequence models for CCG tagging can be improved by using linguistically-motivated prior probability distributions over potential categories. We extend this approach to the task of learning a CCG parser from weak supervision. We present a Bayesian formulation for CCG parser induction that assumes only supervision in the form of an incomplete tag dictionary mapping some word types to sets of potential categories. Our approach outperforms a baseline model trained with uniform priors by exploiting universal, intrinsic properties of the CCG formalism to bias the model toward simpler, more cross-linguistically common categories.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
AAAI1
2015 A Supertag-Context Model for Weakly-Supervised CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that specify the syntactic configurations in which they may occur.We present a novel parsing model with the capacity to capture the associative adjacent-category relationships intrinsic to CCG by parameterizing the relationships between each constituent label and the preterminal categories directly to its left and right, biasing the model toward constituent categories that can combine with their contexts.This builds on the intuitions of Klein and Manning's (2002) "constituentcontext" model, which demonstrated the value of modeling context, but has the advantage of being able to exploit the properties of CCG.Our experiments show that our model outperforms a baseline in which this context information is not captured.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL1
2015 Unsupervised Code-Switching for Multilingual Historical Document Transcription
abstract
Dan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Dan Klein. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Dan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Daniel Klein 0001
HLT-NAACL1
2014 Weakly-Supervised Bayesian Learning of a CCG Supertagger
abstract
We present a Bayesian formulation for weakly-supervised learning of a Combinatory Categorial Grammar (CCG) supertagger with an HMM.We assume supervision in the form of a tag dictionary, and our prior encourages the use of crosslinguistically common category structures as well as transitions between tags that can combine locally according to CCG's combinators.Our prior is theoretically appealing since it is motivated by languageindependent, universal properties of the CCG formalism.Empirically, we show that it yields substantial improvements over previous work that used similar biases to initialize an EM-based learner.Additional gains are obtained by further shaping the prior with corpus-specific information that is extracted automatically from raw text and a tag dictionary.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL1
2013 Real-World Semi-Supervised Learning of POS-Taggers for Low-Resource Languages
Dan Garrette, Jason Mielens, Jason Baldridge
ACL (1)1
2013 Learning a Part-of-Speech Tagger from Two Hours of Annotation
Dan Garrette, Jason Baldridge
HLT-NAACL1
2012 Type-Supervised Hidden Markov Models for Part-of-Speech Tagging with Incomplete Tag Dictionaries
Dan Garrette, Jason Baldridge
EMNLP-CoNLL1