VLDB 2026 Research / reviewers in the wild / expert
Terra Blevins
dblp:184/3734
· DBLP profile ↗
16ranked-venue papers
9as first author
12since 2021 · last 2026
0009-0001-9473-1888ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew 0002, Marek Suppa, Hila Gonen, Shachar Mirkin, Vasile Florian Pais, Kaja Dobrovoljc, Voula Giouli, Jun Kevin, Eugene Jang, Eungseo Kim, Jeongyeon Seo, Xenophon Gialis, Yuval Pinter |
LREC | 1 |
| 2025 | Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language ModelsabstractHila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith |
NAACL (Long Papers) | 2 |
| 2024 | MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingabstractTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, Luke Zettlemoyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, Luke Zettlemoyer |
ACL (1) | 2 |
| 2024 | Translate to Disambiguate: Zero-shot Multilingual Word Sense Disambiguation with Pretrained Language ModelsabstractPretrained Language Models (PLMs) learn rich cross-lingual knowledge and perform well on diverse tasks such as translation and multilingual word sense disambiguation (WSD) when finetuned.However, they often struggle at disambiguating word sense in a zero-shot setting.To better understand this contrast, we present a new study investigating how well PLMs capture cross-lingual word sense with Contextual Word-Level Translation (C-WLT), an extension of word-level translation that prompts the model to translate a given word in context.We find that as the model size increases, PLMs encode more cross-lingual word sense knowledge and better use context to improve WLT performance.Building on C-WLT, we introduce a zero-shot prompting approach for WSD, tested on 18 languages from the XL-WSD dataset.Our method outperforms fully supervised baselines on recall for many evaluation languages without additional training or finetuning.This study presents a first step towards understanding how to best leverage the crosslingual knowledge inside PLMs for robust zeroshot reasoning in any language. Haoqiang Kang, Terra Blevins, Luke Zettlemoyer |
EACL (1) | 2 |
| 2024 | Breaking the Curse of Multilinguality with Cross-lingual Expert Language ModelsabstractTerra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer |
EMNLP | 1 |
| 2024 | Detecting Pretraining Data from Large Language ModelsabstractAlthough large language models (LLMs) are widely deployed, the data used to train them is rarely disclosed. Given the incredible scale of this data, up to trillions of tokens, it is all but certain that it includes potentially problematic text such as copyrighted materials, personally identifiable information, and test data for widely reported reference benchmarks. However, we currently have no way to know which data of these types is included or in what proportions. In this paper, we study the pretraining data detection problem: given a piece of text and black-box access to an LLM without knowing the pretraining data, can we determine if the model was trained on the provided text? To facilitate this study, we introduce a dynamic benchmark WIKIMIA that uses data created before and after model training to support gold truth detection. We also introduce a new detection method MIN-K PROB based on a simple hypothesis: an unseen example is likely to contain a few outlier words with low probabilities under the LLM, while a seen example is less likely to have words with such low probabilities. MIN-K PROB can be applied without any knowledge about the pretrainig corpus or any additional training, departing from previous detection methods that require training a reference model on data that is similar to the pretraining data. Moreover, our experiments demonstrate that MIN-K PROB achieves a 7.4% improvement on WIKIMIA over these previous methods. We apply MIN-K PROB to two real-world scenarios, copyrighted book detection and contaminated downstream example detection, and find that it to be a consistently effective solution. Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen 0001, Luke Zettlemoyer |
ICLR | 6 |
| 2024 | Universal NER: A Gold-Standard Multilingual Named Entity Recognition BenchmarkabstractStephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, Yuval Pinter. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Stephen Mayhew 0002, Terra Blevins, Shuheng Liu 0002, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson 0001, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter |
NAACL-HLT | 2 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 4 |
| 2023 | Prompting Language Models for Linguistic StructureabstractAlthough pretrained language models (PLMs) can be prompted to perform a wide range of language tasks, it remains an open question how much this ability comes from generalizable linguistic understanding versus surface-level lexical patterns.To test this, we present a structured prompting approach for linguistic structured prediction tasks, allowing us to perform zero-and few-shot sequence tagging with autoregressive PLMs.We evaluate this approach on part-of-speech tagging, named entity recognition, and sentence chunking, demonstrating strong few-shot performance in all cases.We also find that while PLMs contain significant prior knowledge of task labels due to task leakage into the pretraining corpus, structured prompting can also retrieve linguistic structure with arbitrary labels.These findings indicate that the in-context learning ability and linguistic knowledge of PLMs generalizes beyond memorization of their training data. Terra Blevins, Hila Gonen, Luke Zettlemoyer |
ACL (1) | 1 |
| 2022 | Analyzing the Mono- and Cross-Lingual Pretraining Dynamics of Multilingual Language ModelsabstractThe emergent cross-lingual transfer seen in multilingual pretrained models has sparked significant interest in studying their behavior.However, because these analyses have focused on fully trained multilingual models, little is known about the dynamics of the multilingual pretraining process.We investigate when these models acquire their in-language and crosslingual abilities by probing checkpoints taken from throughout XLM-R pretraining, using a suite of linguistic tasks.Our analysis shows that the model achieves high in-language performance early on, with lower-level linguistic skills acquired before more complex ones.In contrast, the point in pretraining when the model learns to transfer cross-lingually differs across language pairs.Interestingly, we also observe that, across many languages and tasks, the final model layer exhibits significant performance degradation over time, while linguistic knowledge propagates to lower layers of the network.Taken together, these insights highlight the complexity of multilingual pretraining and the resulting varied behavior for different languages over time. Terra Blevins, Hila Gonen, Luke Zettlemoyer |
EMNLP | 1 |
| 2022 | Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained ModelsabstractEnglish pretrained language models, which make up the backbone of many modern NLP systems, require huge amounts of unlabeled training data.These models are generally presented as being trained only on English text but have been found to transfer surprisingly well to other languages.We investigate this phenomenon and find that common English pretraining corpora actually contain significant amounts of non-English text: even when less than 1% of data is not English (well within the error rate of strong language classifiers), this leads to hundreds of millions of foreign language tokens in large-scale datasets.We then demonstrate that even these small percentages of non-English data facilitate cross-lingual transfer for models trained on them, with target language performance strongly correlated to the amount of in-language data seen during pretraining.In light of these findings, we argue that no model is truly monolingual when pretrained at scale, which should be considered when evaluating cross-lingual transfer. Terra Blevins, Luke Zettlemoyer |
EMNLP | 1 |
| 2021 | FEWS: Large-Scale, Low-Shot Word Sense Disambiguation with the DictionaryabstractCurrent models for Word Sense Disambiguation (WSD) struggle to disambiguate rare senses, despite reaching human performance on global WSD metrics.This stems from a lack of data for both modeling and evaluating rare senses in existing WSD datasets.In this paper, we introduce FEWS (Few-shot Examples of Word Senses), a new low-shot WSD dataset automatically extracted from example sentences in Wiktionary.FEWS has high sense coverage across different natural language domains and provides: (1) a large training set that covers many more senses than previous datasets and (2) a comprehensive evaluation set containing few-and zero-shot examples of a wide variety of senses.We establish baselines on FEWS with knowledgebased and neural WSD approaches and present transfer learning experiments demonstrating that models additionally trained with FEWS better capture rare senses in existing WSD datasets.Finally, we find humans outperform the best baseline models on FEWS, indicating that FEWS will support significant future work on low-shot WSD. Terra Blevins, Mandar Joshi, Luke Zettlemoyer |
EACL | 1 |
| 2020 | Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encodersabstractA major obstacle in Word Sense Disambiguation (WSD) is that word senses are not uniformly distributed, causing existing models to generally perform poorly on senses that are either rare or unseen during training.We propose a bi-encoder model that independently embeds (1) the target word with its surrounding context and (2) the dictionary definition, or gloss, of each sense.The encoders are jointly optimized in the same representation space, so that sense disambiguation can be performed by finding the nearest sense embedding for each target word embedding.Our system outperforms previous state-of-the-art models on English all-words WSD; these gains predominantly come from improved performance on rare senses, leading to a 31.1% error reduction on less frequent senses over prior work.This demonstrates that rare senses can be more effectively disambiguated by modeling their definitions. Terra Blevins, Luke Zettlemoyer |
ACL | 1 |
| 2019 | Better Character Language Modeling through MorphologyabstractWe incorporate morphological supervision into character language models (CLMs) via multitasking and show that this addition improves bits-per-character (BPC) performance across 24 languages, even when the morphology data and language modeling data are disjoint.Analyzing the CLMs shows that inflected words benefit more from explicitly modeling morphology than uninflected words, and that morphological supervision improves performance even as the amount of language modeling data grows.We then transfer morphological supervision across languages to improve language modeling performance in the low-resource setting. Terra Blevins, Luke Zettlemoyer |
ACL (1) | 1 |
| 2016 | Mining Paraphrasal Typed Templates from a Plain Text CorpusabstractFinding paraphrases in text is an important task with implications for generation, summarization and question answering, among other applications.Of particular interest to those applications is the specific formulation of the task where the paraphrases are templated, which provides an easy way to lexicalize one message in multiple ways by simply plugging in the relevant entities.Previous work has focused on mining paraphrases from parallel and comparable corpora, or mining very short sub-sentence synonyms and paraphrases.In this paper we present an approach which combines distributional and KB-driven methods to allow robust mining of sentence-level paraphrasal templates, utilizing a rich type system for the slots, from a plain text corpus. Or Biran, Terra Blevins, Kathy McKeown |
ACL (1) | 2 |
| 2016 | Automatically Processing Tweets from Gang-Involved Youth: Towards Detecting Loss and AggressionabstractViolence is a serious problems for cities like Chicago and has been exacerbated by the use of social media by gang-involved youths for taunting rival gangs. We present a corpus of tweets from a young and powerful female gang member and her communicators, which we have annotated with discourse intention, using a deep read to understand how and what triggered conversations to escalate into aggression. We use this corpus to develop a part-of-speech tagger and phrase table for the variant of English that is used and a classifier for identifying tweets that express grieving and aggression. Terra Blevins, Robert Kwiatkowski, Jamie C. Macbeth, Kathy McKeown, Desmond Upton Patton, Owen Rambow |
COLING | 1 |