VLDB 2026 Research / reviewers in the wild / expert
Emily Allaway
dblp:220/4016
· DBLP profile ↗
14ranked-venue papers
7as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generics are puzzling. Can language models find the missing piece?abstractGeneric sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic framework has proven difficult, in part because speakers use generics to generalise properties with widely different statistical prevalence. In this work, we study the implicit quantification and context-sensitivity of generics by leveraging language models as models of language. We create ConGen, a dataset of 2873 naturally occurring generic and quantified sentences in context, and define p-acceptability, a metric based on surprisal that is sensitive to quantification. Our experiments show generics are more context-sensitive than determiner quantifiers and about 20% of naturally occurring generics we analyze express weak generalisations. We also explore how human biases in stereotypes can be observed in language models. Gustavo Cilleruelo Calderón, Emily Allaway, Barry Haddow, Alexandra Birch |
COLING | 2 |
| 2025 | VISaGE: Understanding Visual Generics and ExceptionsabstractWhile Vision Language Models (VLMs) learn conceptual representations, in the form of generalized knowledge, during training, they are typically used to analyze individual instances.When evaluation instances are atypical, this paradigm results in tension between two priors in the model.The first is a pragmatic prior that the textual and visual input are both relevant, arising from VLM finetuning on congruent inputs; the second is a semantic prior that the conceptual representation is generally true for instances of the category.In order to understand how VLMs trade off these priors, we introduce a new evaluation dataset, VISaGE, consisting of both typical and exceptional images.In carefully balanced experiments, we show that conceptual understanding degrades when the assumption of congruency underlying the pragmatic prior is violated with incongruent images.This effect is stronger than the effect of the semantic prior when querying about individual instances. Stella Frank, Emily Allaway |
EMNLP | 2 |
| 2025 | Evaluating Defeasible Reasoning in LLMs with DEFREASINGabstractEmily Allaway, Kathleen McKeown. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Emily Allaway, Kathy McKeown |
NAACL (Long Papers) | 1 |
| 2024 | Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process GenericsabstractAbstract Large language models (LLMs) have garnered a great deal of attention for their exceptional generative performance on commonsense and reasoning tasks. In this work, we investigate LLMs’ capabilities for generalization using a particularly challenging type of statement: generics. Generics express generalizations (e.g., birds can fly) but do so without explicit quantification. They are notable because they generalize over their instantiations (e.g., sparrows can fly) yet hold true even in the presence of exceptions (e.g., penguins do not). For humans, these generic generalizations play a fundamental role in cognition, concept acquisition, and intuitive reasoning. We investigate how LLMs respond to and reason about generics. To this end, we first propose a framework grounded in pragmatics to automatically generate both exceptions and instantiations – collectively exemplars. We make use of focus—a pragmatic phenomenon that highlights meaning-bearing elements in a sentence—to capture the full range of interpretations of generics across different contexts of use. This allows us to derive precise logical definitions for exemplars and operationalize them to automatically generate exemplars from LLMs. Using our system, we generate a dataset of ∼370kexemplars across ∼17k generics and conduct a human validation of a sample of the generated data. We use our final generated dataset to investigate how LLMs reason about generics. Humans have a documented tendency to conflate universally quantified statements (e.g., all birds can fly) with generics. Therefore, we probe whether LLMs exhibit similar overgeneralization behavior in terms of quantification and in property inheritance. We find that LLMs do show evidence of overgeneralization, although they sometimes struggle to reason about exceptions. Furthermore, we find that LLMs may exhibit similar non-logical behavior to humans when considering property inheritance from generics. Emily Allaway, Chandra Bhagavatula, Jena D. Hwang, Kathy McKeown, Sarah-Jane Leslie |
Comput. Linguistics | 1 |
| 2023 | Penguins Don't Fly: Reasoning about Generics through Instantiations and ExceptionsabstractEmily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathleen McKeown, Doug Downey, Yejin Choi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathy McKeown, Doug Downey, Yejin Choi 0001 |
EACL | 1 |
| 2022 | SafeText: A Benchmark for Exploring Physical Safety in Language ModelsabstractSharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton, Desmond Upton Patton, Kathy McKeown, William Yang Wang |
EMNLP | 2 |
| 2021 | A Unified Feature Representation for Lexical ConnotationsabstractIdeological attitudes and stance are often expressed through subtle meanings of words and phrases.Understanding these connotations is critical to recognizing the cultural and emotional perspectives of the speaker.In this paper, we use distant labeling to create a new lexical resource representing connotation aspects for nouns and adjectives.Our analysis shows that it aligns well with human judgments.Additionally, we present a method for creating lexical representations that capture connotations within the embedding space and show that using the embeddings provides a statistically significant improvement on the task of stance detection when data is limited. Emily Allaway, Kathy McKeown |
EACL | 1 |
| 2021 | Sequential Cross-Document Coreference ResolutionabstractRelating entities and events in text is a key component of natural language understanding.Cross-document coreference resolution, in particular, is important for the growing interest in multi-document analysis tasks.In this work we propose a new model that extends the efficient sequential prediction paradigm for coreference resolution to cross-document settings and achieves competitive results for both entity and event coreference while providing strong evidence of the efficacy of both sequential models and higher-order inference in cross-document settings.Our model incrementally composes mentions into cluster representations and predicts links between a mention and the already constructed clusters, approximating a higher-order model.In addition, we conduct extensive ablation studies that provide new insights into the importance of various inputs and representation types in coreference. Emily Allaway, Miguel Ballesteros |
EMNLP (1) | 1 |
| 2021 | Human Rationales as Attribution Priors for Explainable Stance DetectionabstractAs NLP systems become better at detecting opinions and beliefs from text, it is important to ensure not only that models are accurate but also that they arrive at their predictions in ways that align with human reasoning.In this work, we present a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data.We show that in a data-scarce setting, our approach can improve the reasoning of a state-of-the-art classifierparticularly for inputs containing challenging phenomena such as sarcasm-at no cost in predictive performance.Furthermore, we demonstrate that attention weights surpass a leading attribution method in providing faithful explanations of our model's predictions, thus serving as a computationally cheap and reliable source of attributions for our model. Sahil Jayaram, Emily Allaway |
EMNLP (1) | 2 |
| 2021 | Adversarial Learning for Zero-Shot Stance Detection on Social MediaabstractEmily Allaway, Malavika Srikanth, Kathleen McKeown. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Emily Allaway, Malavika Srikanth, Kathy McKeown |
NAACL-HLT | 1 |
| 2020 | Event-Guided Denoising for Multilingual Relation LearningabstractGeneral purpose relation extraction has recently seen considerable gains in part due to a massively data-intensive distant supervision technique from Soares et al. (2019)that produces stateof-the-art results across many benchmarks.In this work, we present a methodology for collecting high quality training data for relation extraction from unlabeled text that achieves a nearrecreation of their zero-shot and few-shot results at a fraction of the training cost.Our approach exploits the predictable distributional structure of date-marked news articles to build a denoised corpus -the extraction process filters out low quality examples.We show that a smaller multilingual encoder trained on this corpus performs comparably to the current state-of-the-art (when both receive little to no fine-tuning) on few-shot and standard relation benchmarks in English and Spanish despite using many fewer examples (50k vs. 300mil+). Amith Ananthram, Emily Allaway, Kathy McKeown |
COLING | 2 |
| 2020 | Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic RepresentationsabstractStance detection is an important component of understanding hidden influences in everyday life.Since there are thousands of potential topics to take a stance on, most with little to no training data, we focus on zero-shot stance detection: classifying stance from no training examples.In this paper, we present a new dataset for zero-shot stance detection that captures a wider range of topics and lexical variation than in previous datasets.Additionally, we propose a new model for stance detection that implicitly captures relationships between topics using generalized topic representations and show that this model improves performance on a number of challenging linguistic phenomena. Emily Allaway, Kathy McKeown |
EMNLP (1) | 1 |
| 2019 | ATOMIC: An Atlas of Machine Commonsense for If-Then ReasoningabstractWe present ATOMIC, an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge. Compared to existing resources that center around taxonomic knowledge, ATOMIC focuses on inferential knowledge organized as typed if-then relations with variables (e.g., “if X pays Y a compliment, then Y will likely return the compliment”). We propose nine if-then relation types to distinguish causes vs. effects, agents vs. themes, voluntary vs. involuntary events, and actions vs. mental states. By generatively training on the rich inferential knowledge described in ATOMIC, we show that neural models can acquire simple commonsense capabilities and reason about previously unseen events. Experimental results demonstrate that multitask models that incorporate the hierarchical structure of if-then relation types lead to more accurate inference compared to models trained in isolation, as measured by both automatic and human evaluation. Maarten Sap, Ronan Le Bras 0001, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, Yejin Choi 0001 |
AAAI | 3 |
| 2018 | Event2Mind: Commonsense Inference on Events, Intents, and ReactionsabstractWe investigate a new commonsense inference task: given an event described in a short free-form text ("X drinks coffee in the morning"), a system reasons about the likely intents ("X wants to stay awake") and reactions ("X feels alert") of the event's participants.To support this study, we construct a new crowdsourced corpus of 25,000 event phrases covering a diverse range of everyday events and situations.We report baseline performance on this task, demonstrating that neural encoder-decoder models can successfully compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants.In addition, we demonstrate how commonsense inference on people's intents and reactions can help unveil the implicit gender inequality prevalent in modern movie scripts. 1 https://tinyurl.com/event2mind Hannah Rashkin, Maarten Sap, Emily Allaway, Noah A. Smith, Yejin Choi 0001 |
ACL (1) | 3 |