EDBT 2026 Demo / reviewers in the wild / expert
Rachel Rudinger
dblp:136/8740
· DBLP profile ↗
41ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-5506-4701ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksabstractNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 9 |
| 2026 | Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of PlausibilityabstractWe investigate the degree to which human (and LLM) plausibility judgments of multiplechoice commonsense benchmark answers are subject to influence by (im)plausibility arguments for or against an answer, in particular, using rationales generated by LLMs.We collect 3, 000 plausibility judgments from humans and another 13, 600 judgments from LLMs.Overall, we observe increases and decreases in mean human plausibility ratings in the presence of LLM-generated PRO and CON rationales, respectively, suggesting that, on the whole, human judges find these rationales convincing.Experiments with LLMs reveal similar patterns of influence.Our findings demonstrate a novel use of LLMs for studying aspects of human cognition, while also raising practical concerns that, even in domains where humans are "experts" (i.e., common sense), LLMs have the potential to exert considerable influence on people's beliefs. 1 Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger |
ACL (1) | 4 |
| 2025 | On the Mutual Influence of Gender and Occupation in LLM RepresentationsabstractWe examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names in LLMs influence each other mutually.We find that LLMs' first-name gender representations correlate with real-world gender statistics associated with the name, and are influenced by the co-occurrence of stereotypically feminine or masculine occupations.Additionally, we study the influence of firstname gender representations on LLMs in a downstream occupation prediction task and their potential as an internal metric to identify extrinsic model biases.While feminine firstname embeddings often raise the probabilities for female-dominated jobs (and vice versa for male-dominated jobs), reliably using these internal gender representations for bias detection remains challenging.Is Jody male or female? Female MaleJody is a nurse.Is Jody male or female? Female MaleJody is a comedian.Is Jody male or female? Female Male Haozhe An, Connor Baumler, Abhilasha Sancheti, Rachel Rudinger |
ACL (1) | 4 |
| 2025 | Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasabstractNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 0005, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 5 |
| 2025 | Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveabstractMultiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform.We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity; 2) match LLM use cases; and 3) fully test knowledge.We instead advocate for generative formats based on human testing-where LLMs construct and explain answers-better capturing user needs and knowledge while remaining easy to score.We then show even when MCQA is a useful format, its datasets suffer from: leakage; unanswerability; shortcuts; and saturation.In each issue, we give fixes from education, like rubrics to guide MCQ writing; scoring methods to bridle guessing; and Item Response Theory to build harder MCQs.Lastly, we discuss LLM errors in MCQA-robustness, biases, and unfaithful explanations-showing how our prior solutions better measure or address these issues.While we do not need to desert MCQA, we encourage more efforts in refining the task based on educational testing, advancing evaluations. Q1. What's wrong with MCQA's format?A) It doesn't apply to many tasks ( §3.1) B) It's misaligned with LLM use cases ( §3.2) C) It doesn't fully test knowledge ( §3.3) Q2.What's wrong with MCQA datasets?A) Test sets are contaminated ( §5.1) B) They have unanswerable questions ( §5.2) C) They contain shortcuts ( §5.3) D) They're too easy for LLMs ( §5.4) Q3.How do LLMs struggle with MCQA?A) They lack robustness ( §6.1) B) They exhibit biases ( §6.2) C) They give unfaithful explanations ( §6.3) Q4.How can insights from education improve MCQA?A) Improve knowledge testing via generative formats ( §4) B) Combat test set leakage with fresh questions ( §5.1) C) Write MCQs informed by educational rubrics ( §5.2) D) Use calibration scoring to curb guessing ( §5.3.1)E) Find harder MCQs with item response theory ( §5.4.1) Nishant Balepur, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 2 |
| 2025 | Multiple LLM Agents Debate for Equitable Cultural AlignmentabstractLarge Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-turn approaches, we propose to exploit the complementary strengths of multiple LLMs to promote cultural adaptability. We introduce a Multi-Agent Debate framework, where two LLM-based agents debate over a cultural scenario and collaboratively reach a final decision. We propose two variants: one where either LLM agents exclusively debate and another where they dynamically choose between self-reflection and debate during their turns. We evaluate these approaches on 7 open-weight LLMs (and 21 LLM combinations) using the NormAd-ETI benchmark for social etiquette norms in 75 countries. Experiments show that debate improves both overall accuracy and cultural group parity over single-LLM baselines. Notably, multi-agent debate enables relatively small LLMs (7-9B) to achieve accuracies comparable to that of a much larger model (27B parameters). Dayeon Ki, Rachel Rudinger, Tianyi Zhou 0001, Marine Carpuat |
ACL (1) | 2 |
| 2025 | Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat LogsabstractRupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger, Claire Bonial, Philip Resnik. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Rupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger, Claire Bonial, Philip Resnik |
ACL (1) | 4 |
| 2025 | No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking QuestionsabstractQuestions help unlock information to satisfy users' information needs.However, when the question is poorly posed, answerers (whether human or computer) may struggle to answer the question in a way that satisfies the asker, despite possibly knowing everything necessary to address the asker's latent information need.Using Reddit question-answer interactions from r/NoStupidQuestions, we develop a computational framework grounded in linguistic theory to study poorly-posedness of questions by generating spaces of potential interpretations of questions and computing distributions over these spaces based on interpretations chosen by both human answerers in the Reddit question thread, as well as by a suite of large language models.Both humans and models struggle to converge on dominant interpretations when faced with poorly posed questions, but employ different strategies: humans focus on specific interpretations through question negotiation, while models attempt comprehensive coverage by addressing many interpretations simultaneously. Neha Srikanth, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 2 |
| 2025 | A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps UsersabstractNishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng 0005, Fumeng Yang, Rachel Rudinger, Jordan L. Boyd-Graber |
EMNLP | 7 |
| 2025 | 'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?abstractLarge Language Models (LLMs) are increasingly involved in high-stakes domains, yet how they reason about socially-sensitive decisions still remains underexplored.We present a largescale audit of LLMs' treatment of socioeconomic status (SES) in college admissions decisions using a novel dual-process framework inspired by cognitive science.Leveraging a synthetic dataset of 30,000 applicant profiles 1 grounded in real-world correlations, we prompt 4 open-source LLMs (Qwen 2, Mistral v0.3, Gemma 2, Llama 3.1) under 2 modes: a fast, decision-only setup (System 1) and a slower, explanation-based setup (System 2).Results from 5 million prompts reveals that LLMs consistently favor low-SES applicants-even when controlling for academic performance-and that System 2 amplifies this tendency by explicitly invoking SES as compensatory justification, highlighting both their potential and volatility as decision-makers.We then propose DPAF, a dual-process audit framework to probe LLMs' reasoning behaviors in sensitive applications. Huy Nghiem, Phuong-Anh Nguyen-Le, John Prindle, Rachel Rudinger, Hal Daumé III |
EMNLP | 4 |
| 2025 | Natural Language Inference Improves Compositionality in Vision-Language ModelsabstractCompositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (LLMs) to break them down into subsets of questions and answers. However, these methods primarily operate on the surface level, failing to incorporate deeper lexical understanding while introducing incorrect assumptions generated by the LLM. In response to these issues, we present Caption Expansion with Contradictions and Entailments (CECE), a principled approach that leverages Natural Language Inference (NLI) to generate entailments and contradictions from a given premise. CECE produces lexically diverse sentences while maintaining their core meaning. Through extensive experiments, we show that CECE enhances interpretability and reduces overreliance on biased or superficial features. By balancing CECE along the original premise, we achieve significant improvements over previous methods without requiring additional fine-tuning, producing state-of-the-art results on benchmarks that score agreement with human judgments for image-text alignment, and achieving an increase in performance on Winoground of $+19.2\%$ (group score) and $+12.9\%$ on EqBen (group score) over the best prior work (finetuned with targeted data). Paola Cascante-Bonilla, Yang Trista Cao, Hal Daumé III, Rachel Rudinger |
ICLR | 5 |
| 2025 | Language Models Predict Empathy Gaps Between Social In-groups and Out-groupsabstractYu Hou, Hal Daumé Iii, Rachel Rudinger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hal Daumé III, Rachel Rudinger |
NAACL (Long Papers) | 3 |
| 2025 | NLI under the Microscope: What Atomic Hypothesis Decomposition RevealsabstractNeha Srikanth, Rachel Rudinger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Neha Srikanth, Rachel Rudinger |
NAACL (Long Papers) | 2 |
| 2024 | Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?abstractMultiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs).To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices.In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain.To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference.Our key findings are threefold.First, we find no evidence that the choices-only accuracy stems from memorization alone.Second, priors over individual choices do not fully explain choicesonly accuracy, hinting that LLMs use the group dynamics of choices.Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question.Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA.Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making. 1Question: Which of these contains only a solution?Answer: (B) Question: Which can be considered a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Question: Which can be considered a solution?Step 2: Answer the Question from Step 1 Classify Choice (A) Correctness Classify Choice (B) Correctness ... Abductive Question Inference ( §6) Question: Which of these contains only a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Full MCQA Prompt Choices-only Prompt ( §3) LLMs Can Perform MCQA with no Question, but how? Classify Choice (D) Correctness Question: Which of these contains only a solution?Choices: (A) \n (B) \n (C) \n (D) \n Answer: (B) Step 1: Guess the Question No Choices Empty Choices Choice: a can of mixed fruit Answer: False Choice: a bottle of juice Answer: True Choice: a jar of pickles Answer: False Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger |
ACL (1) | 3 |
| 2024 | Assessing Common Ground through Language-based Cultural Consensus in Humans and Large Language Models
Sophie Domanski, Rachel Rudinger, Marine Carpuat, Patrick Shafto, Yi Ting Huang |
CogSci | 2 |
| 2024 | Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the USabstractRecent work has highlighted the culturallycontingent nature of commonsense knowledge (Shen et al., 2024).We introduce AMAMMERE (/A:.mA:.mu:.reI/),from the Akan word meaning 'culture.'This test set of 525 multiple-choice questions is designed to evaluate the commonsense knowledge of English LLMs, relative to the cultural contexts of Ghana and the United States.To create AMAMMERE, we select a set of multiplechoice questions (MCQs) from existing commonsense datasets and rewrite them in a multistage process involving surveys of Ghanaian and U.S. participants.In three rounds of surveys, participants from both pools are solicited to ( 1) write correct and incorrect answer choices, ( 2) rate individual answer choices on a 5-point Likert scale, and (3) select the best answer choice from the newly-constructed MCQ items, in a final validation step.By engaging participants at multiple stages, our procedure ensures that participant perspectives are incorporated both in the creation and validation of test items, resulting in high levels of agreement within each pool.We evaluate several off-the-shelf English LLMs on AMAMMERE. 1 Uniformly, models prefer answers choices that align with the preferences of U.S. annotators over Ghanaian annotators.Additionally, when test items specify a cultural context (Ghana or the U.S.), models exhibit some ability to adapt, but performance is consistently better in U.S. contexts than Ghanaian.As large resources are devoted to the advancement of English LLMs, our findings underscore the need for culturally adaptable models and evaluations to meet the needs of diverse English-speaking populations around the world. Christabel Acquaye, Haozhe An, Rachel Rudinger |
EMNLP | 3 |
| 2024 | On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language ModelsabstractWe study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship prediction.We show that models are less likely to predict romantic relationships for (a) same-gender character pairs than different-gender pairs; and (b) intra/inter-racial character pairs involving Asian names as compared to Black, Hispanic, or White names.We examine the contextualized embeddings of first names and find that gender for Asian names is less discernible than non-Asian names.We discuss the social implications of our findings, underlining the need to prioritize the development of inclusive and equitable technology.2 Except for Hispanic wherein we did not get any names in 5 -10% bin and only 1 name in 25 -50% bin.480 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0 % Female 0-2 2-5 5-10 10-25 25-50 50-75 75-90 90-95 95-98 98-100 % Female 0.56 0.55 0.59 0.58 0.61 0.59 0.69 0.68 0.64 0.72 0.56 0.52 0.57 0.58 0.59 0.58 0.65 0.66 0.63 0.69 0.61 0.58 0.63 0.60 0.64 0.62 0.70 0.69 0.67 0.72 0.58 0.57 0.59 0.60 0.62 0.59 0.67 0.66 0.63 0.70 0.61 0.60 0.64 0.62 0.65 0.62 0.69 0.69 0.66 0.69 0.59 0.56 0.61 0.60 0.61 0.60 0.66 0.64 0.62 0.67 0.66 0.64 0.67 0.66 0.67 0.64 0.69 0.66 0.65 0.67 0.67 0.65 0.68 0.66 0.68 0.64 0.67 0.67 0.64 0.67 0.64 0.61 0.65 0.63 0.65 0.61 0.65 0.63 0.62 0.65 0.70 0.66 0.70 0.68 0.68 0.66 0.68 0.67 0.65 0.68 Male Neutral Female Male Neutral Female Asian (Recall) 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0 Abhilasha Sancheti, Haozhe An, Rachel Rudinger |
EMNLP | 3 |
| 2024 | Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question AnsweringabstractNeha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh Nguyen, Rachel Rudinger, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Neha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh C. Nguyen, Rachel Rudinger, Jordan L. Boyd-Graber |
NAACL-HLT | 6 |
| 2024 | How Often Are Errors in Natural Language Reasoning Due to Paraphrastic Variability?abstractAbstract Large language models have been shown to behave inconsistently in response to meaning-preserving paraphrastic inputs. At the same time, researchers evaluate the knowledge and reasoning abilities of these models with test evaluations that do not disaggregate the effect of paraphrastic variability on performance. We propose a metric, PC, for evaluating the paraphrastic consistency of natural language reasoning models based on the probability of a model achieving the same correctness on two paraphrases of the same problem. We mathematically connect this metric to the proportion of a model’s variance in correctness attributable to paraphrasing. To estimate PC, we collect ParaNlu, a dataset of 7,782 human-written and validated paraphrased reasoning problems constructed on top of existing benchmark datasets for defeasible and abductive natural language inference.1 Using ParaNlu, we measure the paraphrastic consistency of several model classes and show that consistency dramatically increases with pretraining but not fine-tuning. All models tested exhibited room for improvement in paraphrastic consistency. Neha Srikanth, Marine Carpuat, Rachel Rudinger |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning ModelsabstractA common limitation of diagnostic tests for detecting social biases in NLP models is that they may only detect stereotypic associations that are pre-specified by the designer of the test.Since enumerating all possible problematic associations is infeasible, it is likely these tests fail to detect biases that are present in a model but not pre-specified by the designer.To address this limitation, we propose SODAPOP 1 (SOcial bias Discovery from Answers about PeOPle), an approach for automatic social bias discovery in social commonsense question-answering.The SODAPOP pipeline generates modified instances from the Social IQa dataset (Sap et al., 2019b) by ( 1) substituting names associated with different demographic groups, and (2) generating many distractor answers from a masked language model.By using a social commonsense model to score the generated distractors, we are able to uncover the model's stereotypic associations between demographic groups and an open set of words.We also test SODAPOP on debiased models and show the limitations of multiple state-of-the-art debiasing algorithms. Haozhe An, Zongxia Li, Jieyu Zhao 0001, Rachel Rudinger |
EACL | 4 |
| 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and ProhibitionsabstractReviewing and comprehending key obligations, entitlements, and prohibitions in legal contracts can be a tedious task due to their length and domain-specificity.Furthermore, the key rights and duties requiring review vary for each contracting party.In this work, we propose a new task of party-specific extractive summarization for legal contracts to facilitate faster reviewing and improved comprehension of rights and duties.To facilitate this, we curate a dataset comprising of party-specific pairwise importance comparisons annotated by legal experts, covering ∼293K sentence pairs that include obligations, entitlements, and prohibitions extracted from lease agreements.Using this dataset, we train a pairwise importance ranker and propose a pipeline-based extractive summarization system that generates a party-specific contract summary.We establish the need for incorporating domain-specific notion of importance during summarization by comparing our system against various baselines using both automatic and human evaluation methods 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 4 |
| 2022 | Entailment Relation Aware Paraphrase GenerationabstractWe introduce a new task of entailment relation aware paraphrase generation which aims at generating a paraphrase conforming to a given entailment relation (e.g. equivalent, forward entailing, or reverse entailing) with respect to a given input. We propose a reinforcement learning-based weakly-supervised paraphrasing system, ERAP, that can be trained using existing paraphrase and natural language inference (NLI) corpora without an explicit task-specific corpus. A combination of automated and human evaluations show that ERAP generates paraphrases conforming to the specified entailment relation and are of good quality as compared to the baselines and uncontrolled paraphrasing systems. Using ERAP for augmenting training data for downstream textual entailment task improves performance over an uncontrolled paraphrasing system, and introduces fewer training artifacts, indicating the benefit of explicit control during paraphrasing. Abhilasha Sancheti, Balaji Vasan Srinivasan, Rachel Rudinger |
AAAI | 3 |
| 2022 | Agent-Specific Deontic Modality Detection in Legal LanguageabstractLegal documents are typically long and written in legalese, which makes it particularly difficult for laypeople to understand their rights and duties.While natural language understanding technologies can be valuable in supporting such understanding in the legal domain, the limited availability of datasets annotated for deontic modalities in the legal domain, due to the cost of hiring experts and privacy issues, is a bottleneck.To this end, we introduce, LEXDE-MOD, a corpus of English contracts annotated with deontic modality expressed with respect to a contracting party or agent along with the modal triggers.We benchmark this dataset on two tasks: (i) agent-specific multi-label deontic modality classification, and (ii) agent-specific deontic modality and trigger span detection using Transformer-based (Vaswani et al., 2017) language models.Transfer learning experiments show that the linguistic diversity of modal expressions in LEXDEMOD generalizes reasonably from lease to employment and rental agreements.A small case study indicates that a model trained on LEXDEMOD can detect red flags with high recall.We believe our work offers a new research direction for deontic modality detection in the legal domain 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 4 |
| 2022 | Recognition of They/Them as Singular Personal Pronouns in Coreference ResolutionabstractAs using they/them as personal pronouns becomes increasingly common in English, it is important that coreference resolution systems work as well for individuals who use personal "they" as they do for those who use gendered personal pronouns.We introduce a new benchmark for coreference resolution systems which evaluates singular personal "they" recognition.Using these WinoNB schemas, we evaluate a number of publicly available coreference resolution systems and confirm their bias toward resolving "they" pronouns as plural. Connor Baumler, Rachel Rudinger |
NAACL-HLT | 2 |
| 2022 | Theory-Grounded Measurement of U.S. Social Stereotypes in English Language ModelsabstractYang Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, Linda Zou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yang Trista Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, Linda Zou |
NAACL-HLT | 4 |
| 2022 | Partial-input baselines show that NLI models can ignore context, but they don'tabstractWhen strong partial-input baselines reveal artifacts in crowdsourced NLI datasets, the performance of full-input models trained on such datasets is often dismissed as reliance on spurious correlations.We investigate whether stateof-the-art NLI models are capable of overriding default inferences made by a partial-input baseline.We introduce an evaluation set of 600 examples consisting of perturbed premises to examine a RoBERTa model's sensitivity to edited contexts.Our results indicate that NLI models are still capable of learning to condition on context-a necessary component of inferential reasoning-despite being trained on artifact-ridden datasets. Neha Srikanth, Rachel Rudinger |
NAACL-HLT | 2 |
| 2021 | Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionabstractThe black-box nature of neural models has motivated a line of research that aims to generate natural language rationales to explain why a model made certain predictions. Such rationale generation models, to date, have been trained on dataset-specific crowdsourced rationales, but this approach is costly and is not generalizable to new tasks and domains. In this paper, we investigate the extent to which neural models can reason about natural language rationales that explain model predictions, relying only on distant supervision with no additional annotation cost for human-written rationales. We investigate multiple ways to automatically generate rationales using pre-trained language models, neural knowledge models, and distant supervision from related tasks, and train generative models capable of composing explanatory rationales for unseen instances. We demonstrate our approach on the defeasible inference task, a nonmonotonic reasoning task in which an inference may be strengthened or weakened when new information (an update) is introduced. Our model shows promises at generating post-hoc rationales explaining why an inference is more or less likely given the additional information, however, it mostly generates trivial rationales reflecting the fundamental limitations of neural language models. Conversely, the more realistic setup of jointly predicting the update or its type and generating rationale is more challenging, suggesting an important future direction. Faeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin Choi 0001 |
AAAI | 3 |
| 2020 | "You are grounded!": Latent Name Artifacts in Pre-trained Language ModelsabstractPre-trained language models (LMs) may perpetuate biases originating in their training corpus to downstream models.We focus on artifacts associated with the representation of given names (e.g., Donald), which, depending on the corpus, may be associated with specific entities, as indicated by next token prediction (e.g., Trump).While helpful in some contexts, grounding happens also in underspecified or inappropriate contexts.For example, endings generated for 'Donald is a' substantially differ from those of other names, and often have more-than-average negative sentiment.We demonstrate the potential effect on downstream tasks with reading comprehension probes where name perturbation changes the model answers.As a silver lining, our experiments suggest that additional pre-training on different corpora may mitigate this bias. ModelMain Corpus Type Gen. Cls.Named Entities from News Named Entities from History Model Minimal News History Infrml Avg Minimal News History Infrml Avg GPT 0.0 7.0 12.7 1.4 5.3 0.0 21.9 39.1 7.8 17.2 GPT2-small 22.5 63.4 50.7 15.5 38.0 12.5 29.7 56.2 12.5 27.7 GPT2-medium 33.8 64.8 49.3 12.7 40.2 21.9 32.8 62.5 4.7 30.5 GPT2-large 43.7 66.2 47.9 16.9 43.7 29.7 29.7 56.2 12.5 32.0 GPT2-XL 50.7 62.0 45.1 21.1 44.7 28.1 31.2 60.9 14.1 33.6 TransformerXL 14.1 18.3 15.5 12.7 15.2 35.9 43.8 51.6 37.5 42.2 XLNet-base 4.2 33.8 12.7 4.2 13.7 0.0 34.4 23.4 3.1 15.2 XLNet-large 11.3 40.8 23.9 9.9 21.5 6.2 29.7 31.2 7.8 18.7 Average 22.5 44.5 32.2 11.8 27.7 16.8 31.7 47.6 12.5 27.1 Vered Shwartz, Rachel Rudinger, Oyvind Tafjord |
EMNLP (1) | 2 |
| 2020 | Causal Inference of Script KnowledgeabstractWhen does a sequence of events define an everyday scenario and how can this knowledge be induced from text?Prior works in inducing such scripts have relied on, in one form or another, measures of correlation between instances of events in a corpus.We argue from both a conceptual and practical sense that a purely correlation-based approach is insufficient, and instead propose an approach to script induction based on the causal effect between events, formally defined via interventions.Through both human and automatic evaluations, we show that the output of our method based on causal effects better matches the intuition of what a script represents. Noah Weber, Rachel Rudinger, Benjamin Van Durme |
EMNLP (1) | 2 |
| 2020 | The Universal Decompositional Semantics Dataset and Decomp ToolkitabstractWe present the Universal Decompositional Semantics (UDS) dataset (v1.0), which is bundled with the Decomp toolkit (v0.1). UDS1.0 unifies five high-quality, decompositional semantics-aligned annotation sets within a single semantic graph specification—with graph structures defined by the predicative patterns produced by the PredPatt tool and real-valued node and edge attributes constructed using sophisticated normalization procedures. The Decomp toolkit provides a suite of Python 3 tools for querying UDS graphs using SPARQL. Both UDS1.0 and Decomp0.1 are publicly available at http://decomp.io. Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang 0012, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
LREC | 10 |
| 2019 | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine TranslationabstractWe present PARABANK, a large-scale English paraphrase dataset that surpasses prior work in both quantity and quality. Following the approach of PARANMT (Wieting and Gimpel, 2018), we train a Czech-English neural machine translation (NMT) system to generate novel paraphrases of English reference sentences. By adding lexical constraints to the NMT decoding procedure, however, we are able to produce multiple high-quality sentential paraphrases per source sentence, yielding an English paraphrase resource with more than 4 billion generated tokens and exhibiting greater lexical diversity. Using human judgments, we also demonstrate that PARABANK’s paraphrases improve over PARANMT on both semantic similarity and fluency. Finally, we use PARABANK to train a monolingual NMT model with the same support for lexically-constrained decoding for sentence rewriting tasks. Edward J. Hu, Rachel Rudinger, Matt Post, Benjamin Van Durme |
AAAI | 2 |
| 2018 | Collecting Diverse Natural Language Inference Problems for Sentence Representation EvaluationabstractAdam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Adam Poliak, Aparajita Haldar, Rachel Rudinger, Edward J. Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme |
EMNLP | 3 |
| 2018 | Neural-Davidsonian Semantic Proto-role LabelingabstractWe present a model for semantic proto-role labeling (SPRL) using an adapted bidirectional LSTM encoding strategy that we call Neural-Davidsonian: predicate-argument structure is represented as pairs of hidden states corresponding to predicate and argument head tokens of the input sequence.We demonstrate:(1) state-of-the-art results in SPRL, and (2) that our network naturally shares parameters between attributes, allowing for learning new attribute types with limited added supervision. Rachel Rudinger, Adam R. Teichert, Ryan Culkin, Sheng Zhang 0012, Benjamin Van Durme |
EMNLP | 1 |
| 2018 | Lexicosyntactic inference in neural modelsabstractWe investigate neural models' ability to capture lexicosyntactic inferences: inferences triggered by the interaction of lexical and syntactic information.We take the task of event factuality prediction as a case study and build a factuality judgment dataset for all English clause-embedding verbs in various syntactic contexts.We use this dataset, which we make publicly available, to probe the behavior of current state-of-the-art neural systems, showing that these systems make certain systematic errors that are clearly visible through the lens of factuality prediction. Aaron Steven White, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
EMNLP | 2 |
| 2018 | Cross-lingual Decompositional Semantic ParsingabstractWe introduce the task of cross-lingual decompositional semantic parsing: mapping content provided in a source language into a decompositional semantic analysis based on a target language.We present: (1) a form of decompositional semantic analysis designed to allow systems to target varying levels of structural complexity (shallow to deep analysis), (2) an evaluation metric to measure the similarity between system output and reference semantic analysis, (3) an end-to-end model with a novel annotating mechanism that supports intra-sentential coreference, and (4) an evaluation dataset on which our model outperforms strong baselines by at least 1.75 F 1 score. Sheng Zhang 0012, Xutai Ma, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
EMNLP | 3 |
| 2018 | Neural Models of FactualityabstractRachel Rudinger, Aaron Steven White, Benjamin Van Durme. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Rachel Rudinger, Aaron Steven White, Benjamin Van Durme |
NAACL-HLT | 1 |
| 2017 | Ordinal Common-sense InferenceabstractHumans have the capacity to draw common-sense inferences from natural language: various things that are likely but not certain to hold based on established discourse, and are rarely stated explicitly. We propose an evaluation of automated common-sense inference based on an extension of recognizing textual entailment: predicting ordinal human responses on the subjective likelihood of an inference holding in a given context. We describe a framework for extracting common-sense knowledge from corpora, which is then used to construct a dataset for this ordinal entailment task. We train a neural sequence-to-sequence model on this dataset, which we use to score and generate possible inferences. Further, we annotate subsets of previously established datasets via our ordinal annotation protocol in order to then analyze the distinctions between these and what we have constructed. Sheng Zhang 0012, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Universal Decompositional Semantics on Universal DependenciesabstractAaron Steven White, Drew Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. Aaron Steven White, Dee Ann Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang 0012, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
EMNLP | 6 |
| 2015 | Script Induction as Language ModelingabstractThe narrative cloze is an evaluation metric commonly used for work on automatic script induction.While prior work in this area has focused on count-based methods from distributional semantics, such as pointwise mutual information, we argue that the narrative cloze can be productively reframed as a language modeling task.By training a discriminative language model for this task, we attain improvements of up to 27 percent over prior methods on standard narrative cloze metrics. Rachel Rudinger, Pushpendre Rastogi, Francis Ferraro, Benjamin Van Durme |
EMNLP | 1 |
| 2015 | Semantic Proto-RolesabstractWe present the first large-scale, corpus based verification of Dowty’s seminal theory of proto-roles. Our results demonstrate both the need for and the feasibility of a property-based annotation scheme of semantic relationships, as opposed to the currently dominant notion of categorical roles. Dee Ann Reisinger, Rachel Rudinger, Francis Ferraro, Craig Harman, Kyle Rawlins, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | SenseSpotting: Never let your parallel data tie you to an old domain
Marine Carpuat, Hal Daumé III, Katharine Henry, Ann Irvine, Jagadeesh Jagarlamudi, Rachel Rudinger |
ACL (1) | 6 |