Abhilasha Ravichander

dblp:170/4795 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
23since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 11 first-author · 21 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
abstract
Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context.However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming.In this work, we release HALOGEN , a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative models spanning nine domains including programming, scientific attribution, and summarization, and (2) automatic highprecision verifiers for each use case that decompose LLM generations into atomic units, and verify each unit against a high-quality knowledge source.We use this framework to evaluate ∼150,000 generations from 14 language models, finding that even the best-performing models are riddled with hallucinations (sometimes up to 86% of generated atomic facts depending on the domain).We further define a novel error classification for LLM hallucinations based on whether they likely stem from incorrect recollection of training data (Type A errors), or incorrect knowledge in training data (Type B errors), or are fabrication (Type C errors).We hope our framework provides a foundation to enable the principled study of why generative models hallucinate, and advances the development of trustworthy large language models.
Abhilasha Ravichander, Shrusti Ghela, Dave Wadden, Yejin Choi 0001
ACL (1)1
2025 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
abstract
We introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024 tasks carefully selected from over one million human-chatbot conversation logs. For automated evaluation with WildBench, we have developed two metrics, WB-Reward and WB-Score, which are computable using advanced LLMs such as GPT-4-turbo. WildBench evaluation uses task-specific checklists to evaluate model outputs systematically and provides structured explanations that justify the scores and comparisons, resulting in more reliable and interpretable automatic judgments. WB-Reward employs fine-grained pairwise comparisons between model responses, generating five potential outcomes: much better, slightly better, slightly worse, much worse, or a tie. Unlike previous evaluations that employed a single baseline model, we selected three baseline models at varying performance levels to ensure a comprehensive pairwise evaluation. Additionally, we propose a simple method to mitigate length bias, by converting outcomes of “slightly better/worse” to “tie” if the winner response exceeds the loser one by more than K characters. WB-Score evaluates the quality of model outputs individually, making it a fast and cost-efficient evaluation metric. WildBench results demonstrate a strong correlation with the human-voted Elo ratings from Chatbot Arena on hard tasks. Specifically, WB-Reward achieves a Pearson correlation of 0.98 with top-ranking models. Additionally, WB-Score reaches 0.95, surpassing both ArenaHard’s 0.91 and AlpacaEval2.0’s 0.89 for length-controlled win rates, as well as the 0.87 for regular win rates.
Bill Y. Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras 0001, Yejin Choi 0001
ICLR4
2025 Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
abstract
Abhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Yuchen Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Y. Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi 0001
NAACL (Long Papers)1
2025 Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
abstract
Large language models (LLMs) frequently generate hallucinations—content that deviates from factually inaccurate or deviates from provided context—posing challenges for diagnosis. However, diagnosing the causes of hallucination is challenging due to the complex interplay of underlying causes. This paper introduces a framework to systematically understand the sources of hallucination behavior in large language models. Our key insight is that hallucinations arise when more frequent but non-factual associations outweigh faithful ones. Through theoretical and empirical analyses, we demonstrate that decoder-only transformers effectively function as subsequence embedding models, with the fully-connected layers encoding input-output associations. We propose a tracing algorithm that identifies causal subsequences by analyzing hallucination probabilities across randomized input contexts. Experiments show our method outperforms standard attribution techniques in identifying hallucination causes and is supported by evidence from the model’s training corpus. This work provides a unified perspective on hallucinations and a robust framework for their cause and analysis.
Yiyou Sun, Yu Gai, Lijie Chen 0001, Abhilasha Ravichander, Yejin Choi 0001, Nouha Dziri, Dawn Song
NeurIPS4
2024 Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
abstract
Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs).To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices.In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain.To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference.Our key findings are threefold.First, we find no evidence that the choices-only accuracy stems from memorization alone.Second, priors over individual choices do not fully explain choicesonly accuracy, hinting that LLMs use the group dynamics of choices.Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question.Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA.Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making. 1Question: Which of these contains only a solution?Answer: (B) Question: Which can be considered a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Question: Which can be considered a solution?Step 2: Answer the Question from Step 1 Classify Choice (A) Correctness Classify Choice (B) Correctness ... Abductive Question Inference ( §6) Question: Which of these contains only a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Full MCQA Prompt Choices-only Prompt ( §3) LLMs Can Perform MCQA with no Question, but how? Classify Choice (D) Correctness Question: Which of these contains only a solution?Choices: (A) \n (B) \n (C) \n (D) \n Answer: (B) Step 1: Guess the Question No Choices Empty Choices Choice: a can of mixed fruit Answer: False Choice: a bottle of juice Answer: True Choice: a jar of pickles Answer: False
Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger
ACL (1)2
2024 OLMo: Accelerating the Science of Language Models
abstract
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi
ACL (1)28
2024 Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
abstract
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo
ACL (1)23
2024 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
abstract
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, Bill Yuchen Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Raghavi Chandu, Kai-Wei Chang 0001, Yejin Choi 0001, Bill Y. Lin
ACL (1)3
2024 What's In My Big Data?
abstract
Large text corpora are the backbone of language models. However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination). In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node. We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*. Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination. For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE. We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them.
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge
ICLR4
2024 The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
abstract
Alignment tuning has become the de facto standard practice for enabling base large language models (LLMs) to serve as open-domain AI assistants. The alignment tuning process typically involves instruction learning through supervised fine-tuning (SFT) and preference tuning via reinforcement learning from human feedback (RLHF). A recent study, LIMA (Zhou et al., 2023), shows that using merely 1K examples for SFT can achieve significant alignment performance as well, suggesting that the effect of alignment tuning might be "superficial." This raises questions about how exactly the alignment tuning transforms a base LLM. We analyze the effect of alignment tuning by examining the token distribution shift between base LLMs and their aligned counterparts (e.g., Llama-2 and Llama-2-chat). Our findings reveal that base LLMs and their alignment-tuned versions perform nearly identically in decoding on the majority of token positions (i.e., they share the top-ranked tokens). Most distribution shifts occur with stylistic tokens (e.g., discourse markers, safety disclaimers). This direct evidence strongly supports the hypothesis that alignment tuning primarily learns to adopt the language style of AI assistants, and that the knowledge required for answering user queries predominantly comes from the base LLMs themselves. Based on these findings, we rethink the alignment of LLMs by posing the research question: how effectively can we align base LLMs without SFT or RLHF? To address this, we introduce a simple, tuning-free alignment method, URIAL (Untuned LLMs with Restyled In-context Alignment). URIAL achieves effective alignment purely through in-context learning (ICL) with base LLMs, requiring as few as three constant stylistic examples and a system prompt. We conduct a fine-grained and interpretable evaluation on a diverse set of examples, named just-eval-instruct. Results demonstrate that base LLMs with URIAL can match or even surpass the performance of LLMs aligned with SFT (Mistral-7b-Instruct) or SFT+RLHF (Llama-2-70b-chat). We show that the gap between tuning-free and tuning-based alignment methods can be significantly reduced through strategic prompting and ICL. Our findings on the superficial nature of alignment tuning and results with URIAL suggest that deeper analysis and theoretical understanding of alignment is crucial to future LLM research.
Bill Y. Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Raghavi Chandu, Chandra Bhagavatula, Yejin Choi 0001
ICLR2
2024 The Generative AI Paradox: "What It Can Create, It May Not Understand"
abstract
The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models now take only seconds to produce outputs that would challenge or exceed the capabilities even of expert humans. At the same time, models still show basic errors in understanding that would not be expected even in non-expert humans. This presents us with an apparent paradox: how do we reconcile seemingly superhuman capabilities with the persistence of errors that few humans would make? In this work, we posit that this tension reflects a divergence in the configuration of intelligence in today's generative models relative to intelligence in humans. Specifically, we propose and test the **Generative AI Paradox** hypothesis: generative models, having been trained directly to reproduce expert-like outputs, acquire generative capabilities that are not contingent upon---and can therefore exceed---their ability to understand those same types of outputs. This contrasts with humans, for whom basic understanding almost always precedes the ability to generate expert-level outputs. We test this hypothesis through controlled experiments analyzing generation vs.~understanding in generative models, across both language and image modalities. Our results show that although models can outperform humans in generation, they consistently fall short of human capabilities in measures of understanding, as well as weaker correlation between generation and understanding performance, and more brittleness to adversarial inputs. Our findings support the hypothesis that models' generative capability may not be contingent upon understanding capability, and call for caution in interpreting artificial intelligence by analogy to human intelligence.
Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Jena D. Hwang, Jillian Fisher, Abhilasha Ravichander, Khyathi Raghavi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, Yejin Choi 0001
ICLR9
2024 MacGyver: Are Large Language Models Creative Problem Solvers?
abstract
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas Griffiths, Faeze Brahman. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras 0001, Raja Marjieh, Nanyun Peng 0001, Yejin Choi 0001, Thomas L. Griffiths 0001, Faeze Brahman
NAACL-HLT2
2024 The Art of Saying No: Contextual Noncompliance in Language Models
abstract
Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities.
Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi
NeurIPS6
2024 Understanding How to Inform Blind and Low-Vision Users about Data Privacy through Privacy Question Answering Assistants
Yuanyuan Feng, Abhilasha Ravichander, Yaxing Yao, Shikun Zhang, Rex Chen, Shomir Wilson, Norman M. Sadeh
USENIX Security Symposium2
2024 Incorporating Taxonomic Reasoning and Regulatory Knowledge into Automated Privacy Question Answering
Abhilasha Ravichander, Ian Yang, Rex Chen, Shomir Wilson, Thomas B. Norton, Norman M. Sadeh
WISE (1)1
2023 Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
abstract
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Liwei Jiang, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren, Sean Welleck, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Raghavi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Y. Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren 0001, Sean Welleck, Yejin Choi 0001
EMNLP6
2022 CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
abstract
The full power of human language-based communication cannot be realized without negation.All human languages have some form of negation.Despite this, negation remains a challenging phenomenon for current natural language understanding systems.To facilitate the future development of models that can process negation effectively, we present CONDAQA, the first English reading comprehension dataset which requires reasoning about the implications of negated statements in paragraphs.We collect paragraphs with diverse negation cues, then have crowdworkers ask questions about the implications of the negated statement in the passage.We also have workers make three kinds of edits to the passage-paraphrasing the negated statement, changing the scope of the negation, and reversing the negation-resulting in clusters of question-answer pairs that are difficult for models to answer with spurious shortcuts.CONDAQA features 14,182 questionanswer pairs with over 200 unique negation cues and is challenging for current state-ofthe-art models.The best performing model on CONDAQA (UNIFIEDQA-V2-3B) achieves only 42% on our consistency metric, well below human performance which is 81%.We release our dataset, along with fully-finetuned, few-shot, and zero-shot evaluations, to facilitate the development of future NLP methods that work on negated language.
Abhilasha Ravichander, Matt Gardner 0001, Ana Marasovic
EMNLP1
2022 A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus
abstract
Over the past decade, researchers have started to explore the use of NLP to develop tools aimed at helping the public, vendors, and regulators analyze disclosures made in privacy policies. With the introduction of new privacy regulations, the language of privacy policies is also evolving, and disclosures made by the same organization are not always the same in different languages, especially when used to communicate with users who fall under different jurisdictions. This work explores the use of language technologies to capture and analyze these differences at scale. We introduce an annotation scheme designed to capture the nuances of two new landmark privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. We then introduce the first bilingual corpus of mobile app privacy policies consisting of 64 privacy policies in English (292K words) and 91 privacy policies in German (478K words), respectively with manual annotations for 8K and 19K fine-grained data practices. The annotations are used to develop computational methods that can automatically extract “disclosures” from privacy policies. Analysis of a subset of 59 “semi-parallel” policies reveals differences that can be attributed to different regulatory regimes, suggesting that systematic analysis of policies using automated language technologies is indeed a worthwhile endeavor.
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas B. Norton, Thomas Hupperich, Shomir Wilson, Norman M. Sadeh
LREC6
2021 Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy?
abstract
Abhilasha Ravichander, Alan W Black, Thomas Norton, Shomir Wilson, Norman Sadeh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Abhilasha Ravichander, Alan W. Black, Thomas B. Norton, Shomir Wilson, Norman M. Sadeh
ACL/IJCNLP (1)1
2021 Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?
abstract
Although neural models have achieved impressive results on several NLP benchmarks, little is understood about the mechanisms they use to perform language tasks.Thus, much recent attention has been devoted to analyzing the sentence representations learned by neural encoders, through the lens of 'probing' tasks.However, to what extent was the information encoded in sentence representations, as discovered through a probe, actually used by the model to perform its task?In this work, we examine this probing paradigm through a case study in Natural Language Inference, showing that models can learn to encode linguistic properties even if they are not needed for the task on which the model was trained.We further identify that pretrained word embeddings play a considerable role in encoding these properties rather than the training task itself, highlighting the importance of careful controls when designing probing experiments.Finally, through a set of controlled synthetic tasks, we demonstrate models can encode these properties considerably above chance-level even when distributed in the data as random noise, calling into question the interpretation of absolute claims on probing tasks. 1
Abhilasha Ravichander, Yonatan Belinkov, Eduard H. Hovy
EACL1
2021 NoiseQA: Challenge Set Evaluation for User-Centric Question Answering
abstract
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, Alan W Black. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard H. Hovy, Alan W. Black
EACL1
2021 Measuring and Improving Consistency in Pretrained Language Models
abstract
Abstract Consistency of a model—that is, the invariance of its behavior under meaning-preserving alternations in its input—is a highly desirable property in natural language processing. In this paper we study the question: Are Pretrained Language Models (PLMs) consistent with respect to factual knowledge? To this end, we create ParaRel🤘, a high-quality resource of cloze-style query English paraphrases. It contains a total of 328 paraphrases for 38 relations. Using ParaRel🤘, we show that the consistency of all PLMs we experiment with is poor— though with high variance between relations. Our analysis of the representational spaces of PLMs suggests that they have a poor structure and are currently not suitable for representing knowledge robustly. Finally, we propose a method for improving model consistency and experimentally demonstrate its effectiveness.1
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg
Trans. Assoc. Comput. Linguistics4
2021 Erratum: Measuring and Improving Consistency in Pretrained Language Models
abstract
Abstract During production of this paper, an error was introduced to the formula on the bottom of the right column of page 1020. In the last two terms of the formula, the n and m subscripts were swapped. The correct formula is:Lc=∑n=1k∑m=n+1kDKL(Qnri∥Qmri)+DKL(Qmri∥Qnri)The paper has been updated.
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg
Trans. Assoc. Comput. Linguistics4
2019 Exploring Numeracy in Word Embeddings
abstract
Word embeddings are now pervasive across NLP subfields as the de-facto method of forming text representataions.In this work, we show that existing embedding models are inadequate at constructing representations that capture salient aspects of mathematical meaning for numbers, which is important for language understanding.Numbers are ubiquitous and frequently appear in text.Inspired by cognitive studies on how humans perceive numbers, we develop an analysis framework to test how well word embeddings capture two essential properties of numbers: magnitude (e.g.3<4) and numeration (e.g.3=three).Our experiments reveal that most models capture an approximate notion of magnitude, but are inadequate at capturing numeration.We hope that our observations provide a starting point for the development of methods which better capture numeracy in NLP systems.
Aakanksha Naik, Abhilasha Ravichander, Carolyn P. Rosé, Eduard H. Hovy
ACL (1)2
2019 EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference
abstract
Quantitative reasoning is a higher-order reasoning skill that any intelligent natural language understanding system can reasonably be expected to handle.We present EQUATE 1 (Evaluating Quantitative Understanding Aptitude in Textual Entailment), a new framework for quantitative reasoning in textual entailment.We benchmark the performance of 9 published NLI models on EQUATE, and find that on average, state-of-the-art methods do not achieve an absolute improvement over a majority-class baseline, suggesting that they do not implicitly learn to reason with quantities.We establish a new baseline Q-REAS that manipulates quantities symbolically.In comparison to the best performing NLI model, it achieves success on numerical reasoning tests (+24.2%),but has limited verbal reasoning capabilities (-8.1%).We hope our evaluation framework will support the development of models of quantitative reasoning in language understanding.
Abhilasha Ravichander, Aakanksha Naik, Carolyn P. Rosé, Eduard H. Hovy
CoNLL1
2019 Question Answering for Privacy Policies: Combining Computational and Legal Perspectives
abstract
Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, Norman Sadeh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Abhilasha Ravichander, Alan W. Black, Shomir Wilson, Thomas B. Norton, Norman M. Sadeh
EMNLP/IJCNLP (1)1
2019 MAPS: Scaling Privacy Compliance Analysis to a Million Apps
abstract
Abstract The app economy is largely reliant on data collection as its primary revenue model. To comply with legal requirements, app developers are often obligated to notify users of their privacy practices in privacy policies. However, prior research has suggested that many developers are not accurately disclosing their apps’ privacy practices. Evaluating discrepancies between apps’ code and privacy policies enables the identification of potential compliance issues. In this study, we introduce the Mobile App Privacy System (MAPS) for conducting an extensive privacy census of Android apps. We designed a pipeline for retrieving and analyzing large app populations based on code analysis and machine learning techniques. In its first application, we conduct a privacy evaluation for a set of 1,035,853 Android apps from the Google Play Store. We find broad evidence of potential non-compliance. Many apps do not have a privacy policy to begin with. Policies that do exist are often silent on the practices performed by apps. For example, 12.1% of apps have at least one location-related potential compliance issue. We hope that our extensive analysis will motivate app stores, government regulators, and app developers to more effectively review apps for potential compliance issues.
Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha Ravichander, Ziqi Wang 0007, Joel R. Reidenberg, N. Cameron Russell, Norman M. Sadeh
Proc. Priv. Enhancing Technol.4
2018 Stress Test Evaluation for Natural Language Inference
abstract
Natural language inference (NLI) is the task of determining if a natural language hypothesis can be inferred from a given premise in a justifiable manner. NLI was proposed as a benchmark task for natural language understanding. Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text. However, the extent to which these models understand the semantic content of sentences is unclear. In this work, we propose an evaluation methodology consisting of automatically constructed “stress tests” that allow us to examine whether systems have the ability to make real inferential decisions. Our evaluation of six sentence-encoder models on these stress tests reveals strengths and weaknesses of these models with respect to challenging linguistic phenomena, and suggests important directions for future work in this area.
Aakanksha Naik, Abhilasha Ravichander, Norman M. Sadeh, Carolyn P. Rosé, Graham Neubig
COLING2
2018 An Empirical Study of Self-Disclosure in Spoken Dialogue Systems
abstract
Self-disclosure is a key social strategy employed in conversation to build relations and increase conversational depth.It has been heavily studied in psychology and linguistic literature, particularly for its ability to induce self-disclosure from the recipient, a phenomena known as reciprocity.However, we know little about how self-disclosure manifests in conversation with automated dialog systems, especially as any self-disclosure on the part of a dialog system is patently disingenuous.In this work, we run a large-scale quantitative analysis on the effect of selfdisclosure by analyzing interactions between real-world users and a spoken dialog system in the context of social conversation.We find that indicators of reciprocity occur even in human-machine dialog, with far-reaching implications for chatbots in a variety of domains including education, negotiation and social dialog.
Abhilasha Ravichander, Alan W. Black
SIGDIAL Conference1
2017 How Would You Say It? Eliciting Lexically Diverse Dialogue for Supervised Semantic Parsing
abstract
Building dialogue interfaces for realworld scenarios often entails training semantic parsers starting from zero examples.How can we build datasets that better capture the variety of ways users might phrase their queries, and what queries are actually realistic?Wang et al. (2015) proposed a method to build semantic parsing datasets by generating canonical utterances using a grammar and having crowdworkers paraphrase them into natural wording.A limitation of this approach is that it induces bias towards using similar language as the canonical utterances.In this work, we present a methodology that elicits meaningful and lexically diverse queries from users for semantic parsing tasks.Starting from a seed lexicon and a generative grammar, we pair logical forms with mixed text-image representations and ask crowdworkers to paraphrase and confirm the plausibility of the queries that they generated.We use this method to build a semantic parsing dataset from scratch for a dialog agent in a smart-home simulation.We find evidence that this dataset, which we have named SMARTHOME, is demonstrably more lexically diverse and difficult to parse than existing domain-specific semantic parsing datasets.
Abhilasha Ravichander, Thomas Manzini, Matthias Grabmair, Graham Neubig, Jonathan Francis, Eric Nyberg
SIGDIAL Conference1
2015 VISAGE: A Support Vector Machine Approach to Group Dynamic Analysis
abstract
A group is defined as a collective entity usually consisting of two or more individuals each connected by social relationships. The term 'group dynamics' was originally coined by social psychologist Kurt Lewin to describe the positive and negative forces within groups of people. Its study is useful today in a wide variety of applications such as gaining a better understanding of decision making behavior or evaluating the health of workplace environments. Metrics need to be defined to measure the quality of relations in a group by performing an analysis on each individual member of the group. Previous research in the field performs this through the means of surveying (generally through questionnaires) each member of the group. In this project we propose a novel method to analyze group dynamics from a single static image by performing automated facial expression recognition.
Abhilasha Ravichander, Supriya Vijay, Varshini Ramaseshan
ICMLA1