Bhuwan Dhingra

dblp:180/5692 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-6874-9515ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
abstract
Yukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama
ACL (1)4
2025 Real-time Factuality Assessment from Adversarial Feedback
abstract
We show that existing evaluations for assessing the factuality of news from conventional sources, such as claims on fact-checking websites, result in high accuracies over time for LLM-based detectors-even after their knowledge cutoffs.This suggests that recent popular false information from such sources can be easily identified due to its likely presence in pretraining/retrieval corpora or the emergence of salient, yet shallow, patterns in these datasets.Instead, we argue that a proper factuality evaluation dataset should test a model's ability to reason about current events by retrieving and reading related evidence.To this end, we develop a novel pipeline that leverages natural language feedback from a RAG-based detector to iteratively modify real-time news into deceptive variants that challenge LLMs.Our iterative rewrite decreases the binary classification ROC-AUC by an absolute 17.5 percent for a strong RAG-based GPT-4o detector.Our experiments reveal the important role of RAG in both evaluating and generating challenging news examples, as retrieval-free LLM detectors are vulnerable to unseen events and adversarial attacks, while feedback from RAG-based evaluation helps discover more deceitful patterns.
Sanxing Chen, Bhuwan Dhingra
ACL (1)3
2025 Close or Cloze? Assessing the Robustness of Large Language Models to Adversarial Perturbations via Word Recovery
abstract
The current generation of large language models (LLMs) show a surprising degree of robustness to adversarial perturbations, but it is unclear when these models implicitly recover the original text and when they rely on surrounding context. To isolate this recovery faculty of language models, we study a new diagnostic task —Adversarial Word Recovery — an extension of spellchecking where the inputs may be adversarial. We collect a new dataset using 9 popular perturbation attack strategies from the literature and organize them using a taxonomy of phonetic, typo, and visual attacks. We use this dataset to study the word recovery performance of the current generation of LLMs, finding that proprietary models (GPT-4, GPT-3.5 and Palm-2) match or surpass human performance. Conversely, open-source models (Llama-2, Mistral, Falcon) demonstrate a material gap between human performance, especially on visual attacks. For these open models, we show that performance of word recovery without context correlates to word recovery with context, and ultimately affects downstream task performance on a hateful, offensive, and toxic classification task. Finally, to show improving word recovery can improve robustness, we mitigate these attacks with a small Byt5 model tuned to recover visually attacked words.
Luke Moffett, Bhuwan Dhingra
COLING2
2025 To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
abstract
Large Language Models (LLMs) are often augmented with external contexts, such as those used in retrieval-augmented generation (RAG). However, these contexts can be inaccurate or intentionally misleading, leading to conflicts with the model’s internal knowledge. We argue that robust LLMs should demonstrate situated faithfulness, dynamically calibrating their trust in external information based on their confidence in the internal knowledge and the external context to resolve knowledge conflicts. To benchmark this capability, we evaluate LLMs across several QA datasets, including a newly created dataset featuring in-the-wild incorrect contexts sourced from Reddit posts. We show that when provided with both correct and incorrect contexts, both open-source and proprietary models tend to overly rely on external information, regardless of its factual accuracy. To enhance situated faithfulness, we propose two approaches: Self-Guided Confidence Reasoning (SCR) and Rule-Based Confidence Reasoning (RCR). SCR enables models to self-access the confidence of external information relative to their own internal knowledge to produce the most accurate answer. RCR, in contrast, extracts explicit confidence signals from the LLM and determines the final answer using predefined rules. Our results show that for LLMs with strong reasoning capabilities, such as GPT-4o and GPT-4o mini, SCR outperforms RCR, achieving improvements of up to 24.2\% over a direct input augmentation baseline. Conversely, for a smaller model like Llama-3-8B, RCR outperforms SCR. Fine-tuning SCR with our proposed Confidence Reasoning Direct Preference Optimization (CR-DPO) method improves performance on both seen and unseen datasets, yielding an average improvement of 8.9\% on Llama-3-8B. In addition to quantitative results, we offer insights into the relative strengths of SCR and RCR. Our findings highlight promising avenues for improving situated faithfulness in LLMs.
Sanxing Chen, Hongyi Cai, Bhuwan Dhingra
ICLR4
2025 Improving Model Alignment Through Collective Intelligence of Open-Source Models
abstract
Building helpful and harmless large language models (LLMs) requires effective model alignment approach based on human instructions and feedback, which necessitates high-quality human-labeled data. Constructing such datasets is often expensive and hard to scale, and may face potential limitations on diversity and generalization. To address these challenges, we introduce Mixture of Agents Alignment (MoAA), that leverages the collective strengths of various language models to provide high-quality data for model alignment. By employing MoAA, we enhance both supervised fine-tuning and preference optimization, leading to improved performance compared to using a single model alone to generate alignment data (e.g. using GPT-4o alone). Evaluation results show that our approach can improve win rate of LLaMA-3.1-8B-Instruct from 19.5 to 48.3 on Arena-Hard and from 22.33 to 57.23 on AlpacaEval2, highlighting a promising direction for model alignment through this new scalable and diverse synthetic data recipe. Furthermore, we demonstrate that MoAA enables a self-improvement pipeline, where models fine-tuned on MoA-generated data surpass their own initial capabilities, providing evidence that our approach can push the frontier of open-source LLMs without reliance on stronger external supervision. Data and code will be released.
Roy Xie, Shang Zhu, Ben Athiwaratkun, Bhuwan Dhingra, Shuaiwen Song, Ce Zhang 0001, James Zou 0001
ICML6
2025 Evaluating Morphological Compositional Generalization in Large Language Models
abstract
Mete Ismayilzada, Defne Circi, Jonne Sälevä, Hale Sirin, Abdullatif Köksal, Bhuwan Dhingra, Antoine Bosselut, Duygu Ataman, Lonneke Van Der Plas. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mete Ismayilzada, Defne Circi, Jonne Sälevä, Hale Sirin, Abdullatif Köksal, Bhuwan Dhingra, Antoine Bosselut, Duygu Ataman, Lonneke van der Plas
NAACL (Long Papers)6
2025 MatViX: Multimodal Information Extraction from Visually Rich Articles
abstract
Ghazal Khalighinejad, Sharon Scott, Ollie Liu, Kelly L. Anderson, Rickard Stureborg, Aman Tyagi, Bhuwan Dhingra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ghazal Khalighinejad, Sharon Scott, Ollie Liu, Kelly L. Anderson, Rickard Stureborg, Aman Tyagi, Bhuwan Dhingra
NAACL (Long Papers)7
2025 Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
abstract
Contrastive learning (CL) is a prevalent technique for training embedding models, which pulls semantically similar examples (positives) closer in the representation space while pushing dissimilar ones (negatives) further apart. A key source of negatives are "in-batch" examples, i.e., positives from other examples in the batch. Effectiveness of such models is hence strongly influenced by the size and quality of training batches. In this work, we propose *Breaking the Batch Barrier* (B3), a novel batch construction strategy designed to curate high-quality batches for CL. Our approach begins by using a pretrained teacher embedding model to rank all examples in the dataset, from which a sparse similarity graph is constructed. A community detection algorithm is then applied to this graph to identify clusters of examples that serve as strong negatives for one another. The clusters are then used to construct batches that are rich in in-batch negatives. Empirical results on the MMEB multimodal embedding benchmark (36 tasks) demonstrate that our method sets a new state of the art, outperforming previous best methods by +1.3 and +2.9 points at the 7B and 2B model scales, respectively. Notably, models trained with B3 surpass existing state-of-the-art results even with a batch size as small as 64, which is 4–16× smaller than that required by other methods. Moreover, experiments show that B3 generalizes well across domains and tasks, maintaining strong performance even when trained with considerably weaker teachers.
Raghuveer Thirukovalluru, Ye Liu 0006, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen, Bhuwan Dhingra
NeurIPS10
2025 Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals
abstract
Large language models (LLMs) process entire input contexts indiscriminately, which is inefficient when the information required to answer a query is localized within the context. We present dynamic context cutoff, a novel method enabling LLMs to self-terminate processing upon acquiring sufficient task-relevant information. Through analysis of model internals, we discover that specific attention heads inherently encode "sufficiency signals" -- detectable through lightweight classifiers -- that predict when critical information has been processed. This reveals a new efficiency paradigm: models' internal understanding naturally dictates processing needs rather than external compression heuristics. Comprehensive experiments across six QA datasets (up to 40K tokens) with three model families (LLaMA/Qwen/Mistral, 1B-70B) demonstrate 3.4% accuracy improvement while achieving 1.33x token reduction on average. Furthermore, our method demonstrates superior performance compared to other context efficiency methods at equivalent token reduction rates. Additionally, we observe an emergent scaling phenomenon: while smaller models require probing for sufficiency detection, larger models exhibit intrinsic self-assessment capabilities through prompting.
Roy Xie, Paul Rosu, Chunyuan Deng, Bolun Sun, Bhuwan Dhingra
NeurIPS7
2024 Sequence Reducible Holdout Loss for Language Model Pretraining
abstract
Data selection techniques, which adaptively select datapoints inside the training loop, have demonstrated empirical benefits in reducing the number of gradient steps to train neural models. However, these techniques have so far largely been applied to classification. In this work, we study their applicability to language model pretraining, a highly time-intensive task. We propose a simple modification to an existing data selection technique (reducible hold-out loss training) in order to adapt it to the sequence losses typical in language modeling. We experiment on both autoregressive and masked language modelling, and show that applying data selection to pretraining offers notable benefits including a 4.3% reduction in total number of steps, a 21.5% steps reduction in average, to an intermediate target perplexity, over the course of pretraining an autoregressive language model. Further, data selection trained language models demonstrate significantly better generalization ability on out of domain datasets - 7.9% reduction in total number of steps and 23.2% average steps reduction to an intermediate target perplexity.
Raghuveer Thirukovalluru, Nicholas Monath, Bhuwan Dhingra, Sam Wiseman
LREC/COLING3
2024 Atomic Self-Consistency for Better Long Form Generations
abstract
Recent work has aimed to improve LLM generations by filtering out hallucinations, thereby improving the precision of the information in responses.Correctness of a long-form response, however, also depends on the recall of multiple pieces of information relevant to the question.In this paper, we introduce Atomic Self-Consistency (ASC), a technique for improving the recall of relevant information in an LLM response.ASC follows recent work, Universal Self-Consistency (USC) in using multiple stochastic samples from an LLM to improve the long-form response.Unlike USC which only focuses on selecting the best single generation, ASC picks authentic subparts from the samples and merges them into a superior composite answer.Through extensive experiments and ablations, we show that merging relevant subparts of multiple samples performs significantly better than picking a single sample.ASC demonstrates significant gains over USC on multiple factoids and open-ended QA datasets -ASQA, QAMPARI, QUEST, ELI5 with ChatGPT and Llama3.Our analysis also reveals untapped potential for enhancing long-form generations using approach of merging multiple samples.
Raghuveer Thirukovalluru, Bhuwan Dhingra
EMNLP3
2024 ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods
abstract
The rapid scaling of large language models (LLMs) has raised concerns about the transparency and fair use of the data used in their pretraining.Detecting such content is challenging due to the scale of the data and limited exposure of each instance during training.We propose RECALL, (Relative Conditional Log-Likelihood), a novel membership inference attack (MIA) to detect LLMs' pretraining data by leveraging their conditional language modeling capabilities.RECALL examines the relative change in conditional log-likelihoods when prefixing target data points with non-member context.Our empirical findings show that conditioning member data on non-member prefixes induces a larger decrease in log-likelihood compared to non-member data.We conduct comprehensive experiments and show that RE-CALL achieves state-of-the-art performance on WikiMIA dataset, even with random and synthetic prefixes, and can be further improved using an ensemble approach.Moreover, we conduct an in-depth analysis of LLMs' behavior with different membership contexts, providing insights into how LLMs leverage membership information for effective inference at both the sequence and token level.
Roy Xie, Ruomin Huang, Minxing Zhang, Jian Pei 0001, Neil Zhenqiang Gong, Bhuwan Dhingra
EMNLP8
2024 Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language Models
abstract
Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. PPI achieves this by combining small amounts of human-labeled data with larger amounts of data labeled by a reasonably accurate---but potentially biased---automatic system, in a way that results in tighter confidence intervals for certain parameters of interest (e.g., the mean performance of a language model). In this paper, we propose a method called Stratified Prediction-Powered Inference (StratPPI), in which we show that the basic PPI estimates can be considerably improved by employing simple data stratification strategies. Without making any assumptions on the underlying automatic labeling system or data distribution, we derive an algorithm for computing provably valid confidence intervals for parameters of any dimensionality that is based on stratified sampling. In particular, we show both theoretically and empirically that, with appropriate choices of stratification and sample allocation, our approach can provide substantially tighter confidence intervals than unstratified approaches. Specifically, StratPPI is expected to improve in cases where the performance of the autorater varies across different conditional distributions of the target data.
Adam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra, Amir Globerson, William W. Cohen
NeurIPS4
2023 Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
abstract
Human data labeling is an important and expensive task at the heart of supervised learning systems. Hierarchies help humans understand and organize concepts. We ask whether and how concept hierarchies can inform the design of annotation interfaces to improve labeling quality and efficiency. We study this question through annotation of vaccine misinformation, where the labeling task is difficult and highly subjective. We investigate 6 user interface designs for crowdsourcing hierarchical labels by collecting over 18,000 individual annotations. Under a fixed budget, integrating hierarchies into the design improves crowdsource workers’ F1 scores. We attribute this to (1) Grouping similar concepts, improving F1 scores by +0.16 over random groupings, (2) Strong relative performance on high-difficulty examples (relative F1 score difference of +0.40), and (3) Filtering out obvious negatives, increasing precision by +0.07. Ultimately, labeling schemes integrating the hierarchy outperform those that do not — achieving mean F1 of 0.70.
Rickard Stureborg, Bhuwan Dhingra, Jun Yang 0001
CHI2
2023 Salient Span Masking for Temporal Understanding
abstract
Salient Span Masking (SSM) has shown itself to be an effective strategy to improve closedbook question answering performance.SSM extends general masked language model pretraining by creating additional unsupervised training sentences that mask a single entity or date span, thus oversampling factual information.Despite the success of this paradigm, the span types and sampling strategies are relatively arbitrary and not widely studied for other tasks.Thus, we investigate SSM from the perspective of temporal tasks, where learning a good representation of various temporal expressions is important.To that end, we introduce Temporal Span Masking (TSM) intermediate training.First, we find that SSM alone improves the downstream performance on three temporal tasks by an avg.+5.8 points.Further, we are able to achieve additional improvements (avg.+0.29 points) by adding the TSM task.These comprise the new best reported results on the targeted tasks.Our analysis suggests that the effectiveness of SSM stems from the sentences chosen in the training data rather than the mask choice: sentences with entities frequently also contain temporal expressions.Nonetheless, the additional targeted spans of TSM can still improve performance, especially in a zero-shot context.
Jeremy R. Cole, Aditi Chaudhary, Bhuwan Dhingra, Partha Talukdar
EACL3
2023 DiffQG: Generating Questions to Summarize Factual Changes
abstract
Jeremy R. Cole, Palak Jain, Julian Martin Eisenschlos, Michael J.Q. Zhang, Eunsol Choi, Bhuwan Dhingra. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Jeremy R. Cole, Palak Jain 0006, Julian Martin Eisenschlos, Michael J. Q. Zhang, Eunsol Choi, Bhuwan Dhingra
EACL6
2023 Learning the Legibility of Visual Text Perturbations
abstract
Many adversarial attacks in NLP perturb inputs to produce visually similar strings ('ergo' → 'εrgo') which are legible to humans but degrade model performance.Although preserving legibility is a necessary condition for text perturbation, little work has been done to systematically characterize it; instead, legibility is typically loosely enforced via intuitions around the nature and extent of perturbations.Particularly, it is unclear to what extent can inputs be perturbed while preserving legibility, or how to quantify the legibility of a perturbed string.In this work, we address this gap by learning models that predict the legibility of a perturbed string, and rank candidate perturbations based on their legibility.To do so, we collect and release LEGIT, a human-annotated dataset comprising the legibility of visually perturbed text.Using this dataset, we build both text-and vision-based models which achieve up to 0.91 F1 score in predicting whether an input is legible, and an accuracy of 0.86 in predicting which of two given perturbations is more legible.Additionally, we discover that legible perturbations from the LEGIT dataset are more effective at lowering the performance of NLP models than best-known attack strategies, suggesting that current models may be vulnerable to a broad range of perturbations beyond what is captured by existing visual attacks. 1
Dev Seth, Rickard Stureborg, Danish Pruthi, Bhuwan Dhingra
EACL4
2023 Selectively Answering Ambiguous Questions
abstract
Trustworthy language models should abstain from answering questions when they do not know the answer.However, the answer to a question can be unknown for a variety of reasons.Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown.But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context.We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous.In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work.We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning.Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.
Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Martin Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein
EMNLP5
2023 Valla: Standardizing and Benchmarking Authorship Attribution and Verification Through Empirical Evaluation and Comparative Analysis
abstract
Jacob Tyo, Bhuwan Dhingra, Zachary C. Lipton. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jacob Tyo, Bhuwan Dhingra, Zachary C. Lipton
IJCNLP (1)2
2022 ASQA: Factoid Questions Meet Long-Form Answers
abstract
An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA).This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations.The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality.In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA.Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation.Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity.In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question.We use this notion of correctness to define an automated metric of performance for ASQA.Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang
EMNLP3
2022 Time-Aware Language Models as Temporal Knowledge Bases
abstract
Abstract Many facts come with an expiration date, from the name of the President to the basketball team Lebron James plays for. However, most language models (LMs) are trained on snapshots of data collected at a specific moment in time. This can limit their utility, especially in the closed-book setting where the pretraining corpus must contain the facts the model should memorize. We introduce a diagnostic dataset aimed at probing LMs for factual knowledge that changes over time and highlight problems with LMs at either end of the spectrum—those trained on specific slices of temporal data, as well as those trained on a wide range of temporal data. To mitigate these problems, we propose a simple technique for jointly modeling text with its timestamp. This improves memorization of seen facts from the training time period, as well as calibration on predictions about unseen facts from future time periods. We also show that models trained with temporal context can be efficiently “refreshed” as new data arrives, without the need for retraining from scratch.
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen
Trans. Assoc. Comput. Linguistics1
2022 Evaluating Explanations: How Much Do Explanations from the Teacher Aid Students?
abstract
Abstract While many methods purport to explain predictions by highlighting salient features, what aims these explanations serve and how they ought to be evaluated often go unstated. In this work, we introduce a framework to quantify the value of explanations via the accuracy gains that they confer on a student model trained to simulate a teacher model. Crucially, the explanations are available to the student during training, but are not available at test time. Compared with prior proposals, our approach is less easily gamed, enabling principled, automatic, model-agnostic evaluation of attributions. Using our framework, we compare numerous attribution methods for text classification and question answering, and observe quantitative differences that are consistent (to a moderate to high degree) across different student model architectures and learning strategies.1
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio B. Soares, Michael Collins 0001, Zachary C. Lipton, Graham Neubig, William W. Cohen
Trans. Assoc. Comput. Linguistics3
2021 Reasoning Over Virtual Knowledge Bases With Open Predicate Relations
abstract
We present the Open Predicate Query Language (OPQL); a method for constructing a virtual KB (VKB) trained entirely from text. Large Knowledge Bases (KBs) are indispensable for a wide-range of industry applications such as question answering and recommendation. Typically, KBs encode world knowledge in a structured, readily accessible form derived from laborious human annotation efforts. Unfortunately, while they are extremely high precision, KBs are inevitably highly incomplete and automated methods for enriching them are far too inaccurate. Instead, OPQL constructs a VKB by encoding and indexing a set of relation mentions in a way that naturally enables reasoning and can be trained without any structured supervision. We demonstrate that OPQL outperforms prior VKB methods on two different KB reasoning tasks and, additionally, can be used as an external memory integrated into a language model (OPQL-LM) leading to improvements on two open-domain question answering tasks.
Haitian Sun, Patrick Verga, Bhuwan Dhingra, Ruslan Salakhutdinov, William W. Cohen
ICML3
2021 Fool Me Twice: Entailment from Wikipedia Gamification
abstract
Julian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan Boyd-Graber. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Julian Martin Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan L. Boyd-Graber
NAACL-HLT2
2021 Differentiable Open-Ended Commonsense Reasoning
abstract
Bill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, William Cohen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Bill Y. Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren 0001, William W. Cohen
NAACL-HLT3
2020 Learning to Deceive with Attention-Based Explanations
abstract
Attention mechanisms are ubiquitous components in neural architectures applied to natural language processing.In addition to yielding gains in predictive accuracy, attention weights are often claimed to confer interpretability, purportedly useful both for providing insights to practitioners and for explaining why a model makes its decisions to stakeholders.We call the latter use of attention mechanisms into question by demonstrating a simple method for training models to produce deceptive attention masks.Our method diminishes the total weight assigned to designated impermissible tokens, even when the models can be shown to nevertheless rely on these features to drive predictions.Across multiple models and tasks, our approach manipulates attention weights while paying surprisingly little cost in accuracy.Through a human study, we show that our manipulated attention-based explanations deceive people into thinking that predictions from a model biased against gender minorities do not rely on the gender.Consequently, our results cast doubt on attention's reliability as a tool for auditing algorithms in the context of fairness and accountability.1
Danish Pruthi, Bhuwan Dhingra, Graham Neubig, Zachary C. Lipton
ACL3
2020 ToTTo: A Controlled Table-To-Text Generation Dataset
abstract
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Ankur P. Parikh, Xuezhi Wang 0002, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das 0001
EMNLP (1)5
2020 Differentiable Reasoning over a Virtual Knowledge Base
Bhuwan Dhingra, Manzil Zaheer, Vidhisha Balachandran, Graham Neubig, Ruslan Salakhutdinov, William W. Cohen
ICLR1
2019 Handling Divergent Reference Texts when Evaluating Table-to-Text Generation
abstract
Automatically constructed datasets for generating text from semi-structured data (tables), such as WikiBio (Lebret et al., 2016), often contain reference texts that diverge from the information in the corresponding semistructured data.We show that metrics which rely solely on the reference texts, such as BLEU and ROUGE, show poor correlation with human judgments when those references diverge.We propose a new metric, PAR-ENT, which aligns n-grams from the reference and generated texts to the semi-structured data before computing their precision and recall.Through a large scale human evaluation study of table-to-text models for WikiBio, we show that PARENT correlates with human judgments better than existing text generation metrics.We also adapt and evaluate the information extraction based evaluation proposed in Wiseman et al. (2017), and show that PAR-ENT has comparable correlation to it, while being easier to use.We show that PARENT is also applicable when the reference texts are elicited from humans using the data from the WebNLG challenge.1 * Work done during an internship at Google.
Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das 0001, William W. Cohen
ACL (1)1
2019 Combating Adversarial Misspellings with Robust Word Recognition
abstract
To combat adversarial spelling mistakes, we propose placing a word recognition model in front of the downstream classifier.Our word recognition models build upon the RNN semicharacter architecture, introducing several new backoff strategies for handling rare and unseen words.Trained to recognize words corrupted by random adds, drops, swaps, and keyboard mistakes, our method achieves 32% relative (and 3.3% absolute) error reduction over the vanilla semi-character model.Notably, our pipeline confers robustness on the downstream classifier, outperforming both adversarial training and off-the-shelf spell checkers.Against a BERT model fine-tuned for sentiment analysis, a single adversarially-chosen character attack lowers accuracy from 90.3% to 45.8%.Our defense restores accuracy to 75% 1 .Surprisingly, better word recognition does not always entail greater robustness.Our analysis reveals that robustness also depends upon a quantity that we denote the sensitivity.
Danish Pruthi, Bhuwan Dhingra, Zachary C. Lipton
ACL (1)2
2019 PubMedQA: A Dataset for Biomedical Research Question Answering
abstract
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, Xinghua Lu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qiao Jin 0001, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, Xinghua Lu 0001
EMNLP/IJCNLP (1)2
2018 Open Domain Question Answering Using Early Fusion of Knowledge Bases and Text
abstract
Open Domain Question Answering (QA) is evolving from complex pipelined systems to end-to-end deep neural networks.Specialized neural models have been developed for extracting answers from either text alone or Knowledge Bases (KBs) alone.In this paper we look at a more practical setting, namely QA over the combination of a KB and entitylinked text, which is appropriate when an incomplete KB is available with a large text corpus.Building on recent advances in graph representation learning we propose a novel model, GRAFT-Net, for extracting answers from a question-specific subgraph containing text and KB entities and relations.We construct a suite of benchmark tasks for this problem, varying the difficulty of questions, the amount of training data, and KB completeness.We show that GRAFT-Net is competitive with the state-of-the-art when tested using either KBs or text alone, and vastly outperforms existing methods in the combined setting.
Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, William W. Cohen
EMNLP2
2018 GLoMo: Unsupervised Learning of Transferable Relational Graphs
abstract
Modern deep transfer learning approaches have mainly focused on learning generic feature vectors from one task that are transferable to other tasks, such as word embeddings in language and pretrained convolutional features in vision. However, these approaches usually transfer unary features and largely ignore more structured graphical representations. This work explores the possibility of learning generic latent relational graphs that capture dependencies between pairs of data units (e.g., words or pixels) from large-scale unlabeled data and transferring the graphs to downstream tasks. Our proposed transfer learning framework improves performance on various tasks including question answering, natural language inference, sentiment analysis, and image classification. We also show that the learned graphs are generic enough to be transferred to different embeddings on which the graphs have not been trained (including GloVe embeddings, ELMo embeddings, and task-specific RNN hidden units), or embedding-free units such as image pixels.
Zhilin Yang 0001, Junbo Jake Zhao, Bhuwan Dhingra, Kaiming He, William W. Cohen, Ruslan Salakhutdinov, Yann LeCun
NeurIPS3
2017 Bootstrapping Distantly Supervised IE Using Joint Learning and Small Well-Structured Corpora
abstract
We propose a framework to improve the performance of distantly-supervised relation extraction, by jointly learning to solve two related tasks: concept-instance extraction and relation extraction. We further extend this framework to make a novel use of document structure: in some small, well-structured corpora, sections can be identified that correspond to relation arguments, and distantly-labeled examples from such sections tend to have good precision. Using these as seeds we extract additional relation examples by applying label propagation on a graph composed of noisy examples extracted from a large unstructured testing corpus. Combined with the soft constraint that concept examples should have the same type as the second argument of the relation, we get significant improvements over several state-of-the-art approaches to distantly-supervised relation extraction, and reasonable extraction performance even with very small set of distant labels.
Lidong Bing, Bhuwan Dhingra, Kathryn Mazaitis, Jonghyuk Park 0004, William W. Cohen
AAAI2
2017 Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
abstract
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, Li Deng. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Bhuwan Dhingra, Lihong Li 0001, Xiujun Li, Jianfeng Gao 0001, Yun-Nung Chen, Faisal Ahmed 0001, Li Deng 0001
ACL (1)1
2017 Gated-Attention Readers for Text Comprehension
abstract
In this paper we study the problem of answering cloze-style questions over documents.Our model, the Gated-Attention (GA) Reader 1 , integrates a multi-hop architecture with a novel attention mechanism, which is based on multiplicative interactions between the query embedding and the intermediate states of a recurrent neural network document reader.This enables the reader to build query-specific representations of tokens in the document for accurate answer selection.The GA Reader obtains state-of-the-art results on three benchmarks for this task-the CNN & Daily Mail news stories and the Who Did What dataset.The effectiveness of multiplicative interaction is demonstrated by an ablation study, and by comparing to alternative compositional operators for implementing the gated-attention.
Bhuwan Dhingra, Hanxiao Liu, Zhilin Yang 0001, William W. Cohen, Ruslan Salakhutdinov
ACL (1)1
2017 Words or Characters? Fine-grained Gating for Reading Comprehension
Zhilin Yang 0001, Bhuwan Dhingra, Junjie Hu 0001, William W. Cohen, Ruslan Salakhutdinov
ICLR (Poster)2
2017 Using Graphs of Classifiers to Impose Declarative Constraints on Semi-supervised Learning
abstract
We propose a general approach to modeling semi-supervised learning (SSL) algorithms. Specifically, we present a declarative language for modeling both traditional supervised classification tasks and many SSL heuristics, including both well-known heuristics such as co-training and novel domain-specific heuristics. In addition to representing individual SSL heuristics, we show that multiple heuristics can be automatically combined using Bayesian optimization methods. We experiment with two classes of tasks, link-based text classification and relation extraction. We show modest improvements on well-studied link-based classification benchmarks, and state-of-the-art results on relation-extraction tasks for two realistic domains.
Lidong Bing, William W. Cohen, Bhuwan Dhingra
IJCAI3