EDBT 2026 Demo / reviewers in the wild / expert
Pradeep Dasigi
dblp:27/7184
· DBLP profile ↗
24ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0001-7127-1316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackabstractLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lester James V. Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar 0009, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi |
ACL (1) | 9 |
| 2025 | Generalizing Verifiable Instruction FollowingabstractA crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like only answer with yes or no" ormention the word `abracadabra' at least 3 times" that the user adds to craft a more useful answer.Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code. Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert 0001, Hannaneh Hajishirzi |
NeurIPS | 6 |
| 2025 | A Large-Scale Study of Reranker Relevance Feedback at InferenceabstractNeural IR systems often employ a retrieve-and-rerank framework: a bi-encoder retrieves a fixed number of candidates (e.g., 𝐾=100), which a cross-encoder then reranks.Recent studies have indicated that relevance feedback from the reranker at inference time can improve the recall of the retriever.The approach works by updating the retriever's query representations via a distillation process that aligns it with the reranker's predictions.While a powerful idea, the arguably narrow scope of past studies focusing on a small number of specific domains such as english question answering and entity retrieval has left a gap in our understanding of how well it generalizes.In this paper, we study inference-time reranker relevance feedback extensively across multiple retrieval domains, languages, and modalities, while also investigating aspects such as the performance and latency implications of the number of distillation updates and feedback candidates. Revanth Gangi Reddy, Pradeep Dasigi, Md. Arafat Sultan, Arman Cohan, Avirup Sil, Heng Ji 0001, Hannaneh Hajishirzi |
SIGIR | 2 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 35 |
| 2024 | Scalable Data Ablation Approximations for Language Models through Modular Training and MergingabstractTraining data compositions for Large Language Models (LLMs) can significantly affect their downstream performance.However, a thorough data ablation study exploring large sets of candidate data mixtures is typically prohibitively expensive since the full effect is seen only after training the models; this can lead practitioners to settle for sub-optimal data mixtures.We propose an efficient method for approximating data ablations which trains individual models on subsets of a training corpus and reuses them across evaluations of combinations of subsets.In continued pre-training experiments, we find that, given an arbitrary evaluation set, the perplexity score of a single model trained on a candidate set of data is strongly correlated with perplexity scores of parameter averages of models trained on distinct partitions of that data.From this finding, we posit that researchers and practitioners can conduct inexpensive simulations of data ablations by maintaining a pool of models that were each trained on partitions of a large training corpus, and assessing candidate data mixtures by evaluating parameter averages of combinations of these models.This approach allows for substantial improvements in amortized training efficiency -scaling only linearly with respect to new data -by enabling reuse of previous training computation, opening new avenues for improving model performance through rigorous, incremental data assessment and mixing. Clara Na, Ian Magnusson, Ananya Harsh Jha, Tom Sherborne, Emma Strubell, Jesse Dodge, Pradeep Dasigi |
EMNLP | 7 |
| 2024 | TRAM: Bridging Trust Regions and Sharpness Aware MinimizationabstractSharpness-aware minimization (SAM) reports improving domain generalization by
reducing the loss surface curvature in the parameter space. However,
generalization during _fine-tuning_ is often more dependent on the
transferability of _representations_ in the function space. Trust-region
methods (TR) target this goal by regularizing representation curvature to reduce
catastrophic forgetting of pre-trained task-agnostic information while adopting
task-specific skills. We consider unifying these strategies for low curvature in
both parameter space and function space to improve out-of-domain (OOD)
generalization. We propose **Trust Region Aware Minimization** (TRAM), a
SAM algorithm fine-tuning for low parameter sharpness and smooth, informative
representations preserving pre-trained structure. TRAM uses a trust region bound
to inform the SAM adversarial neighborhood, introducing an awareness of function
curvature within optimization for flatter minima. We empirically validate TRAM
in vision (cross-dataset adaptation) and text (OOD language modeling, zero-shot
cross-lingual transfer) tasks where robust domain transfer and representation
generality are critical. TRAM outperforms SAM- and TR-based optimization across
all tasks, notably surpassing competing methods for hard transfer between
_anticorrelated_ domains. TRAM establishes a novel standard in
fine-tuning for domain-generalizable models with minimal additional computation
over previous sharpness-aware methods. Tom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao Peng 0009 |
ICLR | 3 |
| 2024 | Evaluating In-Context Learning of Libraries for Code GenerationabstractArkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Arkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi |
NAACL-HLT | 4 |
| 2024 | The Art of Saying No: Contextual Noncompliance in Language ModelsabstractChat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities. Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 4 |
| 2023 | LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form SummarizationabstractKalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo |
EACL | 5 |
| 2023 | AGRO: Adversarial discovery of error-prone Groups for Robust Optimization
Bhargavi Paranjape, Pradeep Dasigi, Vivek Srikumar, Luke Zettlemoyer, Hannaneh Hajishirzi |
ICLR | 2 |
| 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open ResourcesabstractIn this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research. Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi |
NeurIPS | 3 |
| 2022 | Generating Data to Mitigate Spurious Correlations in Natural Language Inference DatasetsabstractNatural language processing models often exploit spurious correlations between taskindependent features and labels in datasets to perform well only within the distributions they are trained on, while not generalising to different task distributions.We propose to tackle this problem by generating a debiased version of a dataset, which can then be used to train a debiased, off-the-shelf model, by simply replacing its training data.Our approach consists of 1) a method for training data generators to generate high-quality, label-consistent data samples; and 2) a filtering mechanism for removing data points that contribute to spurious correlations, measured in terms of z-statistics.We generate debiased versions of the SNLI and MNLI datasets, 1 and we evaluate on a large suite of debiased, outof-distribution, and adversarial test sets.Results show that models trained on our debiased datasets generalise better than those trained on the original datasets in all settings.On the majority of the datasets, our method outperforms or performs comparably to previous state-ofthe-art debiasing strategies, and when combined with an orthogonal technique, productof-experts, it improves further and outperforms previous best results of SNLI-hard and MNLI-hard.* Work done while at the Allen Institute for AI. 1 All our code and the generated datasets are available at https://github.com/jimmycode/ gen-debiased-nli.Generator (Section 2 & Section 4.1) sample z-filter (Section 3 & Section 4.2) Yuxiang Wu, Matt Gardner 0001, Pontus Stenetorp, Pradeep Dasigi |
ACL (1) | 4 |
| 2021 | Learning with Instance Bundles for Reading ComprehensionabstractWhen training most modern reading comprehension models, all the questions associated with a context are treated as being independent from each other.However, closely related questions and their corresponding answers are not independent, and leveraging these relationships could provide a strong supervision signal to a model.Drawing on ideas from contrastive estimation, we introduce several new supervision losses that compare question-answer scores across multiple related instances.Specifically, we normalize these scores across various neighborhoods of closely contrasting questions and/or answers, adding a cross entropy loss term in addition to traditional maximum likelihood estimation.Our techniques require bundles of related question-answer pairs, which we either mine from within existing data or create using automated heuristics.We empirically demonstrate the effectiveness of training with instance bundles on two datasets-HotpotQA and ROPES-showing up to 9% absolute gains in accuracy. Dheeru Dua, Pradeep Dasigi, Sameer Singh 0001, Matt Gardner 0001 |
EMNLP (1) | 2 |
| 2021 | Mitigating False-Negative Contexts in Multi-document Question Answering with Retrieval MarginalizationabstractQuestion Answering (QA) tasks requiring information from multiple documents often rely on a retrieval model to identify relevant information for reasoning.The retrieval model is typically trained to maximize the likelihood of the labeled supporting evidence.However, when retrieving from large text corpora such as Wikipedia, the correct answer can often be obtained from multiple evidence candidates.Moreover, not all such candidates are labeled as positive during annotation, rendering the training signal weak and noisy.This problem is exacerbated when the questions are unanswerable or when the answers are Boolean, since the model cannot rely on lexical overlap to make a connection between the answer and supporting evidence.We develop a new parameterization of set-valued retrieval that handles unanswerable queries, and we show that marginalizing over this set during training allows a model to mitigate false negatives in supporting evidence annotations.We test our method on two multi-document QA datasets, IIRC and HotpotQA.On IIRC, we show that joint modeling with marginalization improves model performance by 5.5 F1 points and achieves a new state-of-the-art performance of 50.5 F1.We also show that retrieval marginalization results in 4.1 QA F1 improvement over a non-marginalized baseline on HotpotQA in the fullwiki setting. 1 * Majority of the work done as an intern at AI2. 1 Code available at https://github.com/ niansong1996/retrieval_marginalization.An Example in IIRC: Q: How many other Cardinals participated in the 2005 papal conclave with Policarpo? Ansong Ni, Matt Gardner 0001, Pradeep Dasigi |
EMNLP (1) | 3 |
| 2021 | A Dataset of Information-Seeking Questions and Answers Anchored in Research PapersabstractPradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner 0001 |
NAACL-HLT | 1 |
| 2020 | IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsabstractHumans often have to read multiple documents to address their information needs.However, most existing reading comprehension (RC) tasks only focus on questions for which the contexts provide all the information required to answer them, thus not evaluating a system's performance at identifying a potential lack of sufficient information and locating sources for that information.To fill this gap, we present a dataset, IIRC, with more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.The questions were written by crowd workers who did not have access to any of the linked documents, leading to questions that have little lexical overlap with the contexts where the answers appear.This process also gave many questions without answers, and those that require discrete reasoning, increasing the difficulty of the task.We follow recent modeling work on various reading comprehension datasets to construct a baseline model for this dataset, finding that it achieves 31.1% F1 on this task, while estimated human performance is 88.4%.The dataset, code for the baseline system, and a leaderboard can be found at https://allennlp.org/iirc. James Ferguson, Matt Gardner 0001, Hannaneh Hajishirzi, Tushar Khot, Pradeep Dasigi |
EMNLP (1) | 5 |
| 2019 | Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential ReasoningabstractPradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, Matt Gardner 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2017 | Ontology-Aware Token Embeddings for Prepositional Phrase AttachmentabstractType-level word embeddings use the same set of parameters to represent all instances of a word regardless of its context, ignoring the inherent lexical ambiguity in language.Instead, we embed semantic concepts (or synsets) as defined in WordNet and represent a word token in a particular context by estimating a distribution over relevant semantic concepts.We use the new, context-sensitive embeddings in a model for predicting prepositional phrase (PP) attachments and jointly learn the concept embeddings and model parameters.We show that using context-sensitive embeddings improves the accuracy of the PP attachment model by 5.4% absolute points, which amounts to a 34.4% relative reduction in errors. Pradeep Dasigi, Waleed Ammar, Chris Dyer, Eduard H. Hovy |
ACL (1) | 1 |
| 2017 | Neural Semantic Parsing with Type Constraints for Semi-Structured TablesabstractWe present a new semantic parsing model for answering compositional questions on semi-structured Wikipedia tables.Our parser is an encoder-decoder neural network with two key technical innovations:(1) a grammar for the decoder that only generates well-typed logical forms; and(2) an entity embedding and linking module that identifies entity mentions while generalizing across tables.We also introduce a novel method for training our neural model with question-answer supervision.On the WIKITABLEQUESTIONS data set, our parser achieves a state-of-theart accuracy of 43.3% for a single model and 45.9% for a 5-model ensemble, improving on the best prior score of 38.7% set by a 15-model ensemble.These results suggest that type constraints and entity linking are valuable components to incorporate in neural semantic parsers. Jayant Krishnamurthy, Pradeep Dasigi, Matt Gardner 0001 |
EMNLP | 2 |
| 2014 | Modeling Newswire Events using Neural Networks for Anomaly Detection
Pradeep Dasigi, Eduard H. Hovy |
COLING | 1 |
| 2014 | Tharwa: A Large Scale Dialectal Arabic - Standard Arabic - English Lexicon
Mona T. Diab, Mohamed Al-Badrashiny, Maryam Aminian, Heba Elfardy, Nizar Habash, Abdelati Hawwari, Wael Salloum, Pradeep Dasigi, Ramy Eskander |
LREC | 9 |
| 2012 | Subgroup Detection in Ideological Discussions
Amjad Abu-Jbara, Pradeep Dasigi, Mona T. Diab, Dragomir R. Radev |
ACL (1) | 2 |
| 2011 | CODACT: Towards Identifying Orthographic Variants in Dialectal Arabic
Pradeep Dasigi, Mona T. Diab |
IJCNLP | 1 |
| 2009 | Experiments in CLIR using fuzzy string search based on surface similarityabstractCross Language Information Retrieval (CLIR) between languages of the same origin is an interesting topic of research. The similarity of the writing systems used for these languages can be used effectively to not only improve CLIR, but to overcome the problems of textual variations, textual errors, and even the lack of linguistic resources like stemmers to an extent. We have conducted CLIR experiments between three languages which use writing systems (scripts) of Brahmi-origin, namely Hindi, Bengali and Marathi. We found significant improvements for all the six language pairs using a method for fuzzy text search based on Surface Similarity. In this paper we report these results and compare them with a baseline CLIR system and a CLIR system that uses Scaled Edit Distance (SED) for fuzzy string matching. Sethuramalingam Subramaniam, Anil Kumar Singh 0001, Pradeep Dasigi, Vasudeva Varma |
SIGIR | 3 |