Prafulla Kumar Choubey

dblp:203/8260 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 8 first-author · 11 since 2021
YearPublicationVenuePosition
2026 GTA: Generating Long-horizon Tasks for Web Agents at Scale
abstract
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen 0001, Jonathan May, Chien-Sheng Wu
ACL (1)3
2025 Unanswerability Evaluation for Retrieval Augmented Generation
abstract
Existing evaluation frameworks for retrievalaugmented generation (RAG) systems focus on answerable queries, but they overlook the importance of appropriately rejecting unanswerable requests.In this paper, we introduce UAEval4RAG, a comprehensive evaluation framework designed to evaluate whether RAG systems effectively handle unanswerable queries specific to a given knowledge base.We first define a taxonomy with six unanswerable categories, and UAEval4RAG automatically synthesizes diverse and challenging queries for any given knowledge base and evaluate the RAG systems with unanswered ratio and acceptable ratio metrics.We also conduct experiments with various RAG components and prompting strategies across four datasets, which reveals that due to varying knowledge distribution across datasets, no single configuration consistently delivers optimal performance on both answerable and unanswerable requests across different knowledge bases.Our findings highlight the critical role of component selection and prompt design in optimizing RAG systems to balance the accuracy of answerable queries with high rejection rates of unanswerable ones.UAEval4RAG provides valuable insights and tools for developing more robust and reliable RAG systems.
Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu
ACL (1)2
2025 SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
abstract
Indexing is an important step towards strong performance in retrieval-augmented generation (RAG) systems. However, existing methods organize data based on either semantic similarity (similarity) or related information (relatedness), but do not cover both perspectives comprehensively. Our analysis reveals that modeling only one perspective results in insufficient knowledge synthesis, leading to suboptimal performance on complex tasks requiring multihop reasoning. In this paper, we propose SiReRAG, a novel RAG indexing approach that explicitly considers both similar and related information. On the similarity side, we follow existing work and explore some variances to construct a similarity tree based on recursive summarization. On the relatedness side, SiReRAG extracts propositions and entities from texts, groups propositions via shared entities, and generates recursive summaries to construct a relatedness tree. We index and flatten both similarity and relatedness trees into a unified retrieval pool. Our experiments demonstrate that SiReRAG consistently outperforms state-of-the-art indexing methods on three multihop datasets (MuSiQue, 2WikiMultiHopQA, and HotpotQA), with an average 1.9% improvement in F1 scores. As a reasonably efficient solution, SiReRAG enhances existing reranking methods significantly, with up to 7.8% improvement in average F1 scores. Our code is available at https://github.com/SalesforceAIResearch/SiReRAG.
Prafulla Kumar Choubey, Alexander R. Fabbri, Gabriel Bernadett-Shapiro, Rui Zhang 0037, Prasenjit Mitra 0001, Caiming Xiong, Chien-Sheng Wu
ICLR2
2025 Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
abstract
Kaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu
NAACL (Long Papers)3
2024 Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
abstract
Kung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kung-Hsiang Huang, Philippe Laban, Alexander R. Fabbri, Prafulla Kumar Choubey, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu
NAACL-HLT4
2023 Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuning
Chen Xing, Prafulla Kumar Choubey, Chien-Sheng Wu, Caiming Xiong
ICLR3
2022 Conformal Predictor for Improving Zero-Shot Text Classification Efficiency
abstract
Pre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification.0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text pair separately.This increases the computational cost to make inferences linearly in the number of labels.In this work, we improve the efficiency of such cross-encoder-based 0shot models by restricting the number of likely labels using another fast base classifier-based conformal predictor (CP) calibrated on samples labeled by the 0shot model.Since a CP generates prediction sets with coverage guarantees, it reduces the number of target labels without excluding the most probable label based on the 0shot model.We experiment with three intent and two topic classification datasets.With a suitable CP for each dataset, we reduce the average inference time for NLI-and NSP-based models by 25.6% and 22.2% respectively, without dropping performance below the predefined error rate of 1%.
Prafulla Kumar Choubey, Yu Bai 0017, Chien-Sheng Wu, Wenhao Liu 0003, Nazneen Fatema Rajani
EMNLP1
2022 Improving Factual Consistency in Summarization with Compression-Based Post-Editing
abstract
State-of-the-art summarization models still struggle to be factually consistent with the input text.A model-agnostic way to address this problem is post-editing the generated summaries.However, existing approaches typically fail to remove entity errors if a suitable input entity replacement is not available or may insert erroneous content.In our work, we focus on removing extrinsic entity errors, or entities not in the source, to improve consistency while retaining the summary's essential information and form.We propose to use sentence-compression data to train the post-editing model to take a summary with extrinsic entity errors marked with special tokens and output a compressed, well-formed summary with those errors removed.We show that this model improves factual consistency while maintaining ROUGE, improving entity precision by up to 30% on XSum, and that this model can be applied on top of another post-editor, improving entity precision by up to a total of 38%.We perform an extensive comparison of post-editing approaches that demonstrate trade-offs between factual consistency, informativeness, and grammaticality, and we analyze settings where posteditors show the largest improvements.
Alexander R. Fabbri, Prafulla Kumar Choubey, Jesse Vig, Chien-Sheng Wu, Caiming Xiong
EMNLP2
2022 P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts
Benjamin Newman, Prafulla Kumar Choubey, Nazneen Fatema Rajani
ICLR2
2021 Automatic Data Acquisition for Event Coreference Resolution
abstract
We propose to leverage lexical paraphrases and high precision rules informed by news discourse structure to automatically collect coreferential and non-coreferential event pairs from unlabeled English news articles.We perform both manual validation and empirical evaluation on multiple evaluation datasets with different event domains and text genres to assess the quality of our acquired event pairs.We found that a model trained on our acquired event pairs performs comparably as the supervised model when applied to new data out of the training data domains.Further, augmenting human-annotated data with the acquired event pairs provides empirical performance gains on both in-domain and out-of-domain evaluation datasets.
Prafulla Kumar Choubey, Ruihong Huang
EACL1
2021 GFST: Gender-Filtered Self-Training for More Accurate Gender in Translation
abstract
Targeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context.Furthermore, these incorrectly gendered translations have the potential to reflect or amplify social biases.We propose gender-filtered self-training (GFST) to improve gender translation accuracy on unambiguously gendered inputs.Our GFST approach uses a source monolingual corpus and an initial model to generate gender-specific pseudo-parallel corpora which are then filtered and added to the training data.We evaluate GFST on translation from English into five languages, finding that it improves gender accuracy without damaging generic quality.We also show the viability of GFST on several experimental settings, including re-training from scratch, fine-tuning, controlling the gender balance of the data, forward translation, and back-translation. 1
Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana Dinu
EMNLP (1)1
2020 Discourse as a Function of Event: Profiling Discourse Structure in News Articles around the Main Event
abstract
Understanding discourse structures of news articles is vital to effectively contextualize the occurrence of a news event.To enable computational modeling of news structures, we apply an existing theory of functional discourse structure for news articles that revolves around the main event and create a human-annotated corpus of 802 documents spanning over four domains and three media sources.Next, we propose several documentlevel neural-network models to automatically construct news content structures.Finally, we demonstrate that incorporating system predicted news structures yields new state-of-theart performance for event coreference resolution.The news documents we annotated are openly available and the annotations are publicly released for future research 1 .
Prafulla Kumar Choubey, Aaron Lee, Ruihong Huang, Lu Wang 0008
ACL1
2020 One Classifier for All Ambiguous Words: Overcoming Data Sparsity by Utilizing Sense Correlations Across Words
abstract
Most supervised word sense disambiguation (WSD) systems build word-specific classifiers by leveraging labeled data. However, when using word-specific classifiers, the sparseness of annotations leads to inferior sense disambiguation performance on less frequently seen words. To combat data sparsity, we propose to learn a single model that derives sense representations and meanwhile enforces congruence between a word instance and its right sense by using both sense-annotated data and lexical resources. The model is shared across words that allows utilizing sense correlations across words, and therefore helps to transfer common disambiguation rules from annotation-rich words to annotation-lean words. Empirical evaluation on benchmark datasets shows that the proposed shared model outperforms the equivalent classifier-based models by 1.7%, 2.5% and 3.8% in F1-score when using GloVe, ELMo and BERT word embeddings respectively.
Prafulla Kumar Choubey, Ruihong Huang
LREC1
2019 In Plain Sight: Media Bias Through the Lens of Factual Reporting
abstract
Lisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, Lu Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, Lu Wang 0008
EMNLP/IJCNLP (1)5
2018 Improving Event Coreference Resolution by Modeling Correlations between Event Coreference Chains and Document Topic Structures
abstract
This paper proposes a novel approach for event coreference resolution that models correlations between event coreference chains and document topical structures through an Integer Linear Programming formulation.We explicitly model correlations between the main event chains of a document with topic transition sentences, inter-coreference chain correlations, event mention distributional characteristics and sub-event structure, and use them with scores obtained from a local coreference relation classifier for jointly resolving multiple event chains in a document.Our experiments across KBP 2016 and 2017 datasets suggest that each of the structures contribute to improving event coreference resolution performance.
Prafulla Kumar Choubey, Ruihong Huang
ACL (1)1
2017 A Sequential Model for Classifying Temporal Relations between Intra-Sentence Events
abstract
We present a sequential model for temporal relation classification between intrasentence events.The key observation is that the overall syntactic structure and compositional meanings of the multi-word context between events are important for distinguishing among fine-grained temporal relations.Specifically, our approach first extracts a sequence of context words that indicates the temporal relation between two events, which well align with the dependency path between two event mentions.The context word sequence, together with a parts-of-speech tag sequence and a dependency relation sequence that are generated corresponding to the word sequence, are then provided as input to bidirectional recurrent neural network (LSTM) models.The neural nets learn compositional syntactic and semantic representations of contexts surrounding the two events and predict the temporal relation between them.Evaluation of the proposed approach on TimeBank corpus shows that sequential modeling is capable of accurately recognizing temporal relations between events, which outperforms a neural net model using various discrete features as input that imitates previous feature based models.
Prafulla Kumar Choubey, Ruihong Huang
EMNLP1
2017 Event Coreference Resolution by Iteratively Unfolding Inter-dependencies among Events
abstract
We introduce a novel iterative approach for event coreference resolution that gradually builds event clusters by exploiting inter-dependencies among event mentions within the same chain as well as across event chains.Among event mentions in the same chain, we distinguish within-and cross-document event coreference links by using two distinct pairwise classifiers, trained separately to capture differences in feature distributions of within-and crossdocument event clusters.Our event coreference approach alternates between WD and CD clustering and combines arguments from both event clusters after every merge, continuing till no more merge can be made.And then it performs further merging between event chains that are both closely related to a set of other chains of events.Experiments on the ECB+ corpus show that our model outperforms state-of-the-art methods in joint task of WD and CD event coreference resolution.
Prafulla Kumar Choubey, Ruihong Huang
EMNLP1