EDBT 2026 Demo / reviewers in the wild / expert
Kalpesh Krishna
dblp:207/8485
· DBLP profile ↗
22ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0001-6574-0817ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented GenerationabstractSatyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui |
NAACL (Long Papers) | 2 |
| 2024 | PostMark: A Robust Blackbox Watermark for Large Language ModelsabstractThe most effective techniques to detect LLMgenerated text rely on inserting a detectable signature-or watermark-during the model's decoding process.Most existing watermarking methods require access to the underlying LLM's logits, which LLM API providers are loath to share due to fears of model distillation.As such, these watermarks must be implemented independently by each LLM provider.In this paper, we develop POSTMARK, a modular post-hoc watermarking procedure in which an input-dependent set of words (determined via a semantic embedding) is inserted into the text after the decoding process has completed.Critically, POSTMARK does not require logit access, which means it can be implemented by a third party.We also show that POST-MARK is more robust to paraphrasing attacks than existing watermarking methods: our experiments cover eight baseline algorithms, five base LLMs, and three datasets.Finally, we evaluate the impact of POSTMARK on text quality using both automated and human assessments, highlighting the trade-off between quality and robustness to paraphrasing.We release our code, outputs, and annotations at https://github.com/lilakk/PostMark. Yapei Chang, Kalpesh Krishna, Amir Houmansadr, John Wieting, Mohit Iyyer |
EMNLP | 2 |
| 2024 | Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationabstractAs large language models (LLMs) evolve, evaluating their output reliably becomes increasingly difficult due to the high cost of human evaluation.To address this, we introduce FLAMe, a family of Foundational Large Autorater Models.FLAMe is trained on a diverse set of over 100 quality assessment tasks, incorporating 5M+ human judgments curated from publicly released human evaluations.FLAMe outperforms models like GPT-4 and Claude-3 on various held-out tasks, and serves as a powerful starting point for finetuning, as shown in our reward model evaluation case study (FLAMe-RM).On Reward-Bench, FLAMe-RM-24B achieves 87.8% accuracy, surpassing GPT-4-0125 (85.9%) and GPT-4o (84.7%).Additionally, we introduce FLAMe-Opt-RM, an efficient tail-patch finetuning approach that offers competitive Re-wardBench performance using 25× fewer training datapoints.Our FLAMe variants outperform popular proprietary LLM-as-a-Judge models on 8 of 12 autorater benchmarks, covering 53 quality assessment tasks, including RewardBench and LLM-AggreFact.Finally, our analysis shows that FLAMe is significantly less biased than other LLM-as-a-Judge models on the CoBBLEr autorater bias benchmark.1 * Tu Vu and Kalpesh Krishna contributed equally to the project leadership, design, and implementation of the work.† Work done while at UMass Amherst.‡ Equal contribution as senior advisors.1 The FLAMe collection is available at https:// huggingface.co/datasets/google/flame-collection."""Input format.""" INSTRUCTIONS:"""Task definition and evaluation instructions.""" Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, Yun-Hsuan Sung |
EMNLP | 2 |
| 2023 | Stealing the Decoding Algorithms of Language ModelsabstractA key component of generating text from modern language models (LM) is the selection and tuning of decoding algorithms. These algorithms determine how to generate text from the internal probability distribution generated by the LM. The process of choosing a decoding algorithm and tuning its hyperparameters takes significant time, manual effort, and computation, and it also requires extensive human evaluation. Therefore, the identity and hyperparameters of such decoding algorithms are considered to be extremely valuable to their owners. In this work, we show, for the first time, that an adversary with typical API access to an LM can steal the type and hyperparameters of its decoding algorithms at very low monetary costs. Our attack is effective against popular LMs used in text generation APIs, including GPT-2, GPT-3 and GPT-Neo. We demonstrate the feasibility of stealing such information with only a few dollars, e.g., 0.8, 1, 4, and 40 for the four versions of GPT-3. Ali Naseh, Kalpesh Krishna, Mohit Iyyer, Amir Houmansadr |
CCS | 2 |
| 2023 | LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form SummarizationabstractKalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo |
EACL | 1 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 2 |
| 2023 | Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseabstractThe rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier detection. However, the robustness of these detection algorithms to paraphrases of AI-generated text remains unclear. To stress test these detectors, we build a 11B parameter paraphrase generation model (DIPPER) that can paraphrase paragraphs, condition on surrounding context, and control lexical diversity and content reordering. Paraphrasing text generated by three large language models (including GPT3.5-davinci-003) with DIPPER successfully evades several detectors, including watermarking, GPTZero, DetectGPT, and OpenAI's text classifier. For example, DIPPER drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%), without appreciably modifying the input semantics.
To increase the robustness of AI-generated text detection to paraphrase attacks, we introduce a simple defense that relies on retrieving semantically-similar generations and must be maintained by a language model API provider. Given a candidate text, our algorithm searches a database of sequences previously generated by the API, looking for sequences that match the candidate text within a certain threshold. We empirically verify our defense using a database of 15M generations from a fine-tuned T5-XXL model and find that it can detect 80% to 97% of paraphrased generations across different settings while only classifying 1% of human-written sequences as AI-generated. We open-source our models, code and data. Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer |
NeurIPS | 1 |
| 2022 | Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsabstractStyle transfer is the task of rewriting a sentence into a target style while approximately preserving content.While most prior literature assumes access to a large style-labelled corpus, recent work (Riley et al., 2021) has attempted "few-shot" style transfer using just 3-10 sentences at inference for style extraction.In this work, we study a relevant low-resource setting: style transfer for languages where no style-labelled corpora are available.We notice that existing few-shot methods perform this task poorly, often copying inputs verbatim.We push the state-of-the-art for few-shot style transfer with a new method modeling the stylistic difference between paraphrases.When compared to prior work, our model achieves 2-3x better performance in formality transfer and code-mixing addition across seven languages.Moreover, our method is better at controlling the style transfer magnitude using an input scalar knob.We report promising qualitative results for several attribute transfer tasks (sentiment transfer, simplification, gender neutralization, text anonymization) all without retraining the model.Finally, we find model evaluation to be difficult due to the lack of datasets and metrics for many languages.To facilitate future research we crowdsource formality annotations for 4000 sentence pairs in four Indic languages, and use this data to design our automatic evaluations. 1 Kalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta, Partha Talukdar |
ACL (1) | 1 |
| 2022 | RELiC: Retrieving Evidence for Literary ClaimsabstractHumanities scholars commonly provide evidence for claims that they make about a work of literature (e.g., a novel) in the form of quotations from the work.We collect a large-scale dataset (RELiC) of 78K literary quotations and surrounding critical analysis and use it to formulate the novel task of literary evidence retrieval, in which models are given an excerpt of literary analysis surrounding a masked quotation and asked to retrieve the quoted passage from the set of all passages in the work.Solving this retrieval task requires a deep understanding of complex literary and linguistic phenomena, which proves challenging to methods that overwhelmingly rely on lexical and semantic similarity matching.We implement a RoBERTa-based dense passage retriever for this task that outperforms existing pretrained information retrieval baselines; however, experiments and analysis by human domain experts indicate that there is substantial room for improvement over our dense retriever. Katherine Thai, Yapei Chang, Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 3 |
| 2022 | RankGen: Improving Text Generation with Large Ranking ModelsabstractGiven an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts.To address these issues we present RANKGEN, a 1.2B parameter encoder model for English that scores model generations given a prefix.RANKGEN can be flexibly incorporated as a scoring function in beam search and used to decode from any pretrained language model.We train RANKGEN using large-scale contrastive learning to map a prefix close to the ground-truth sequence that follows it and far away from two types of negatives:(1) random sequences from the same document as the prefix, and (2) sequences generated from a large language model conditioned on the prefix.Experiments across four different language models (345M-11B parameters) and two domains show that RANKGEN significantly outperforms decoding algorithms like nucleus, top-k, and typical sampling on both automatic metrics (85.0 vs 77.3 MAUVE) as well as human evaluations with English writers (74.5% human preference over nucleus sampling).Analysis reveals that RANKGEN outputs are more relevant to the prefix and improve continuity and coherence compared to baselines.We release our model checkpoints, code, and human preference data with explanations to facilitate future research.1 Kalpesh Krishna, Yapei Chang, John Wieting, Mohit Iyyer |
EMNLP | 1 |
| 2022 | SLING: Sino Linguistic Evaluation of Large Language ModelsabstractTo understand what kinds of linguistic knowledge are encoded by pretrained Chinese language models (LMs), we introduce the benchmark of Sino LINGuistics (SLING), which consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.Each pair demonstrates the acceptability contrast of a specific syntactic or semantic phenomenon (e.g., The keys are lost vs.The keys is lost), and an LM should assign lower perplexity to the acceptable sentence.In contrast to the CLiMP dataset (Xiang et al., 2021), which also contains Chinese minimal pairs and was created by translating the vocabulary of the English BLiMP dataset, the minimal pairs in SLING are derived primarily by applying syntactic and lexical transformations to naturally-occurring, linguist-annotated sentences from the Chinese Treebank 9.0, thus addressing severe issues in CLiMP's data generation process.We test 18 publicly available pretrained monolingual (e.g., BERT-base-zh, CPM) and multi-lingual (e.g., mT5, XLM) language models on SLING.Our experiments show that the average accuracy for LMs is far below human performance (69.7% vs. 97.1%),while BERT-base-zh achieves the highest accuracy (84.8%) of all tested LMs, even much larger ones.Additionally, we find that most LMs have a strong gender and number (singular/plural) bias, and they perform better on local phenomena than hierarchical ones. 1 Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit Iyyer |
EMNLP | 2 |
| 2022 | Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureabstractLiterary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works published around the world.Machine translation (MT) holds potential to complement the work of human translators by improving both training procedures and their overall efficiency.Literary translation is less constrained than more traditional MT settings since translators must balance meaning equivalence, readability, and critical interpretability in the target language.This property, along with the complex discourse-level context present in literary texts, also makes literary MT more challenging to computationally model and evaluate.To explore this task, we collect a dataset (PAR3) of non-English language novels in the public domain, each aligned at the paragraph level to both human and automatic English translations.Using PAR3, we discover that expert literary translators prefer reference human translations over machinetranslated paragraphs at a rate of 84%, while state-of-the-art automatic MT metrics do not correlate with those preferences.The experts note that MT outputs contain not only mistranslations, but also discourse-disrupting errors and stylistic inconsistencies.To address these problems, we train a post-editing model whose output is preferred over normal MT output at a rate of 69% by experts.We publicly release PAR3 to spur future research into literary MT. 1 Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, Mohit Iyyer |
EMNLP | 3 |
| 2021 | Weakly-Supervised Open-Retrieval Conversational Question Answering
Chen Qu 0001, Liu Yang 0005, Cen Chen 0001, W. Bruce Croft, Kalpesh Krishna, Mohit Iyyer |
ECIR (1) | 5 |
| 2021 | Do Long-Range Language Models Actually Use Long-Range Context?abstractLanguage models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions.Recent efforts to improve the efficiency of self-attention have led to a proliferation of long-range Transformer language models, which can process much longer sequences than models of the past.However, the ways in which such models take advantage of the longrange context remain unclear.In this paper, we perform a fine-grained analysis of two longrange Transformer language models (including the Routing Transformer, which achieves state-of-the-art perplexity on the PG-19 longsequence LM benchmark dataset) that accept input sequences of up to 8K tokens.Our results reveal that providing long-range context (i.e., beyond the previous 2K tokens) to these models only improves their predictions on a small set of tokens (e.g., those that can be copied from the distant context) and does not help at all for sentence-level prediction tasks.Finally, we discover that PG-19 contains a variety of different document types and domains, and that long-range context helps most for literary novels (as opposed to textbooks or magazines). Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer |
EMNLP (1) | 2 |
| 2021 | Hurdles to Progress in Long-form Question AnsweringabstractThe task of long-form question answering (LFQA) involves retrieving documents relevant to a given question and using them to generate a paragraph-length answer.While many models have recently been proposed for LFQA, we show in this paper that the task formulation raises fundamental challenges regarding evaluation and dataset creation that currently preclude meaningful modeling progress.To demonstrate these challenges, we first design a new system that relies on sparse attention and contrastive retriever learning to achieve state-of-the-art performance on the ELI5 LFQA dataset.While our system tops the public leaderboard, a detailed analysis reveals several troubling trends: (1) our system's generated answers are not actually grounded in the documents that it retrieves; (2) ELI5 contains significant train / validation overlap, as at least 81% of ELI5 validation questions occur in paraphrased form in the training set; (3) ROUGE-L is not an informative metric of generated answer quality and can be easily gamed; and (4) human evaluations used for other text generation tasks are unreliable for LFQA.We offer suggestions to mitigate each of these issues, which we hope will lead to more rigorous LFQA research and meaningful progress in the future. 1 * Work done during an internship at Google Research. 1 Resources accompanying our paper can be found in https://github.com/martiansideofthemoon/ hurdles-longform-qa Val Q: Can you protect electronics from EMPs/solar flares?If so, how?Generation using random retrievals, 24.8 ROUGE-L Yes, you absolutely can, in fact you can build a Faraday cage around your electronics, and protect them from solar flares...This is what is done with the Faraday cage around your electronics, which is the problem.The reason it is expensive is because it requires a huge amount of power and is expensive to replace... designed to shield your electronics from solar flares, you will have to pay for the protection.This is because you have to buy a piece of equipment that is designed to shield your electronics from solar flares, and that is expensive.... This is also expensive, but not as expensive as the protection you need to shield your electronics from solar flares... designed to be as cheap as possible... Generation using predicted retrievals, 19.0 ROUGE-L Yes, you can shield them.But it's a slow process... Also, the equipment that's powered by them is a lot more expensive than you'd think, so it's hard to make sure that you're not just shielding them from your remote control.Train Q1: How does an EMP ruin electronics?What does it do?How would they be fixed?Can It be protected against?How?Train Q2: If Earth were hit with a massive EMP, would all of our currently technology be completely unusable permanently?Train Q3: Whenever a electromagnetic pulse (EMP) is released what does it do to electronics to disable them?Train Q4: If earth was hit with an EMP, could we ever restore electricity?If not, why?Train Q5: What are solar flares and why does it impact our electronics?Train Q6.When an EMP goes off, can the electronics affected be replaced?Gold Answer, 18.6 ROUGE-L I'll start with the grounding question, because that's the easiest to answer: Doesn't help a bit.All that matters is that the metal container is conductive and doesn't have gaps...completely seal your Faraday cage.Consider soldering the lid on to that paint can... look at little baggie it comes in.Sealed mylar.That protected that chip from air travel at 35,000 feet, land travel through rural, urban, and suburban areas, and all the electromagnetic radiation that the trip entails...No lead shielding.No safes.... Random Train Ans, 19.4 ROUGE-LThe fast lane/slow lane is a bit of a misnomer.It gives the impression that new, faster lanes are being built.In reality, normal speed will be... Kalpesh Krishna, Aurko Roy, Mohit Iyyer |
NAACL-HLT | 1 |
| 2020 | Reformulating Unsupervised Style Transfer as Paraphrase GenerationabstractModern NLP defines the task of style transfer as modifying the style of a given sentence without appreciably changing its semantics, which implies that the outputs of style transfer systems should be paraphrases of their inputs.However, many existing systems purportedly designed for style transfer inherently warp the input's meaning through attribute transfer, which changes semantic properties such as sentiment.In this paper, we reformulate unsupervised style transfer as a paraphrase generation problem, and present a simple methodology based on fine-tuning pretrained language models on automatically generated paraphrase data.Despite its simplicity, our method significantly outperforms state-of-the-art style transfer systems on both human and automatic evaluations.We also survey 23 style transfer papers and discover that existing automatic metrics can be easily gamed and propose fixed variants.Finally, we pivot to a more real-world style transfer setting by collecting a large dataset of 15M sentences in 11 diverse styles, which we use for an in-depth analysis of our system. Kalpesh Krishna, John Wieting, Mohit Iyyer |
EMNLP (1) | 1 |
| 2020 | Thieves on Sesame Street! Model Extraction of BERT-based APIs
Kalpesh Krishna, Gaurav Tomar, Ankur P. Parikh, Nicolas Papernot, Mohit Iyyer |
ICLR | 1 |
| 2019 | Syntactically Supervised Transformers for Faster Neural Machine TranslationabstractStandard decoders for neural machine translation autoregressively generate a single target token per time step, which slows inference especially for long outputs.While architectural advances such as the Transformer fully parallelize the decoder computations at training time, inference still proceeds sequentially.Recent developments in nonand semiautoregressive decoding produce multiple tokens per time step independently of the others, which improves inference speed but deteriorates translation quality.In this work, we propose the syntactically supervised Transformer (SynST), which first autoregressively predicts a chunked parse tree before generating all of the target tokens in one shot conditioned on the predicted parse.A series of controlled experiments demonstrates that SynST decodes sentences ∼ 5× faster than the baseline autoregressive Transformer while achieving higher BLEU scores than most competing methods on En-De and En-Fr datasets. Nader Akoury, Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 2 |
| 2019 | Generating Question-Answer HierarchiesabstractThe process of knowledge acquisition can be viewed as a question-answer game between a student and a teacher in which the student typically starts by asking broad, open-ended questions before drilling down into specifics (Hintikka, 1981;Hakkarainen and Sintonen, 2002).This pedagogical perspective motivates a new way of representing documents.In this paper, we present SQUASH (Specificity-controlled Question-Answer Hierarchies), a novel and challenging text generation task that converts an input document into a hierarchy of question-answer pairs.Users can click on high-level questions (e.g., "Why did Frodo leave the Fellowship?") to reveal related but more specific questions (e.g., "Who did Frodo leave with?").Using a question taxonomy loosely based on Lehnert (1978), we classify questions in existing reading comprehension datasets as either GENERAL or SPECIFIC.We then use these labels as input to a pipelined system centered around a conditional neural language model.We extensively evaluate the quality of the generated QA hierarchies through crowdsourced experiments and report strong empirical results. Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 1 |
| 2019 | Trick or TReAT : Thematic Reinforcement for Artistic Typography
Purva Tendulkar, Kalpesh Krishna, Ramprasaath R. Selvaraju, Devi Parikh |
ICCC | 2 |
| 2018 | Revisiting the Importance of Encoding Logic Rules in Sentiment ClassificationabstractWe analyze the performance of different sentiment classification models on syntacticallycomplex inputs like A-but-B sentences.The first contribution of this analysis addresses reproducible research: to meaningfully compare different models, their accuracies must be averaged over far more random seeds than what has traditionally been reported.With proper averaging in place, we notice that the distillation model described in Hu et al. (2016), which incorporates explicit logic rules for sentiment classification, is ineffective.In contrast, using contextualized ELMo embeddings (Peters et al., 2018a) instead of logic rules yields significantly better performance.Additionally, we provide analysis and visualizations that demonstrate ELMo's ability to implicitly learn logic rules.Finally, a crowdsourced analysis reveals how ELMo outperforms baseline models even on sentences with ambiguous sentiment labels. Kalpesh Krishna, Preethi Jyothi, Mohit Iyyer |
EMNLP | 1 |
| 2018 | A Study of All-Convolutional Encoders for Connectionist Temporal ClassificationabstractConnectionist temporal classification (CTC) is a popular sequence prediction approach for automatic speech recognition that is typically used with models based on recurrent neural networks (RNNs). We explore whether deep convolutional neural networks (CNNs) can be used effectively instead of RNNs as the “encoder” in CTC. CNNs lack an explicit representation of the entire sequence, but have the advantage that they are much faster to train. We present an exploration of CNN s as encoders for CTC models, in the context of character-based (lexicon-free) automatic speech recognition. In particular, we explore a range of one-dimensional convolutionallayers, which are particularly efficient. We compare the performance of our CNN-based models against typical RNN-based models in terms of training time, decoding time, model size and word error rate (WER) on the Switchboard Eva12000 corpus. We find that our CNN-based models are close in performance to LSTMs, while not matching them, and are much faster to train and decode. Kalpesh Krishna, Liang Lu 0001, Kevin Gimpel, Karen Livescu |
ICASSP | 1 |