EDBT 2026 Demo / reviewers in the wild / expert
Danish Pruthi
dblp:192/7349
· DBLP profile ↗
19ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0002-1789-6803ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond World Models: Rethinking Understanding in AI ModelsabstractWorld models have garnered substantial interest in the AI community. These are internal representations that simulate aspects of the external world, track entities and states, capture causal relationships, and enable prediction of consequences. This contrasts with representations based solely on statistical correlations. A key motivation behind this research direction is that humans possess such mental world models, and finding evidence of similar representations in AI models might indicate that these models "understand" the world in a human-like way. In this paper, we use case studies from the philosophy of science literature to critically examine whether the world model framework adequately characterizes human-level understanding. We focus on specific philosophical analyses where the distinction between world model capabilities and human understanding is most pronounced. While these represent particular views of understanding rather than universal definitions, they help us explore the limits of world models. Danish Pruthi |
AAAI | 2 |
| 2026 | TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated StoriesabstractMillions of users across the globe turn to AI chatbots for their creative needs, inviting widespread interest in understanding how they represent diverse cultures. However, evaluating cultural representations in open-ended tasks remains challenging and underexplored. In this work, we present TALES, an evaluation of cultural misrepresentations in LLM-generated stories for diverse Indian cultural identities. First, we develop TALES-Tax, a taxonomy of cultural misrepresentations by collating insights from participants with lived experiences in India through focus groups (N=9) and individual surveys (N=15). Using TALES-Tax, we evaluate 6 models through a large-scale annotation study spanning 2,925 annotations from 108 annotators with lived experience and native language proficiency from across 71 regions in India and 14 languages. Concerningly, we find that 88% of the generated stories contain misrepresentations, and such errors are more prevalent in mid- and low-resourced languages and stories based in peri-urban regions in India. We also transform the annotations into TALES-QA, a standalone question bank to evaluate the cultural knowledge of models. Kirti Bhagat, Shaily Bhatt, Athul Velagapudi, Aditya Vashistha, Shachi Dave, Danish Pruthi |
CHI | 6 |
| 2025 | All That Glitters is Not Novel: Plagiarism in AI Generated ResearchabstractAutomating scientific research is considered the final frontier of science.Recently, several papers claim autonomous research agents can generate novel research ideas.Amidst the prevailing optimism, we document a critical concern: a considerable fraction of such research documents are smartly plagiarized.Unlike past efforts where experts evaluate the novelty and feasibility of research ideas, we request 13 experts to operate under a different situational logic: to identify similarities between LLM-generated research documents and existing work.Concerningly, the experts identify 24% of the 50 evaluated research documents to be either paraphrased (with one-to-one methodological mapping), or significantly borrowed from existing work.These reported instances are cross-verified by authors of the source papers.Experts find an additional 32% ideas to partially overlap with prior work, and a small fraction to be completely original.Problematically, these LLM-generated research documents do not acknowledge original sources, and bypass inbuilt plagiarism detectors.Lastly, through controlled experiments we show that automated plagiarism detectors are inadequate at catching plagiarized ideas from such systems.We recommend a careful assessment of LLMgenerated research, and discuss the implications of our findings on academic publishing. Danish Pruthi |
ACL (1) | 2 |
| 2025 | FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and StereotypesabstractJanki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta, Danish Pruthi, Mitesh M Khapra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Janki Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Danish Pruthi, Mitesh M. Khapra |
ACL (1) | 5 |
| 2025 | Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on TwitchabstractTo meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement (\textit{e.g.}, users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalence, relatively little is known about the effectiveness of these systems. In this paper, we conduct an audit of Twitch’s automated moderation tool (\texttt{AutoMod}) to investigate its effectiveness in flagging hateful content. For our audit, we create streaming accounts to act as siloed test beds, and interface with the live chat using Twitch’s APIs to send over 107,000 comments collated from 4 datasets. We measure \texttt{AutoMod}‘s accuracy in flagging blatantly hateful content containing misogyny, racism, ableism and homophobia. Our experiments reveal that a large fraction of hateful messages, up to 94% on some datasets, \text{\textit{bypass moderation}}. Contextual addition of slurs to these messages results in 100% removal, revealing \texttt{AutoMod}‘s reliance on slurs as a hate signal. We also find that contrary to Twitch’s community guidelines, \texttt{AutoMod} blocks up to 89.5% of benign examples that use sensitive words in pedagogical or empowering contexts. Overall, our audit points to large gaps in \texttt{AutoMod}‘s capabilities and underscores the importance for such systems to understand context effectively. Prarabdh Shukla, Wei Yin Chong, Brennan Schaffner, Danish Pruthi, Arjun Nitin Bhagoji |
ACL (1) | 5 |
| 2025 | STAMP Your Content: Proving Dataset Membership via Watermarked RephrasingsabstractGiven how large parts of publicly available text are crawled to pretrain large language models (LLMs), data creators increasingly worry about the inclusion of their proprietary data for model training without attribution or licensing. Their concerns are also shared by benchmark curators whose test-sets might be compromised. In this paper, we present STAMP, a framework for detecting dataset membership—i.e., determining the inclusion of a dataset in the pretraining corpora of LLMs. Given an original piece of content, our proposal involves first generating multiple rephrases, each embedding a watermark with a unique secret key. One version is to be released publicly, while others are to be kept private. Subsequently, creators can compare model likelihoods between public and private versions using paired statistical tests to prove membership. We show that our framework can successfully detect contamination across four benchmarks which appear only once in the training data and constitute less than 0.001% of the total tokens, outperforming several contamination detection and dataset inference baselines. We verify that STAMP preserves both the semantic meaning and utility of the original data. We apply STAMP to two real-world scenarios to confirm the inclusion of paper abstracts and blog articles in the pretraining corpora. Saksham Rastogi, Pratyush Maini, Danish Pruthi |
ICML | 3 |
| 2025 | Knowledge Graph Guided Evaluation of Abstention TechniquesabstractKinshuk Vasisht, Navreet Kaur, Danish Pruthi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kinshuk Vasisht, Navreet Kaur 0002, Danish Pruthi |
NAACL (Long Papers) | 3 |
| 2024 | Revisiting the Robustness of Watermarking to Paraphrasing AttacksabstractAmidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model.Many recent watermarking techniques slightly modify the output probabilities of LMs to embed a signal in the generated output that can later be detected.Since early proposals for text watermarking, questions about their robustness to paraphrasing have been prominently discussed.Lately, some techniques are deliberately designed and claimed to be robust to paraphrasing.However, such watermarking schemes do not adequately account for the ease with which they can be reverse-engineered.We show that with access to only a limited number of generations from a black-box watermarked model, we can drastically increase the effectiveness of paraphrasing attacks to evade watermark detection, thereby rendering the watermark ineffective. 1 Saksham Rastogi, Danish Pruthi |
EMNLP | 2 |
| 2023 | Learning the Legibility of Visual Text PerturbationsabstractMany adversarial attacks in NLP perturb inputs to produce visually similar strings ('ergo' → 'εrgo') which are legible to humans but degrade model performance.Although preserving legibility is a necessary condition for text perturbation, little work has been done to systematically characterize it; instead, legibility is typically loosely enforced via intuitions around the nature and extent of perturbations.Particularly, it is unclear to what extent can inputs be perturbed while preserving legibility, or how to quantify the legibility of a perturbed string.In this work, we address this gap by learning models that predict the legibility of a perturbed string, and rank candidate perturbations based on their legibility.To do so, we collect and release LEGIT, a human-annotated dataset comprising the legibility of visually perturbed text.Using this dataset, we build both text-and vision-based models which achieve up to 0.91 F1 score in predicting whether an input is legible, and an accuracy of 0.86 in predicting which of two given perturbations is more legible.Additionally, we discover that legible perturbations from the LEGIT dataset are more effective at lowering the performance of NLP models than best-known attack strategies, suggesting that current models may be vulnerable to a broad range of perturbations beyond what is captured by existing visual attacks. 1 Dev Seth, Rickard Stureborg, Danish Pruthi, Bhuwan Dhingra |
EACL | 3 |
| 2023 | Model-tuning Via Prompts Makes NLP Models Adversarially RobustabstractIn recent years, NLP practitioners have converged on the following practice: (i) import an off-the-shelf pretrained (masked) language model; (ii) append a multilayer perceptron atop the CLS token's hidden representation (with randomly initialized weights); and (iii) finetune the entire model on a downstream task (MLP-FT).This procedure has produced massive gains on standard NLP benchmarks, but these models remain brittle, even to mild adversarial perturbations.In this work, we demonstrate surprising gains in adversarial robustness enjoyed by Model-tuning Via Prompts (MVP), an alternative method of adapting to downstream tasks.Rather than appending an MLP head to make output prediction, MVP appends a prompt template to the input, and makes prediction via text infilling/completion. Across 5 NLP datasets, 4 adversarial attacks, and 3 different models, MVP improves performance against adversarial substitutions by an average of 8% over standard methods and even outperforms adversarial training-based state-of-art defenses by 3.5%.By combining MVP with adversarial training, we achieve further improvements in adversarial robustness while maintaining performance on unperturbed examples.Finally, we conduct ablations to investigate the mechanism underlying these gains.Notably, we find that the main causes of vulnerability of MLP-FT can be attributed to the misalignment between pre-training and fine-tuning tasks, and the randomly initialized MLP parameters. 1 Mrigank Raman, Pratyush Maini, J. Zico Kolter, Zachary C. Lipton, Danish Pruthi |
EMNLP | 5 |
| 2023 | Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsabstractRecent progress in generative models has resulted in models that produce both realistic as well as relevant images for most textual inputs. These models are being used to generate millions of images everyday, and hold the potential to drastically impact areas such as generative art, digital marketing and data augmentation. Given their outsized impact, it is important to ensure that the generated content reflects the artifacts and surroundings across the globe, rather than over-representing certain parts of the world. In this paper, we measure the geographical representativeness of common nouns (e.g., a house) generated through DALL•E 2 and Stable Diffusion models using a crowdsourced study comprising 540 participants across 27 countries. For deliberately underspecified inputs without country names, the generated images most reflect the surroundings of the United States followed by India, and the top generations rarely reflect surroundings from all other countries (average score less than 3 out of 5). Specifying the country names in the input increases the representativeness by 1.44 points on average on a 5 − point Likert scale for DALL•E 2 and 0.75 for Stable Diffusion, however, the overall scores for many countries still remain low, highlighting the need for future models to be more geographically inclusive. Lastly, we examine the feasibility of quantifying the geographical representativeness of generated images without conducting user studies.1 Abhipsa Basu, Venkatesh Babu Radhakrishnan, Danish Pruthi |
ICCV | 3 |
| 2022 | Explain, Edit, and Understand: Rethinking User Study Design for Evaluating Model ExplanationsabstractIn attempts to "explain" predictions of machine learning models, researchers have proposed hundreds of techniques for attributing predictions to features that are deemed important. While these attributions are often claimed to hold the potential to improve human "understanding" of the models, surprisingly little work explicitly evaluates progress towards this aspiration. In this paper, we conduct a crowdsourcing study, where participants interact with deception detection models that have been trained to distinguish between genuine and fake hotel reviews. They are challenged both to simulate the model on fresh reviews, and to edit reviews with the goal of lowering the probability of the originally predicted class. Successful manipulations would lead to an adversarial example. During the training (but not the test) phase, input spans are highlighted to communicate salience. Through our evaluation, we observe that for a linear bag-of-words model, participants with access to the feature coefficients during training are able to cause a larger reduction in model confidence in the testing phase when compared to the no-explanation control. For the BERT-based classifier, popular local explanations do not improve their ability to reduce the model confidence over the no-explanation case. Remarkably, when the explanation for the BERT model is given by the (global) attributions of a linear model trained to imitate the BERT model, people can effectively manipulate the model. Siddhant Arora, Danish Pruthi, Norman M. Sadeh, William W. Cohen, Zachary C. Lipton, Graham Neubig |
AAAI | 2 |
| 2022 | Measures of Information Reflect Memorization PatternsabstractNeural networks are known to exploit spurious artifacts (or shortcuts) that co-occur with a target label, exhibiting heuristic memorization. On the other hand, networks have been shown to memorize training examples, resulting in example-level memorization. These kinds of memorization impede generalization of networks beyond their training distributions. Detecting such memorization could be challenging, often requiring researchers to curate tailored test sets. In this work, we hypothesize—and subsequently show—that the diversity in the activation patterns of different neurons is reflective of model generalization and memorization. We quantify the diversity in the neural activations through information-theoretic measures and find support for our hypothesis in experiments spanning several natural language and vision tasks. Importantly, we discover that information organization points to the two forms of memorization, even for neural activations computed on unlabeled in-distribution examples. Lastly, we demonstrate the utility of our findings for the problem of model selection. Rachit Bansal, Danish Pruthi, Yonatan Belinkov |
NeurIPS | 2 |
| 2022 | Learning to Scaffold: Optimizing Model Explanations for TeachingabstractModern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior. However, what is the precise goal of providing such explanations, and how can we demonstrate that explanations achieve this goal? Some research argues that explanations should help teach a student (either human or machine) to simulate the model being explained, and that the quality of explanations can be measured by the simulation accuracy of students on unexplained examples. In this work, leveraging meta-learning techniques, we extend this idea to improve the quality of the explanations themselves, specifically by optimizing explanations such that student models more effectively learn to simulate the original model. We train models on three natural language processing and computer vision tasks, and find that students trained with explanations extracted with our framework are able to simulate the teacher significantly more effectively than ones produced with previous methods. Through human annotations and a user study, we further find that these learned explanations more closely align with how humans would explain the required decisions in these tasks. Our code is available at https://github.com/coderpat/learning-scaffold. Patrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins, Graham Neubig |
NeurIPS | 3 |
| 2022 | Evaluating Explanations: How Much Do Explanations from the Teacher Aid Students?abstractAbstract While many methods purport to explain predictions by highlighting salient features, what aims these explanations serve and how they ought to be evaluated often go unstated. In this work, we introduce a framework to quantify the value of explanations via the accuracy gains that they confer on a student model trained to simulate a teacher model. Crucially, the explanations are available to the student during training, but are not available at test time. Compared with prior proposals, our approach is less easily gamed, enabling principled, automatic, model-agnostic evaluation of attributions. Using our framework, we compare numerous attribution methods for text classification and question answering, and observe quantitative differences that are consistent (to a moderate to high degree) across different student model architectures and learning strategies.1 Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio B. Soares, Michael Collins 0001, Zachary C. Lipton, Graham Neubig, William W. Cohen |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Do Context-Aware Translation Models Pay the Right Attention?abstractKayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig |
ACL/IJCNLP (1) | 3 |
| 2020 | Learning to Deceive with Attention-Based ExplanationsabstractAttention mechanisms are ubiquitous components in neural architectures applied to natural language processing.In addition to yielding gains in predictive accuracy, attention weights are often claimed to confer interpretability, purportedly useful both for providing insights to practitioners and for explaining why a model makes its decisions to stakeholders.We call the latter use of attention mechanisms into question by demonstrating a simple method for training models to produce deceptive attention masks.Our method diminishes the total weight assigned to designated impermissible tokens, even when the models can be shown to nevertheless rely on these features to drive predictions.Across multiple models and tasks, our approach manipulates attention weights while paying surprisingly little cost in accuracy.Through a human study, we show that our manipulated attention-based explanations deceive people into thinking that predictions from a model biased against gender minorities do not rely on the gender.Consequently, our results cast doubt on attention's reliability as a tool for auditing algorithms in the context of fairness and accountability.1 Danish Pruthi, Bhuwan Dhingra, Graham Neubig, Zachary C. Lipton |
ACL | 1 |
| 2019 | Combating Adversarial Misspellings with Robust Word RecognitionabstractTo combat adversarial spelling mistakes, we propose placing a word recognition model in front of the downstream classifier.Our word recognition models build upon the RNN semicharacter architecture, introducing several new backoff strategies for handling rare and unseen words.Trained to recognize words corrupted by random adds, drops, swaps, and keyboard mistakes, our method achieves 32% relative (and 3.3% absolute) error reduction over the vanilla semi-character model.Notably, our pipeline confers robustness on the downstream classifier, outperforming both adversarial training and off-the-shelf spell checkers.Against a BERT model fine-tuned for sentiment analysis, a single adversarially-chosen character attack lowers accuracy from 90.3% to 45.8%.Our defense restores accuracy to 75% 1 .Surprisingly, better word recognition does not always entail greater robustness.Our analysis reveals that robustness also depends upon a quantity that we denote the sensitivity. Danish Pruthi, Bhuwan Dhingra, Zachary C. Lipton |
ACL (1) | 1 |
| 2018 | SPINE: SParse Interpretable Neural EmbeddingsabstractPrediction without justification has limited utility. Much of the success of neural models can be attributed to their ability to learn rich, dense and expressive representations. While these representations capture the underlying complexity and latent trends in the data, they are far from being interpretable. We propose a novel variant of denoising k-sparse autoencoders that generates highly efficient and interpretable distributed word representations (word embeddings), beginning with existing word representations from state-of-the-art methods like GloVe and word2vec. Through large scale human evaluation, we report that our resulting word embedddings are much more interpretable than the original GloVe and word2vec embeddings. Moreover, our embeddings outperform existing popular word embeddings on a diverse suite of benchmark downstream tasks. Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Eduard H. Hovy |
AAAI | 2 |