Ameet Deshpande

dblp:220/4337 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0001-9885-0385ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 37% Reinforcement learning · 31% Trustworthy machine learning · 15%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 78% Data mining · 22%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.812024
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs · ICLR 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs · ICLR 2024
Machine learning › Reinforcement learning
off-policy evaluation
0.812024
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation · NeurIPS 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction
0.812024
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation · NeurIPS 2024
Information retrieval
generative engine optimization
0.812024
GEO: Generative Engine Optimization · KDD 2024
Information retrieval
retrieval models
0.812024
GEO: Generative Engine Optimization · KDD 2024
Information retrieval
search engines
0.812024
GEO: Generative Engine Optimization · KDD 2024
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.712023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023
Data mining › predictive modeling › classification › multi-label classification
extreme classification
0.712023
SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme Classification · ICML 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme Classification · ICML 2023
Natural language and speech › Language models and text generation
natural language understanding
0.212023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

semantic class descriptions · 1.3hybrid matching · 1.3contrastive learning · 1.3state abstraction · 0.8persona prompting · 0.8large language model · 0.8de-biasing prompt · 0.8black-box optimization · 0.8abstract reward process · 0.8dataset construction · 0.7
YearPublicationVenuePosition
2025 Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
abstract
Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B. Sai, John J Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya Sai, John J. Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran
ACL (1)3
2024 Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
abstract
Recent works have showcased the ability of large-scale language models (LLMs) to embody diverse personas in their responses, exemplified by prompts like ‘_You are Yoda. Explain the Theory of Relativity._’ While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs’ capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform _basic reasoning tasks_. Our study covers 24 reasoning datasets (spanning mathematics, law, medicine, morals, and more), 4 LLMs (2 versions of ChatGPT-3.5, GPT-4-Turbo, and Llama-2-70b-chat), and 19 diverse personas (e.g., ‘an Asian person’) spanning 5 socio-demographic groups: race, gender, religion, disability, and political affiliation. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked (‘_Are Black people less skilled at mathematics?_’), they manifest stereotypical and often erroneous presumptions when prompted to answer questions while adopting a persona. These can be observed as abstentions in the model’s response, e.g., ‘_As a Black person, I am unable to answer this question as it requires math knowledge_’, and generally result in a substantial drop in performance on reasoning tasks. Our experiments with ChatGPT-3.5 show that this bias is _ubiquitous_—80% of our personas demonstrate bias; it is _significant_—some datasets show performance drops of 70%+; and can be especially _harmful for certain groups_—some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all four LLMs exhibit persona-induced bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern as they do not always manifest as explicit abstentions, and can also be hard-to-avoid—we find de-biasing prompts to have minimal to no effect. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs—a trend on the rise—can surface their deep-rooted biases and have unforeseeable and detrimental side-effects.
Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, Tushar Khot
ICLR3
2024 GEO: Generative Engine Optimization
abstract
The advent of large language models (LLMs) has ushered in a new paradigm of search engines that use generative models to gather and summarize information to answer user queries. This emerging technology, which we formalize under the unified framework of generative engines (GEs), can generate accurate and personalized responses, rapidly replacing traditional search engines like Google and Bing. Generative Engines typically satisfy queries by synthesizing information from multiple sources and summarizing them using LLMs. While this shift significantly improvesuser utility and generative search engine traffic, it poses a huge challenge for the third stakeholder -- website and content creators. Given the black-box and fast-moving nature of generative engines, content creators have little to no control over when and how their content is displayed. With generative engines here to stay, we must ensure the creator economy is not disadvantaged. To address this, we introduce Generative Engine Optimization (GEO), the first novel paradigm to aid content creators in improving their content visibility in generative engine responses through a flexible black-box optimization framework for optimizing and defining visibility metrics. We facilitate systematic evaluation by introducing GEO-bench, a large-scale benchmark of diverse user queries across multiple domains, along with relevant web sources to answer these queries. Through rigorous evaluation, we demonstrate that GEO can boost visibility by up to 40% in generative engine responses. Moreover, we show the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods. Our work opens a new frontier in information discovery systems, with profound implications for both developers of generative engines and content creators.
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande
KDD6
2024 QualEval: Qualitative Evaluation for Model Improvement
abstract
Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan
NAACL-HLT2
2024 Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
abstract
Evaluating policies using off-policy data is crucial for applying reinforcement learning to real-world problems such as healthcare and autonomous driving. Previous methods for *off-policy evaluation* (OPE) generally suffer from high variance or irreducible bias, leading to unacceptably high prediction errors. In this work, we introduce STAR, a framework for OPE that encompasses a broad range of estimators -- which include existing OPE methods as special cases -- that achieve lower mean squared prediction errors. STAR leverages state abstraction to distill complex, potentially continuous problems into compact, discrete models which we call *abstract reward processes* (ARPs). Predictions from ARPs estimated from off-policy data are provably consistent (asymptotically correct). Rather than proposing a specific estimator, we present a new framework for OPE and empirically demonstrate that estimators within STAR outperform existing methods. The best STAR estimator outperforms baselines in all twelve cases studied, and even the median STAR estimator surpasses the baselines in seven out of the twelve cases.
Shreyas Chaudhari, Ameet Deshpande, Bruno C. da Silva 0001, Philip S. Thomas
NeurIPS2
2023 C-STS: Conditional Semantic Textual Similarity
abstract
Ameet Deshpande, Carlos Jimenez, Howard Chen, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen, Karthik Narasimhan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Ameet Deshpande, Carlos E. Jimenez, Howard Chen 0003, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen 0001, Karthik Narasimhan
EMNLP1
2023 SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme Classification
abstract
Extreme classification (XC) involves predicting over large numbers of classes (thousands to millions), with real-world applications like news article classification and e-commerce product tagging. The zero-shot version of this task requires generalization to novel classes without additional supervision. In this paper, we develop SemSup-XC, a model that achieves state-of-the-art zero-shot and few-shot performance on three XC datasets derived from legal, e-commerce, and Wikipedia data. To develop SemSup-XC, we use automatically collected semantic class descriptions to represent classes and facilitate generalization through a novel hybrid matching module that matches input instances to class descriptions using a combination of semantic and lexical similarity. Trained with contrastive learning, SemSup-XC significantly outperforms baselines and establishes state-of-the-art performance on all three datasets considered, gaining up to 12 precision points on zero-shot and more than 10 precision points on one-shot tests, with similar gains for recall@10. Our ablation studies highlight the relative importance of our hybrid matching module and automatically collected class descriptions.
Pranjal Aggarwal, Ameet Deshpande, Karthik Narasimhan
ICML2
2022 When is BERT Multilingual? Isolating Crucial Ingredients for Cross-lingual Transfer
abstract
Ameet Deshpande, Partha Talukdar, Karthik Narasimhan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ameet Deshpande, Partha Talukdar, Karthik Narasimhan
NAACL-HLT1
2019 FigureNet : A Deep Learning model for Question-Answering on Scientific Plots
abstract
Deep Learning has managed to push boundaries in a wide variety of tasks. One area of interest is to tackle problems in reasoning and understanding, with an aim to emulate human intelligence. In this work, we describe a deep learning model that addresses the reasoning task of question-answering on categorical plots. We introduce a novel architecture FigureNet, that learns to identify various plot elements, quantify the represented values and determine a relative ordering of these statistical values. We test our model on the FigureQA dataset which provides images and accompanying questions for scientific plots like bar graphs and pie charts, augmented with rich annotations. Our approach outperforms the state-of-the-art Relation Networks baseline by approximately 7% on this dataset, with a training time that is over an order of magnitude lesser.
Revanth Reddy, Rahul Ramesh, Ameet Deshpande, Mitesh M. Khapra
IJCNN3