Tanmay Rajpurohit

dblp:119/9714 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-9302-4244ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 40% Representation and self-supervised learning · 26% Vision and language · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
mathematical reasoning
1.222023
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning · ICLR 2023
LILA: A Unified Benchmark for Mathematical Reasoning · EMNLP 2022
Information retrieval
generative engine optimization
0.812024
GEO: Generative Engine Optimization · KDD 2024
Information retrieval
retrieval models
0.812024
GEO: Generative Engine Optimization · KDD 2024
Information retrieval
search engines
0.812024
GEO: Generative Engine Optimization · KDD 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Hyperbolic Image-text Representations · ICML 2023
Machine learning › Representation and self-supervised learning
hierarchical representation
0.712023
Hyperbolic Image-text Representations · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
0.712023
Hyperbolic Image-text Representations · ICML 2023
Computer vision › Vision and language › multimodal representation
image-text representation
0.712023
Hyperbolic Image-text Representations · ICML 2023
Natural language and speech › Question answering and dialogue systems
math word problem solving
0.712023
Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation · EMNLP 2023
Computer vision › Vision and language › vision-language model
prompt learning
0.712023
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning · ICLR 2023
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.712023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023
Natural language and speech › Language models and text generation › evaluation of language models › reasoning evaluation
mathematical reasoning benchmark
0.612022
LILA: A Unified Benchmark for Mathematical Reasoning · EMNLP 2022
Natural language and speech › Language models and text generation
natural language understanding
0.212023
C-STS: Conditional Semantic Textual Similarity · EMNLP 2023
Computing education
intelligent tutoring systems
0.212023
Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation · EMNLP 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.212022
LILA: A Unified Benchmark for Mathematical Reasoning · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

personalized learning · 1.3knowledge tracing · 1.3knowledge distillation · 1.3large language model · 0.8black-box optimization · 0.8reinforcement learning · 0.7policy gradient · 0.7hyperbolic geometry · 0.7dynamic prompting · 0.7dataset construction · 0.7contrastive learning · 0.7benchmarking · 0.6
YearPublicationVenuePosition
2025 Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
abstract
Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B. Sai, John J Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya Sai, John J. Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran
ACL (1)6
2024 GEO: Generative Engine Optimization
abstract
The advent of large language models (LLMs) has ushered in a new paradigm of search engines that use generative models to gather and summarize information to answer user queries. This emerging technology, which we formalize under the unified framework of generative engines (GEs), can generate accurate and personalized responses, rapidly replacing traditional search engines like Google and Bing. Generative Engines typically satisfy queries by synthesizing information from multiple sources and summarizing them using LLMs. While this shift significantly improvesuser utility and generative search engine traffic, it poses a huge challenge for the third stakeholder -- website and content creators. Given the black-box and fast-moving nature of generative engines, content creators have little to no control over when and how their content is displayed. With generative engines here to stay, we must ensure the creator economy is not disadvantaged. To address this, we introduce Generative Engine Optimization (GEO), the first novel paradigm to aid content creators in improving their content visibility in generative engine responses through a flexible black-box optimization framework for optimizing and defining visibility metrics. We facilitate systematic evaluation by introducing GEO-bench, a large-scale benchmark of diverse user queries across multiple domains, along with relevant web sources to answer these queries. Through rigorous evaluation, we demonstrate that GEO can boost visibility by up to 40% in generative engine responses. Moreover, we show the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods. Our work opens a new frontier in information discovery systems, with profound implications for both developers of generative engines and content creators.
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande
KDD3
2024 QualEval: Qualitative Evaluation for Model Improvement
abstract
Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan
NAACL-HLT4
2023 C-STS: Conditional Semantic Textual Similarity
abstract
Ameet Deshpande, Carlos Jimenez, Howard Chen, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen, Karthik Narasimhan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Ameet Deshpande, Carlos E. Jimenez, Howard Chen 0003, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen 0001, Karthik Narasimhan
EMNLP6
2023 Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation
abstract
In this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principles, such as knowledge tracing and personalized learning. Concretely, we let GPT-3 be a math tutor and run two steps iteratively: 1) assessing the student model's current learning status on a GPT-generated exercise book, and 2) improving the student model by training it with tailored exercise samples generated by GPT-3. Experimental results reveal that our approach outperforms LLMs (e.g., GPT-3 and PaLM) in accuracy across three distinct benchmarks while employing significantly fewer parameters. Furthermore, we provide a comprehensive analysis of the various components within our methodology to substantiate their efficacy.
Zhenwen Liang, Wenhao Yu 0002, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang 0001, Ashwin Kalyan
EMNLP3
2023 Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Pan Lu, Liang Qiu 0001, Kai-Wei Chang 0001, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, Ashwin Kalyan
ICLR6
2023 Hyperbolic Image-text Representations
abstract
Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrastive model that yields hyperbolic representations of images and text. Hyperbolic spaces have suitable geometric properties to embed tree-like data, so MERU can better capture the underlying hierarchy in image-text datasets. Our results show that MERU learns a highly interpretable and structured representation space while being competitive with CLIP's performance on standard multi-modal tasks like image classification and image-text retrieval.
Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson 0001, Ramakrishna Vedantam
ICML3
2022 LILA: A Unified Benchmark for Mathematical Reasoning
abstract
Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan
EMNLP7