Keshav Santhanam

dblp:221/1812 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-5939-7944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 62% Cloud and datacenter computing · 38%
Artificial intelligence
5 papers
Language models and text generation · 44% Efficient and distributed learning · 30% Transfer learning and domain adaptation · 25%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel query processing
intra-query parallelism
0.912025
PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025
Cloud and datacenter computing › inference serving
LLM serving
0.912025
PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025
Parallel and multicore computing
parallel query processing
0.912025
PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025
Natural language and speech › Language models and text generation › neural language model
autoregressive transformer
0.712023
Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer Models · NeurIPS 2023
Machine learning › Efficient and distributed learning
inference efficiency
0.712023
Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer Models · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.712023
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023
Information retrieval › document retrieval
passage retrieval
0.712023
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023
Information retrieval
reranking
0.712023
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads · OSDI 2020
Parallel and multicore computing › task scheduling
heterogeneity-aware scheduling
0.412020
Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads · OSDI 2020
Machine learning › Efficient and distributed learning
distributed training
0.112020
Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads · OSDI 2020

Methods — techniques the papers use, named apart from their topics

rule-based multilingual validation · 1.7large language model prompting · 1.7prompt optimization · 1.5language model · 1.5compiler · 1.5knowledge distillation · 1.3LLM prompting · 1.3scheduling policy · 0.9cost model · 0.7
YearPublicationVenuePosition
2025 ColBERT-Serve: Efficient Multi-stage Memory-Mapped Scoring
Kaili Huang, Thejas Venkatesh, Uma Dingankar, Antonio Mallia, Daniel Campos, Christopher Potts, Matei Zaharia, Kwabena Boahen 0001, Omar Khattab, Saarthak Sarup, Keshav Santhanam
ECIR (4)12
2025 PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
abstract
LLM serving systems typically treat user prompts as monolithic inputs, optimizing inference through decoding tricks or inter-query batching. However, many real-world prompts contain *latent semantic parallelism*—decomposable structures where subtasks can be executed independently to reduce latency while preserving meaning.We introduce PARALLELPROMPT, the first benchmark for measuring intra-query parallelism in natural user prompts. Our dataset comprises over 37,000 real-world prompts from public LLM chat logs, each annotated with a structured schema capturing task templates, shared context, and iteration inputs. These schemas are extracted using LLM-assisted prompting with rule-based multilingual validation.To evaluate the benefits of decomposition, we provide an execution suite that benchmarks serial vs. parallel strategies, measuring latency, structural adherence, and semantic fidelity. Our results show that intra-query parallelism can be successfully parsed in over 75\% of curated datasets, unlocking up to *$5\times$ speedups* on tasks like translation, comprehension, and comparative analysis, with minimal quality degradation.By releasing this benchmark, curation pipeline, and evaluation suite, we provide the first standardized testbed for studying structure-aware execution in LLM serving pipelines.
Steven Kolawole, Keshav Santhanam, Virginia Smith, Pratiksha Thaker
NeurIPS2
2024 DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines
abstract
The ML community is rapidly exploring techniques for prompting language models (LMs) and for stacking them into pipelines that solve complex tasks. Unfortunately, existing LM pipelines are typically implemented using hard-coded “prompt templates”, i.e. lengthy strings discovered via trial and error. Toward a more systematic approach for developing and optimizing LM pipelines, we introduce DSPy, a programming model that abstracts LM pipelines as text transformation graphs, or imperative computational graphs where LMs are invoked through declarative modules. DSPy modules are parameterized, meaning they can learn how to apply compositions of prompting, finetuning, augmentation, and reasoning techniques. We design a compiler that will optimize any DSPy pipeline to maximize a given metric, by creating and collecting demonstrations. We conduct two case studies, showing that succinct DSPy programs can express and optimize pipelines that reason about math word problems, tackle multi-hop retrieval, answer complex questions, and control agent loops. Within minutes of compiling, DSPy can automatically produce pipelines that outperform out-of-the-box few-shot prompting as well as expert-created demonstrations for GPT-3.5 and Llama2-13b-chat. On top of that, DSPy programs compiled for relatively small LMs like 770M parameter T5 and Llama2-13b-chat are competitive with many approaches that rely on large and proprietary LMs like GPT-3.5 and on expert-written prompt chains. DSPy is available at https://github.com/stanfordnlp/dspy
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma 0005, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, Christopher Potts
ICLR5
2023 UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers
abstract
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, Christopher Potts. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts
EMNLP3
2023 Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer Models
abstract
Large language models (LLMs) are highly capable but also computationally expensive. Characterizing the _fundamental tradeoff_ between inference efficiency and model capabilities is thus important, but requires an efficiency metric that is comparable across models from different providers. Unfortunately, raw runtimes measured through black-box APIs do not satisfy this property: model providers can implement software and hardware optimizations orthogonal to the model, and shared infrastructure introduces performance contention. We propose a new metric for inference efficiency called _idealized runtime_, that puts models on equal footing as though they were served on uniform hardware and software without performance contention, and a cost model to efficiently estimate this metric for autoregressive Transformer models. We also propose variants of the idealized runtime that incorporate the number and type of accelerators needed to serve the model. Using these metrics, we compare ten LLMs developed in 2022 to provide the first analysis of inference efficiency-capability tradeoffs; we make several observations from this analysis, including the fact that the superior inference runtime performance of certain APIs is often a byproduct of optimizations within the API rather than the underlying model. Our code is open sourced at https://github.com/stanford-crfm/helm-efficiency.
Deepak Narayanan, Keshav Santhanam, Peter Henderson 0002, Rishi Bommasani, Percy Liang
NeurIPS2
2022 PLAID: An Efficient Engine for Late Interaction Retrieval
abstract
Pre-trained language models are increasingly important components across multiple information retrieval (IR) paradigms. Late interaction, introduced with the ColBERT model and recently refined in ColBERTv2, is a popular paradigm that holds state-of-the-art status across many benchmarks. To dramatically speed up the search latency of late interaction, we introduce the Performance-optimized Late Interaction Driver (PLAID) engine. Without impacting quality, PLAID swiftly eliminates low-scoring passages using a novel centroid interaction mechanism that treats every passage as a lightweight bag of centroids. PLAID uses centroid interaction as well as centroid pruning, a mechanism for sparsifying the bag of centroids, within a highly-optimized engine to reduce late interaction search latency by up to 7x on a GPU and 45x on a CPU against vanilla ColBERTv2, while continuing to deliver state-of-the-art retrieval quality. This allows the PLAID engine with ColBERTv2 to achieve latency of tens of milliseconds on a GPU and tens or just few hundreds of milliseconds on a CPU at large scale, even at the largest scales we evaluate with 140M passages.
Keshav Santhanam, Omar Khattab, Christopher Potts, Matei Zaharia
CIKM1
2022 ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
abstract
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia
NAACL-HLT1
2020 Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, Matei Zaharia
OSDI2
2018 ROLA: A New Distributed Transaction Protocol and Its Formal Analysis
abstract
Designers of distributed database systems face the choice between stronger consistency guarantees and better performance. A number of applications only require read atomicity (RA) and prevention of lost updates (PLU). Existing distributed database systems that meet these requirements also provide additional stronger consistency guarantees (such as causal consistency ), and therefore incur lower performance. In this paper we define a new distributed transaction protocol, ROLA, that targets applications where only RA and PLU are needed. We formally model ROLA in Maude. We then perform model checking to analyze both the correctness and the performance of ROLA. For correctness , we use standard model checking to analyze ROLA’s satisfaction of RA and PLU. To analyze performance we: (a) use statistical model checking to analyze key performance properties; and (b) compare these performance results with those obtained by analyzing in Maude the well-known protocol Walter. Our results show that ROLA outperforms Walter.
Si Liu 0003, Peter Csaba Ölveczky, Keshav Santhanam, Qi Wang 0017, Indranil Gupta, José Meseguer 0001
FASE3