VLDB 2026 Research / reviewers in the wild / expert
Keyon Vafa
dblp:241/7140
· DBLP profile ↗
10ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Trustworthy machine learning · 44% Language models and text generation · 29% Generative modeling · 16% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% |
Topics — the 15 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.2 | 3 | 2025 | What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models · ICML 2025 Potemkin Understanding in Large Language Models · ICML 2025 Rationales for Sequential Predictions · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.6 | 2 | 2025 | Potemkin Understanding in Large Language Models · ICML 2025 Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function · ICML 2024 |
Machine learning › Trustworthy machine learning
benchmark validity |
0.9 | 1 | 2025 | Potemkin Understanding in Large Language Models · ICML 2025 |
Machine learning › Learning theory
inductive bias |
0.9 | 1 | 2025 | What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.7 | 1 | 2023 | An Invariant Learning Characterization of Controlled Text Generation · ACL (1) 2023 |
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning |
0.7 | 1 | 2023 | An Invariant Learning Characterization of Controlled Text Generation · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | An Invariant Learning Characterization of Controlled Text Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text generation |
0.7 | 1 | 2023 | An Invariant Learning Characterization of Controlled Text Generation · ACL (1) 2023 |
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction |
0.5 | 1 | 2021 | Rationales for Sequential Predictions · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
topic model |
0.4 | 1 | 2020 | Text-Based Ideal Points · ACL 2020 |
Computational social science and digital humanities › political science
political text analysis |
0.4 | 1 | 2020 | Text-Based Ideal Points · ACL 2020 |
Machine learning › Generative modeling
autoregressive model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Machine learning › Generative modeling › normalizing flow
discrete flow model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Machine learning › Generative modeling › generative model
discrete generative model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.1 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
user study · 1.7image-based steering · 1.7NLP-based prediction of human generalization · 1.5lower bound estimation · 0.9inductive bias probe · 0.9benchmark design · 0.9myhill-nerode theorem · 0.8automaton recovery metrics · 0.8invariant learning · 0.7attribute classifier · 0.7randomization · 0.5causal inference · 0.5attenuation bias correction · 0.5probabilistic topic model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Potemkin Understanding in Large Language ModelsabstractLarge language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM’s capabilities based on its answers to a curated set of questions? This paper first introduces a formal framework to address this question. The key is to note that the benchmarks used to test LLMs—such as AP exams—are also those used to test people. However, this raises an implication: such benchmarks are only valid tests if LLMs misunderstand concepts in ways that mirror human misunderstandings. Otherwise, success on benchmarks only demonstrates potemkin understanding: the illusion of understanding driven by answers irreconcilable with how any human would interpret a concept. We present two procedures for quantifying the existence of potemkins: one using a specially designed benchmark in three domains, the other using a general procedure that provides a lower-bound on their prevalence. We find that potemkins are ubiquitous across models, tasks, and domains. We also find that these failures reflect not just incorrect understanding, but deeper internal incoherence in concept representations. Marina Mancoridis, Bec Weeks, Keyon Vafa, Sendhil Mullainathan |
ICML | 3 |
| 2025 | What Has a Foundation Model Found? Using Inductive Bias to Probe for World ModelsabstractFoundation models are premised on the idea that sequence prediction can uncover deeper domain understanding, much like how Kepler's predictions of planetary motion later led to the discovery of Newtonian mechanics. However, evaluating whether these models truly capture deeper structure remains a challenge. We develop a technique for evaluating foundation models that examines how they adapt to synthetic datasets generated from some postulated world model. Our technique measures whether the foundation model's inductive bias aligns with the world model, and so we refer to it as an inductive bias probe. Across multiple domains, we find that foundation models can excel at their training tasks yet fail to develop inductive biases towards the underlying world model when adapted to new tasks. We particularly find that foundation models trained on orbital trajectories consistently fail to apply Newtonian mechanics when adapted to new physics tasks. Further analysis reveals that these models behave as if they develop task-specific heuristics that fail to generalize. Keyon Vafa, Peter G. Chang, Ashesh Rambachan, Sendhil Mullainathan |
ICML | 1 |
| 2025 | What's Producible May Not Be Reachable: Measuring the Steerability of Generative ModelsabstractHow should we evaluate the quality of generative models? Many existing metrics focus on a model's producibility, i.e. the quality and breadth of outputs it can generate. However, the actual value from using a generative model stems not just from what it can produce but whether a user with a specific goal can produce an output that satisfies that goal. We refer to this property as steerability. In this paper, we first introduce a mathematical decomposition for quantifying steerability independently from producibility. Steerability is more challenging to evaluate than producibility because it requires knowing a user's goals. We address this issue by creating a benchmark task that relies on one key idea: sample an output from a generative model and ask users to reproduce it. We implement this benchmark in user studies of text-to-image and large language models. Despite the ability of these models to produce high-quality outputs, they all perform poorly on steerability. These results suggest that we need to focus on improving the steerability of generative models. We show such improvements are indeed possible: simple image-based steering mechanisms achieve more than 2x improvement on this benchmark. Keyon Vafa, Sarah Bentley, Jon M. Kleinberg, Sendhil Mullainathan |
NeurIPS | 1 |
| 2024 | Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization FunctionabstractWhat makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they will be used for. We consider a setting where these deployment decisions are made by people, and in particular, people’s beliefs about where an LLM will perform well. We model such beliefs as the consequence of a human generalization function: having seen what an LLM gets right or wrong, people generalize to where else it might succeed. We collect a dataset of 19K examples of how humans make generalizations across 79 tasks from the MMLU and BIG-Bench benchmarks. We show that the human generalization function can be predicted using NLP methods: people have consistent structured ways to generalize. We then evaluate LLM alignment with the human generalization function. Our results show that – especially for cases where the cost of mistakes is high – more capable models (e.g. GPT-4) can do worse on the instances people choose to use them for, exactly because they are not aligned with the human generalization function. Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan |
ICML | 1 |
| 2024 | Evaluating the World Model Implicit in a Generative ModelabstractRecent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlying reality is governed by a deterministic finite automaton. This includes problems as diverse as simple logical reasoning, geographic navigation, game-playing, and chemistry. We propose new evaluation metrics for world model recovery inspired by the classic Myhill-Nerode theorem from language theory. We illustrate their utility in three domains: game playing, logic puzzles, and navigation. In all domains, the generative models we consider do well on existing diagnostics for assessing world models, but our evaluation metrics reveal their world models to be far less coherent than they appear. Such incoherence creates fragility: using a generative model to solve related but subtly different tasks can lead to failures. Building generative models that meaningfully capture the underlying logic of the domains they model would be immensely valuable; our results suggest new ways to assess how close a given model is to that goal. Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg, Sendhil Mullainathan |
NeurIPS | 1 |
| 2023 | An Invariant Learning Characterization of Controlled Text GenerationabstractControlled generation refers to the problem of creating text that contains stylistic or semantic attributes of interest.Many approaches reduce this problem to training a predictor of the desired attribute.For example, researchers hoping to deploy a large language model to produce non-toxic content may use a toxicity classifier to filter generated text.In practice, the generated text to classify, which is determined by user prompts, may come from a wide range of distributions.In this paper, we show that the performance of controlled generation may be poor if the distributions of text in response to user prompts differ from the distribution the predictor was trained on.To address this problem, we cast controlled generation under distribution shift as an invariant learning problem: the most effective predictor should be invariant across multiple text environments.We then discuss a natural solution that arises from this characterization and propose heuristics for selecting natural environments.We study this characterization and the proposed method empirically using both synthetic and real data.Experiments demonstrate both the challenge of distribution shift in controlled generation and the potential of invariance methods in this setting. Carolina Zheng, Claudia Shi, Keyon Vafa, Amir Feder, David M. Blei |
ACL (1) | 3 |
| 2021 | Rationales for Sequential PredictionsabstractSequence models are a critical component of modern NLP systems, but their predictions are difficult to explain.We consider model explanations though rationales, subsets of context that can explain individual model predictions.We find sequential rationales by solving a combinatorial optimization: the best rationale is the smallest subset of input tokens that would predict the same output as the full sequence.Enumerating all subsets is intractable, so we propose an efficient greedy algorithm to approximate this objective.The algorithm, which is called greedy rationalization, applies to any model.For this approach to be effective, the model should form compatible conditional distributions when making predictions on incomplete subsets of the context.This condition can be enforced with a short finetuning step.We study greedy rationalization on language modeling and machine translation.Compared to existing baselines, greedy rationalization is best at optimizing the sequential objective and provides the most faithful rationales.On a new dataset of annotated sequential rationales, greedy rationales are most similar to human rationales. Keyon Vafa, Yuntian Deng, David M. Blei, Alexander M. Rush |
EMNLP (1) | 1 |
| 2021 | Assessing the Effects of Friend-to-Friend Texting onTurnout in the 2018 US Midterm ElectionsabstractRecent mobile app technology lets people systematize the process of messaging their friends to urge them to vote. Prior to the most recent US midterm elections in 2018, the mobile app Outvote randomized an aspect of their system, hoping to unobtrusively assess the causal effect of their users’ messages on voter turnout. However, properly assessing this causal effect is hindered by multiple statistical challenges, including attenuation bias due to mismeasurement of subjects’ outcomes and low precision due to two-sided non-compliance with subjects’ assignments. We address these challenges, which are likely to impinge upon any study that seeks to randomize authentic friend-to-friend interactions, by tailoring the statistical analysis to make use of additional data about both users and subjects. Using meta-data of users’ in-app behavior, we reconstruct subjects’ positions in users’ queues. We use this information to refine the study population to more compliant subjects who were higher in the queues, and we do so in a systematic way which optimizes a proxy for the study’s power. To mitigate attenuation bias, we then use ancillary data of subjects’ matches to the voter rolls that lets us refine the study population to one with low rates of outcome mismeasurement. Our analysis reveals statistically significant treatment effects from friend-to-friend mobilization efforts ( 8.3, CI = (1.2, 15.3)) that are among the largest reported in the get-out-the-vote (GOTV) literature. While social pressure from friends has long been conjectured to play a role in effective GOTV treatments, the present study is among the first to assess these effects experimentally. Aaron Schein, Keyon Vafa, Dhanya Sridhar, Victor Veitch, Jeffrey Quinn, James Moffet, David M. Blei, Donald P. Green |
WWW | 2 |
| 2020 | Text-Based Ideal PointsabstractIdeal point models analyze lawmakers' votes to quantify their political positions, or ideal points.But votes are not the only way to express a political position.Lawmakers also give speeches, release press statements, and post tweets.In this paper, we introduce the text-based ideal point model (tbip), an unsupervised probabilistic topic model that analyzes texts to quantify the political positions of its authors.We demonstrate the tbip with two types of politicized text data: U.S. Senate speeches and senator tweets.Though the model does not analyze their votes or political affiliations, the tbip separates lawmakers by party, learns interpretable politicized topics, and infers ideal points close to the classical vote-based ideal points.One benefit of analyzing texts, as opposed to votes, is that the tbip can estimate ideal points of anyone who authors political texts, including non-voting actors.To this end, we use it to study tweets from the 2020 Democratic presidential candidates.Using only the texts of their tweets, it identifies them along an interpretable progressive-tomoderate spectrum. Keyon Vafa, Suresh Naidu, David M. Blei |
ACL | 1 |
| 2019 | Discrete Flows: Invertible Generative Models of Discrete DataabstractWhile normalizing flows have led to significant advances in modeling high-dimensional continuous distributions, their applicability to discrete distributions remains unknown. In this paper, we show that flows can in fact be extended to discrete events---and under a simple change-of-variables formula not requiring log-determinant-Jacobian computations. Discrete flows have numerous applications. We consider two flow architectures: discrete autoregressive flows that enable bidirectionality, allowing, for example, tokens in text to depend on both left-to-right and right-to-left contexts in an exact language model; and discrete bipartite flows that enable efficient non-autoregressive generation as in RealNVP. Empirically, we find that discrete autoregressive flows outperform autoregressive baselines on synthetic discrete distributions, an addition task, and Potts models; and bipartite flows can obtain competitive performance with autoregressive baselines on character-level language modeling for Penn Tree Bank and text8. Dustin Tran, Keyon Vafa, Kumar Krishna Agrawal, Laurent Dinh, Ben Poole |
NeurIPS | 2 |