VLDB 2026 Research / reviewers in the wild / expert
Taylor Lundy
dblp:243/2600
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
3 papers |
Algorithmic game theory and mechanism design · 76% Mathematical optimization · 24% | |
| Artificial intelligence
3 papers |
Language models and text generation · 76% Deep learning architectures and training · 24% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational social science and digital humanities · 61% Computational finance and economics · 39% | |
| Network and information security
1 paper |
Blockchain and cryptocurrency security · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models · NeurIPS 2025 |
Blockchain and cryptocurrency security › blockchain applications
non-fungible tokens |
0.9 | 1 | 2025 | NFTs as a Data-Rich Test Bed: Conspicuous Consumption and its Determinants · WWW 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | STEER: Assessing the Economic Rationality of Large Language Models · ICML 2024 |
Algorithmic game theory and mechanism design
pricing |
0.8 | 1 | 2024 | Pay to (Not) Play: Monetizing Impatience in Mobile Games · AAAI 2024 |
Algorithmic game theory and mechanism design
revenue maximization |
0.8 | 1 | 2024 | Pay to (Not) Play: Monetizing Impatience in Mobile Games · AAAI 2024 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end learning |
0.6 | 1 | 2022 | The Perils of Learning Before Optimizing · AAAI 2022 |
Mathematical optimization › optimization for machine learning
predict-then-optimize |
0.6 | 1 | 2022 | The Perils of Learning Before Optimizing · AAAI 2022 |
Algorithmic game theory and mechanism design › mechanism design
incentive compatibility |
0.4 | 1 | 2020 | Limitations of Incentive Compatibility on Discrete Type Spaces · AAAI 2020 |
Algorithmic game theory and mechanism design
mechanism design |
0.4 | 1 | 2020 | Limitations of Incentive Compatibility on Discrete Type Spaces · AAAI 2020 |
Natural language and speech › Language models and text generation
LLM agents |
0.2 | 1 | 2024 | STEER: Assessing the Economic Rationality of Large Language Models · ICML 2024 |
Mathematical optimization
stochastic optimization |
0.2 | 1 | 2022 | The Perils of Learning Before Optimizing · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
vision transformer embeddings · 1.7template adaptation · 1.7econometric modeling · 1.7LLM-assisted data generation · 1.7utility modeling · 1.5optimization · 1.5price of correlation · 1.1end-to-end differentiation · 1.1chain-of-thought prompting · 0.8benchmark distribution · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multidimensional Bayesian Utility Maximization: Tight Approximations to WelfareabstractWe initiate the study of multidimensional Bayesian utility maximization, focusing on the unit-demand setting where values are i.i.d. across both items and buyers. The seminal result of Hartline and Roughgarden '08 studies simple, information-robust mechanisms that maximize utility for $n$ i.i.d. agents and $m$ identical items via an approximation to social welfare as an upper bound, and they prove this gap between optimal utility and social welfare is $\Theta(1+\log{n/m})$ in this setting. We extend these results to the multidimensional setting. To do so, we develop simple, prior-independent, approximately-optimal mechanisms, targeting the simplest benchmark of optimal welfare. We give a $(1-1/e)$-approximation when there are more items than buyers, and a $\Theta(\log{n/m})$-approximation when there are more buyers than items, and we prove that this bound is tight in both $n$ and $m$ by reducing the i.i.d. unit-demand setting to the identical items setting. Finally, we include an extensive discussion section on why Bayesian utility maximization is a promising research direction. In particular, we characterize complexities in this setting that defy our intuition from the welfare and revenue literature, and motivate why coming up with a better benchmark than welfare is a hard problem itself. Kira Goldner, Taylor Lundy |
NeurIPS | 2 |
| 2025 | STEER-ME: Assessing the Microeconomic Reasoning of Large Language ModelsabstractLarge language models (LLMs) are increasingly being asked to make economically rational decisions and indeed are already being applied to economic tasks like stock picking and financial analysis. Existing LLM benchmarks tend to focus on specific applications, making them insufficient for characterizing economic reasoning more broadly. In previous work, we offered a blueprint for comprehensively benchmarking $\textit{strategic}$ decision-making Raman et al. 2024. However, this work did not engage with the even larger microeconomic literature on $\textit{non-strategic}$ settings. We address this gap here, taxonomizing microeconomic reasoning into $58$ distinct elements, each grounded in up to $10$ distinct domains, $5$ perspectives, and $3$ types. The generation of benchmark data across this combinatorial space is powered by a novel LLM-assisted data generation protocol that we dub auto-STEER, which generates a set of questions by adapting handwritten templates to target new domains and perspectives. By generating fresh questions for each element, auto-STEER induces diversity which could help to reduce the risk of data contamination. We use this benchmark to evaluate $27$ LLMs spanning a range of scales and adaptation strategies, comparing performance across multiple formats—multiple-choice and free-text question answering—and scoring schemes. Our results surface systematic limitations in current LLMs' ability to generalize economic reasoning across types, formats, and textual perturbations, and establish a foundation for evaluating and improving economic competence in foundation models. Narun K. Raman, Taylor Lundy, Thiago Amin, Kevin Leyton-Brown, Jesse Perla |
NeurIPS | 2 |
| 2025 | NFTs as a Data-Rich Test Bed: Conspicuous Consumption and its DeterminantsabstractConspicuous consumption occurs when a consumer derives value from a good based on its social meaning as a signal of wealth, taste, and/or community affiliation. Common conspicuous goods include designer footwear, country club memberships, and artwork; conspicuous goods also exist in the digital sphere, with non-fungible tokens (NFTs) as a prominent example. The NFT market merits deeper study for two key reasons: first, it is poorly understood relative to its economic scale; and second, it is unusually amenable to analysis because NFT transactions are publicly available on the blockchain, making them useful as a test bed for conspicuous consumption dynamics. This paper introduces a model that incorporates two previously identified elements of conspicuous consumption: the bandwagon effect (goods increase in value as they become more popular) and the snob effect (goods increase in value as they become rarer). Our model resolves the apparent tension between these two effects, exhibiting net complementarity between others' and one's own conspicuous consumption. We also introduce a novel dataset combining NFT transactions with embeddings of the corresponding NFT images computed using an off-the-shelf vision transformer architecture. We use our dataset to validate the model, showing that the bandwagon effect raises an NFT collection's value as more consumers join, while the snob effect drives consumers to seek rarer NFTs within a given collection. Taylor Lundy, Narun K. Raman, Scott Duke Kominers, Kevin Leyton-Brown |
WWW | 1 |
| 2024 | Pay to (Not) Play: Monetizing Impatience in Mobile GamesabstractMobile gaming is a rapidly growing and incredibly profitable sector; having grown seven-fold over the past 10 years, it now grosses over $100 billion annually. This growth was due in large part to a shift in monetization strategies: rather than charging players an upfront cost ("pay-to-play"), games often request optional microtransactions throughout gameplay ("free-to-play"). We focus on a common scenario in which games include wait times---gating either items or game progression---that players can pay to skip. Game designers typically say that they optimize for player happiness rather than revenue; however, prices for skips are typically set at levels that few players are willing to pay, leading to low purchase rates. Under a traditional analysis, it would seem that game designers fail at their stated goal if few players buy what they are selling. We argue that an alternate model can better explain this dynamic: players value tasks more highly as they are perceived to be more difficult. While skips can increase players' utilities by providing instant gratification, pricing skips too cheaply can lower players' utilities by decreasing the perceived amount of work needed to complete a task. We show that high revenue, high player utility, and low purchase rates can all coexist under this model, particularly under a realistic distribution of players having few buyers but a few big-spending "whales." We also investigate how a game designer should optimize prices under our model. An appendix of the paper with proofs, more comprehensive results and visualizations can be found at https://arxiv.org/abs/2312.10205. Taylor Lundy, Narun K. Raman, Hu Fu 0001, Kevin Leyton-Brown |
AAAI | 1 |
| 2024 | UNSAT Solver Synthesis via Monte Carlo Forest Search
Chris Cameron, Jason S. Hartford, Taylor Lundy, Tuan Truong, Alan Milligan, Rex Chen, Kevin Leyton-Brown |
CPAIOR (1) | 3 |
| 2024 | STEER: Assessing the Economic Rationality of Large Language ModelsabstractThere is increasing interest in using LLMs as decision-making "agents". Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions---and more broadly, determining whether an LLM agent is reliable enough to be trusted---requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "rationality report card". Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior. Narun K. Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz |
ICML | 2 |
| 2022 | The Perils of Learning Before OptimizingabstractFormulating real-world optimization problems often begins with making predictions from historical data (e.g., an optimizer that aims to recommend fast routes relies upon travel-time predictions). Typically, learning the prediction model used to generate the optimization problem and solving that problem are performed in two separate stages. Recent work has showed how such prediction models can be learned end-to-end by differentiating through the optimization task. Such methods often yield empirical improvements, which are typically attributed to end-to-end making better error tradeoffs than the standard loss function used in a two-stage solution. We refine this explanation and more precisely characterize when end-to-end can improve performance. When prediction targets are stochastic, a two-stage solution must make an a priori choice about which statistics of the target distribution to model---we consider expectations over prediction targets---while an end-to-end solution can make this choice adaptively. We show that the performance gap between a two-stage and end-to-end approach is closely related to the \emph{price of correlation} concept in stochastic optimization and show the implications of some existing POC results for the predict-then-optimize problem. We then consider a novel and particularly practical setting, where multiple prediction targets are combined to obtain each of the objective function’s coefficients. We give explicit constructions where (1) two-stage performs unboundedly worse than end-to-end; and (2) two-stage is optimal. We use simulations to experimentally quantify performance gaps and identify a wide range of real-world applications from the literature whose objective functions rely on multiple prediction targets, suggesting that end-to-end learning could yield significant improvements. Chris Cameron, Jason S. Hartford, Taylor Lundy, Kevin Leyton-Brown |
AAAI | 3 |
| 2020 | Limitations of Incentive Compatibility on Discrete Type Spaces
Taylor Lundy, Hu Fu 0001 |
AAAI | 1 |