Taylor Lundy

dblp:243/2600 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 76% Mathematical optimization · 24%
Artificial intelligence
3 papers
Language models and text generation · 76% Deep learning architectures and training · 24%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational social science and digital humanities · 61% Computational finance and economics · 39%
Network and information security
1 paper
Blockchain and cryptocurrency security · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
0.912025
STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models · NeurIPS 2025
Blockchain and cryptocurrency security › blockchain applications
non-fungible tokens
0.912025
NFTs as a Data-Rich Test Bed: Conspicuous Consumption and its Determinants · WWW 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
STEER: Assessing the Economic Rationality of Large Language Models · ICML 2024
Algorithmic game theory and mechanism design
pricing
0.812024
Pay to (Not) Play: Monetizing Impatience in Mobile Games · AAAI 2024
Algorithmic game theory and mechanism design
revenue maximization
0.812024
Pay to (Not) Play: Monetizing Impatience in Mobile Games · AAAI 2024
Machine learning › Deep learning architectures and training › neural network training
end-to-end learning
0.612022
The Perils of Learning Before Optimizing · AAAI 2022
Mathematical optimization › optimization for machine learning
predict-then-optimize
0.612022
The Perils of Learning Before Optimizing · AAAI 2022
Algorithmic game theory and mechanism design › mechanism design
incentive compatibility
0.412020
Limitations of Incentive Compatibility on Discrete Type Spaces · AAAI 2020
Algorithmic game theory and mechanism design
mechanism design
0.412020
Limitations of Incentive Compatibility on Discrete Type Spaces · AAAI 2020
Natural language and speech › Language models and text generation
LLM agents
0.212024
STEER: Assessing the Economic Rationality of Large Language Models · ICML 2024
Mathematical optimization
stochastic optimization
0.212022
The Perils of Learning Before Optimizing · AAAI 2022

Methods — techniques the papers use, named apart from their topics

vision transformer embeddings · 1.7template adaptation · 1.7econometric modeling · 1.7LLM-assisted data generation · 1.7utility modeling · 1.5optimization · 1.5price of correlation · 1.1end-to-end differentiation · 1.1chain-of-thought prompting · 0.8benchmark distribution · 0.8
YearPublicationVenuePosition
2025 Multidimensional Bayesian Utility Maximization: Tight Approximations to Welfare
abstract
We initiate the study of multidimensional Bayesian utility maximization, focusing on the unit-demand setting where values are i.i.d. across both items and buyers. The seminal result of Hartline and Roughgarden '08 studies simple, information-robust mechanisms that maximize utility for $n$ i.i.d. agents and $m$ identical items via an approximation to social welfare as an upper bound, and they prove this gap between optimal utility and social welfare is $\Theta(1+\log{n/m})$ in this setting. We extend these results to the multidimensional setting. To do so, we develop simple, prior-independent, approximately-optimal mechanisms, targeting the simplest benchmark of optimal welfare. We give a $(1-1/e)$-approximation when there are more items than buyers, and a $\Theta(\log{n/m})$-approximation when there are more buyers than items, and we prove that this bound is tight in both $n$ and $m$ by reducing the i.i.d. unit-demand setting to the identical items setting. Finally, we include an extensive discussion section on why Bayesian utility maximization is a promising research direction. In particular, we characterize complexities in this setting that defy our intuition from the welfare and revenue literature, and motivate why coming up with a better benchmark than welfare is a hard problem itself.
Kira Goldner, Taylor Lundy
NeurIPS2
2025 STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
abstract
Large language models (LLMs) are increasingly being asked to make economically rational decisions and indeed are already being applied to economic tasks like stock picking and financial analysis. Existing LLM benchmarks tend to focus on specific applications, making them insufficient for characterizing economic reasoning more broadly. In previous work, we offered a blueprint for comprehensively benchmarking $\textit{strategic}$ decision-making Raman et al. 2024. However, this work did not engage with the even larger microeconomic literature on $\textit{non-strategic}$ settings. We address this gap here, taxonomizing microeconomic reasoning into $58$ distinct elements, each grounded in up to $10$ distinct domains, $5$ perspectives, and $3$ types. The generation of benchmark data across this combinatorial space is powered by a novel LLM-assisted data generation protocol that we dub auto-STEER, which generates a set of questions by adapting handwritten templates to target new domains and perspectives. By generating fresh questions for each element, auto-STEER induces diversity which could help to reduce the risk of data contamination. We use this benchmark to evaluate $27$ LLMs spanning a range of scales and adaptation strategies, comparing performance across multiple formats—multiple-choice and free-text question answering—and scoring schemes. Our results surface systematic limitations in current LLMs' ability to generalize economic reasoning across types, formats, and textual perturbations, and establish a foundation for evaluating and improving economic competence in foundation models.
Narun K. Raman, Taylor Lundy, Thiago Amin, Kevin Leyton-Brown, Jesse Perla
NeurIPS2
2025 NFTs as a Data-Rich Test Bed: Conspicuous Consumption and its Determinants
abstract
Conspicuous consumption occurs when a consumer derives value from a good based on its social meaning as a signal of wealth, taste, and/or community affiliation. Common conspicuous goods include designer footwear, country club memberships, and artwork; conspicuous goods also exist in the digital sphere, with non-fungible tokens (NFTs) as a prominent example. The NFT market merits deeper study for two key reasons: first, it is poorly understood relative to its economic scale; and second, it is unusually amenable to analysis because NFT transactions are publicly available on the blockchain, making them useful as a test bed for conspicuous consumption dynamics. This paper introduces a model that incorporates two previously identified elements of conspicuous consumption: the bandwagon effect (goods increase in value as they become more popular) and the snob effect (goods increase in value as they become rarer). Our model resolves the apparent tension between these two effects, exhibiting net complementarity between others' and one's own conspicuous consumption. We also introduce a novel dataset combining NFT transactions with embeddings of the corresponding NFT images computed using an off-the-shelf vision transformer architecture. We use our dataset to validate the model, showing that the bandwagon effect raises an NFT collection's value as more consumers join, while the snob effect drives consumers to seek rarer NFTs within a given collection.
Taylor Lundy, Narun K. Raman, Scott Duke Kominers, Kevin Leyton-Brown
WWW1
2024 Pay to (Not) Play: Monetizing Impatience in Mobile Games
abstract
Mobile gaming is a rapidly growing and incredibly profitable sector; having grown seven-fold over the past 10 years, it now grosses over $100 billion annually. This growth was due in large part to a shift in monetization strategies: rather than charging players an upfront cost ("pay-to-play"), games often request optional microtransactions throughout gameplay ("free-to-play"). We focus on a common scenario in which games include wait times---gating either items or game progression---that players can pay to skip. Game designers typically say that they optimize for player happiness rather than revenue; however, prices for skips are typically set at levels that few players are willing to pay, leading to low purchase rates. Under a traditional analysis, it would seem that game designers fail at their stated goal if few players buy what they are selling. We argue that an alternate model can better explain this dynamic: players value tasks more highly as they are perceived to be more difficult. While skips can increase players' utilities by providing instant gratification, pricing skips too cheaply can lower players' utilities by decreasing the perceived amount of work needed to complete a task. We show that high revenue, high player utility, and low purchase rates can all coexist under this model, particularly under a realistic distribution of players having few buyers but a few big-spending "whales." We also investigate how a game designer should optimize prices under our model. An appendix of the paper with proofs, more comprehensive results and visualizations can be found at https://arxiv.org/abs/2312.10205.
Taylor Lundy, Narun K. Raman, Hu Fu 0001, Kevin Leyton-Brown
AAAI1
2024 UNSAT Solver Synthesis via Monte Carlo Forest Search
Chris Cameron, Jason S. Hartford, Taylor Lundy, Tuan Truong, Alan Milligan, Rex Chen, Kevin Leyton-Brown
CPAIOR (1)3
2024 STEER: Assessing the Economic Rationality of Large Language Models
abstract
There is increasing interest in using LLMs as decision-making "agents". Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions---and more broadly, determining whether an LLM agent is reliable enough to be trusted---requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "rationality report card". Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.
Narun K. Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz
ICML2
2022 The Perils of Learning Before Optimizing
abstract
Formulating real-world optimization problems often begins with making predictions from historical data (e.g., an optimizer that aims to recommend fast routes relies upon travel-time predictions). Typically, learning the prediction model used to generate the optimization problem and solving that problem are performed in two separate stages. Recent work has showed how such prediction models can be learned end-to-end by differentiating through the optimization task. Such methods often yield empirical improvements, which are typically attributed to end-to-end making better error tradeoffs than the standard loss function used in a two-stage solution. We refine this explanation and more precisely characterize when end-to-end can improve performance. When prediction targets are stochastic, a two-stage solution must make an a priori choice about which statistics of the target distribution to model---we consider expectations over prediction targets---while an end-to-end solution can make this choice adaptively. We show that the performance gap between a two-stage and end-to-end approach is closely related to the \emph{price of correlation} concept in stochastic optimization and show the implications of some existing POC results for the predict-then-optimize problem. We then consider a novel and particularly practical setting, where multiple prediction targets are combined to obtain each of the objective function’s coefficients. We give explicit constructions where (1) two-stage performs unboundedly worse than end-to-end; and (2) two-stage is optimal. We use simulations to experimentally quantify performance gaps and identify a wide range of real-world applications from the literature whose objective functions rely on multiple prediction targets, suggesting that end-to-end learning could yield significant improvements.
Chris Cameron, Jason S. Hartford, Taylor Lundy, Kevin Leyton-Brown
AAAI3
2020 Limitations of Incentive Compatibility on Discrete Type Spaces
Taylor Lundy, Hu Fu 0001
AAAI1