Tyler Lu

dblp:77/2582 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0002-7433-8421ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 11 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorTheory of computation · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 52% Probabilistic and Bayesian machine learning · 15% Multi-agent systems · 12%
Theoretical computer science
13 papers
Algorithmic game theory and mechanism design · 85% Mathematical optimization · 11% Algorithms and data structures · 3%
Databases, data mining, and information retrieval
4 papers
Recommender systems · 53% Information retrieval · 41% Web and social media mining · 6%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% Energy-efficient computing · 50%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
social choice
1.982025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Optimal social choice functions: A utilitarian view · Artif. Intell. 2015
Multi-Winner Social Choice with Incomplete Preferences · IJCAI 2013
Information retrieval › ranking › text ranking
comment ranking
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Information retrieval
ranking
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Algorithmic game theory and mechanism design › social choice › proportional representation
justified representation
0.912025
Representative Ranking for Deliberation in the Public Sphere · ICML 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.822020
ConQUR: Mitigating Delusional Bias in Deep Q-Learning · ICML 2020
Non-delusional Q-learning and value-iteration · NeurIPS 2018
Machine learning › Reinforcement learning
value-based reinforcement learning
0.822020
ConQUR: Mitigating Delusional Bias in Deep Q-Learning · ICML 2020
Non-delusional Q-learning and value-iteration · NeurIPS 2018
Recommender systems › interactive recommendation › conversational recommendation
critiquing-based recommendation
0.612022
Discovering Personalized Semantics for Soft Attributes in Recommender Systems using Concept Activation Vectors · WWW 2022
Recommender systems
interactive recommendation
0.612022
Discovering Personalized Semantics for Soft Attributes in Recommender Systems using Concept Activation Vectors · WWW 2022
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.412020
Gradient-Based Optimization for Bayesian Preference Elicitation · AAAI 2020
Knowledge, reasoning and agents › Multi-agent systems › social choice
computational social choice
0.412020
Preference elicitation and robust winner determination for single- and multi-winner social choice · Artif. Intell. 2020
Machine learning › Reinforcement learning › deep reinforcement learning
deep q-learning
0.412020
ConQUR: Mitigating Delusional Bias in Deep Q-Learning · ICML 2020
Knowledge, reasoning and agents › Multi-agent systems › social choice › computational social choice
preference elicitation
0.412020
Preference elicitation and robust winner determination for single- and multi-winner social choice · Artif. Intell. 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information
0.412020
Gradient-Based Optimization for Bayesian Preference Elicitation · AAAI 2020
Recommender systems
preference elicitation
0.412020
Gradient-Based Optimization for Bayesian Preference Elicitation · AAAI 2020
Machine learning › Reinforcement learning
dynamic programming
0.312018
Non-delusional Q-learning and value-iteration · NeurIPS 2018
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312018
Data center cooling using model-predictive control · NeurIPS 2018
Robotics › Motion planning and robot control › robot control
model predictive control
0.312018
Data center cooling using model-predictive control · NeurIPS 2018
Machine learning › Reinforcement learning › dynamic programming
value iteration
0.312018
Non-delusional Q-learning and value-iteration · NeurIPS 2018
Energy-efficient computing › thermal management
datacenter cooling
0.312018
Data center cooling using model-predictive control · NeurIPS 2018
Cloud and datacenter computing › resource management
datacenter resource management
0.312018
Data center cooling using model-predictive control · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › structured prediction › ranking model
mallows model
0.322014
Effective sampling and learning for mallows models with pairwise-preference data · J. Mach. Learn. Res. 2014
Learning Mallows Models with Pairwise Preferences · ICML 2011
Machine learning › Reinforcement learning
preference learning
0.322014
Effective sampling and learning for mallows models with pairwise-preference data · J. Mach. Learn. Res. 2014
Learning Mallows Models with Pairwise Preferences · ICML 2011
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
dynamic bayesian network
0.312017
Logistic Markov Decision Processes · IJCAI 2017
Machine learning › Reinforcement learning
markov decision process
0.312017
Logistic Markov Decision Processes · IJCAI 2017
Algorithmic game theory and mechanism design › pricing
price competition
0.222014
On the Value of Using Group Discounts under Price Competition · AAAI 2013
On the value of using group discounts under price competition · Artif. Intell. 2014
Mathematical optimization › combinatorial optimization
assignment problem
0.212015
Value-Directed Compression of Large-Scale Assignment Problems · AAAI 2015
Mathematical optimization
integer programming
0.212015
Value-Directed Compression of Large-Scale Assignment Problems · AAAI 2015
Mathematical optimization
linear programming
0.212015
Value-Directed Compression of Large-Scale Assignment Problems · AAAI 2015
Robotics › Robot manipulation › robot design
mechanism design
0.212014
On the value of using group discounts under price competition · Artif. Intell. 2014
Algorithms and data structures › randomized algorithms
sampling
0.212014
Effective sampling and learning for mallows models with pairwise-preference data · J. Mach. Learn. Res. 2014

Methods — techniques the papers use, named apart from their topics

civility classification · 1.7concept activation vectors · 1.1monte carlo estimation · 0.9gradient methods · 0.9reinforcement learning · 0.7model predictive control · 0.7search framework · 0.4q-approximator training · 0.4penalization scheme · 0.4local backup · 0.3information sets · 0.3logistic regression · 0.3constraint generation · 0.3approximate linear programming · 0.3greedy algorithm · 0.2value-directed compression · 0.2distributed data processing · 0.2column generation · 0.2
YearPublicationVenuePosition
2025 Representative Ranking for Deliberation in the Public Sphere
abstract
Online comment sections, such as those on news sites or social media, have the potential to foster informal public deliberation, However, this potential is often undermined by the frequency of toxic or low-quality exchanges that occur in these settings. To combat this, platforms increasingly leverage algorithmic ranking to facilitate higher-quality discussions, e.g., by using civility classifiers or forms of prosocial ranking. Yet, these interventions may also inadvertently reduce the visibility of legitimate viewpoints, undermining another key aspect of deliberation: representation of diverse views. We seek to remedy this problem by introducing guarantees of representation into these methods. In particular, we adopt the notion of *justified representation* (JR) from the social choice literature and incorporate a JR constraint into the comment ranking setting. We find that enforcing JR leads to greater inclusion of diverse viewpoints while still being compatible with optimizing for user engagement or other measures of conversational quality.
Manon Revel, Smitha Milli, Tyler Lu, Jamelle Watson-Daniels, Maximilian Nickel
ICML3
2024 Discovering Personalized Semantics for Soft Attributes in Recommender Systems Using Concept Activation Vectors
abstract
Interactive recommender systems have emerged as a promising paradigm to overcome the limitations of the primitive user feedback used by traditional recommender systems (e.g., clicks, item consumption, ratings). They allow users to express intent, preferences, constraints, and contexts in a richer fashion, often using natural language (including faceted search and dialogue). Yet more research is needed to find the most effective ways to use this feedback. One challenge is inferring a user’s semantic intent from the open-ended terms or attributes often used to describe a desired item. This is critical for recommender systems that wish to support users in their everyday, intuitive use of natural language to refine recommendation results. Leveraging concept activation vectors (CAVs) [ 26 ], a recently developed approach for model interpretability in machine learning, we develop a framework to learn a representation that captures the semantics of such attributes and connects them to user preferences and behaviors in recommender systems. One novel feature of our approach is its ability to distinguish objective and subjective attributes (both subjectivity of degree and of sense ) and associate different senses of subjective attributes with different users. We demonstrate on both synthetic and real-world datasets that our CAV representation not only accurately interprets users’ subjective semantics but also can be used to improve recommendations through interactive item critiquing .
Christina Göpfert, Alex Haig, Yinlam Chow, Ivan Vendrov, Tyler Lu, Deepak Ramachandran, Hubert Pham, Mohammad Ghavamzadeh, Craig Boutilier
Trans. Recomm. Syst.6
2022 Discovering Personalized Semantics for Soft Attributes in Recommender Systems using Concept Activation Vectors
abstract
Interactive recommender systems (RSs) allow users to express intent, preferences and contexts in a rich fashion, often using natural language. One challenge in using such feedback is inferring a user’s semantic intent from the open-ended terms used to describe an item, and using it to refine recommendation results. Leveraging concept activation vectors (CAVs) [21], we develop a framework to learn a representation that captures the semantics of such attributes and connects them to user preferences and behaviors in RSs. A novel feature of our approach is its ability to distinguish objective and subjective attributes and associate different senses with different users. Using synthetic and real-world datasets, we show that our CAV representation accurately interprets users’ subjective semantics, and can improve recommendations via interactive critiquing.
Christina Göpfert, Yinlam Chow, Ivan Vendrov, Tyler Lu, Deepak Ramachandran, Craig Boutilier
WWW5
2020 Gradient-Based Optimization for Bayesian Preference Elicitation
abstract
Effective techniques for eliciting user preferences have taken on added importance as recommender systems (RSs) become increasingly interactive and conversational. A common and conceptually appealing Bayesian criterion for selecting queries is expected value of information (EVOI). Unfortunately, it is computationally prohibitive to construct queries with maximum EVOI in RSs with large item spaces. We tackle this issue by introducing a continuous formulation of EVOI as a differentiable network that can be optimized using gradient methods available in modern machine learning computational frameworks (e.g., TensorFlow, PyTorch). We exploit this to develop a novel Monte Carlo method for EVOI optimization, which is much more scalable for large item spaces than methods requiring explicit enumeration of items. While we emphasize the use of this approach for pairwise (or k-wise) comparisons of items, we also demonstrate how our method can be adapted to queries involving subsets of item attributes or “partial items,” which are often more cognitively manageable for users. Experiments show that our gradient-based EVOI technique achieves state-of-the-art performance across several domains while scaling to large item spaces.
Ivan Vendrov, Tyler Lu, Craig Boutilier
AAAI2
2020 ConQUR: Mitigating Delusional Bias in Deep Q-Learning
abstract
Delusional bias is a fundamental source of error in approximate Q-learning. To date, the only techniques that explicitly address delusion require comprehensive search using tabular value estimates. In this paper, we develop efficient methods to mitigate delusional bias by training Q-approximators with labels that are "consistent" with the underlying greedy policy class. We introduce a simple penalization scheme that encourages Q-labels used across training batches to remain (jointly) consistent with the expressible policy class. We also propose a search framework that allows multiple Q-approximators to be generated and tracked, thus mitigating the effect of premature (implicit) policy commitments. Experimental results demonstrate that these methods can improve the performance of Q-learning in a variety of Atari games, sometimes dramatically.
DiJia Su, Jayden Ooi, Tyler Lu, Dale Schuurmans, Craig Boutilier
ICML3
2020 Preference elicitation and robust winner determination for single- and multi-winner social choice
Tyler Lu, Craig Boutilier
Artif. Intell.1
2018 Data center cooling using model-predictive control
abstract
Despite impressive recent advances in reinforcement learning (RL), its deployment in real-world physical systems is often complicated by unexpected events, limited data, and the potential for expensive failures. In this paper, we describe an application of RL “in the wild” to the task of regulating temperatures and airflow inside a large-scale data center (DC). Adopting a data-driven, model-based approach, we demonstrate that an RL agent with little prior knowledge is able to effectively and safely regulate conditions on a server floor after just a few hours of exploration, while improving operational efficiency relative to existing PID controllers.
Nevena Lazic, Craig Boutilier, Tyler Lu, Eehern Wong, Binz Roy, M. K. Ryu, Greg Imwalle
NeurIPS3
2018 Non-delusional Q-learning and value-iteration
abstract
We identify a fundamental source of error in Q-learning and other forms of dynamic programming with function approximation. Delusional bias arises when the approximation architecture limits the class of expressible greedy policies. Since standard Q-updates make globally uncoordinated action choices with respect to the expressible policy class, inconsistent or even conflicting Q-value estimates can result, leading to pathological behaviour such as over/under-estimation, instability and even divergence. To solve this problem, we introduce a new notion of policy consistency and define a local backup process that ensures global consistency through the use of information sets---sets that record constraints on policies consistent with backed-up Q-values. We prove that both the model-based and model-free algorithms using this backup remove delusional bias, yielding the first known algorithms that guarantee optimal results under general conditions. These algorithms furthermore only require polynomially many information sets (from a potentially exponential support). Finally, we suggest other practical heuristics for value-iteration and Q-learning that attempt to reduce delusional bias.
Tyler Lu, Dale Schuurmans, Craig Boutilier
NeurIPS1
2017 Logistic Markov Decision Processes
abstract
User modeling in advertising and recommendation has typically focused on myopic predictors of user responses. In this work, we consider the long-term decision problem associated with user interaction. We propose a concise specification of long-term interaction dynamics by combining factored dynamic Bayesian networks with logistic predictors of user responses, allowing state-of-the-art prediction models to be seamlessly extended. We show how to solve such models at scale by providing a constraint generation approach for approximate linear programming that overcomes the variable coupling and non-linearity induced by the logistic regression predictor. The efficacy of the approach is demonstrated on advertising domains with up to 2^54 states and 2^39 actions.
Martin Mladenov, Craig Boutilier, Dale Schuurmans, Ofer Meshi, Gal Elidan, Tyler Lu
IJCAI6
2016 Budget Allocation using Weakly Coupled, Constrained Markov Decision Processes
Craig Boutilier, Tyler Lu
UAI2
2015 Value-Directed Compression of Large-Scale Assignment Problems
abstract
Data-driven analytics — in areas ranging from consumer marketing to public policy — often allow behavior prediction at the level of individuals rather than population segments, offering the opportunity to improve decisions that impact large populations. Modeling such (generalized) assignment problems as linear programs, we propose a general value-directed compression technique for solving such problems at scale. We dynamically segment the population into cells using a form of column generation, constructing groups of individuals who can provably be treated identically in the optimal solution. This compression allows problems, unsolvable using standard LP techniques, to be solved effectively. Indeed, once a compressed LP is constructed, problems can solved in milliseconds. We provide a theoretical analysis of themethods, outline the distributed implementation of the requisite data processing, and show how a single compressed LP can be used to solve multiple variants of the original LP near-optimally in real-time (e.g., tosupport scenario analysis). We also show how the method can be leveraged in integer programming models. Experimental results on marketing contact optimization and political legislature problems validate the performance of our technique.
Tyler Lu, Craig Boutilier
AAAI1
2015 Optimal social choice functions: A utilitarian view
Craig Boutilier, Ioannis Caragiannis, Simi Haber, Tyler Lu, Ariel D. Procaccia, Or Sheffet
Artif. Intell.4
2014 On the value of using group discounts under price competition
Reshef Meir, Tyler Lu, Moshe Tennenholtz, Craig Boutilier
Artif. Intell.2
2014 Effective sampling and learning for mallows models with pairwise-preference data
Tyler Lu, Craig Boutilier
J. Mach. Learn. Res.1
2013 On the Value of Using Group Discounts under Price Competition
abstract
The increasing use of group discounts has provided opportunities for buying groups with diverse preferences to coordinate their behavior in order to exploit the best offers from multiple vendors. We analyze this problem from the viewpoint of the vendors, asking under what conditions a vendor should adopt a volume-based price schedule rather than posting a fixed price, either as a monopolist or when competing with other vendors. When vendors have uncertainty about buyers' valuations specified by a known distribution, we show that a vendor is always better off posting a fixed price, provided that buyers' types are i.i.d. and that other vendors also use fixed prices. We also show that these assumptions cannot be relaxed: if buyers are not i.i.d., or other vendors post discount schedules, then posting a schedule may yield higher profit for the vendor. We provide similar results under a distribution-free uncertainty model, where vendors minimize their maximum regret over all type realizations.
Reshef Meir, Tyler Lu, Moshe Tennenholtz, Craig Boutilier
AAAI2
2013 Multi-Winner Social Choice with Incomplete Preferences
Tyler Lu, Craig Boutilier
IJCAI1
2012 Optimal social choice functions: a utilitarian view
abstract
We adopt a utilitarian perspective on social choice, assuming that agents have (possibly latent) utility functions over some space of alternatives. For many reasons one might consider mechanisms, or social choice functions, that only have access to the ordinal rankings of alternatives by the individual agents rather than their utility functions. In this context, one possible objective for a social choice function is the maximization of (expected) social welfare relative to the information contained in these rankings. We study such optimal social choice functions under three different models, and underscore the important role played by scoring functions. In our worst-case model, no assumptions are made about the underlying distribution and we analyze the worst-case distortion---or degree to which the selected alternative does not maximize social welfare---of optimal social choice functions. In our average-case model, we derive optimal functions under neutral (or impartial culture) distributional models. Finally, a very general learning-theoretic model allows for the computation of optimal social choice functions (i.e., that maximize expected social welfare) under arbitrary, sampleable distributions. In the latter case, we provide both algorithms and sample complexity results for the class of scoring functions, and further validate the approach empirically.
Craig Boutilier, Ioannis Caragiannis, Simi Haber, Tyler Lu, Ariel D. Procaccia, Or Sheffet
EC4
2012 Matching models for preference-sensitive group purchasing
abstract
Matching buyers and sellers is one of the most fundamental problems in economics and market design. An interesting variant of the matching problem arises when self-interested buyers come together in order to induce sellers to offer quantity or volume discounts, as is common in buying consortia, and more recently in the consumer group couponing space (e.g., Groupon). We consider a general model of this problem in which a group or buying consortium is faced with volume discount offers from multiple vendors, but group members have distinct preferences for different vendor offerings. Unlike some recent formulations of matching games that involve quantity discounts, the combination of varying preferences and discounts can render the core of the matching game empty, in both the transferable and nontransferable utility sense. Thus, instead of coalitional stability, we propose several forms of Nash stability under various knowledge and transfer/payment assumptions. We investigate the computation of buyer-welfare maximizing matchings and show the existence of transfers (subsidized prices) of a particularly desirable form that support stable matchings. We also study a nontransferable utility model, showing that stable matchings exist; we develop a further variant of the problem in which buyers provide a simple preference ordering over "deals" rather than specific valuations---a model that is especially attractive in the consumer space---and also show the existence of stable matchings. Finally, computational experiments with buyer-welfare maximization demonstrate the value of our approach.
Tyler Lu, Craig Boutilier
EC1
2012 Bayesian Vote Manipulation: Optimal Strategies and Impact on Welfare
Tyler Lu, Pingzhong Tang, Ariel D. Procaccia, Craig Boutilier
UAI1
2011 Learning Mallows Models with Pairwise Preferences
Tyler Lu, Craig Boutilier
ICML1
2011 Budgeted Social Choice: From Consensus to Personalized Decision Making
abstract
We develop a general framework for social choice problems in which a limited number of alternatives can be recommended to an agent population. In our budgeted social choice model, this limit is determined by a budget, capturing problems that arise naturally in a variety of contexts, and spanning the continuum from pure consensus decision making (i.e., standard social choice) to fully personalized recommendation. Our approach applies a form of segmentation to social choice problems— requiring the selection of diverse options tailored to different agent types—and generalizes certain multi-winner election schemes. We show that standard rank aggregation methods perform poorly, and that optimization in our model is NP-complete; but we develop fast greedy algorithms with some theoretical guarantees. Experiments on real-world datasets demonstrate the effectiveness of our algorithms.
Tyler Lu, Craig Boutilier
IJCAI1
2011 Robust Approximation and Incremental Elicitation in Voting Protocols
abstract
While voting schemes provide an effective means for aggregating preferences, methods for the effective elicitation of voter preferences have received little attention. We address this problem by first considering approximate winner determination when incomplete voter preferences are provided. Exploiting natural scoring metrics, we use max regret to measure the quality or robustness of proposed winners, and develop polynomial time algorithms for computing the alternative with minimax regret for several popular voting rules. We then show how minimax regret can be used to effectively drive incremental preference/vote elicitation and devise several heuristics for this process. Despite worst-case theoretical results showing that most voting protocols require nearly complete voter preferences to determine winners, we demonstrate the practical effectiveness of regret-based elicitation for determining both approximate and exact winners on several real-world data sets. 1
Tyler Lu, Craig Boutilier
IJCAI1
2010 The unavailable candidate model: a decision-theoretic view of social choice
abstract
One of the fundamental problems in the theory of social choice is aggregating the rankings of a set of agents (or voters) into a consensus ranking. Rank aggregation has found application in a variety of computational contexts. However, the goal of constructing a consensus ranking rather than, say, a single outcome (or winner) is often left unjustified, calling into question the suitability of classical rank aggregation methods. We introduce a novel model which offers a decision-theoretic motivation for constructing a consensus ranking. Our unavailable candidate model assumes that a consensus choice must be made, but that candidates may become unavailable after voters express their preferences. Roughly speaking, a consensus ranking serves as a compact, easily communicable representation of a decision policy that can be used to make choices in the face of uncertain candidate availability. We use this model to define a principled aggregation method that minimizes expected voter dissatisfaction with the chosen candidate. We give exact and approximation algorithms for computing optimal rankings and provide computational evidence for the effectiveness of a simple greedy scheme. We also describe strong connections to popular voting protocols such as the plurality rule and the Kemeny consensus, showing specifically that Kemeny produces optimal rankings in the unavailable candidate model under certain conditions.
Tyler Lu, Craig Boutilier
EC1
2008 Does Unlabeled Data Provably Help? Worst-case Analysis of the Sample Complexity of Semi-Supervised Learning
Shai Ben-David, Tyler Lu, Dávid Pál
COLT2