EDBT 2026 Demo / reviewers in the wild / expert
Tom Yan
dblp:213/7323
· DBLP profile ↗
12ranked-venue papers
8as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Reinforcement learning · 35% Probabilistic and Bayesian machine learning · 15% Trustworthy machine learning · 14% | |
| Theoretical computer science
6 papers |
Algorithmic game theory and mechanism design · 78% Mathematical optimization · 12% Computational complexity · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 29 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Algorithmic game theory and mechanism design
cooperative game theory |
0.9 | 2 | 2021 | If You Like Shapley Then You'll Love the Core · AAAI 2021 Evaluating and Rewarding Teamwork Using Cooperative Game Abstractions · NeurIPS 2020 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Stackelberg Learning with Outcome-based Payment · NeurIPS 2025 |
Algorithmic game theory and mechanism design
stackelberg game |
0.9 | 1 | 2025 | Stackelberg Learning with Outcome-based Payment · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
active learning |
0.8 | 1 | 2024 | The Human-AI Substitution game: active learning from a strategic labeler · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.8 | 1 | 2024 | Foundations of Testing for Finite-Sample Causal Discovery · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | Foundations of Testing for Finite-Sample Causal Discovery · ICML 2024 |
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
0.8 | 1 | 2024 | A theoretical case-study of Scalable Oversight in Hierarchical Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.8 | 1 | 2024 | A theoretical case-study of Scalable Oversight in Hierarchical Reinforcement Learning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
scalable oversight |
0.8 | 1 | 2024 | A theoretical case-study of Scalable Oversight in Hierarchical Reinforcement Learning · NeurIPS 2024 |
Mathematical optimization
multiple hypothesis testing |
0.8 | 1 | 2024 | Foundations of Testing for Finite-Sample Causal Discovery · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.6 | 1 | 2022 | Active fairness auditing · ICML 2022 |
Machine learning › Trustworthy machine learning › fairness
fairness auditing |
0.6 | 1 | 2022 | Active fairness auditing · ICML 2022 |
Computational complexity
query complexity |
0.6 | 1 | 2022 | Active fairness auditing · ICML 2022 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.5 | 1 | 2021 | Inverse Reinforcement Learning From Like-Minded Teachers · AAAI 2021 |
Machine learning › Reinforcement learning
markov decision process |
0.5 | 1 | 2021 | Inverse Reinforcement Learning From Like-Minded Teachers · AAAI 2021 |
Machine learning and data management
data valuation |
0.5 | 1 | 2021 | If You Like Shapley Then You'll Love the Core · AAAI 2021 |
Algorithmic game theory and mechanism design › cooperative game theory › solution concepts
core |
0.5 | 1 | 2021 | If You Like Shapley Then You'll Love the Core · AAAI 2021 |
Algorithmic game theory and mechanism design
revenue maximization |
0.5 | 1 | 2021 | Revenue maximization via machine learning with noisy data · NeurIPS 2021 |
Algorithmic game theory and mechanism design › cooperative game theory › solution concepts
shapley value |
0.5 | 1 | 2021 | If You Like Shapley Then You'll Love the Core · AAAI 2021 |
Algorithmic game theory and mechanism design › cooperative game theory
solution concepts |
0.5 | 1 | 2021 | If You Like Shapley Then You'll Love the Core · AAAI 2021 |
Computer vision › Video understanding and tracking
action recognition |
0.4 | 1 | 2020 | Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Natural language and speech › Information extraction and text analysis › event analysis
event understanding |
0.4 | 1 | 2020 | Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Algorithmic game theory and mechanism design › cooperative game theory › solution concepts › shapley value
shapley value approximation |
0.4 | 1 | 2020 | Evaluating and Rewarding Teamwork Using Cooperative Game Abstractions · NeurIPS 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2018 | Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANs · AAAI 2018 |
Machine learning › Learning paradigms
semi-supervised learning |
0.3 | 1 | 2018 | Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANs · AAAI 2018 |
Bioinformatics and computational biology
biomedical data analysis |
0.3 | 1 | 2018 | Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANs · AAAI 2018 |
Machine learning › Trustworthy machine learning
strategic behavior |
0.2 | 1 | 2024 | The Human-AI Substitution game: active learning from a strategic labeler · ICLR 2024 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.1 | 1 | 2021 | Inverse Reinforcement Learning From Like-Minded Teachers · AAAI 2021 |
Computer vision › Video understanding and tracking › action recognition
multimodal action recognition |
0.1 | 1 | 2020 | Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Methods — techniques the papers use, named apart from their topics
stackelberg learning · 1.7mechanism design · 1.7structured testing · 1.5query complexity analysis · 1.5multi-constraint bandits · 1.5deterministic algorithm design · 1.5query-based auditing · 1.1cooperative game theory · 1.0approximation algorithm · 1.0sub-MDP reward design · 0.8hierarchical experimental design · 0.8randomized algorithms · 0.6randomized algorithm · 0.6sub-gaussian noise · 0.5machine learning · 0.5cycle-consistency loss · 0.3cycle Wasserstein regression GAN · 0.3adversarial learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stackelberg Learning with Outcome-based PaymentabstractWith businesses starting to deploy agents to act on their behalf, an emerging challenge that businesses have to contend with is how to incentivize other agents with differing interests to work alongside its own agent. In present day commerce, payment is a common way that different parties use to \emph{economically} align their interests. In this paper, we study how one could analogously learn such payment schemes for aligning agents in the decentralized multi-agent setting. We model this problem as a Stackelberg Markov game, in which the leader can commit to a policy and also designate a set of outcome-based payments. We are interested in answering the question: when do efficient learning algorithms exist? To this end, we characterize the computational and statistical complexity of planning and learning in general-sum and cooperative games. In general-sum games, we find that planning is computationally intractable. In cooperative games, we show that learning can be statistically hard without payment and efficient with payment, showing that payment is necessary for learning even with aligned rewards. Altogether, our work aims to consolidate our theoretical understanding of outcome-based payment algorithms that can economically align decentralized agents. Tom Yan, Chicheng Zhang |
NeurIPS | 1 |
| 2024 | The Human-AI Substitution game: active learning from a strategic labelerabstractThe standard active learning setting assumes a willing labeler, who provides labels on informative examples to speed up learning. However, if the labeler wishes to be compensated for as many labels as possible before learning finishes, the labeler may benefit from actually slowing down learning. This incentive arises for instance if the labeler is to be replaced by the ML model once it is trained. In this paper, we initiate the study of learning from a strategic labeler, who may abstain from labeling to slow down learning. We first prove that strategic abstention can prolong learning, and propose a novel complexity measure and representation to analyze the query complexity of the learning game. Next, we develop a near-optimal deterministic algorithm, prove its robustness to strategic labeling, and contrast it with other active learning algorithms. We also analyze extensions that encompass more general learning goals and labeler assumptions. Finally, we characterize the query cost of multi-task active learning, with and without abstention. Our first exploration of strategic labeling aims to consolidate our theoretical understanding of the \emph{imitative} nature of ML in human-AI interaction. Tom Yan, Chicheng Zhang |
ICLR | 1 |
| 2024 | Foundations of Testing for Finite-Sample Causal DiscoveryabstractDiscovery of causal relationships is a fundamental goal of science and vital for sound decision making. As such, there has been considerable interest in causal discovery methods with provable guarantees. Existing works have thus far largely focused on discovery under hard intervention and infinite-samples, in which intervening on a node readily reveals the orientation of every edge incident to the node. This setup however overlooks the stochasticity inherent in real-world, finite-sample settings. Our work takes a step towards studying finite-sample causal discovery, wherein multiple interventions on a node are now needed for edge orientation. In this work, we study the canonical setup in theoretical causal discovery literature, where one assumes causal sufficiency and access to the graph skeleton. Our key observation is that discovery may be viewed as structured, multiple testing, and we develop a novel testing framework to this end. Crucially, our framework allows for anytime valid testing as multiple tests are needed to conclude an edge orientation. It also allows for flexible combination of structured test-statistics (enabling one to use Meek rules to propagate edge orientation) as well as robust testing. Through empirical simulations, we confirm the usefulness of our framework. In closing, using this testing framework, we show how one may efficiently verify graph structure by drawing a connection to multi-constraint bandits and designing a novel algorithm to this end. Tom Yan, Ziyu Xu 0001, Zachary C. Lipton |
ICML | 1 |
| 2024 | A theoretical case-study of Scalable Oversight in Hierarchical Reinforcement LearningabstractA key source of complexity in next-generation AI models is the size of model outputs, making it time-consuming to parse and provide reliable feedback on. To ensure such models are aligned, we will need to bolster our understanding of scalable oversight and how to scale up human feedback. To this end, we study the challenges of scalable oversight in the context of goal-conditioned hierarchical reinforcement learning. Hierarchical structure is a promising entrypoint into studying how to scale up human feedback, which in this work we assume can only be provided for model outputs below a threshold size. In the cardinal feedback setting, we develop an apt sub-MDP reward and algorithm that allows us to acquire and scale up low-level feedback for learning with sublinear regret. In the ordinal feedback setting, we show the necessity of both high- and low-level feedback, and develop a hierarchical experimental design algorithm that efficiently acquires both types of feedback for learning. Altogether, our work aims to consolidate the foundations of scalable oversight, formalizing and studying the various challenges thereof. Tom Yan, Zachary C. Lipton |
NeurIPS | 1 |
| 2022 | Margin-distancing for safe model explanationabstractThe growing use of machine learning models in consequential settings has highlighted an important and seemingly irreconcilable tension between transparency and vulnerability to gaming. While this has sparked sizable debate in legal literature, there has been comparatively less technical study of this contention. In this work, we propose a clean-cut formulation of this tension and a way to make the tradeoff between transparency and gaming. We identify the source of gaming as being points close to the decision boundary of the model. And we initiate an investigation on how to provide example-based explanations that are expansive and yet consistent with a version space that is sufficiently uncertain with respect to the boundary points’ labels. Finally, we furnish our theoretical results with empirical investigations of this tradeoff on real-world datasets. Tom Yan, Chicheng Zhang |
AISTATS | 1 |
| 2022 | Active fairness auditingabstractThe fast spreading adoption of machine learning (ML) by companies across industries poses significant regulatory challenges. One such challenge is scalability: how can regulatory bodies efficiently audit these ML models, ensuring that they are fair? In this paper, we initiate the study of query-based auditing algorithms that can estimate the demographic parity of ML models in a query-efficient manner. We propose an optimal deterministic algorithm, as well as a practical randomized, oracle-efficient algorithm with comparable guarantees. Furthermore, we make inroads into understanding the optimal query complexity of randomized active fairness estimation algorithms. Our first exploration of active fairness estimation aims to put AI governance on firmer theoretical foundations. Tom Yan, Chicheng Zhang |
ICML | 1 |
| 2021 | Inverse Reinforcement Learning From Like-Minded TeachersabstractWe study the problem of learning a policy in a Markov decision process (MDP) based on observations of the actions taken by multiple teachers. We assume that the teachers are like-minded in that their reward functions -- while different from each other -- are random perturbations of an underlying reward function. Under this assumption, we demonstrate that inverse reinforcement learning algorithms that satisfy a certain property -- that of matching feature expectations -- yield policies that are approximately optimal with respect to the underlying reward function, and that no algorithm can do better in the worst case. We also show how to efficiently recover the optimal policy when the MDP has one state -- a setting that is akin to multi-armed bandits. Ritesh Noothigattu, Tom Yan, Ariel D. Procaccia |
AAAI | 2 |
| 2021 | If You Like Shapley Then You'll Love the CoreabstractThe prevalent approach to problems of credit assignment in machine learning -- such as feature and data valuation -- is to model the problem at hand as a cooperative game and apply the Shapley value. But cooperative game theory offers a rich menu of alternative solution concepts, which famously includes the core and its variants. Our goal is to challenge the machine learning community's current consensus around the Shapley value, and make a case for the core as a viable alternative. To that end, we prove that arbitrarily good approximations to the least core -- a core relaxation that is always feasible -- can be computed efficiently (but prove an impossibility for a more refined solution concept, the nucleolus). We also perform experiments that corroborate these theoretical results and shed light on settings where the least core may be preferable to the Shapley value. Tom Yan, Ariel D. Procaccia |
AAAI | 1 |
| 2021 | Revenue maximization via machine learning with noisy dataabstractIncreasingly, copious amounts of consumer data are used to learn high-revenue mechanisms via machine learning. Existing research on mechanism design via machine learning assumes that there is a distribution over the buyers' values for the items for sale and that the learning algorithm's input is a training set sampled from this distribution. This setup makes the strong assumption that no noise is introduced during data collection. In order to help place mechanism design via machine learning on firm foundations, we investigate the extent to which this learning process is robust to noise. Optimizing revenue using noisy data is challenging because revenue functions are extremely volatile: an infinitesimal change in the buyers' values can cause a steep drop in revenue. Nonetheless, we provide guarantees when arbitrarily correlated noise is added to the training set; we only require that the noise has bounded magnitude or is sub-Gaussian. We conclude with an application of our guarantees to multi-task mechanism design, where there are multiple distributions over buyers' values and the goal is to learn a high-revenue mechanism per distribution. To our knowledge, we are the first to study mechanism design via machine learning with noisy data as well as multi-task mechanism design. Ellen Vitercik, Tom Yan |
NeurIPS | 2 |
| 2020 | Evaluating and Rewarding Teamwork Using Cooperative Game AbstractionsabstractCan we predict how well a team of individuals will perform together? How should individuals be rewarded for their contributions to the team performance? Cooperative game theory gives us a powerful set of tools for answering these questions: the characteristic function and solution concepts like the Shapley Value. There are two major difficulties in applying these techniques to real world problems: first, the characteristic function is rarely given to us and needs to be learned from data. Second, the Shapley Value is combinatorial in nature. We introduce a parametric model called cooperative game abstractions (CGAs) for estimating characteristic functions from data. CGAs are easy to learn, readily interpretable, and crucially allows linear-time computation of the Shapley Value. We provide identification results and sample complexity bounds for CGA models as well as error bounds in the estimation of the Shapley Value using CGAs. We apply our methods to study teams of artificial RL agents as well as real world teams from professional sports. Tom Yan, Christian Kroer, Alexander Peysakhovich |
NeurIPS | 1 |
| 2020 | Moments in Time Dataset: One Million Videos for Event UnderstandingabstractWe present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions occurring in 3 second videos poses many challenges: meaningful events do not include only people, but also objects, animals, and natural phenomena; visual and auditory events can be symmetrical in time ("opening" is "closing" in reverse), and either transient or sustained. We describe the annotation process of our dataset (each video is tagged with one action or activity label among 339 different classes), analyze its scale and diversity in comparison to other large-scale video datasets for action recognition, and report results of several baseline models addressing separately, and jointly, three modalities: spatial, temporal and auditory. The Moments in Time dataset, designed to have a large coverage and diversity of events in both visual and auditory modalities, can serve as a new challenge to develop models that scale to the level of complexity and abstract reasoning that a human processes on a daily basis. Mathew Monfort, Carl Vondrick, Aude Oliva, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa M. Brown, Quanfu Fan, Dan Gutfreund |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2018 | Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANsabstractThe biomedical field offers many learning tasks that share unique challenges: large amounts of unpaired data, and a high cost to generate labels. In this work, we develop a method to address these issues with semi-supervised learning in regression tasks (e.g., translation from source to target). Our model uses adversarial signals to learn from unpaired datapoints, and imposes a cycle-loss reconstruction error penalty to regularize mappings in either direction against one another. We first evaluate our method on synthetic experiments, demonstrating two primary advantages of the system: 1) distribution matching via the adversarial loss and 2) regularization towards invertible mappings via the cycle loss. We then show a regularization effect and improved performance when paired data is supplemented by additional unpaired data on two real biomedical regression tasks: estimating the physiological effect of medical treatments, and extrapolating gene expression (transcriptomics) signals. Our proposed technique is a promising initial step towards more robust use of adversarial signals in semi-supervised regression, and could be useful for other tasks (e.g., causal inference or modality translation) in the biomedical field. Matthew B. A. McDermott, Tom Yan, Tristan Naumann, Nathan Hunt, Harini Suresh, Peter Szolovits, Marzyeh Ghassemi |
AAAI | 2 |