EDBT 2026 Demo / reviewers in the wild / expert
Dung Daniel T. Ngo
dblp:296/8379
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 72% Reinforcement learning · 28% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.3 | 3 | 2025 | Discretization-free Multicalibration through Loss Minimization over Tree Ensembles · NeurIPS 2025 Reconciling Model Multiplicity for Downstream Decision Making · ICLR 2025 Strategic Instrumental Variable Regression: Recovering Causal Relationships From Strategic Responses · ICML 2022 |
Machine learning › Trustworthy machine learning
calibration |
1.7 | 2 | 2025 | Discretization-free Multicalibration through Loss Minimization over Tree Ensembles · NeurIPS 2025 Reconciling Model Multiplicity for Downstream Decision Making · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness › fairness criteria
multicalibration |
1.7 | 2 | 2025 | Discretization-free Multicalibration through Loss Minimization over Tree Ensembles · NeurIPS 2025 Reconciling Model Multiplicity for Downstream Decision Making · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
1.1 | 2 | 2022 | Incentivizing Combinatorial Bandit Exploration · NeurIPS 2022 Improved Regret for Differentially Private Exploration in Linear MDP · ICML 2022 |
Machine learning › Trustworthy machine learning
model multiplicity |
0.9 | 1 | 2025 | Reconciling Model Multiplicity for Downstream Decision Making · ICLR 2025 |
Machine learning › Reinforcement learning › multi-armed bandit
incentivized exploration |
0.6 | 1 | 2022 | Incentivizing Combinatorial Bandit Exploration · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
strategic behavior |
0.6 | 1 | 2022 | Strategic Instrumental Variable Regression: Recovering Causal Relationships From Strategic Responses · ICML 2022 |
Algorithmic game theory and mechanism design › mechanism design
incentive compatibility |
0.6 | 1 | 2022 | Incentivizing Combinatorial Bandit Exploration · NeurIPS 2022 |
Algorithmic game theory and mechanism design › mechanism design
dynamic mechanism design |
0.5 | 1 | 2021 | Incentivizing Compliance with Algorithmic Instruments · ICML 2021 |
Machine learning › Reinforcement learning
bandit |
0.2 | 1 | 2022 | Incentivizing Combinatorial Bandit Exploration · NeurIPS 2022 |
Machine learning › Reinforcement learning › bandit
combinatorial semi-bandits |
0.2 | 1 | 2022 | Incentivizing Combinatorial Bandit Exploration · NeurIPS 2022 |
Privacy and data protection
differential privacy |
0.2 | 1 | 2022 | Improved Regret for Differentially Private Exploration in Linear MDP · ICML 2022 |
Machine learning › Reinforcement learning
regret minimization |
0.1 | 1 | 2021 | Incentivizing Compliance with Algorithmic Instruments · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
thompson sampling · 1.1linear MDP · 1.1adaptive policy update schedule · 1.1randomized recommendation mechanism · 1.0tree ensemble learning · 0.9multicalibration · 0.9empirical risk minimization · 0.9calibration algorithm · 0.9instrumental variable regression · 0.6instrumental variables · 0.5instrumental variable · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reconciling Model Multiplicity for Downstream Decision MakingabstractWe consider the problem of model multiplicity in downstream decision-making, a setting where two predictive models of equivalent accuracy cannot agree on what action to take for a downstream decision-making problem. Prior work attempts to address model multiplicity by resolving prediction disagreement between models. However, we show that even when the two predictive models approximately agree on their individual predictions almost everywhere, these models can lead the downstream decision-maker to take actions with substantially higher losses. We address this issue by proposing a framework that calibrates the predictive models with respect to both a finite set of downstream decision-making problems and the individual probability prediction. Specifically, leveraging tools from multi-calibration, we provide an algorithm that, at each time-step, first reconciles the differences in individual probability prediction, then calibrates the updated models such that they are indistinguishable from the true probability distribution to the decision-makers. We extend our results to the setting where one does not have direct access to the true probability distribution and instead relies on a set of i.i.d data to be the empirical distribution. Furthermore, we generalize our results to the settings where one has more than two predictive models and an infinitely large downstream action set. Finally, we provide a set of experiments to evaluate our methods empirically. Compared to existing work, our proposed algorithm creates a pair of predictive models with improved downstream decision-making losses and agrees on their best-response actions almost everywhere. Ally Yalei Du, Dung Daniel T. Ngo, Steven Z. Wu |
ICLR | 2 |
| 2025 | Discretization-free Multicalibration through Loss Minimization over Tree EnsemblesabstractIn recent years, multicalibration has emerged as a desirable learning objective for ensuring that a predictor is calibrated across a rich collection of overlapping subpopulations. Existing approaches typically achieve multicalibration by discretizing the predictor's output space and iteratively adjusting its output values. However, this discretization approach departs from the standard empirical risk minimization (ERM) pipeline, introduces rounding error and an additional sensitive hyperparameter, and may distort the predictor’s outputs in ways that hinder downstream decision-making.
In this work, we propose a discretization-free multicalibration method that directly optimizes an empirical risk objective over an ensemble of depth-two decision trees. Our ERM approach can be implemented using off-the-shelf tree ensemble learning methods such as LightGBM. Our algorithm provably achieves multicalibration, provided that the data distribution satisfies a technical condition we term as loss saturation. Across multiple datasets, our empirical evaluation shows that this condition is always met in practice. Our discretization-free algorithm consistently matches or outperforms existing multicalibration approaches—even when evaluated using a discretization-based multicalibration metric that shares its discretization granularity with the baselines. Code to replicate the results in this work is available at https://github.com/hjenryin/Discretization-free-MC. Hongyi Henry Jin, Zijun Ding, Dung Daniel T. Ngo, Steven Z. Wu |
NeurIPS | 3 |
| 2022 | Strategic Instrumental Variable Regression: Recovering Causal Relationships From Strategic ResponsesabstractIn settings where Machine Learning (ML) algorithms automate or inform consequential decisions about people, individual decision subjects are often incentivized to strategically modify their observable attributes to receive more favorable predictions. As a result, the distribution the assessment rule is trained on may differ from the one it operates on in deployment. While such distribution shifts, in general, can hinder accurate predictions, our work identifies a unique opportunity associated with shifts due to strategic responses: We show that we can use strategic responses effectively to recover causal relationships between the observable features and outcomes we wish to predict, even under the presence of unobserved confounding variables. Specifically, our work establishes a novel connection between strategic responses to ML models and instrumental variable (IV) regression by observing that the sequence of deployed models can be viewed as an instrument that affects agents’ observable features but does not directly influence their outcomes. We show that our causal recovery method can be utilized to improve decision-making across several important criteria: individual fairness, agent outcomes, and predictive risk. In particular, we show that if decision subjects differ in their ability to modify non-causal attributes, any decision rule deviating from the causal coefficients can lead to (potentially unbounded) individual-level unfairness. . Keegan Harris, Dung Daniel T. Ngo, Logan Stapleton, Hoda Heidari, Steven Z. Wu |
ICML | 2 |
| 2022 | Improved Regret for Differentially Private Exploration in Linear MDPabstractWe study privacy-preserving exploration in sequential decision-making for environments that rely on sensitive data such as medical records. In particular, we focus on solving the problem of reinforcement learning (RL) subject to the constraint of (joint) differential privacy in the linear MDP setting, where both dynamics and rewards are given by linear functions. Prior work on this problem due to (Luyo et al., 2021) achieves a regret rate that has a dependence of O(K^{3/5}) on the number of episodes K. We provide a private algorithm with an improved regret rate with an optimal dependence of O($\sqrt{}$K) on the number of episodes. The key recipe for our stronger regret guarantee is the adaptivity in the policy update schedule, in which an update only occurs when sufficient changes in the data are detected. As a result, our algorithm benefits from low switching cost and only performs O(log(K)) updates, which greatly reduces the amount of privacy noise. Finally, in the most prevalent privacy regimes where the privacy parameter \epsilon is a constant, our algorithm incurs negligible privacy cost{—}in comparison with the existing non-private regret bounds, the additional regret due to privacy appears in lower-order terms. Dung Daniel T. Ngo, Giuseppe Vietri, Steven Z. Wu |
ICML | 1 |
| 2022 | Incentivizing Combinatorial Bandit ExplorationabstractConsider a bandit algorithm that recommends actions to self-interested users in a recommendation system. The users are free to choose other actions and need to be incentivized to follow the algorithm's recommendations. While the users prefer to exploit, the algorithm can incentivize them to explore by leveraging the information collected from the previous users. All published work on this problem, known as incentivized exploration, focuses on small, unstructured action sets and mainly targets the case when the users' beliefs are independent across actions. However, realistic exploration problems often feature large, structured action sets and highly correlated beliefs. We focus on a paradigmatic exploration problem with structure: combinatorial semi-bandits. We prove that Thompson Sampling, when applied to combinatorial semi-bandits, is incentive-compatible when initialized with a sufficient number of samples of each arm (where this number is determined in advance by the Bayesian prior). Moreover, we design incentive-compatible algorithms for collecting the initial samples. Xinyan Hu, Dung Daniel T. Ngo, Aleksandrs Slivkins, Steven Z. Wu |
NeurIPS | 2 |
| 2021 | Incentivizing Compliance with Algorithmic InstrumentsabstractRandomized experiments can be susceptible to selection bias due to potential non-compliance by the participants. While much of the existing work has studied compliance as a static behavior, we propose a game-theoretic model to study compliance as dynamic behavior that may change over time. In rounds, a social planner interacts with a sequence of heterogeneous agents who arrive with their unobserved private type that determines both their prior preferences across the actions (e.g., control and treatment) and their baseline rewards without taking any treatment. The planner provides each agent with a randomized recommendation that may alter their beliefs and their action selection. We develop a novel recommendation mechanism that views the planner’s recommendation as a form of instrumental variable (IV) that only affects an agents’ action selection, but not the observed rewards. We construct such IVs by carefully mapping the history –the interactions between the planner and the previous agents– to a random recommendation. Even though the initial agents may be completely non-compliant, our mechanism can incentivize compliance over time, thereby enabling the estimation of the treatment effect of each treatment, and minimizing the cumulative regret of the planner whose goal is to identify the optimal treatment. Dung Daniel T. Ngo, Logan Stapleton, Vasilis Syrgkanis, Steven Z. Wu |
ICML | 1 |