VLDB 2026 Research / reviewers in the wild / expert
Xiaojie Mao
dblp:222/3283
· DBLP profile ↗
10ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Probabilistic and Bayesian machine learning · 46% Reinforcement learning · 24% Trustworthy machine learning · 14% | |
| Theoretical computer science
4 papers |
Mathematical optimization · 52% Algorithmic game theory and mechanism design · 32% Computational complexity · 16% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
3.3 | 5 | 2025 | Learning with Selectively Labeled Data from Multiple Decision-makers · ICML 2025 Localized Debiased Machine Learning: Efficient Inference on Quantile Treatment Effects and Beyond · J. Mach. Learn. Res. 2024 Minimax Instrumental Variable Regression and L2 Convergence Guarantees without Identification or Closedness · COLT 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable regression |
1.3 | 2 | 2023 | Minimax Instrumental Variable Regression and L2 Convergence Guarantees without Identification or Closedness · COLT 2023 Inference on Strongly Identified Functionals of Weakly Identified Functions · COLT 2023 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.9 | 2 | 2022 | Doubly Robust Distributionally Robust Off-Policy Evaluation and Learning · ICML 2022 Causal Inference with Noisy and Missing Covariates via Matrix Factorization · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable |
0.9 | 1 | 2025 | Learning with Selectively Labeled Data from Multiple Decision-makers · ICML 2025 |
Machine learning › Learning theory
online learning |
0.9 | 1 | 2025 | Online Strategic Classification With Noise and Partial Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning
regret minimization |
0.9 | 1 | 2025 | Online Strategic Classification With Noise and Partial Feedback · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Learning with Selectively Labeled Data from Multiple Decision-makers · ICML 2025 |
Machine learning › Trustworthy machine learning › dataset bias
selection bias |
0.9 | 1 | 2025 | Learning with Selectively Labeled Data from Multiple Decision-makers · ICML 2025 |
Algorithmic game theory and mechanism design › strategic behavior
strategic classification |
0.9 | 1 | 2025 | Online Strategic Classification With Noise and Partial Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning › bandit
bandit feedback |
0.8 | 1 | 2024 | Contextual Linear Optimization with Bandit Feedback · NeurIPS 2024 |
Natural language and speech › Language models and text generation › prompt tuning
context optimization |
0.8 | 1 | 2024 | Contextual Linear Optimization with Bandit Feedback · NeurIPS 2024 |
Mathematical optimization
continuous optimization |
0.8 | 1 | 2024 | Localized Debiased Machine Learning: Efficient Inference on Quantile Treatment Effects and Beyond · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory › statistical estimation
nonparametric estimation |
0.7 | 1 | 2023 | Minimax Instrumental Variable Regression and L2 Convergence Guarantees without Identification or Closedness · COLT 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
proximal causal inference |
0.7 | 1 | 2023 | Inference on Strongly Identified Functionals of Weakly Identified Functions · COLT 2023 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.6 | 1 | 2022 | Doubly Robust Distributionally Robust Off-Policy Evaluation and Learning · ICML 2022 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.4 | 1 | 2020 | Smooth Contextual Bandits: Bridging the Parametric and Non-differentiable Regret Regimes · COLT 2020 |
Computational complexity
learning theory |
0.4 | 1 | 2020 | Smooth Contextual Bandits: Bridging the Parametric and Non-differentiable Regret Regimes · COLT 2020 |
Mathematical optimization › online optimization
regret bounds |
0.4 | 1 | 2020 | Smooth Contextual Bandits: Bridging the Parametric and Non-differentiable Regret Regimes · COLT 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
deconfounding |
0.3 | 1 | 2018 | Causal Inference with Noisy and Missing Covariates via Matrix Factorization · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation |
0.3 | 1 | 2018 | Causal Inference with Noisy and Missing Covariates via Matrix Factorization · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.3 | 1 | 2025 | Online Strategic Classification With Noise and Partial Feedback · NeurIPS 2025 |
Mathematical optimization
stochastic optimization |
0.2 | 1 | 2024 | Contextual Linear Optimization with Bandit Feedback · NeurIPS 2024 |
Recommender systems › collaborative filtering
matrix factorization |
0.1 | 1 | 2018 | Causal Inference with Noisy and Missing Covariates via Matrix Factorization · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
massart noise model · 1.7surrogate loss · 1.5regret bounds · 1.5localized debiased machine learning · 1.5induced empirical risk minimization · 1.5penalized estimation · 1.3minimax estimation · 1.3partial identification · 0.9online learning algorithms · 0.9online learning algorithm · 0.9cost-sensitive learning · 0.9semiparametric efficiency · 0.8nuisance function estimation · 0.8nonparametric estimation · 0.4hölder class · 0.4matrix factorization · 0.3exponential family matrix completion · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning with Selectively Labeled Data from Multiple Decision-makersabstractWe study the problem of classification with selectively labeled data, whose distribution may differ from the full population due to historical decision-making. We exploit the fact that in many applications historical decisions were made by multiple decision-makers, each with different decision rules. We analyze this setup under a principled instrumental variable (IV) framework and rigorously study the identification of classification risk. We establish conditions for the exact identification of classification risk and derive tight partial identification bounds when exact identification fails. We further propose a unified cost-sensitive learning (UCL) approach to learn classifiers robust to selection bias in both identification settings. Finally, we theoretically and numerically validate the efficacy of our proposed method. Xiaojie Mao |
ICML | 3 |
| 2025 | Online Strategic Classification With Noise and Partial FeedbackabstractIn this paper, we study an online strategic classification problem, where a principal aims to learn an accurate binary linear classifier from sequentially arriving agents. For each agent, the principal announces a classifier. The agent can strategically exercise costly manipulations on his features to be classified as the favorable positive class. The principal is unaware of the true feature-label distribution, but observes all reported features and only labels of positively classified agents. We assume that the true feature-label distribution is given by a halfspace model subject to arbitrary feature-dependent bounded noise (i.e., Massart Noise). This problem faces the combined challenges of agents' strategic feature manipulations, partial label observations, and label noises. We tackle these challenges by a novel learning algorithm. We show that the proposed algorithm yields classifiers that converge to the clairvoyant optimal one and attains a regret rate of $ O(\sqrt{T})$ up to poly-logarithmic and constant factors over $T$ cycles. Tianrun Zhao, Xiaojie Mao |
NeurIPS | 2 |
| 2024 | Contextual Linear Optimization with Bandit FeedbackabstractContextual linear optimization (CLO) uses predictive contextual features to reduce uncertainty in random cost coefficients and thereby improve average-cost performance. An example is the stochastic shortest path problem with random edge costs (e.g., traffic) and contextual features (e.g., lagged traffic, weather). Existing work on CLO assumes the data has fully observed cost coefficient vectors, but in many applications, we can only see the realized cost of a historical decision, that is, just one projection of the random cost coefficient vector, to which we refer as bandit feedback. We study a class of offline learning algorithms for CLO with bandit feedback, which we term induced empirical risk minimization (IERM), where we fit a predictive model to directly optimize the downstream performance of the policy it induces. We show a fast-rate regret bound for IERM that allows for misspecified model classes and flexible choices of the optimization estimate, and we develop computationally tractable surrogate losses. A byproduct of our theory of independent interest is fast-rate regret bound for IERM with full feedback and misspecified policy class. We compare the performance of different modeling choices numerically using a stochastic shortest path example and provide practical insights from the empirical results. Yichun Hu, Nathan Kallus, Xiaojie Mao, Yanchen Wu |
NeurIPS | 3 |
| 2024 | Localized Debiased Machine Learning: Efficient Inference on Quantile Treatment Effects and BeyondabstractWe consider estimating a low-dimensional parameter in an estimating equation involving high-dimensional nuisance functions that depend on the target parameter as an input. A central example is the efficient estimating equation for the (local) quantile treatment effect ((L)QTE) in causal inference, which involves the covariate-conditional cumulative distribution function evaluated at the quantile to be estimated. Existing approaches based on flexibly estimating the nuisances and plugging in the estimates, such as debiased machine learning (DML), require we learn the nuisance at all possible inputs. For (L)QTE, DML requires we learn the whole covariate-conditional cumulative distribution function. We instead propose localized debiased machine learning (LDML), which avoids this burdensome step and needs only estimate nuisances at a single initial rough guess for the target parameter. For (L)QTE, LDML involves learning just two regression functions, a standard task for machine learning methods. We prove that under lax rate conditions our estimator has the same favorable asymptotic behavior as the infeasible estimator that uses the unknown true nuisances. Thus, LDML notably enables practically-feasible and theoretically-grounded efficient estimation of important quantities in causal inference such as (L)QTEs when we must control for many covariates and/or flexible relationships, as we demonstrate in empirical studies. Nathan Kallus, Xiaojie Mao, Masatoshi Uehara |
J. Mach. Learn. Res. | 2 |
| 2023 | Inference on Strongly Identified Functionals of Weakly Identified FunctionsabstractIn a variety of applications, including nonparametric instrumental variable (NPIV) analysis, proximal causal inference under unmeasured confounding, and missing-not-at-random data with shadow variables, we are interested in inference on a continuous linear functional (e.g., average causal effects) of nuisance function (e.g., NPIV regression) defined by conditional moment restrictions. These nuisance functions are generally weakly identified, in that the conditional moment restrictions can be severely ill-posed as well as admit multiple solutions. This is sometimes resolved by imposing strong conditions that imply the function can be estimated at rates that make inference on the functional possible. In this paper, we study a novel condition for the functional to be strongly identified even when the nuisance function is not; that is, the functional is amenable to asymptotically-normal estimation at root-n-rates. The condition implies the existence of debiasing nuisance functions, and we propose penalized minimax estimators for both the primary and debiasing nuisance functions. The proposed nuisance estimators can accommodate flexible function classes, and importantly they can converge to fixed limits determined by the penalization regardless of the identifiability of the nuisances. We use the penalized nuisance estimators to form a debiased estimator for the functional of interest and prove its asymptotic normality under generic high-level conditions, which provide for asymptotically valid confidence intervals. We also illustrate our method in a novel partially linear proximal causal inference problem and a partially linear instrumental variable regression problem. Andrew Bennett, Nathan Kallus, Xiaojie Mao, Whitney Newey, Vasilis Syrgkanis, Masatoshi Uehara |
COLT | 3 |
| 2023 | Minimax Instrumental Variable Regression and L2 Convergence Guarantees without Identification or ClosednessabstractIn this paper, we study nonparametric estimation of instrumental variable (IV) regressions. Recently, many flexible machine learning methods have been developed for instrumental variable estimation. However, these methods have at least one of the following limitations: (1) restricting the IV regression to be uniquely identified; (2) only obtaining estimation error rates in terms weak metrics (e.g., projected norm) rather than strong metrics (e.g., L_2 norm); or (3) imposing the so-called closedness condition that requires a certain conditional expectation operator to be sufficiently smooth. In this paper, we present the first method and analysis that can avoid all three limitations, while still permitting general function approximation. Specifically, we propose a new penalized minimax estimator that can converge to a fixed IV solution even when there are multiple solutions, and we derive a strong L_2 error rate for our estimator under lax conditions. Notably, this guarantee only needs a widely-used source condition and realizability assumptions, but not the so-called closedness condition. We argue that the source condition and the closedness condition are inherently conflicting, so relaxing the latter significantly improves upon the existing literature that requires both conditions. Our estimator can achieve this improvement because it builds on a novel formulation of the IV estimation problem as a constrained optimization problem. Andrew Bennett, Nathan Kallus, Xiaojie Mao, Whitney Newey, Vasilis Syrgkanis, Masatoshi Uehara |
COLT | 3 |
| 2022 | Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningabstractOff-policy evaluation and learning (OPE/L) use offline observational data to make better decisions, which is crucial in applications where online experimentation is limited. However, depending entirely on logged data, OPE/L is sensitive to environment distribution shifts — discrepancies between the data-generating environment and that where policies are deployed. Si et al., (2020) proposed distributionally robust OPE/L (DROPE/L) to address this, but the proposal relies on inverse-propensity weighting, whose estimation error and regret will deteriorate if propensities are nonparametrically estimated and whose variance is suboptimal even if not. For standard, non-robust, OPE/L, this is solved by doubly robust (DR) methods, but they do not naturally extend to the more complex DROPE/L, which involves a worst-case expectation. In this paper, we propose the first DR algorithms for DROPE/L with KL-divergence uncertainty sets. For evaluation, we propose Localized Doubly Robust DROPE (LDR$^2$OPE) and show that it achieves semiparametric efficiency under weak product rates conditions. Thanks to a localization technique, LDR$^2$OPE only requires fitting a small number of regressions, just like DR methods for standard OPE. For learning, we propose Continuum Doubly Robust DROPL (CDR$^2$OPL) and show that, under a product rate condition involving a continuum of regressions, it enjoys a fast regret rate of $O(N^{-1/2})$ even when unknown propensities are nonparametrically estimated. We empirically validate our algorithms in simulations and further extend our results to general $f$-divergence uncertainty sets. Nathan Kallus, Xiaojie Mao, Zhengyuan Zhou |
ICML | 2 |
| 2020 | Smooth Contextual Bandits: Bridging the Parametric and Non-differentiable Regret RegimesabstractWe study a nonparametric contextual bandit problem where the expected reward functions belong to a Hölder class with smoothness parameter $\beta$. We show how this interpolates between two extremes that were previously studied in isolation: non-differentiable bandits ($\beta\leq1$), where rate-optimal regret is achieved by running separate non-contextual bandits in different context regions, and parametric-response bandits (satisfying $\beta=\infty$), where rate-optimal regret can be achieved with minimal or no exploration due to infinite extrapolatability. We develop a novel algorithm that carefully adjusts to all smoothness settings and we prove its regret is rate-optimal by establishing matching upper and lower bounds, recovering the existing results at the two extremes. In this sense, our work bridges the gap between the existing literature on parametric and non-differentiable contextual bandit problems and between bandit algorithms that exclusively use global or local information, shedding light on the crucial interplay of complexity and regret in contextual bandits. Yichun Hu, Nathan Kallus, Xiaojie Mao |
COLT | 3 |
| 2019 | Interval Estimation of Individual-Level Causal Effects Under Unobserved ConfoundingabstractWe study the problem of learning conditional average treatment effects (CATE) from observational data with unobserved confounders. The CATE function maps baseline covariates to individual causal effect predictions and is key for personalized assessments. Recent work has focused on how to learn CATE under unconfoundedness, i.e., when there are no unobserved confounders. Since CATE may not be identified when unconfoundedness is violated, we develop a functional interval estimator that predicts bounds on the individual causal effects under realistic violations of unconfoundedness. Our estimator takes the form of a weighted kernel estimator with weights that vary adversarially. We prove that our estimator is sharp in that it converges exactly to the tightest bounds possible on CATE when there may be unobserved confounders. Further, we study personalized decision rules derived from our estimator and prove that they achieve optimal minimax regret asymptotically. We assess our approach in a simulation study as well as demonstrate its application in the case of hormone replacement therapy by comparing conclusions from a real observational study and clinical trial. Nathan Kallus, Xiaojie Mao, Angela Zhou |
AISTATS | 2 |
| 2018 | Causal Inference with Noisy and Missing Covariates via Matrix FactorizationabstractValid causal inference in observational studies often requires controlling for confounders. However, in practice measurements of confounders may be noisy, and can lead to biased estimates of causal effects. We show that we can reduce bias induced by measurement noise using a large number of noisy measurements of the underlying confounders. We propose the use of matrix factorization to infer the confounders from noisy covariates. This flexible and principled framework adapts to missing values, accommodates a wide variety of data types, and can enhance a wide variety of causal inference methods. We bound the error for the induced average treatment effect estimator and show it is consistent in a linear regression setting, using Exponential Family Matrix Completion preprocessing. We demonstrate the effectiveness of the proposed procedure in numerical experiments with both synthetic data and real clinical data. Nathan Kallus, Xiaojie Mao, Madeleine Udell |
NeurIPS | 2 |