VLDB 2026 Research / reviewers in the wild / expert
Christina J. Yuan
dblp:305/7681
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 86% Probabilistic and Bayesian machine learning · 11% Learning theory · 3% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
off-policy evaluation |
1.3 | 2 | 2024 | OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024 SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Reinforcement learning › off-policy evaluation
estimator selection |
0.8 | 1 | 2024 | OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024 |
Machine learning › Reinforcement learning
policy evaluation |
0.8 | 1 | 2024 | OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024 |
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation |
0.5 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling |
0.5 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff |
0.1 | 1 | 2021 | SOPE: Spectrum of Off-Policy Estimators · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
statistical estimation · 0.8reweighting · 0.8trajectory importance sampling · 0.5state-action visitation distribution · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsabstractOffline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from other policies. Evaluating a new policy online without a confident estimate of its performance can lead to costly, unsafe, or hazardous outcomes, especially in education and healthcare. Several OPE estimators have been proposed in the last decade, many of which have hyperparameters and require training. Unfortunately, choosing the best OPE algorithm for each task and domain is still unclear. In this paper, we propose a new algorithm that adaptively blends a set of OPE estimators given a dataset without relying on an explicit selection using a statistical procedure. We prove that our estimator is consistent and satisfies several desirable properties for policy evaluation. Additionally, we demonstrate that when compared to alternative approaches, our estimator can be used to select higher-performing policies in healthcare and robotics. Our work contributes to improving ease of use for a general-purpose, estimator-agnostic, off-policy evaluation framework for offline RL. Allen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath, Yannis Flet-Berliac, Emma Brunskill |
NeurIPS | 3 |
| 2021 | SOPE: Spectrum of Off-Policy EstimatorsabstractMany sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the most common OPE techniques that provides unbiased estimates is trajectory based importance sampling (IS). However, due to the high variance of trajectory IS estimates, importance sampling methods based on state-action visitation distributions (SIS) have recently been adopted. Unfortunately, while SIS often provides lower variance estimates for long horizons, estimating the state-action distribution ratios can be challenging and lead to biased estimates. In this paper, we present a new perspective on this bias-variance trade-off and show the existence of a spectrum of estimators whose endpoints are SIS and IS. Additionally, we also establish a spectrum for doubly-robust and weighted version of these estimators. We provide empirical evidence that estimators in this spectrum can be used to trade-off between the bias and variance of IS and SIS and can achieve lower mean-squared error than both IS and SIS. Christina J. Yuan, Yash Chandak, Stephen Giguere 0001, Philip S. Thomas, Scott Niekum |
NeurIPS | 1 |