EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyan Hu 0003
dblp:66/1557-3
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-5766-1059ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 70% Learning theory · 25% Language models and text generation · 3% | |
| Theoretical computer science
1 paper |
Information theory · 50% Logic in computer science · 50% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.9 | 1 | 2025 | PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs · ICML 2025 |
Machine learning › Learning theory › online learning
online model selection |
0.9 | 1 | 2025 | PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs · ICML 2025 |
Machine learning › Reinforcement learning
function approximation |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Learning theory
information-theoretic analysis |
0.8 | 1 | 2024 | An Information Theoretic Approach to Interaction-Grounded Learning · ICML 2024 |
Machine learning › Reinforcement learning › imitation learning › inverse reinforcement learning
inverse optimal control |
0.8 | 1 | 2024 | An Information Theoretic Approach to Interaction-Grounded Learning · ICML 2024 |
Machine learning › Reinforcement learning › markov decision process
low-rank MDP |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Reinforcement learning
reinforcement learning theory |
0.8 | 1 | 2024 | An Information Theoretic Approach to Interaction-Grounded Learning · ICML 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Learning theory
sample complexity |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.7 | 1 | 2023 | Provably (More) Sample-Efficient Offline RL with Options · NeurIPS 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.7 | 1 | 2023 | Provably (More) Sample-Efficient Offline RL with Options · NeurIPS 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework |
0.7 | 1 | 2023 | Provably (More) Sample-Efficient Offline RL with Options · NeurIPS 2023 |
Logic in computer science › knowledge representation and reasoning
conditional independence |
0.2 | 1 | 2024 | An Information Theoretic Approach to Interaction-Grounded Learning · ICML 2024 |
Information theory › information measures
mutual information |
0.2 | 1 | 2024 | An Information Theoretic Approach to Interaction-Grounded Learning · ICML 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
information gathering |
0.2 | 1 | 2023 | Provably (More) Sample-Efficient Offline RL with Options · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
upper confidence bound · 1.6variational mutual information estimation · 1.5min-max optimization · 1.5f-information measures · 1.5random fourier features · 0.9kernel methods · 0.9maximum likelihood estimation · 0.8least-squares value iteration · 0.8pessimistic value iteration · 0.7information-theoretic lower bounds · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative ModelsabstractExisting frameworks for evaluating and comparing generative models consider an offline setting, where the evaluator has access to large batches of data produced by the models. However, in practical scenarios, the goal is often to identify and select the best model using the fewest possible generated samples to minimize the costs of querying data from the sub-optimal models. In this work, we propose an online evaluation and selection framework to find the generative model that maximizes a standard assessment score among a group of available models. We view the task as a multi-armed bandit (MAB) and propose upper confidence bound (UCB) bandit algorithms to identify the model producing data with the best evaluation score that quantifies the quality and diversity of generated data. Specifically, we develop the MAB-based selection of generative models considering the Fr{é}chet Distance (FD) and Inception Score (IS) metrics, resulting in the FD-UCB and IS-UCB algorithms. We prove regret bounds for these algorithms and present numerical results on standard image datasets. Our empirical results suggest the efficacy of MAB approaches for the sample-efficient evaluation and selection of deep generative models. The project code is available at \url{https://github.com/yannxiaoyanhu/dgm-online-eval}. Xiaoyan Hu 0003, Ho-fung Leung, Farzan Farnia |
AISTATS | 1 |
| 2025 | PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMsabstractSelecting a sample generation scheme from multiple prompt-based generative models, including large language models (LLMs) and prompt-guided image and video generation models, is typically addressed by choosing the model that maximizes an averaged evaluation score. However, this score-based selection overlooks the possibility that different models achieve the best generation performance for different types of text prompts. An online identification of the best generation model for various input prompts can reduce the costs associated with querying sub-optimal models. In this work, we explore the possibility of varying rankings of text-based generative models for different text prompts and propose an online learning framework to predict the best data generation model for a given input prompt. The proposed PAK-UCB algorithm addresses a contextual bandit (CB) setting with shared context variables across the arms, utilizing the generated data to update kernel-based functions that predict the score of each model available for unseen text prompts. Additionally, we leverage random Fourier features (RFF) to accelerate the online learning process of PAK-UCB. Our numerical experiments on real and simulated text-to-image and image-to-text generative models show that RFF-UCB performs successfully in identifying the best generation model across different sample types. The code is available at: github.com/yannxiaoyanhu/dgm-online-select. Xiaoyan Hu 0003, Ho-fung Leung, Farzan Farnia |
ICML | 1 |
| 2024 | Provably Efficient CVaR RL in Low-rank MDPsabstractWe study risk-sensitive Reinforcement Learning (RL), where we aim to maximize
the Conditional Value at Risk (CVaR) with a fixed risk tolerance $\tau$.
Prior theoretical work studying risk-sensitive RL focuses on the tabular Markov Decision Processes (MDPs) setting.
To extend CVaR RL to settings where state space is large, function approximation must be deployed.
We study CVaR RL in low-rank MDPs with nonlinear function approximation. Low-rank MDPs assume the underlying transition kernel admits a low-rank decomposition, but unlike prior linear models, low-rank MDPs do not assume the feature or state-action representation is known.
We propose a novel Upper Confidence Bound (UCB) bonus-driven algorithm to carefully balance the interplay between exploration, exploitation, and representation learning in CVaR RL.
We prove that our algorithm achieves a sample complexity of $\tilde{O}\left(\frac{H^7 A^2 d^4}{\tau^2 \epsilon^2}\right)$ to yield an $\epsilon$-optimal CVaR, where $H$ is the length of each episode, $A$ is the capacity of action space, and $d$ is the dimension of representations.
Computational-wise, we design a novel discretized Least-Squares Value Iteration (LSVI) algorithm for the CVaR objective as the planning oracle and show that we can find the near-optimal policy in a polynomial running time with a Maximum Likelihood Estimation oracle.
To our knowledge, this is the first provably efficient CVaR RL algorithm in low-rank MDPs. Yulai Zhao 0002, Wenhao Zhan, Xiaoyan Hu 0003, Ho-fung Leung, Farzan Farnia, Wen Sun 0002, Jason D. Lee |
ICLR | 3 |
| 2024 | An Information Theoretic Approach to Interaction-Grounded LearningabstractReinforcement learning (RL) problems where the learner attempts to infer an unobserved reward from some feedback variables have been studied in several recent papers. The setting of Interaction-Grounded Learning (IGL) is an example of such feedback-based reinforcement learning tasks where the learner optimizes the return by inferring latent binary rewards from the interaction with the environment. In the IGL setting, a relevant assumption used in the RL literature is that the feedback variable $Y$ is conditionally independent of the context-action $(X,A)$ given the latent reward $R$. In this work, we propose *Variational Information-based IGL (VI-IGL)* as an information-theoretic method to enforce the conditional independence assumption in the IGL-based RL problem. The VI-IGL framework learns a reward decoder using an information-based objective based on the conditional mutual information (MI) between the context-action $(X,A)$ and the feedback variable $Y$ observed from the environment. To estimate and optimize the information-based terms for the continuous random variables in the RL problem, VI-IGL leverages the variational representation of mutual information and results in a min-max optimization problem. Theoretical analysis shows that the optimization problem can be sample-efficiently solved. Furthermore, we extend the VI-IGL framework to general $f$-Information measures in the information theory literature, leading to the generalized $f$-VI-IGL framework to address the RL problem under the IGL condition. Finally, the empirical results on several reinforcement learning settings indicate an improved performance in comparison to the previous IGL-based RL algorithm. Xiaoyan Hu 0003, Farzan Farnia, Ho-fung Leung |
ICML | 1 |
| 2023 | A Tighter Problem-Dependent Regret Bound for Risk-Sensitive Reinforcement LearningabstractWe study the regret for risk-sensitive reinforcement learning (RL) with the exponential utility in the episodic MDP. Recent works establish both a lower bound $\Omega((e^{|\beta|(H-1)/2}-1)\sqrt{SAT}/|\beta|)$ and the best known (upper) bound $\tilde{O}((e^{|\beta|H}-1)\sqrt{H^2SAT}/|\beta|)$, where $H$ is the length of the episode, $S$ the size of state space, $A$ the size of action space, $T$ the total number of timesteps, and $\beta$ the risk parameter. The gap between the upper and the lower bound is exponential and hence is unsatisfactory. In this paper, we show that a variant of UCB-Advantage algorithm reduces a factor of $\sqrt{H}$ from the best previously known bound in any arbitrary MDP. To further sharpen the regret bound, we introduce a brand new mechanism of regret analysis and derive a problem-dependent regret bound without prior knowledge of the MDP from the algorithm. This bound is much tighter in MDPs with special structures. Particularly, we show that a regret that matches the information-theoretic lower bound up to logarithmic factors can be attained within a rich class of MDPs, which improves an exponential factor over the best previously known bound. Further, we derive a novel information-theoretic lower bound of $\Omega(\max_{h\in[H]} c_{v,h+1}^*\sqrt{SAT}/|\beta|)$, where $\max_{h\in[H]} c_{v,h+1}^*$ is a problem-dependent statistic. This lower bound shows that the problem-dependent regret bound achieved by the algorithm is optimal in its dependence on $\max_{h\in[H]} c_{v,h+1}^*$. Xiaoyan Hu 0003, Ho-fung Leung |
AISTATS | 1 |
| 2023 | Provably (More) Sample-Efficient Offline RL with OptionsabstractThe options framework yields empirical success in long-horizon planning problems of reinforcement learning (RL). Recent works show that options help improve the sample efficiency in online RL. However, these results are no longer applicable to scenarios where exploring the environment online is risky, e.g., automated driving and healthcare. In this paper, we provide the first analysis of the sample complexity for offline RL with options, where the agent learns from a dataset without further interaction with the environment. We derive a novel information-theoretic lower bound, which generalizes the one for offline learning with actions. We propose the PEssimistic Value Iteration for Learning with Options (PEVIO) algorithm and establish near-optimal suboptimality bounds for two popular data-collection procedures, where the first one collects state-option transitions and the second one collects state-action transitions. We show that compared to offline RL with actions, using options not only enjoys a faster finite-time convergence rate (to the optimal value) but also attains a better performance when either the options are carefully designed or the offline data is limited. Based on these results, we analyze the pros and cons of the data-collection procedures. Xiaoyan Hu 0003, Ho-fung Leung |
NeurIPS | 1 |