VLDB 2026 Research / reviewers in the wild / expert
Yulai Zhao 0002
dblp:64/6357-2
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6930-3590ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 51% Reinforcement learning · 28% Optimization for machine learning · 8% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
3.3 | 4 | 2025 | Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025 Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025 Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
diffusion model conditioning |
0.9 | 1 | 2025 | Adding Conditional Control to Diffusion Models with Reinforcement Learning · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
0.9 | 1 | 2025 | Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
guided sampling |
0.9 | 1 | 2025 | Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
inference-time optimization |
0.9 | 1 | 2025 | Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025 |
Machine learning › Generative modeling
iterative refinement |
0.9 | 1 | 2025 | Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025 |
Machine learning › Reinforcement learning
reinforcement learning for generative models |
0.9 | 1 | 2025 | Adding Conditional Control to Diffusion Models with Reinforcement Learning · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation |
0.9 | 1 | 2025 | Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024 |
Machine learning › Reinforcement learning
function approximation |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Reinforcement learning › markov decision process
low-rank MDP |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Optimization for machine learning
model-based optimization |
0.8 | 1 | 2024 | Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024 |
Machine learning › Optimization for machine learning › model-based optimization
offline model-based optimization |
0.8 | 1 | 2024 | Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024 |
Machine learning › Learning paradigms › incremental learning
online fine-tuning |
0.8 | 1 | 2024 | Feedback Efficient Online Fine-Tuning of Diffusion Models · ICML 2024 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning |
0.8 | 1 | 2024 | Feedback Efficient Online Fine-Tuning of Diffusion Models · ICML 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Learning theory
sample complexity |
0.8 | 1 | 2024 | Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.7 | 1 | 2023 | Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.7 | 1 | 2023 | Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning
policy optimization |
0.7 | 1 | 2023 | Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
classifier guidance · 1.7reinforcement learning · 1.6value function · 0.9theoretical guarantee · 0.9soft value-based decoding · 0.9noising-denoising · 0.9classifier-free guidance · 0.9upper confidence bound · 0.8maximum likelihood estimation · 0.8least-squares value iteration · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adding Conditional Control to Diffusion Models with Reinforcement LearningabstractDiffusion models are powerful generative models that allow for precise control over the characteristics of the generated samples. While these diffusion models trained on large datasets have achieved success, there is often a need to introduce additional controls in downstream fine-tuning processes, treating these powerful models as pre-trained diffusion models. This work presents a novel method based on reinforcement learning (RL) to add such controls using an offline dataset comprising inputs and labels. We formulate this task as an RL problem, with the classifier learned from the offline dataset and the KL divergence against pre-trained models serving as the reward functions. Our method, **CTRL** (**C**onditioning pre-**T**rained diffusion models with **R**einforcement **L**earning), produces soft-optimal policies that maximize the abovementioned reward functions. We formally demonstrate that our method enables sampling from the conditional distribution with additional controls during inference.
Our RL-based approach offers several advantages over existing methods. Compared to classifier-free guidance,
it improves sample efficiency and can greatly simplify dataset construction by leveraging conditional independence between the inputs and additional controls. Additionally, unlike classifier guidance, it eliminates the need to train classifiers from intermediate states to additional controls.
The code is available at https://github.com/zhaoyl18/CTRL. Yulai Zhao 0002, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung, Tommaso Biancalani, Sergey Levine, Ehsan Hajiramezanali |
ICLR | 1 |
| 2025 | Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA DesignabstractTo fully leverage the capabilities of diffusion models, we are often interested in optimizing downstream reward functions during inference. While numerous algorithms for reward-guided generation have been recently proposed due to their significance, current approaches predominantly focus on single-shot generation, transitioning from fully noised to denoised states. We propose a novel framework for inference-time reward optimization with diffusion models. Our approach employs an iterative refinement process consisting of two steps in each iteration: noising and reward-guided denoising. This sequential refinement allows for the gradual correction of errors introduced during reward optimization. Finally, we provide a theoretical guarantee for our framework. Finally, we demonstrate its superior empirical performance in protein and DNA design. Masatoshi Uehara, Xingyu Su, Yulai Zhao 0002, Xiner Li, Aviv Regev, Shuiwang Ji, Sergey Levine, Tommaso Biancalani |
ICML | 3 |
| 2025 | Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based DecodingabstractDiffusion models excel at capturing the natural design spaces of images, molecules, DNA, RNA, and protein sequences. However, rather than merely generating designs that are natural, we often aim to optimize downstream reward functions while preserving the naturalness of these design spaces. Existing methods for achieving this goal often require differentiable proxy models (e.g., classifier guidance or DPS) or involve computationally expensive fine-tuning of diffusion models (e.g., classifier-free guidance, RL-based fine-tuning). In our work, we propose a new method to address these challenges. Our algorithm is an iterative sampling method that integrates soft value functions, which looks ahead to how intermediate noisy states lead to high rewards in the future, into the standard inference procedure of pre-trained diffusion models. Notably, our approach avoids fine-tuning generative models and eliminates the need to construct differentiable models. This enables us to (1) directly utilize non-differentiable features/reward feedback, commonly used in many scientific domains, and (2) apply our method to recent discrete diffusion models in a principled way. Finally, we demonstrate the effectiveness of our algorithm across several domains, including image generation, molecule generation, and DNA/RNA sequence generation. Xiner Li, Yulai Zhao 0002, Chenyu Wang 0003, Gabriele Scalia, Gökcen Eraslan, Surag Nair, Tommaso Biancalani, Shuiwang Ji, Aviv Regev, Sergey Levine, Masatoshi Uehara |
NeurIPS | 2 |
| 2024 | Provably Efficient CVaR RL in Low-rank MDPsabstractWe study risk-sensitive Reinforcement Learning (RL), where we aim to maximize
the Conditional Value at Risk (CVaR) with a fixed risk tolerance $\tau$.
Prior theoretical work studying risk-sensitive RL focuses on the tabular Markov Decision Processes (MDPs) setting.
To extend CVaR RL to settings where state space is large, function approximation must be deployed.
We study CVaR RL in low-rank MDPs with nonlinear function approximation. Low-rank MDPs assume the underlying transition kernel admits a low-rank decomposition, but unlike prior linear models, low-rank MDPs do not assume the feature or state-action representation is known.
We propose a novel Upper Confidence Bound (UCB) bonus-driven algorithm to carefully balance the interplay between exploration, exploitation, and representation learning in CVaR RL.
We prove that our algorithm achieves a sample complexity of $\tilde{O}\left(\frac{H^7 A^2 d^4}{\tau^2 \epsilon^2}\right)$ to yield an $\epsilon$-optimal CVaR, where $H$ is the length of each episode, $A$ is the capacity of action space, and $d$ is the dimension of representations.
Computational-wise, we design a novel discretized Least-Squares Value Iteration (LSVI) algorithm for the CVaR objective as the planning oracle and show that we can find the near-optimal policy in a polynomial running time with a Maximum Likelihood Estimation oracle.
To our knowledge, this is the first provably efficient CVaR RL algorithm in low-rank MDPs. Yulai Zhao 0002, Wenhao Zhan, Xiaoyan Hu 0003, Ho-fung Leung, Farzan Farnia, Wen Sun 0002, Jason D. Lee |
ICLR | 1 |
| 2024 | Feedback Efficient Online Fine-Tuning of Diffusion ModelsabstractDiffusion models excel at modeling complex data distributions, including those of images, proteins, and small molecules. However, in many cases, our goal is to model parts of the distribution that maximize certain properties: for example, we may want to generate images with high aesthetic quality, or molecules with high bioactivity. It is natural to frame this as a reinforcement learning (RL) problem, in which the objective is to finetune a diffusion model to maximize a reward function that corresponds to some property. Even with access to online queries of the ground-truth reward function, efficiently discovering high-reward samples can be challenging: they might have a low probability in the initial distribution, and there might be many infeasible samples that do not even have a well-defined reward (e.g., unnatural images or physically impossible molecules). In this work, we propose a novel reinforcement learning procedure that efficiently explores on the manifold of feasible samples. We present a theoretical analysis providing a regret guarantee, as well as empirical validation across three domains: images, biological sequences, and molecules. Masatoshi Uehara, Yulai Zhao 0002, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel Diamant, Alex M. Tseng, Sergey Levine, Tommaso Biancalani |
ICML | 2 |
| 2024 | Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsabstractAI-driven design problems, such as DNA/protein sequence design, are commonly tackled from two angles: generative modeling, which efficiently captures the feasible design space (e.g., natural images or biological sequences), and model-based optimization, which utilizes reward models for extrapolation. To combine the strengths of both approaches, we adopt a hybrid method that fine-tunes cutting-edge diffusion models by optimizing reward models through RL. Although prior work has explored similar avenues, they primarily focus on scenarios where accurate reward models are accessible. In contrast, we concentrate on an offline setting where a reward model is unknown, and we must learn from static offline datasets, a common scenario in scientific domains. In offline scenarios, existing approaches tend to suffer from overoptimization, as they may be misled by the reward model in out-of-distribution regions. To address this, we introduce a conservative fine-tuning approach, BRAID, by optimizing a conservative reward model, which includes additional penalization outside of offline data distributions. Through empirical and theoretical analysis, we demonstrate the capability of our approach to outperform the best designs in offline data, leveraging the extrapolation capabilities of reward models while avoiding the generation of invalid designs through pre-trained diffusion models. Masatoshi Uehara, Yulai Zhao 0002, Ehsan Hajiramezanali, Gabriele Scalia, Gökcen Eraslan, Avantika Lal, Sergey Levine, Tommaso Biancalani |
NeurIPS | 2 |
| 2023 | Blessing of Class Diversity in Pre-trainingabstractThis paper presents a new statistical analysis aiming to explain the recent superior achievements of the pre-training techniques in natural language processing (NLP). We prove that when the classes of the pre-training task (e.g., different words in the masked language model task) are sufficiently diverse, in the sense that the least singular value of the last linear layer in pre-training (denoted as $\tilde{\nu}$) is large, then pre-training can significantly improve the sample efficiency of downstream tasks. Specially, we show the transfer learning excess risk enjoys an $O\left(\frac{1}{\tilde{\nu} \sqrt{n}}\right)$ rate, in contrast to the $O\left(\frac{1}{\sqrt{m}}\right)$ rate in the standard supervised learning. Here, $n$ is the number of pre-training data and $m$ is the number of data in the downstream task, and typically $n \gg m$. Our proof relies on a vector-form Rademacher complexity chain rule for disassembling composite function classes and a modified self-concordance condition. These techniques can be of independent interest. Yulai Zhao 0002, Jianshu Chen, Simon S. Du |
AISTATS | 1 |
| 2023 | Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement LearningabstractPolicy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statistical guarantees. Leveraging a multi-agent performance difference lemma that characterizes the landscape of multi-agent policy optimization, we find that the localized action value function serves as an ideal descent direction for each local policy. Motivated by the observation, we present a multi-agent PPO algorithm in which the local policy of each agent is updated similarly to vanilla PPO. We prove that with standard regularity conditions on the Markov game and problem-dependent quantities, our algorithm converges to the globally optimal policy at a sublinear rate. We extend our algorithm to the off-policy setting and introduce pessimism to policy evaluation, which aligns with experiments. To our knowledge, this is the first provably convergent multi-agent PPO algorithm in cooperative Markov games. Yulai Zhao 0002, Zhuoran Yang, Zhaoran Wang 0001, Jason D. Lee |
ICML | 1 |
| 2022 | Provably Efficient Policy Optimization for Two-Player Zero-Sum Markov GamesabstractPolicy-based methods with function approximation are widely used for solving two-player zero-sum games with large state and/or action spaces. However, it remains elusive how to obtain optimization and statistical guarantees for such algorithms. We present a new policy optimization algorithm with function approximation and prove that under standard regularity conditions on the Markov game and the function approximation class, our algorithm finds a near-optimal policy within a polynomial number of samples and iterations. To our knowledge, this is the first provably efficient policy optimization algorithm with function approximation that solves two-player zero-sum Markov games. Yulai Zhao 0002, Yuandong Tian, Jason D. Lee, Simon S. Du |
AISTATS | 1 |