Yulai Zhao 0002

dblp:64/6357-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6930-3590ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 51% Reinforcement learning · 28% Optimization for machine learning · 8%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.342025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025
Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
diffusion model conditioning
0.912025
Adding Conditional Control to Diffusion Models with Reinforcement Learning · ICLR 2025
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.912025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
guided sampling
0.912025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
inference-time optimization
0.912025
Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025
Machine learning › Generative modeling
iterative refinement
0.912025
Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning for generative models
0.912025
Adding Conditional Control to Diffusion Models with Reinforcement Learning · ICLR 2025
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation
0.912025
Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design · ICML 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024
Machine learning › Reinforcement learning
function approximation
0.812024
Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024
Machine learning › Reinforcement learning › markov decision process
low-rank MDP
0.812024
Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024
Machine learning › Optimization for machine learning
model-based optimization
0.812024
Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024
Machine learning › Optimization for machine learning › model-based optimization
offline model-based optimization
0.812024
Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models · NeurIPS 2024
Machine learning › Learning paradigms › incremental learning
online fine-tuning
0.812024
Feedback Efficient Online Fine-Tuning of Diffusion Models · ICML 2024
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning
0.812024
Feedback Efficient Online Fine-Tuning of Diffusion Models · ICML 2024
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
0.812024
Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024
Machine learning › Learning theory
sample complexity
0.812024
Provably Efficient CVaR RL in Low-rank MDPs · ICLR 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.712023
Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
policy optimization
0.712023
Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning · ICML 2023

Methods — techniques the papers use, named apart from their topics

classifier guidance · 1.7reinforcement learning · 1.6value function · 0.9theoretical guarantee · 0.9soft value-based decoding · 0.9noising-denoising · 0.9classifier-free guidance · 0.9upper confidence bound · 0.8maximum likelihood estimation · 0.8least-squares value iteration · 0.8
YearPublicationVenuePosition
2025 Adding Conditional Control to Diffusion Models with Reinforcement Learning
abstract
Diffusion models are powerful generative models that allow for precise control over the characteristics of the generated samples. While these diffusion models trained on large datasets have achieved success, there is often a need to introduce additional controls in downstream fine-tuning processes, treating these powerful models as pre-trained diffusion models. This work presents a novel method based on reinforcement learning (RL) to add such controls using an offline dataset comprising inputs and labels. We formulate this task as an RL problem, with the classifier learned from the offline dataset and the KL divergence against pre-trained models serving as the reward functions. Our method, **CTRL** (**C**onditioning pre-**T**rained diffusion models with **R**einforcement **L**earning), produces soft-optimal policies that maximize the abovementioned reward functions. We formally demonstrate that our method enables sampling from the conditional distribution with additional controls during inference. Our RL-based approach offers several advantages over existing methods. Compared to classifier-free guidance, it improves sample efficiency and can greatly simplify dataset construction by leveraging conditional independence between the inputs and additional controls. Additionally, unlike classifier guidance, it eliminates the need to train classifiers from intermediate states to additional controls. The code is available at https://github.com/zhaoyl18/CTRL.
Yulai Zhao 0002, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung, Tommaso Biancalani, Sergey Levine, Ehsan Hajiramezanali
ICLR1
2025 Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design
abstract
To fully leverage the capabilities of diffusion models, we are often interested in optimizing downstream reward functions during inference. While numerous algorithms for reward-guided generation have been recently proposed due to their significance, current approaches predominantly focus on single-shot generation, transitioning from fully noised to denoised states. We propose a novel framework for inference-time reward optimization with diffusion models. Our approach employs an iterative refinement process consisting of two steps in each iteration: noising and reward-guided denoising. This sequential refinement allows for the gradual correction of errors introduced during reward optimization. Finally, we provide a theoretical guarantee for our framework. Finally, we demonstrate its superior empirical performance in protein and DNA design.
Masatoshi Uehara, Xingyu Su, Yulai Zhao 0002, Xiner Li, Aviv Regev, Shuiwang Ji, Sergey Levine, Tommaso Biancalani
ICML3
2025 Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding
abstract
Diffusion models excel at capturing the natural design spaces of images, molecules, DNA, RNA, and protein sequences. However, rather than merely generating designs that are natural, we often aim to optimize downstream reward functions while preserving the naturalness of these design spaces. Existing methods for achieving this goal often require differentiable proxy models (e.g., classifier guidance or DPS) or involve computationally expensive fine-tuning of diffusion models (e.g., classifier-free guidance, RL-based fine-tuning). In our work, we propose a new method to address these challenges. Our algorithm is an iterative sampling method that integrates soft value functions, which looks ahead to how intermediate noisy states lead to high rewards in the future, into the standard inference procedure of pre-trained diffusion models. Notably, our approach avoids fine-tuning generative models and eliminates the need to construct differentiable models. This enables us to (1) directly utilize non-differentiable features/reward feedback, commonly used in many scientific domains, and (2) apply our method to recent discrete diffusion models in a principled way. Finally, we demonstrate the effectiveness of our algorithm across several domains, including image generation, molecule generation, and DNA/RNA sequence generation.
Xiner Li, Yulai Zhao 0002, Chenyu Wang 0003, Gabriele Scalia, Gökcen Eraslan, Surag Nair, Tommaso Biancalani, Shuiwang Ji, Aviv Regev, Sergey Levine, Masatoshi Uehara
NeurIPS2
2024 Provably Efficient CVaR RL in Low-rank MDPs
abstract
We study risk-sensitive Reinforcement Learning (RL), where we aim to maximize the Conditional Value at Risk (CVaR) with a fixed risk tolerance $\tau$. Prior theoretical work studying risk-sensitive RL focuses on the tabular Markov Decision Processes (MDPs) setting. To extend CVaR RL to settings where state space is large, function approximation must be deployed. We study CVaR RL in low-rank MDPs with nonlinear function approximation. Low-rank MDPs assume the underlying transition kernel admits a low-rank decomposition, but unlike prior linear models, low-rank MDPs do not assume the feature or state-action representation is known. We propose a novel Upper Confidence Bound (UCB) bonus-driven algorithm to carefully balance the interplay between exploration, exploitation, and representation learning in CVaR RL. We prove that our algorithm achieves a sample complexity of $\tilde{O}\left(\frac{H^7 A^2 d^4}{\tau^2 \epsilon^2}\right)$ to yield an $\epsilon$-optimal CVaR, where $H$ is the length of each episode, $A$ is the capacity of action space, and $d$ is the dimension of representations. Computational-wise, we design a novel discretized Least-Squares Value Iteration (LSVI) algorithm for the CVaR objective as the planning oracle and show that we can find the near-optimal policy in a polynomial running time with a Maximum Likelihood Estimation oracle. To our knowledge, this is the first provably efficient CVaR RL algorithm in low-rank MDPs.
Yulai Zhao 0002, Wenhao Zhan, Xiaoyan Hu 0003, Ho-fung Leung, Farzan Farnia, Wen Sun 0002, Jason D. Lee
ICLR1
2024 Feedback Efficient Online Fine-Tuning of Diffusion Models
abstract
Diffusion models excel at modeling complex data distributions, including those of images, proteins, and small molecules. However, in many cases, our goal is to model parts of the distribution that maximize certain properties: for example, we may want to generate images with high aesthetic quality, or molecules with high bioactivity. It is natural to frame this as a reinforcement learning (RL) problem, in which the objective is to finetune a diffusion model to maximize a reward function that corresponds to some property. Even with access to online queries of the ground-truth reward function, efficiently discovering high-reward samples can be challenging: they might have a low probability in the initial distribution, and there might be many infeasible samples that do not even have a well-defined reward (e.g., unnatural images or physically impossible molecules). In this work, we propose a novel reinforcement learning procedure that efficiently explores on the manifold of feasible samples. We present a theoretical analysis providing a regret guarantee, as well as empirical validation across three domains: images, biological sequences, and molecules.
Masatoshi Uehara, Yulai Zhao 0002, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel Diamant, Alex M. Tseng, Sergey Levine, Tommaso Biancalani
ICML2
2024 Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models
abstract
AI-driven design problems, such as DNA/protein sequence design, are commonly tackled from two angles: generative modeling, which efficiently captures the feasible design space (e.g., natural images or biological sequences), and model-based optimization, which utilizes reward models for extrapolation. To combine the strengths of both approaches, we adopt a hybrid method that fine-tunes cutting-edge diffusion models by optimizing reward models through RL. Although prior work has explored similar avenues, they primarily focus on scenarios where accurate reward models are accessible. In contrast, we concentrate on an offline setting where a reward model is unknown, and we must learn from static offline datasets, a common scenario in scientific domains. In offline scenarios, existing approaches tend to suffer from overoptimization, as they may be misled by the reward model in out-of-distribution regions. To address this, we introduce a conservative fine-tuning approach, BRAID, by optimizing a conservative reward model, which includes additional penalization outside of offline data distributions. Through empirical and theoretical analysis, we demonstrate the capability of our approach to outperform the best designs in offline data, leveraging the extrapolation capabilities of reward models while avoiding the generation of invalid designs through pre-trained diffusion models.
Masatoshi Uehara, Yulai Zhao 0002, Ehsan Hajiramezanali, Gabriele Scalia, Gökcen Eraslan, Avantika Lal, Sergey Levine, Tommaso Biancalani
NeurIPS2
2023 Blessing of Class Diversity in Pre-training
abstract
This paper presents a new statistical analysis aiming to explain the recent superior achievements of the pre-training techniques in natural language processing (NLP). We prove that when the classes of the pre-training task (e.g., different words in the masked language model task) are sufficiently diverse, in the sense that the least singular value of the last linear layer in pre-training (denoted as $\tilde{\nu}$) is large, then pre-training can significantly improve the sample efficiency of downstream tasks. Specially, we show the transfer learning excess risk enjoys an $O\left(\frac{1}{\tilde{\nu} \sqrt{n}}\right)$ rate, in contrast to the $O\left(\frac{1}{\sqrt{m}}\right)$ rate in the standard supervised learning. Here, $n$ is the number of pre-training data and $m$ is the number of data in the downstream task, and typically $n \gg m$. Our proof relies on a vector-form Rademacher complexity chain rule for disassembling composite function classes and a modified self-concordance condition. These techniques can be of independent interest.
Yulai Zhao 0002, Jianshu Chen, Simon S. Du
AISTATS1
2023 Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning
abstract
Policy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statistical guarantees. Leveraging a multi-agent performance difference lemma that characterizes the landscape of multi-agent policy optimization, we find that the localized action value function serves as an ideal descent direction for each local policy. Motivated by the observation, we present a multi-agent PPO algorithm in which the local policy of each agent is updated similarly to vanilla PPO. We prove that with standard regularity conditions on the Markov game and problem-dependent quantities, our algorithm converges to the globally optimal policy at a sublinear rate. We extend our algorithm to the off-policy setting and introduce pessimism to policy evaluation, which aligns with experiments. To our knowledge, this is the first provably convergent multi-agent PPO algorithm in cooperative Markov games.
Yulai Zhao 0002, Zhuoran Yang, Zhaoran Wang 0001, Jason D. Lee
ICML1
2022 Provably Efficient Policy Optimization for Two-Player Zero-Sum Markov Games
abstract
Policy-based methods with function approximation are widely used for solving two-player zero-sum games with large state and/or action spaces. However, it remains elusive how to obtain optimization and statistical guarantees for such algorithms. We present a new policy optimization algorithm with function approximation and prove that under standard regularity conditions on the Markov game and the function approximation class, our algorithm finds a near-optimal policy within a polynomial number of samples and iterations. To our knowledge, this is the first provably efficient policy optimization algorithm with function approximation that solves two-player zero-sum Markov games.
Yulai Zhao 0002, Yuandong Tian, Jason D. Lee, Simon S. Du
AISTATS1