VLDB 2026 Research / reviewers in the wild / expert
Jiale Zha
dblp:387/1718
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 60% Generative modeling · 40% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning
maximum entropy reinforcement learning |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Methods — techniques the papers use, named apart from their topics
ratio estimator · 0.9q-learning · 0.9actor-critic · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reward-Directed Score-Based Diffusion Models via q-LearningabstractWe propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions. Different from most existing studies, ours does not involve any pretrained model for the unknown score functions of the noise-perturbed data distributions, nor does it attempt to learn the score functions. Instead, we formulate the problem as entropy-regularized continuous-time RL and show that the optimal stochastic policy has a Gaussian distribution with a known covariance matrix. Based on this result, we parameterize the mean of Gaussian policies and develop an actor--critic type (little) $q$-learning algorithm to solve the RL problem. A key ingredient in our algorithm design is to obtain noisy observations from the unknown score function via a ratio estimator. Our formulation can also be adapted to solve pure score-matching and fine-tuning pretrained models. Numerically, we show the effectiveness of our approach by comparing its performance with two state-of-the-art RL methods that fine-tune pretrained models on several generative tasks including high-dimensional image generations. Finally, we discuss extensions of our RL formulation to probability flow ODE implementation of diffusion models and to conditional diffusion models. Jiale Zha, Xun Yu Zhou |
J. Mach. Learn. Res. | 2 |