EDBT 2026 Demo / reviewers in the wild / expert
Xun Yu Zhou
dblp:31/289
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-9908-5697ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 85% Generative modeling · 15% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 98% Information theory · 2% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
actor-critic methods |
2.1 | 3 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 q-Learning in Continuous Time · J. Mach. Learn. Res. 2023 Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
1.5 | 2 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 q-Learning in Continuous Time · J. Mach. Learn. Res. 2023 |
Machine learning › Reinforcement learning
maximum entropy reinforcement learning |
1.3 | 2 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020 |
Machine learning › Reinforcement learning
policy evaluation |
1.1 | 2 | 2022 | Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022 Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation |
0.9 | 1 | 2025 | Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.8 | 2 | 2023 | Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022 q-Learning in Continuous Time · J. Mach. Learn. Res. 2023 |
Machine learning › Reinforcement learning › value-based reinforcement learning › q-learning
continuous-time q-learning |
0.7 | 1 | 2023 | q-Learning in Continuous Time · J. Mach. Learn. Res. 2023 |
Machine learning › Reinforcement learning
temporal difference learning |
0.6 | 1 | 2022 | Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning
value function approximation |
0.6 | 1 | 2022 | Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning
continuous-time reinforcement learning |
0.4 | 1 | 2020 | Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.4 | 1 | 2020 | Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020 |
Mathematical optimization › control theory › optimal control
stochastic control |
0.1 | 2 | 2020 | Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020 Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997 |
Mathematical optimization › control theory › optimal control
linear quadratic control |
0.1 | 1 | 2020 | Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020 |
Information theory
asymptotic analysis |
0.0 | 1 | 1997 | Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997 |
Mathematical optimization
singular perturbation |
0.0 | 1 | 1997 | Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997 |
Methods — techniques the papers use, named apart from their topics
stochastic approximation · 1.8martingale theory · 1.8ratio estimator · 0.9q-learning · 0.9actor-critic · 0.9stochastic control · 0.9entropy regularization · 0.9value function convergence · 0.0asymptotic optimality · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reward-Directed Score-Based Diffusion Models via q-LearningabstractWe propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions. Different from most existing studies, ours does not involve any pretrained model for the unknown score functions of the noise-perturbed data distributions, nor does it attempt to learn the score functions. Instead, we formulate the problem as entropy-regularized continuous-time RL and show that the optimal stochastic policy has a Gaussian distribution with a known covariance matrix. Based on this result, we parameterize the mean of Gaussian policies and develop an actor--critic type (little) $q$-learning algorithm to solve the RL problem. A key ingredient in our algorithm design is to obtain noisy observations from the unknown score function via a ratio estimator. Our formulation can also be adapted to solve pure score-matching and fine-tuning pretrained models. Numerically, we show the effectiveness of our approach by comparing its performance with two state-of-the-art RL methods that fine-tune pretrained models on several generative tasks including high-dimensional image generations. Finally, we discuss extensions of our RL formulation to probability flow ODE implementation of diffusion models and to conditional diffusion models. Jiale Zha, Xun Yu Zhou |
J. Mach. Learn. Res. | 3 |
| 2023 | q-Learning in Continuous TimeabstractWe study the continuous-time counterpart of Q-learning for reinforcement learning (RL) under the entropy-regularized, exploratory diffusion process formulation introduced by Wang et al. (2020). As the conventional (big) Q-function collapses in continuous time, we consider its first-order approximation and coin the term “(little) q-function". This function is related to the instantaneous advantage rate function as well as the Hamiltonian. We develop a “q-learning" theory around the q-function that is independent of time discretization. Given a stochastic policy, we jointly characterize the associated q-function and value function by martingale conditions of certain stochastic processes, in both on-policy and off-policy settings. We then apply the theory to devise different actor--critic algorithms for solving underlying RL problems, depending on whether or not the density function of the Gibbs measure generated from the q-function can be computed explicitly. One of our algorithms interprets the well-known Q-learning algorithm SARSA, and another recovers a policy gradient (PG) based continuous-time algorithm proposed in Jia and Zhou (2022b). Finally, we conduct simulation experiments to compare the performance of our algorithms with those of PG-based algorithms in Jia and Zhou (2022b) and time-discretized conventional Q-learning algorithms. Yanwei Jia, Xun Yu Zhou |
J. Mach. Learn. Res. | 2 |
| 2022 | Asset selection via correlation blockmodel clustering
Wenpin Tang, Xun Yu Zhou |
Expert Syst. Appl. | 3 |
| 2022 | Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale ApproachabstractWe propose a unified framework to study policy evaluation (PE) and the associated temporal difference (TD) methods for reinforcement learning in continuous time and space. We show that PE is equivalent to maintaining the martingale condition of a process. From this perspective, we find that the mean-square TD error approximates the quadratic variation of the martingale and thus is not a suitable objective for PE. We present two methods to use the martingale characterization for designing PE algorithms. The first one minimizes a “martingale loss function", whose solution is proved to be the best approximation of the true value function in the mean--square sense. This method interprets the classical gradient Monte-Carlo algorithm. The second method is based on a system of equations called the “martingale orthogonality conditions" with test functions. Solving these equations in different ways recovers various classical TD algorithms, such as TD($\lambda$), LSTD, and GTD. Different choices of test functions determine in what sense the resulting solutions approximate the true value function. Moreover, we prove that any convergent time-discretized algorithm converges to its continuous-time counterpart as the mesh size goes to zero, and we provide the convergence rate. We demonstrate the theoretical results and corresponding algorithms with numerical experiments and applications. Yanwei Jia, Xun Yu Zhou |
J. Mach. Learn. Res. | 2 |
| 2022 | Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and AlgorithmsabstractWe study policy gradient (PG) for reinforcement learning in continuous time and space under the regularized exploratory formulation developed by Wang et al. (2020). We represent the gradient of the value function with respect to a given parameterized stochastic policy as the expected integration of an auxiliary running reward function that can be evaluated using samples and the current value function. This representation effectively turns PG into a policy evaluation (PE) problem, enabling us to apply the martingale approach recently developed by Jia and Zhou (2022a) for PE to solve our PG problem. Based on this analysis, we propose two types of actor-critic algorithms for RL, where we learn and update value functions and policies simultaneously and alternatingly. The first type is based directly on the aforementioned representation, which involves future trajectories and is offline. The second type, designed for online learning, employs the first-order condition of the policy gradient and turns it into martingale orthogonality conditions. These conditions are then incorporated using stochastic approximation when updating policies. Finally, we demonstrate the algorithms by simulations in two concrete examples. Yanwei Jia, Xun Yu Zhou |
J. Mach. Learn. Res. | 2 |
| 2020 | Reinforcement Learning in Continuous Time and Space: A Stochastic Control ApproachabstractWe consider reinforcement learning (RL) in continuous time with continuous feature and action spaces. We motivate and devise an exploratory formulation for the feature dynamics that captures learning under exploration, with the resulting optimization problem being a revitalization of the classical relaxed stochastic control. We then study the problem of achieving the best trade-off between exploration and exploitation by considering an entropy-regularized reward function. We carry out a complete analysis of the problem in the linear--quadratic (LQ) setting and deduce that the optimal feedback control distribution for balancing exploitation and exploration is Gaussian. This in turn interprets the widely adopted Gaussian exploration in RL, beyond its simplicity for sampling. Moreover, the exploitation and exploration are captured respectively by the mean and variance of the Gaussian distribution. We characterize the cost of exploration, which, for the LQ case, is shown to be proportional to the entropy regularization weight and inversely proportional to the discount rate. Finally, as the weight of exploration decays to zero, we prove the convergence of the solution of the entropy-regularized LQ problem to the one of the classical LQ problem. Thaleia Zariphopoulou, Xun Yu Zhou |
J. Mach. Learn. Res. | 3 |
| 2003 | Indefinite Stochastic Linear Quadratic Control with Markovian Jumps in Infinite Time Horizon
Xun Li 0002, Xun Yu Zhou, Mustapha Ait Rami |
J. Glob. Optim. | 2 |
| 1997 | Hierarchical production controls in a stochastic two-machine flowshop with a finite internal bufferabstractThe authors present an asymptotic analysis of hierarchical production planning in a manufacturing system with two tandem machines that are subject to breakdown and repair. The buffer between the two machines is assumed finite. Therefore, the number of parts in that buffer needs to be nonnegative and bounded above by the buffer size. As the rate of change in machine state approaches infinity, the analysis results in a limiting problem in which the stochastic machine capacity is replaced by the equilibrium mean capacity. The value function for the original problem is proved to converge to the value function of the limiting problem. Controls for the original problem are constructed from near optimal controls of the limiting problem in a way which guarantees their asymptotic optimality. The convergence rate of the value function of the original problem to that of the limiting problem together with the error estimate for the constructed asymptotically optimal controls are obtained.> Suresh P. Sethi, Qing Zhang 0003, Xun Yu Zhou |
IEEE Trans. Robotics Autom. | 3 |