Xun Yu Zhou

dblp:31/289 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-9908-5697ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 85% Generative modeling · 15%
Theoretical computer science
2 papers
Mathematical optimization · 98% Information theory · 2%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
actor-critic methods
2.132025
Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025
q-Learning in Continuous Time · J. Mach. Learn. Res. 2023
Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
1.522025
Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025
q-Learning in Continuous Time · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning
maximum entropy reinforcement learning
1.322025
Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025
Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020
Machine learning › Reinforcement learning
policy evaluation
1.122022
Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022
Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022
Machine learning › Generative modeling
diffusion model
0.912025
Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation
0.912025
Reward-Directed Score-Based Diffusion Models via q-Learning · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.822023
Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms · J. Mach. Learn. Res. 2022
q-Learning in Continuous Time · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning › value-based reinforcement learning › q-learning
continuous-time q-learning
0.712023
q-Learning in Continuous Time · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning
temporal difference learning
0.612022
Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
value function approximation
0.612022
Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
continuous-time reinforcement learning
0.412020
Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412020
Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020
Mathematical optimization › control theory › optimal control
stochastic control
0.122020
Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020
Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997
Mathematical optimization › control theory › optimal control
linear quadratic control
0.112020
Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach · J. Mach. Learn. Res. 2020
Information theory
asymptotic analysis
0.011997
Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997
Mathematical optimization
singular perturbation
0.011997
Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer · IEEE Trans. Robotics Autom. 1997

Methods — techniques the papers use, named apart from their topics

stochastic approximation · 1.8martingale theory · 1.8ratio estimator · 0.9q-learning · 0.9actor-critic · 0.9stochastic control · 0.9entropy regularization · 0.9value function convergence · 0.0asymptotic optimality · 0.0
YearPublicationVenuePosition
2025 Reward-Directed Score-Based Diffusion Models via q-Learning
abstract
We propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions. Different from most existing studies, ours does not involve any pretrained model for the unknown score functions of the noise-perturbed data distributions, nor does it attempt to learn the score functions. Instead, we formulate the problem as entropy-regularized continuous-time RL and show that the optimal stochastic policy has a Gaussian distribution with a known covariance matrix. Based on this result, we parameterize the mean of Gaussian policies and develop an actor--critic type (little) $q$-learning algorithm to solve the RL problem. A key ingredient in our algorithm design is to obtain noisy observations from the unknown score function via a ratio estimator. Our formulation can also be adapted to solve pure score-matching and fine-tuning pretrained models. Numerically, we show the effectiveness of our approach by comparing its performance with two state-of-the-art RL methods that fine-tune pretrained models on several generative tasks including high-dimensional image generations. Finally, we discuss extensions of our RL formulation to probability flow ODE implementation of diffusion models and to conditional diffusion models.
Jiale Zha, Xun Yu Zhou
J. Mach. Learn. Res.3
2023 q-Learning in Continuous Time
abstract
We study the continuous-time counterpart of Q-learning for reinforcement learning (RL) under the entropy-regularized, exploratory diffusion process formulation introduced by Wang et al. (2020). As the conventional (big) Q-function collapses in continuous time, we consider its first-order approximation and coin the term “(little) q-function". This function is related to the instantaneous advantage rate function as well as the Hamiltonian. We develop a “q-learning" theory around the q-function that is independent of time discretization. Given a stochastic policy, we jointly characterize the associated q-function and value function by martingale conditions of certain stochastic processes, in both on-policy and off-policy settings. We then apply the theory to devise different actor--critic algorithms for solving underlying RL problems, depending on whether or not the density function of the Gibbs measure generated from the q-function can be computed explicitly. One of our algorithms interprets the well-known Q-learning algorithm SARSA, and another recovers a policy gradient (PG) based continuous-time algorithm proposed in Jia and Zhou (2022b). Finally, we conduct simulation experiments to compare the performance of our algorithms with those of PG-based algorithms in Jia and Zhou (2022b) and time-discretized conventional Q-learning algorithms.
Yanwei Jia, Xun Yu Zhou
J. Mach. Learn. Res.2
2022 Asset selection via correlation blockmodel clustering
Wenpin Tang, Xun Yu Zhou
Expert Syst. Appl.3
2022 Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach
abstract
We propose a unified framework to study policy evaluation (PE) and the associated temporal difference (TD) methods for reinforcement learning in continuous time and space. We show that PE is equivalent to maintaining the martingale condition of a process. From this perspective, we find that the mean-square TD error approximates the quadratic variation of the martingale and thus is not a suitable objective for PE. We present two methods to use the martingale characterization for designing PE algorithms. The first one minimizes a “martingale loss function", whose solution is proved to be the best approximation of the true value function in the mean--square sense. This method interprets the classical gradient Monte-Carlo algorithm. The second method is based on a system of equations called the “martingale orthogonality conditions" with test functions. Solving these equations in different ways recovers various classical TD algorithms, such as TD($\lambda$), LSTD, and GTD. Different choices of test functions determine in what sense the resulting solutions approximate the true value function. Moreover, we prove that any convergent time-discretized algorithm converges to its continuous-time counterpart as the mesh size goes to zero, and we provide the convergence rate. We demonstrate the theoretical results and corresponding algorithms with numerical experiments and applications.
Yanwei Jia, Xun Yu Zhou
J. Mach. Learn. Res.2
2022 Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms
abstract
We study policy gradient (PG) for reinforcement learning in continuous time and space under the regularized exploratory formulation developed by Wang et al. (2020). We represent the gradient of the value function with respect to a given parameterized stochastic policy as the expected integration of an auxiliary running reward function that can be evaluated using samples and the current value function. This representation effectively turns PG into a policy evaluation (PE) problem, enabling us to apply the martingale approach recently developed by Jia and Zhou (2022a) for PE to solve our PG problem. Based on this analysis, we propose two types of actor-critic algorithms for RL, where we learn and update value functions and policies simultaneously and alternatingly. The first type is based directly on the aforementioned representation, which involves future trajectories and is offline. The second type, designed for online learning, employs the first-order condition of the policy gradient and turns it into martingale orthogonality conditions. These conditions are then incorporated using stochastic approximation when updating policies. Finally, we demonstrate the algorithms by simulations in two concrete examples.
Yanwei Jia, Xun Yu Zhou
J. Mach. Learn. Res.2
2020 Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach
abstract
We consider reinforcement learning (RL) in continuous time with continuous feature and action spaces. We motivate and devise an exploratory formulation for the feature dynamics that captures learning under exploration, with the resulting optimization problem being a revitalization of the classical relaxed stochastic control. We then study the problem of achieving the best trade-off between exploration and exploitation by considering an entropy-regularized reward function. We carry out a complete analysis of the problem in the linear--quadratic (LQ) setting and deduce that the optimal feedback control distribution for balancing exploitation and exploration is Gaussian. This in turn interprets the widely adopted Gaussian exploration in RL, beyond its simplicity for sampling. Moreover, the exploitation and exploration are captured respectively by the mean and variance of the Gaussian distribution. We characterize the cost of exploration, which, for the LQ case, is shown to be proportional to the entropy regularization weight and inversely proportional to the discount rate. Finally, as the weight of exploration decays to zero, we prove the convergence of the solution of the entropy-regularized LQ problem to the one of the classical LQ problem.
Thaleia Zariphopoulou, Xun Yu Zhou
J. Mach. Learn. Res.3
2003 Indefinite Stochastic Linear Quadratic Control with Markovian Jumps in Infinite Time Horizon
Xun Li 0002, Xun Yu Zhou, Mustapha Ait Rami
J. Glob. Optim.2
1997 Hierarchical production controls in a stochastic two-machine flowshop with a finite internal buffer
abstract
The authors present an asymptotic analysis of hierarchical production planning in a manufacturing system with two tandem machines that are subject to breakdown and repair. The buffer between the two machines is assumed finite. Therefore, the number of parts in that buffer needs to be nonnegative and bounded above by the buffer size. As the rate of change in machine state approaches infinity, the analysis results in a limiting problem in which the stochastic machine capacity is replaced by the equilibrium mean capacity. The value function for the original problem is proved to converge to the value function of the limiting problem. Controls for the original problem are constructed from near optimal controls of the limiting problem in a way which guarantees their asymptotic optimality. The convergence rate of the value function of the original problem to that of the limiting problem together with the error estimate for the constructed asymptotically optimal controls are obtained.>
Suresh P. Sethi, Qing Zhang 0003, Xun Yu Zhou
IEEE Trans. Robotics Autom.3