VLDB 2026 Research / reviewers in the wild / expert
Dazi Li
dblp:89/7174
· DBLP profile ↗
25ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-1610-6558ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement learning control for once-through boiler-turbine units based on mechanistic directed multi-subgraph integration
Chiqiang Liu, Dazi Li |
Neurocomputing | 3 |
| 2026 | STAGE: Spatio-Temporal Aggregation via Graph Embedding for Multi-Agent Reinforcement Learning in Industrial Optimization
Chiqiang Liu, Dazi Li, Xin Xu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | HYGMA: Hypergraph Coordination Networks with Dynamic Grouping for Multi-Agent Reinforcement LearningabstractCooperative multi-agent reinforcement learning faces significant challenges in effectively organizing agent relationships and facilitating information exchange, particularly when agents need to adapt their coordination patterns dynamically. This paper presents a novel framework that integrates dynamic spectral clustering with hypergraph neural networks to enable adaptive group formation and efficient information processing in multi-agent systems. The proposed framework dynamically constructs and updates hypergraph structures through spectral clustering on agents' state histories, enabling higher-order relationships to emerge naturally from agent interactions. The hypergraph structure is enhanced with attention mechanisms for selective information processing, providing an expressive and efficient way to model complex agent relationships. This architecture can be implemented in both value-based and policy-based paradigms through a unified objective combining task performance with structural regularization. Extensive experiments on challenging cooperative tasks demonstrate that our method significantly outperforms state-of-the-art approaches in both sample efficiency and final performance. The code is available at: https://github.com/mysteryelder/HYGMA. Chiqiang Liu, Dazi Li |
ICML | 2 |
| 2025 | A novel label-aware global graph construction method and spiking-coded graph neural network for intelligent process fault diagnosis
Dazi Li, Yurui Zhu, Hamid Reza Karimi |
Neurocomputing | 1 |
| 2025 | A novel dynamic nonlinear non-Gaussian approach for fault detection and diagnosis
Yihan Ma 0002, Dazi Li |
Neurocomputing | 3 |
| 2025 | Adaptive generative adversarial maximum entropy inverse reinforcement learning
Li Song 0003, Dazi Li, Xin Xu 0001 |
Inf. Sci. | 2 |
| 2025 | Enhancing Graph Reconstruction: Uniting Dual-Level Graph Structure With Graph Reinforcement LearningabstractA combinatorial optimization problem is typically regarded as a 1-D sorting problem in most existing research. The representation ignores some information about the problem because of dimension compression. When applying reinforcement learning (RL) to this problem, convolutional neural networks (CNNs) used in conventional RL cannot directly extract the connection information between two elements in the feature matrix. A typical class of combinatorial optimization problems, the job shop scheduling problem (JSSP), is used in this article as an example. Considering the limitations in previous research, this article reexamines the task from the perspective of graph reconstruction and proposes a graph RL (GRL) method that combines a double deep Q-network (DDQN) and graph attention network (GAT) to achieve breakthroughs beyond the constraints of CNN performance. Moreover, a dual-level graph representation structure is constructed to comprehensively learn the features of scheduling information and overcome the difficulty of learning dynamic graphs. Experiments show that the quality of the obtained solution and generalization performance are both improved compared with models based on original deep RL (DRL) algorithms. Dazi Li, Yanyang Bao, Xin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Bioinspired actor-critic algorithm for reinforcement learning interpretation with Levy-Brown hybrid exploration strategy
Xiao Wang 0034, Dazi Li |
Neurocomputing | 2 |
| 2024 | Adaptive Evolutionary Reinforcement Learning with Policy DirectionabstractAbstract Evolutionary Reinforcement Learning (ERL) has garnered widespread attention in recent years due to its inherent robustness and parallelism. However, the integration of Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) remains relatively rudimentary and lacks dynamism, which can impact the convergence performance of ERL algorithms. In this study, a dynamic adaptive module is introduced to balance the Evolution Strategies (ES) and RL training within ERL. By incorporating elite strategies, this module leverages advantageous individuals to elevate the overall population's performance. Additionally, RL strategy updates often lack guidance from the population. To address this, we incorporate the strategies of the best individuals from the population, providing valuable policy direction. This is achieved through the formulation of a loss function that employs either L1 or L2 regularization to facilitate RL training. The proposed framework is referred to as Adaptive Evolutionary Reinforcement Learning (AERL). The effectiveness of our framework is evaluated by adopting Soft Actor-Critic (SAC) as the RL algorithm and comparing it with other algorithms in the MuJoCo environment. The results underscore the outstanding convergence performance of our proposed Adaptive Evolutionary Soft Actor-Critic (AESAC) algorithm. Furthermore, ablation experiments are conducted to emphasize the necessity of these two improvements. It is worth noting that the enhancements in AESAC are realized at the population level, enabling broader exploration and effectively reducing the risk of falling into local optima. Caibo Dong, Dazi Li |
Neural Process. Lett. | 2 |
| 2022 | SLMS-SSD: Improving the balance of semantic and spatial information in object detection
Kunfeng Wang, Shuqin Zhang, Yonglin Tian, Dazi Li |
Expert Syst. Appl. | 5 |
| 2022 | AdaBoost maximum entropy deep inverse reinforcement learning with truncated gradient
Li Song 0003, Dazi Li, Xiao Wang 0002, Xin Xu 0001 |
Inf. Sci. | 2 |
| 2022 | Sparse online maximum entropy inverse reinforcement learning via proximal optimization and truncated gradient
Li Song 0003, Dazi Li, Xin Xu 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Online Sparse Temporal Difference Learning Based on Nested Optimization and Regularized Dual AveragingabstractIn policy evaluation of reinforcement learning tasks, the temporal difference (TD) learning with value function approximation has been widely studied. However, feature representation has a decisive influence on both accuracy of value function approximation and convergence rate. Therefore, it is important to develop the feature selection theory and methods that can efficiently prevent overfitting and improve estimation accuracy in TD learning algorithms. In this article, we propose an online sparse TD learning algorithm for policy evaluation by using$\ell _{1}$-regualrization for feature selection. The per-step-time runtime computational complexity of the proposed algorithm is linear with respect to feature dimension. The loss function is defined as a nested optimization with$\ell _{1}$-regularization penalty, and the solver minimizes two suboptimization problems by running stochastic gradient descent and regularized dual averaging method, alternately. The convergence results for the fixed points are also established. The experiments on benchmarks with high-dimensional features show the abilities of learning and generalization of the proposed algorithms. Tianheng Song, Dazi Li, Xin Xu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Speckle noise removal based on structural convolutional neural networks with feature fusion for medical image
Dazi Li, Kunfeng Wang, Daozhong Jiang, Qibing Jin |
Signal Process. Image Commun. | 1 |
| 2021 | Recursive Least-Squares Temporal Difference With Gradient CorrectionabstractSince the late 1980s, temporal difference (TD) learning has dominated the research area of policy evaluation algorithms. However, the demand for the avoidance of TD defects, such as low data-efficiency and divergence in off-policy learning, has inspired the studies of a large number of novel TD-based approaches. Gradient-based and least-squares-based algorithms comprise the major part of these new approaches. This paper aims to combine advantages of these two categories to derive an efficient policy evaluation algorithm with O( n2) per-time-step runtime complexity. The least-squares-based framework is adopted, and the gradient correction is used to improve convergence performance. This paper begins with the revision of a previous O( n3) batch algorithm, least-squares TD with a gradient correction (LS-TDC) to regularize the parameter vector. Based on the recursive least-squares technique, an O( n2) counterpart of LS-TDC called RC is proposed. To increase data efficiency, we generalize RC with eligibility traces. An off-policy extension is also proposed based on importance sampling. In addition, the convergence analysis for RC as well as LS-TDC is given. The empirical results in both on-policy and off-policy benchmarks show that RC has a higher estimation accuracy than that of RLSTD and a significantly lower runtime complexity than that of LSTDC. Tianheng Song, Dazi Li, Weimin Yang, Kotaro Hirasawa |
IEEE Trans. Cybern. | 2 |
| 2021 | Actor-Critic Learning Control With Regularization and Feature Selection in Policy Gradient EstimationabstractActor-critic (AC) learning control architecture has been regarded as an important framework for reinforcement learning (RL) with continuous states and actions. In order to improve learning efficiency and convergence property, previous works have been mainly devoted to solve regularization and feature learning problem in the policy evaluation. In this article, we propose a novel AC learning control method with regularization and feature selection for policy gradient estimation in the actor network. The main contribution is that ℓ1-regularization is used on the actor network to achieve the function of feature selection. In each iteration, policy parameters are updated by the regularized dual-averaging (RDA) technique, which solves a minimization problem that involves two terms: one is the running average of the past policy gradients and the other is the ℓ1-regularization term of policy parameters. Our algorithm can efficiently calculate the solution of the minimization problem, and we call the new adaptation of policy gradient RDApolicy gradient (RDA-PG). The proposed RDA-PG can learn stochastic and deterministic near-optimal policies. The convergence of the proposed algorithm is established based on the theory of two-timescale stochastic approximation. The simulation and experimental results show that RDA-PG performs feature selection successfully in the actor and learns sparse representations of the actor both in stochastic and deterministic cases. RDA-PG performs better than existing AC algorithms on standard RL benchmark problems with irrelevant features or redundant features. Luntong Li, Dazi Li, Tianheng Song, Xin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Multi-Kernel Online Reinforcement Learning for Path Tracking Control of Intelligent VehiclesabstractPath tracking control of intelligent vehicles has to deal with the difficulties of model uncertainties and nonlinearities. As a class of adaptive optimal control methods, reinforcement learning (RL) has received increasing attention in solving difficult control problems. However, feature representation and online learning ability are two major problems to be solved for learning control of uncertain dynamic systems. In this article, we propose a multi-kernel online RL approach for path tracking control of intelligent vehicles. In the proposed approach, a multiple kernel feature learning framework is designed for online learning control based on dual heuristic programming (DHP) and the new online learning control algorithm is called multi-kernel DHP (MKDHP). In MKDHP, instead of the expert knowledge for selecting and fine-tuning of a suitable kernel function, only a set of basic kernel functions is required to be predefined and the multi-kernel features can be learned for value function approximation in the critic. The simulation studies on path tracking control for intelligent vehicles have been conducted under$S$-curve and urban road conditions. The results demonstrated that compared with other typical path tracking controllers for intelligent vehicles, such as the linear quadratic regulator (LQR), the pure pursuit controller and the ribbon-based controller, the proposed multi-kernel learning controller can achieve better performance in terms of tracking precision and smoothness. Zhenhua Huang 0004, Xin Xu 0001, Shiliang Sun, Dazi Li |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2020 | Sparse Proximal Reinforcement Learning via Nested OptimizationabstractWe consider the tasks of feature selection and policy evaluation based on linear value function approximation in reinforcement learning problems. High-dimension feature vectors and limited number of samples can easily cause over-fitting and computation expensive. To prevent this problem, ℓ1-regularized method obtains sparse solutions and thus improves generalization performance. We propose an efficient ℓ1-regularized recursive least squares-based online algorithm with O(n2) complexity per time-step, termed ℓ1-RC. With the help of nested optimization decomposition, ℓ1-RC solves a series of standard optimization problems and avoids minimizing mean squares projected Bellman error with ℓ1-regularization directly. In ℓ1-RC, we propose RC with iterative refinement to minimize the operator error, and we propose an alternating direction method of multipliers with proximal operator to minimize the fixed-point error. The convergence of ℓ1-RC is established based on ordinary differential equation method and some extensions are also given. In empirical computations, some state-of-the-art ℓ1-regularized methods are chosen as the baselines, and ℓ1-RC are tested in both policy evaluation and learning control benchmarks. The empirical results show the effectiveness and advantages of ℓ1-RC. Tianheng Song, Dazi Li, Qibing Jin, Kotaro Hirasawa |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Regularization in DQN for Parameter-Varying Control Learning Tasks
Dazi Li, Chengjia Lei, Qibing Jin, Min Han 0001 |
ISNN (2) | 1 |
| 2018 | Quasi-Linear Recurrent Neural Network based Identification and Predictive ControlabstractIn this paper, aiming at the cumbersome solution of control law in neural network predictive control algorithm, a quasi-linear neural network identification and predictive control algorithm is proposed. The recurrent neural network is embedded into the quasi-linear model, which can be viewed as a quasi-ARX model macroscopically. In the quasi-linear recurrent neural network predictive control, the solution of the control law only need one-step derivation, which can greatly simplify the solution process of control law. At the same time, the quasi-linear recurrent neural network can effectively restrain the over-fitting problem in the identification process. Theoretical analysis and simulations are given to prove the simplicity and effectiveness of the proposed computing method. Dazi Li, Tianjiao Kang, Jinglu Hu, Min Han 0001, Qibing Jin |
IJCNN | 1 |
| 2018 | Actor-Critic Learning Control Based on ℓ2-Regularized Temporal-Difference Prediction With Gradient CorrectionabstractActor-critic based on the policy gradient (PG-based AC) methods have been widely studied to solve learning control problems. In order to increase the data efficiency of learning prediction in the critic of PG-based AC, studies on how to use recursive least-squares temporal difference (RLS-TD) algorithms for policy evaluation have been conducted in recent years. In such contexts, the critic RLS-TD evaluates an unknown mixed policy generated by a series of different actors, but not one fixed policy generated by the current actor. Therefore, this AC framework with RLS-TD critic cannot be proved to converge to the optimal fixed point of learning problem. To address the above problem, this paper proposes a new AC framework named critic-iteration PG (CIPG), which learns the state-value function of current policy in an on-policy way and performs gradient ascent in the direction of improving discounted total reward. During each iteration, CIPG keeps the policy parameters fixed and evaluates the resulting fixed policy by -regularized RLS-TD critic. Our convergence analysis extends previous convergence analysis of PG with function approximation to the case of RLS-TD critic. The simulation results demonstrate that the -regularization term in the critic of CIPG is undamped during the learning process, and CIPG has better learning efficiency and faster convergence rate than conventional AC learning control methods. Luntong Li, Dazi Li, Tianheng Song, Xin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Learning-Based Predictive Control for Discrete-Time Nonlinear Systems With Stochastic DisturbancesabstractIn this paper, a learning-based predictive control (LPC) scheme is proposed for adaptive optimal control of discrete-time nonlinear systems under stochastic disturbances. The proposed LPC scheme is different from conventional model predictive control (MPC), which uses open-loop optimization or simplified closed-loop optimal control techniques in each horizon. In LPC, the control task in each horizon is formulated as a closed-loop nonlinear optimal control problem and a finite-horizon iterative reinforcement learning (RL) algorithm is developed to obtain the closed-loop optimal/suboptimal solutions. Therefore, in LPC, RL and adaptive dynamic programming (ADP) are used as a new class of closed-loop learning-based optimization techniques for nonlinear predictive control with stochastic disturbances. Moreover, LPC also decomposes the infinite-horizon optimal control problem in previous RL and ADP methods into a series of finite horizon problems, so that the computational costs are reduced and the learning efficiency can be improved. Convergence of the finite-horizon iterative RL algorithm in each prediction horizon and the Lyapunov stability of the closed-loop control system are proved. Moreover, by using successive policy updates between adjoint time horizons, LPC also has lower computational costs than conventional MPC which has independent optimization procedures between two different prediction horizons. Simulation results illustrate that compared with conventional nonlinear MPC as well as ADP, the proposed LPC scheme can obtain a better performance both in terms of policy optimality and computational efficiency. Xin Xu 0001, Hong Chen 0003, Chuanqiang Lian, Dazi Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Sustainable ℓ2-regularized actor-critic based on recursive least-squares temporal difference learningabstractLeast-squares temporal difference learning (LSTD) has been used mainly for improving the data efficiency of the critic in actor-critic (AC). However, convergence analysis of the resulted algorithms is difficult when policy is changing. In this paper, a new AC method is proposed based on LSTD under discount criterion. The method comprises two components as the contribution: (1) LSTD works in an on-policy way to achieve a good convergence property of AC. (2) A sustainable ℓ2-regularization version of recursive LSTD, which is termed as RRLSTD, is proposed to solve the ℓ2-regularization problem of the critic in AC. To reduce the computation complexity of RRLSTD, we propose a fast version that is termed as FRRLSTD. Simulation results show that RRLSTD/FRRLSTD-based AC methods have better learning efficiency and faster convergence rate than conventional AC methods. Luntong Li, Dazi Li, Tianheng Song |
SMC | 2 |
| 2016 | A sequential method using multiplicative extreme learning machine for epileptic seizure detection
Dazi Li, Qianwen Xie, Qibing Jin, Kotaro Hirasawa |
Neurocomputing | 1 |
| 2016 | Kernel-Based Least Squares Temporal Difference With Gradient CorrectionabstractA least squares temporal difference with gradient correction (LS-TDC) algorithm and its kernel-based version kernel-based LS-TDC (KLS-TDC) are proposed as policy evaluation algorithms for reinforcement learning (RL). LS-TDC is derived from the TDC algorithm. Attributed to TDC derived by minimizing the mean-square projected Bellman error, LS-TDC has better convergence performance. The least squares technique is used to omit the size-step tuning of the original TDC and enhance robustness. For KLS-TDC, since the kernel method is used, feature vectors can be selected automatically. The approximate linear dependence analysis is performed to realize kernel sparsification. In addition, a policy iteration strategy motivated by KLS-TDC is constructed to solve control learning problems. The convergence and parameter sensitivities of both LS-TDC and KLS-TDC are tested through on-policy learning, off-policy learning, and control learning problems. Experimental results, as compared with a series of corresponding RL algorithms, demonstrate that both LS-TDC and KLS-TDC have better approximation and convergence performance, higher efficiency for sample usage, smaller burden of parameter tuning, and less sensitivity to parameters. Tianheng Song, Dazi Li, Liulin Cao, Kotaro Hirasawa |
IEEE Trans. Neural Networks Learn. Syst. | 2 |