Quan Liu 0004

dblp:67/6917-4 · DBLP profile ↗
← Back
56ranked-venue papers
2as first author
41since 2021 · last 2026
0000-0002-8710-1810ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Dynamic masked attention-based contrastive learning for Multi-Agent Reinforcement Learning
Quan Liu 0004, Renyang You
Eng. Appl. Artif. Intell.2
2026 LECMARL: A cooperative multi-agent reinforcement learning method based on lazy mechanisms and efficient exploration
Quan Liu 0004, Renyang You
Neurocomputing2
2026 Reinforcement learning-based hierarchical control of multi-agent systems under uncertain semi-Markovian switching networks
Renyang You, Quan Liu 0004, Shenghui Guo, Zhiming Cui 0002
Inf. Sci.2
2026 CTG-MARL: Cooperative Transformer Graph Reinforcement Learning with Adaptive Policy Modularization
Quan Liu 0004, Renyang You
Knowl. Based Syst.2
2025 Improved Techniques for Offline Reinforcement Learning: Advantage Value Estimation and Layernorm
abstract
Offline reinforcement learning, which aims to learn an optimal policy from a previously collected static datasets. Due to the overestimation caused by extrapolation error, offline algorithms adopt overly pessimistic approaches, which compromise the generalization ability of the learned policy. To address these issues, we propose a method that adopts a mild constraint learning approach, comprising two components: advantage value estimation and layernorm. The first component estimates the value of actions and then selects valuable actions for imitation, which moderately relaxes the strict conservative learning and the second applies layernorm in the value function network, effectively addressing the overestimation of out-of-distribution(OOD) actions and stabilizing the training process. Experimental results on various tasks in the D4RL MuJoCo benchmark show that, compared to baseline methods, our algorithm achieves better performance in most tasks. Especially, our algorithm exhibits well generalization ability on random, medium-replay, and full-replay datasets.
Xiaosong Liu, Quan Liu 0004
ICASSP2
2025 Diverse Collaboration in Multi-Agent Reinforcement Learning via Self-Adaptive Method
abstract
Multi-Agent Reinforcement Learning (MARL) has shown significant promise in tackling complex cooperative tasks, largely due to parameter sharing among agents. However, while this sharing facilitates teamwork, it can also result in agent homogenization, which limits individualized behaviors. To address this issue, we introduce a novel method called Diverse Collaboration in Multi-Agent Reinforcement Learning via Self-Adaptive Method (DC-SA). DC-SA advances individualized behaviors by maximizing the mutual information between agents’ representations and their trajectories to enhance diversity collaboration. The method employs adaptive weights to balance collaboration and individualization, particularly in scenarios where collaboration is challenging. Our empirical results demonstrate that DC-SA outperforms five baselines on the StarCraft II micromanagement tasks.
Xiang Xue, Quan Liu 0004, Meilong Shi, Yuchao Jin
ICASSP2
2025 Offline-to-Online: Case-Based Knowledge Distillation with Large Language Models for Reinforcement Learning
Quan Liu 0004, Meilong Shi, Zhiming Cui 0002
ICCBR2
2025 A Lyapunov-Based Convex Optimal Control Approach via Input Convex Transformer
Renyang You, Quan Liu 0004, Shenghui Guo, Zhiming Cui 0002
ICIC (10)2
2025 LinFa-Q: Accurate Q-learning with linear function approximation
Zhechao Wang, Qiming Fu 0001, Quan Liu 0004, You Lu 0004, Hongjie Wu, Fuyuan Hu
Neurocomputing4
2025 A Comprehensive Review of Multiagent Reinforcement Learning in Video Games
abstract
Recent advancements in multi-agent reinforcement learning (MARL) have demonstrated its application potential in modern games. Beginning with foundational work and progressing to landmark achievements such as AlphaStar inStarCraft IIand OpenAI Five inDota 2, MARL has proven capable of achieving superhuman performance across diverse game environments through techniques like self-play, supervised learning, and deep reinforcement learning. With its growing impact, a comprehensive review has become increasingly important in this field. This paper aims to provide a thorough examination of MARL's application from turn-based two-agent games to real-time multi-agent video games including popular genres such as Sports games, First-Person Shooter (FPS) games, Real-Time Strategy (RTS) games and Multiplayer Online Battle Arena (MOBA) games. We further analyze critical challenges posed by MARL in video games, including nonstationary, partial observability, sparse rewards, team coordination, and scalability, and highlight successful implementations in games likeRocket League,Minecraft,Quake III Arena,StarCraft II, Dota 2, Honor of Kings,etc. This paper offers insights into MARL in video game AI systems, proposes a novel method to estimate game complexity, and suggests future research directions to advance MARL and its applications in game development, inspiring further innovation in this rapidly evolving field.
Qijin Ji, Xinghong Ling, Quan Liu 0004
IEEE Trans. Games4
2024 Deep Deterministic Strategy Gradient Method Using Plot Experience Playback
abstract
The research on continuous control in reinforcement learning has been a hot topic in recent years. The Deep Deterministic Policy Gradient (DDPG) algorithm performs well in continuous control tasks. DDPG algorithm uses experience replay mechanism to train the network model, in order to further improve the efficiency of experience replay mechanism in the DDPG algorithm, the cumulative reward is used as the transiton classification basis, a Deep Deterministic Policy Gradient with Episodic Experience Replay (EER-DDPG) algorithm is proposed. First of all, the transitions are stored in the unit of episode, and two replay buffers are introduced respectively to classify the transitions according to the cumulative reward. Then, the quality of policy can be improved in network model training period by random samping of the episodes with large cumulative rewards. In the continuous control tasks, this algorithm is verified by experiments, and compared with DDPG algorithm, Trust Region Policy Optimization (TRPO) algorithm and Proximal Policy Optization (PPO) algorithm. The experimental results show that EER-DDPG algorithm has better performance.
Quan Liu 0004, Jianxing Zhang, Bofeng Zhang
ICIS1
2024 Offline Reinforcement Learning with Policy Guidance and Uncertainty Estimation
abstract
Offline reinforcement learning is an approach for transforming static datasets into powerful decision engines. It cannot interact with the environment online, which leads to distribution shifts. Previous approaches addressed this problem by making the current policy as close as possible to the behavior policy. However, this type of approach severely limits the generalization ability of Q-functions. To address the above concerns, offline reinforcement learning with policy guidance and uncertainty estimation (PGUE) is proposed. PGUE proposes a fine-grained adjustment approach that improves the generalization ability of Q-functions using a perturbation model. The enhancement of the out-of-distribution generalization of Q-functions is achieved through the implicit guidance of the state space via a deterministic latent policy. Meanwhile, integrating uncertainty estimation into the loss function improves the in-distribution generalization of Q-functions. On the D4RL benchmark, PGUE has better performance than baselines. Moreover, we verify the state distribution and its out-of-distribution generalization ability.
Quan Liu 0004, Lihua Zhang 0006
ICASSP2
2024 Offline Reinforcement Learning with Generative Adversarial Networks and Uncertainty Estimation
abstract
In recent years, offline reinforcement learning has attracted considerable attention in artificial intelligence. By generating a static dataset through a behavior policy, it is unable to engage in online interactions with the environment. However, this inevitably leads to states or actions undergoing inherent distribution shifts. To address the above concerns, offline reinforcement learning with generative adversarial networks and uncertainty estimation (GANUE) is proposed. This approach parameterizes the conditional generative adversarial networks and avoids the over-constraint problem that would be introduced when using a distance metric. Estimating the Q-value through ensemble uncertainty not only relaxes the policy constraint strength but also enhances the out-of-distribution generalization ability. The experimental results show that GANUE significantly performs better on the Maze2D and Adroit tasks than multiple baselines. In addition, we perform sensitivity analysis experiments for the parameters of the perturbation model.
Quan Liu 0004, Lihua Zhang 0006
ICASSP2
2024 Offline Reinforcement Learning Based on Next State Supervision
abstract
Offline reinforcement learning aims to maximize the use of static offline datasets to train agents without interacting with the environment. For the problem of distribution shifts, most approaches avoid out-of-distribution actions through strong constraints, and do not consider generalization and learning within these domains. Based on this problem, we propose a novel method offline reinforcement learning based on next state supervision (NSS), which consists of two main components, the guidance policy and an adaptive coefficient. The guidance policy outputs the next-state with the highest value within a certain range around the current state and an adaptive coefficient regulates the weight of the penalty term in the learned policy. Empirical studies show that the method improves the performance of the baseline method with constraints while having some generalization ability.
Quan Liu 0004, Lihua Zhang 0006
ICASSP2
2024 Multi-Agent Self-Motivated Learning via Role Representation
abstract
In collaborative multi-agent reinforcement learning (MARL), agents need to reach good collaboration in an evolving environment. The introduction of the concept of role is an effective method to recognize the complex relationship between agents and invisibly guide collaboration, but current methods make it difficult to learn excellent policies in dynamic environments. Therefore, we propose a Self-Motivated learning via Role representations (SMR) frame. Firstly, we generate compact role representations based on the trajectories of agents and achieve rational dynamic role assignments by maximizing mutual information. Second, we use an attention mechanism to make action predictions based on observations and role representations and then generate intrinsic rewards based on their uncertainty about the dynamic environment, which in turn induces implicit collaboration. We conduct experiments and visualizations on different challenging benchmark platforms, and the experimental results show the superiority of our method.
Yuchao Jin, Quan Liu 0004
IJCNN2
2024 Balanced Subgoals Generation in Hierarchical Reinforcement Learning
abstract
Hierarchical Reinforcement Learning(HRL) has exhibited potential in scaling reinforcement learning methods. However, it encounters difficulties to expand in tasks with sparse external rewards. Among them, the issues of exploration inefficiency and non-stationarity are serious. In this paper, we propose a novel HRL approach in which we design two measures for subgoals: balance and potentiality. By generating a new subgoal and strategically extending it toward more promising regions, we enhance exploration efficiency by striking a balance in the process. Lastly, leveraging the two measures, we employ an active exploration policy to prevent the introduction of low-level intrinsic rewards, which naturally avoids non-stationary problems. Experimental results demonstrate that our approach surpasses the current state-of-the-art HRL baselines in MuJoCo tasks with sparse rewards.
Sifeng Tong, Quan Liu 0004
SMC2
2024 Hierarchical reinforcement learning with unlimited option scheduling for sparse rewards in continuous spaces
Quan Liu 0004, Fei Zhu 0003, Lihua Zhang 0006
Expert Syst. Appl.2
2024 Taking complementary advantages: Improving exploration via double self-imitation learning in procedurally-generated environments
Fanzhang Li, Quan Liu 0004, Bangjun Wang, Fei Zhu 0003
Expert Syst. Appl.4
2024 RLUC: Strengthening robustness by attaching constraint considerations to policy network
Jianmin Tang, Quan Liu 0004, Fanzhang Li, Fei Zhu 0003
Expert Syst. Appl.2
2024 Personalized federated reinforcement learning: Balancing personalization and experience sharing via distance constraint
Weicheng Xiong, Quan Liu 0004, Fanzhang Li, Bangjun Wang, Fei Zhu 0003
Expert Syst. Appl.2
2024 Deep Luenberger observer-based consistency tracking for nonlinear heterogeneous multi-agent systems with uncertain drift dynamics
Renyang You, Quan Liu 0004
Knowl. Based Syst.2
2023 Learning Unbiased Rewards with Mutual Information in Adversarial Imitation Learning
abstract
A powerful method for automated decision systems is Adversarial Imitation Learning (AIL). It is based on a generative adversarial framework that alternately optimizes a generator (learner) and a discriminator (reward function). In the popular mind, a high-accuracy discriminator results in informative rewards, thus AIL must balance the performance between a generator and a discriminator. However, we find that the AIL reward function is biased and also results in informative rewards. Thus, we theoretically analyze the bias in the AIL reward function and find that balancing the performance of a generator and a discriminator is not necessary when we recover an unbiased reward function. Further, we propose a mutual information based auxiliary reward function. Experiments on continuous control tasks indicate that MI-GAIL is able to address the bias problem of the AIL reward function and further improve sample efficiency and training stability compared with up-to-date algorithms.
Lihua Zhang 0006, Quan Liu 0004
ICASSP2
2023 A Perturbation-Based Policy Distillation Framework with Generative Adversarial Nets
abstract
We study the problem of imitation learning in automated decision systems, in which a learner is trained to imitate an expert demonstrator. A widely used method is adversarial imitation learning that alternately optimizes a generator (learner) and a discriminator (reward function). However, the discriminator is biased during the initial and intermediate training stages. Consequently, the gradient descent direction of the learner is misguided, which leads to unstable training and sample complexity. In this paper, we propose deep imitation learning through a guidance-based policy distillation (GIL) algorithm. First, GIL proposes a teacher model, the guidance-based variational autoencoder, which is pre-trained with expert demonstrations. Then, GIL proposes a perturbation-based policy distillation method that uses the teacher model to guide the learner in the correct optimization direction, enabling the learner to imitate the expert policy with fewer detours. The experimental results show that our approach achieve higher sample efficiency compared with multiple baselines.
Lihua Zhang 0006, Quan Liu 0004, Xiongzhen Zhang, Yapeng Xu
ICASSP2
2023 Efficient Collaboration via Interaction Information in Multi-agent System
Meilong Shi, Quan Liu 0004
ICONIP (2)2
2023 Self-adaptive Inverse Soft-Q Learning for Imitation
Quan Liu 0004, Xiongzhen Zhang
ICONIP (9)2
2023 Cosine Similarity Based Representation Learning for Adversarial Imitation Learning
abstract
Adversarial imitation learning (AIL) aims to recover the reward signal from expert demonstrations and learn expert policy by employing reward and reinforcement learning. However, the raw state-action features of the demonstrations usually have redundant information for a particular control task, and therefore the reward learned from the raw features is often biased, which eventually results in low sample efficiency and instability in AIL. To address these issues, we present CSAIL: Cosine Similarity based Adversarial Imitation Learning. CSAIL extracts expert policy representations from demonstrations via a novel cosine similarity based loss and recovers a robust and unbiased reward function by the learned representations. Based on the reward, CSAIL mimics the expert policy by the Wasserstein distance optimization method. Experimental results show that CSAIL outperforms existing state-of-the-art AIL methods on challenging Mujoco robot control and autonomous driving tasks.
Xiongzhen Zhang, Quan Liu 0004, Lihua Zhang 0006
SMC2
2023 Temporal-difference emphasis learning with regularized correction for off-policy evaluation and control
Jiaqing Cao, Quan Liu 0004, Qiming Fu 0001
Appl. Intell.2
2023 Hierarchical reinforcement learning with adaptive scheduling for robot control
Quan Liu 0004, Fei Zhu 0003
Eng. Appl. Artif. Intell.2
2023 Learning fair representations for accuracy parity
Tangkun Quan, Fei Zhu 0003, Quan Liu 0004, Fanzhang Li
Eng. Appl. Artif. Intell.3
2023 A stable actor-critic algorithm for solving robotic tasks with multiple constraints
Fei Zhu 0003, Quan Liu 0004, Xinghong Ling
Frontiers Comput. Sci.3
2023 Generalized gradient emphasis learning for off-policy evaluation and control with function approximation
Jiaqing Cao, Quan Liu 0004, Qiming Fu 0001
Neural Comput. Appl.2
2023 Addressing implicit bias in adversarial imitation learning with mutual information
Lihua Zhang 0006, Quan Liu 0004, Fei Zhu 0003
Neural Networks2
2022 Master-Slave Policy Collaboration for Actor-Critic Methods
abstract
Actor-critic methods of deep reinforcement learning are widely used to address continuous control tasks. However, the difficulty in balancing exploration and exploitation, as well as the limitation in the learning efficiency of actors, slow down the overall learning progress, and lead to suboptimal policies. To alleviate these problems, we introduce a policy collaboration mechanism using the master-slave architecture (MSPC). At different stages of training, actors are divided into master actors and slave actors, where the master actors with better performance dominate the training, and the slave actors extract the knowledge of the master actors through policy distillation, thus improving the learning and collaborating efficiency for all actors. Moreover, we propose a new experience replay mechanism, called CER, to further improve the exploration ability and performance of master actors. Finally, we demonstrate empirically the advantages of MSPC by applying it to existing state-of-the-art actor-critic methods in Mujoco robot simulation tasks. We also provide a demonstration showing that MSPC+CER improves over MSPC for sample efficiency and learning speed.
Xiaomu Li, Quan Liu 0004
IJCNN2
2022 Vital node searcher: find out critical node measure with deep reinforcement learning
abstract
How to find the critical nodes in the network structure quickly and accurately is a topic of network science. Various algorithms for critical nodes already exist, of which, however, some are with high time complexity and the rest are limited in application range. To solve this problem, an algorithm, referred to as Vital Node Searcher (VNS), is proposed, which discovers critical nodes from a network based on deep reinforcement learning. The VNS method first takes advantage of the Graph Embedding to downscale the feature information of the target network, and then uses the deep Q network method to extract the critical node sequence. A Long-Short Term network module is designed and applied to fully exploit historical information that is contained in the sequence data. Moreover, a duelling Q network module is developed to enhance the precision of prediction. Both in terms of time complexity and performance, the VNS method is superior compared with other methods, which are validated by experiments of real world datasets. Moreover, VNS method has strong generalisation performance and can be applied to different types of critical node problems. The VNS method performed experiments on four datasets and obtained ANC scores that outperformed the other models respectively. The experiment results demonstrated that the VNS method had a stable and effective performance on finding out the critical node sequence.
Guanting Du, Fei Zhu 0003, Quan Liu 0004
Connect. Sci.3
2022 Improving deep reinforcement learning by safety guarding model via hazardous experience planning
Fei Zhu 0003, Xinghong Ling, Quan Liu 0004
Frontiers Comput. Sci.5
2022 Learning fair representations by separating the relevance of potential information
Tangkun Quan, Fei Zhu 0003, Xinghong Ling, Quan Liu 0004
Inf. Process. Manag.4
2022 Best-in-class imitation: Non-negative positive-unlabeled imitation learning from imperfect demonstrations
Fei Zhu 0003, Xinghong Ling, Quan Liu 0004
Inf. Sci.4
2021 Improving exploration efficiency of deep reinforcement learning through samples produced by generative model
Dayong Xu, Fei Zhu 0003, Quan Liu 0004
Expert Syst. Appl.3
2021 Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
Jiaqing Cao, Quan Liu 0004, Fei Zhu 0003, Qiming Fu 0001
Inf. Sci.2
2021 ARAIL: Learning to rank from incomplete demonstrations
Dayong Xu, Fei Zhu 0003, Quan Liu 0004
Inf. Sci.3
2021 Self-guided deep deterministic policy gradient with multi-actor
Quan Liu 0004
Neural Comput. Appl.2
2019 Efficient reinforcement learning in continuous state and action spaces with Dyna and policy approximation
Quan Liu 0004, Zongzhang Zhang, Qiming Fu 0001
Frontiers Comput. Sci.2
2018 Accurate Q-Learning
Yubin Jiang, Xinghong Ling, Quan Liu 0004
ICONIP (3)4
2018 Deep Deterministic Policy Gradient with Clustered Prioritized Sampling
Fei Zhu 0003, Yuchen Fu, Quan Liu 0004
ICONIP (2)4
2018 Policy Space Noise in Deep Deterministic Policy Gradient
Quan Liu 0004
ICONIP (2)2
2018 Single Trajectory Learning: Exploration Versus Exploitation
abstract
In reinforcement learning (RL), the exploration/exploitation (E/E) dilemma is a very crucial issue, which can be described as searching between the exploration of the environment to find more profitable actions, and the exploitation of the best empirical actions for the current state. We focus on the single trajectory RL problem where an agent is interacting with a partially unknown MDP over single trajectories, and try to deal with the E/E in this setting. Given the reward function, we try to find a good E/E strategy to address the MDPs under some MDP distribution. This is achieved by selecting the best strategy in mean over a potential MDP distribution from a large set of candidate strategies, which is done by exploiting single trajectories drawn from plenty of MDPs. In this paper, we mainly make the following contributions: (1) We discuss the strategy-selector algorithm based on formula set and polynomial function. (2) We provide the theoretical and experimental regret analysis of the learned strategy under an given MDP distribution. (3) We compare these methods with the “state-of-the-art” Bayesian RL method experimentally.
Qiming Fu 0001, Quan Liu 0004, Hongjie Wu
Int. J. Pattern Recognit. Artif. Intell.2
2016 Sparse Kernel-Based Least Squares Temporal Difference with Prioritized Sweeping
Cijia Sun, Xinghong Ling, Yuchen Fu, Quan Liu 0004, Haijun Zhu, Jianwei Zhai
ICONIP (3)4
2016 Deep Q-Learning with Prioritized Sampling
Jianwei Zhai, Quan Liu 0004, Zongzhang Zhang, Haijun Zhu, Cijia Sun
ICONIP (1)2
2016 A Kernel-Based Sarsa( \lambda ) Algorithm with Clustering-Based Sample Sparsification
Haijun Zhu, Fei Zhu 0003, Yuchen Fu, Quan Liu 0004, Jianwei Zhai, Cijia Sun
ICONIP (3)4
2016 Reasoning and predicting POMDP planning complexity via covering numbers
Zongzhang Zhang, Qiming Fu 0001, Quan Liu 0004
Frontiers Comput. Sci.4
2015 A Bayesian Sarsa Learning Algorithm with Bandit-Based Method
Shuhua You, Quan Liu 0004, Qiming Fu 0001, Fei Zhu 0003
ICONIP (1)2
2015 Intelligent Model Learning Based on Variance for Bayesian Reinforcement Learning
abstract
We consider a modular method to reinforcement learning that represents uncertainty of model parameters by maintaining probability distributions over them. The algorithm we call MBDP (model-based Bayesian dynamic programming) can be decomposed into two parallel types of inference: model learning and policy learning. During learning a model, we update posterior distributions of a model over observations after taking an action in each state. During learning a policy, we solve MDPs by dynamic programming with greedy approximation to make an agent choose behaviors which maximize return under the estimated model. Furthermore, we propose a principled method which utilizes the variance of Dirichlet distributions for determining when to learn and relearn the model. We demonstrate that MBDP can find near optimal policies with high probability by sufficient model learning and experimental results show that MBDP performs better compared with current state-of-the-art methods in reinforcement learning.
Shuhua You, Quan Liu 0004, Zongzhang Zhang
ICTAI2
2015 Learning topic of dynamic scene using belief propagation and weighted visual words approach
Chunping Liu, Shengrong Gong, Yi Ji 0001, Quan Liu 0004
Soft Comput.5
2014 Protein-protein interaction network constructing based on text mining and reinforcement learning with application to prostate cancer
abstract
As a notoriously lethal human disease, cancer has obtained much concern for a long time. There have accumulated huge amounts of literature and experimental data on cancer-related research. It is impossible for people to deal with these texts manually to discover novel information and knowledge. However, text mining has an advantage of extracting previously unknown and understandable knowledge from large amounts of texts, and forming well-defined knowledge, providing the possibility to fully taking use of the existed texts. With the proceeding of biomedical research, people have gradually realized that complex biological functions and the phenomenon of life are the results of complex interactions among a variety of biological entities, such as protein. Deeply studying protein interaction network is essential to understand life. We, adopting reinforcement learning idea, put forward an algorithm for protein interaction network constructing. With the algorithm, nodes are used to represent proteins and edges denote interactions. During the evolutionary process, a node selects with which nodes in the network it tends to interact. Keep selecting and carrying on iteration, until eventually attaining an optimal network. The network is the result of the dynamic nature of learning behavior. As a malignancy, prostate cancer has been concerned for a long time. We attain biological texts from PubMed and establish a prostate cancer protein interaction networks by the proposed methods. The results show that our proposed method is pretty good. Network topology analysis results also show that the network node degree distribution is scale-free.
Fei Zhu 0003, Quan Liu 0004, Bairong Shen
BIBM2
2013 The second order temporal difference error for Sarsa(λ)
abstract
Traditional reinforcement learning algorithms, such as Q-learning, Q(λ), Sarsa, and Sarsa(λ), update the action value function using temporal difference (TD) error, which is computed by the last action value function. From the perspective of the TD error, and with respect to the problems of low efficiency and slow convergence of the traditional Sarsa(λ) algorithm, this paper defines the nthorder TD Error, applies it in the traditional Sarsa(λ) algorithm, and develops a fast Sarsa(λ) algorithm based on the 2ndorder TD Error. The algorithm adjusts the Q value with the second-order TD Error and broadcasts the TD Error into the whole state-action space, which speeds up the convergence of the algorithm. This paper also analyzes the convergence rate, and under the condition of one-step update, the results show that the number of iteration depends primarily on γ, ε. Finally, using the proposed algorithm on the traditional reinforcement learning problems, the results show that the algorithm has both a faster convergence rate and better convergence performance.
Qiming Fu 0001, Quan Liu 0004, Guixin Chen
ADPRL2
2012 A parallel scheduling algorithm for reinforcement learning in large state space
Quan Liu 0004, Ling Jing
Frontiers Comput. Sci.1