Ke Lin 0001

dblp:65/3313-1 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-3429-5877ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Sample-efficient backtrack temporal difference deep reinforcement learning
Qi Liu 0027, Pengbin Chen, Ke Lin 0001, Kaidong Zhao, Jinliang Ding
Knowl. Based Syst.3
2025 Aggressive and robust low-level control and trajectory tracking for quadrotors with deep reinforcement learning
Yunjiang Lou, Ke Lin 0001
Neural Comput. Appl.4
2025 Distributional Policy Gradient With Distributional Value Function
abstract
In this article, we propose a distributional policy-gradient method based on distributional reinforcement learning (RL) and policy gradient. Conventional RL algorithms typically estimate the expectation of return, given a state-action pair. Furthermore, distributional RL algorithms consider the return as a random variable and estimate the return distribution that can characterize the probability of different returns resulted by environmental uncertainties. Thus, the return distribution provides more valuable information than its expectation, leading to superior policies in general. Although distributional RL has been investigated widely in value-based RL methods, very few policy-gradient methods take advantage of distributional RL. To bridge this research gap, we propose a distributional policy-gradient method by introducing a distributional value function to the policy gradient (DVDPG). We estimate the distribution of policy gradient instead of the expectation estimated in conventional policy-gradient RL methods. Furthermore, we propose two policy-gradient value sampling mechanisms to do policy improvement. First, we propose a distribution-probability-sampling method that samples the policy-gradient value according to the quantile probability of return distribution. Second, a uniform sample mechanism is proposed. With our sample mechanisms, the proposed distributional policy-gradient method enhances the stochasticity of the policy gradient, improving the exploration efficiency and benefiting to avoid falling into local optimal solutions. In sparse-reward tasks, the distribution-probability-sampling method outperforms the uniform sample mechanism. In dense-reward tasks, the two sample mechanisms perform similarly. Moreover, we show that the conventional policy-gradient method is a special case of the proposed method. Experimental results on various sparse-reward and dense-reward OpenAI-gym tasks illustrate the efficiency of the proposed method, outperforming baselines in almost environments.
Qi Liu 0027, Yanjie Li 0004, Xiongtao Shi, Ke Lin 0001, Yuecheng Liu, Yunjiang Lou
IEEE Trans. Neural Networks Learn. Syst.4
2024 Almost surely safe exploration and exploitation for deep reinforcement learning with state safety estimation
Ke Lin 0001, Qi Liu 0027, Duantengchuan Li, Xiongtao Shi
Inf. Sci.1
2024 A review of graph-based multi-agent pathfinding solvers: From classical to beyond classical
abstract
Multi-agent pathfinding (MAPF) is a well-studied abstract model for navigation in a multi-robot system, where every robot finds the path to its goal position without any collision. Due to its numerous practical applications of multi-robot systems, MAPF has steadily emerged as a research hotspot. The optimal solution for MAPF is NP-hard. In this paper, we offer a comprehensive analysis of different MAPF solvers. First, we review the cutting-edge solvers of classical MAPF, including optimal, bounded sub-optimal, and unbounded sub-optimal. The performance of some representative classical MAPF solvers is quantitatively compared. In the next part, we summarize the beyond classical MAPF solvers, which try to use the classical MAPF solvers in real-world scenarios. Last, we conclude some challenges that MAPF is experiencing in detail, review recent research on these issues, and make some suggestions for further work.
Yanjie Li 0002, Kejian Yan, Ke Lin 0001, Xinyu Wu 0001
Knowl. Based Syst.5
2024 FHCPL: An Intelligent Fixed-Horizon Constrained Policy Learning System for Risk-Sensitive Industrial Scenario
abstract
In many industrial scenarios, safety is a crucial factor to consider. In this article, we focus on the safe reinforcement learning problem that maximizes total rewards while enabling agents to avoid risks. We propose an intelligent fixed-horizon constrained policy learning (FHCPL) system, which allows agents to obtain high returns while maintaining risk avoidance behaviors. For discrete cases, a two-stage policy iteration algorithm, named fixed-horizon constrained policy iteration, is proposed, in which the safety of the learned policy is guaranteed. In the first stage, a policy that satisfies the safety constraint is obtained. In the second stage, a final learned policy that can get high returns while satisfying the safety constraint is reached. For continuous cases, we present the fixed-horizon constrained policy optimization algorithm. Empirical results demonstrate that, with the advantage of the fixed-horizon risk, the FHCPL achieves superior performance in terms of reward maximization and risk avoidance.
Ke Lin 0001, Duantengchuan Li, Yanjie Li 0004, Xindong Wu 0001
IEEE Trans. Ind. Informatics1
2024 Safe Reinforcement Learning in Autonomous Driving With Epistemic Uncertainty Estimation
abstract
Safety is one of the critical challenges in the autonomous driving task. Recent works address the safety by implementing a safe reinforcement learning (safe RL) mechanism. However, most approaches make conservative decisions without knowing the confidence of the actions, which ultimately causes traffic congestion and low travel efficiency. This paper proposes an uncertainty-augmented Lagrangian safe reinforcement algorithm (Lag-U) to improve exploration and safety performance for autonomous driving. First, epistemic uncertainty is introduced into safe RL by using deep ensemble. We use the estimated epistemic uncertainty to encourage exploration and to learn a risk-sensitive policy by adaptively modifying safety constraints. Second, we facilitate an intervention assurance to choose safer actions based on the quantified epistemic uncertainty during deployment. Experimental results prove that the proposed method outperforms other safe RL baselines. The trained vehicle can make a decent trade-off between high efficiency and avoiding risks, thus preventing ultra-conservative policy.
Zheng Zhang 0051, Qi Liu 0027, Yanjie Li 0004, Ke Lin 0001, Linyu Li 0013
IEEE Trans. Intell. Transp. Syst.4
2024 TAG: Teacher-Advice Mechanism With Gaussian Process for Reinforcement Learning
abstract
Reinforcement learning (RL) still suffers from the problem of sample inefficiency and struggles with the exploration issue, particularly in situations with long-delayed rewards, sparse rewards, and deep local optimum. Recently, learning from demonstration (LfD) paradigm was proposed to tackle this problem. However, these methods usually require a large number of demonstrations. In this study, we present a sample efficient teacher-advice mechanism with Gaussian process (TAG) by leveraging a few expert demonstrations. In TAG, a teacher model is built to provide both an advice action and its associated confidence value. Then, a guided policy is formulated to guide the agent in the exploration phase via the defined criteria. Through the TAG mechanism, the agent is capable of exploring the environment more intentionally. Moreover, with the confidence value, the guided policy can guide the agent precisely. Also, due to the strong generalization ability of Gaussian process, the teacher model can utilize the demonstrations more effectively. Therefore, substantial improvement in performance and sample efficiency can be attained. Considerable experiments on sparse reward environments demonstrate that the TAG mechanism can help typical RL algorithms achieve significant performance gains. In addition, the TAG mechanism with soft actor-critic algorithm (TAG-SAC) attains the state-of-the-art performance over other LfD counterparts on several delayed reward and complicated continuous control environments.
Ke Lin 0001, Duantengchuan Li, Yanjie Li 0004, Qi Liu 0027, Yanrui Jin
IEEE Trans. Neural Networks Learn. Syst.1
2023 Distributional reinforcement learning with epistemic and aleatoric uncertainty estimation
Qi Liu 0027, Ke Lin 0001, Xiongtao Shi, Yunjiang Lou
Inf. Sci.4
2023 A fully distributed adaptive event-triggered control for output regulation of multi-agent systems with directed network
Xiongtao Shi, Qi Liu 0027, Ke Lin 0001
Inf. Sci.4
2022 Multi-perspective social recommendation method with graph representation learning
Hai Liu 0004, Duantengchuan Li, Zhaoli Zhang, Ke Lin 0001, Xiaoxuan Shen, Naixue Xiong, Jiazhang Wang
Neurocomputing5
2022 EDMF: Efficient Deep Matrix Factorization With Review Feature Learning for Industrial Recommender System
abstract
Recommendation accuracy is a fundamental problem in the quality of the recommendation system. In this article, we propose an efficient deep matrix factorization (EDMF) with review feature learning for the industrial recommender system. Two characteristics in user’s review are revealed. First, interactivity between the user and the item, which can also be considered as the former’s scoring behavior on the latter, is exploited in a review. Second, the review is only a partial description of the user’s preferences for the item, which is revealed as the sparsity property. Specifically, in the first characteristic, EDMF extracts the interactive features of onefold review by convolutional neural networks with word-attention mechanism. Subsequently,${L}_{0}$norm is leveraged to constrain the review considering that the review information is a sparse feature, which is the second characteristic. Furthermore, the loss function is constructed by maximuma posterioriestimation theory, where the interactivity and sparsity property are converted as two prior probability functions. Finally, the alternative minimization algorithm is introduced to optimize the loss functions. Experimental results on several datasets demonstrate that the proposed methods, which show good industrial conversion application prospects, outperform the state-of-the-art methods in terms of effectiveness and efficiency.
Hai Liu 0004, Duantengchuan Li, Xiaoxuan Shen, Ke Lin 0001, Jiazhang Wang, Zhaoli Zhang, Naixue Xiong
IEEE Trans. Ind. Informatics5
2022 MFDNet: Collaborative Poses Perception and Matrix Fisher Distribution for Head Pose Estimation
abstract
Head pose estimation suffers from several problems, including low pose tolerance under different disturbances and ambiguity arising from common head pose representation. In this study, a robust three-branch model with triplet module and matrix Fisher distribution module is proposed to address these problems. Based on metric learning, the triplet module employs triplet architecture and triplet loss. It is implemented to maximize the distance between embeddings with different pose pairs and minimize the distance between embeddings with same pose pairs. It can learn a highly discriminate and robust embedding related to head pose. Moreover, the rotation matrix instead of Euler angle and unit quaternion is utilized to represent head pose. An exponential probability density model based on the rotation matrix (referred to as the matrix Fisher distribution) is developed to model head rotation uncertainty. The matrix Fisher distribution can further analyze the head pose, and its maximum likelihood obtained using singular value decomposition provides enhanced accuracy. Extensive experiments executed over AFLW2000 and BIWI datasets demonstrate that the proposed model achieves state-of-the-art performance in comparison with traditional methods.
Hai Liu 0004, Shuai Fang, Zhaoli Zhang, Duantengchuan Li, Ke Lin 0001, Jiazhang Wang
IEEE Trans. Multim.5
2021 Efficient Nodes Representation Learning with Residual Feature Propagation
Duantengchuan Li, Ke Lin 0001
PAKDD (2)3
2021 CARM: Confidence-aware recommender model via review representation learning and historical rating behavior in the online platforms
Duantengchuan Li, Hai Liu 0004, Zhaoli Zhang, Ke Lin 0001, Shuai Fang, Zhifei Li 0009, Naixue Xiong
Neurocomputing4
2020 A novel Domain Adaptive Residual Network for automatic Atrial Fibrillation Detection
Yanrui Jin, Chengjin Qin, Jinlei Liu 0001, Ke Lin 0001, Chengliang Liu 0001
Knowl. Based Syst.4