Liangpeng Zhang

dblp:165/3160 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
0.622019
Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019
Increasingly Cautious Optimism for Practical PAC-MDP Exploration · IJCAI 2015
Machine learning › Reinforcement learning › exploration
efficient exploration
0.412019
Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning
markov decision process
0.412019
Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning › dynamic programming
value iteration
0.412019
Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning
bellman equation
0.312017
Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning › value function estimation
overestimation bias
0.312017
Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
value-based reinforcement learning
0.312017
Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
value function estimation
0.312017
Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning › exploration
optimistic exploration
0.212015
Increasingly Cautious Optimism for Practical PAC-MDP Exploration · IJCAI 2015

Methods — techniques the papers use, named apart from their topics

explicit planning · 0.4augmented MDP · 0.4log-normal distribution analysis · 0.3PAC-MDP · 0.2
YearPublicationVenuePosition
2025 A large language model-driven reward design framework via dynamic feedback for reinforcement learning
Shengjie Sun 0002, Runze Liu 0002, Jiafei Lyu, Liangpeng Zhang, Xiu Li 0001
Knowl. Based Syst.5
2019 Explicit Planning for Efficient Exploration in Reinforcement Learning
abstract
Efficient exploration is crucial to achieving good performance in reinforcement learning. Existing systematic exploration strategies (R-MAX, MBIE, UCRL, etc.), despite being promising theoretically, are essentially greedy strategies that follow some predefined heuristics. When the heuristics do not match the dynamics of Markov decision processes (MDPs) well, an excessive amount of time can be wasted in travelling through already-explored states, lowering the overall efficiency. We argue that explicit planning for exploration can help alleviate such a problem, and propose a Value Iteration for Exploration Cost (VIEC) algorithm which computes the optimal exploration scheme by solving an augmented MDP. We then present a detailed analysis of the exploration behaviour of some popular strategies, showing how these strategies can fail and spend O(n^2 md) or O(n^2 m + nmd) steps to collect sufficient data in some tower-shaped MDPs, while the optimal exploration scheme, which can be obtained by VIEC, only needs O(nmd), where n, m are the numbers of states and actions and d is the data demand. The analysis not only points out the weakness of existing heuristic-based strategies, but also suggests a remarkable potential in explicit planning for exploration.
Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001
NeurIPS1
2017 Relief R-CNN: Utilizing Convolutional Features for Fast Object Detection
Guiying Li 0002, Junlong Liu, Chunhui Jiang, Liangpeng Zhang, Minlong Lin, Ke Tang 0001
ISNN (1)4
2017 Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning
abstract
Under/overestimation of state/action values are harmful for reinforcement learning agents. In this paper, we show that a state/action value estimated using the Bellman equation can be decomposed to a weighted sum of path-wise values that follow log-normal distributions. Since log-normal distributions are skewed, the distribution of estimated state/action values can also be skewed, leading to an imbalanced likelihood of under/overestimation. The degree of such imbalance can vary greatly among actions and policies within a single problem instance, making the agent prone to select actions/policies that have inferior expected return and higher likelihood of overestimation. We present a comprehensive analysis to such skewness, examine its factors and impacts through both theoretical and empirical results, and discuss the possible ways to reduce its undesirable effects.
Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001
NIPS1
2015 Increasingly Cautious Optimism for Practical PAC-MDP Exploration
Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001
IJCAI1