VLDB 2026 Research / reviewers in the wild / expert
Liangpeng Zhang
dblp:165/3160
· DBLP profile ↗
5ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
0.6 | 2 | 2019 | Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019 Increasingly Cautious Optimism for Practical PAC-MDP Exploration · IJCAI 2015 |
Machine learning › Reinforcement learning › exploration
efficient exploration |
0.4 | 1 | 2019 | Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
markov decision process |
0.4 | 1 | 2019 | Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
0.4 | 1 | 2019 | Explicit Planning for Efficient Exploration in Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
bellman equation |
0.3 | 1 | 2017 | Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017 |
Machine learning › Reinforcement learning › value function estimation
overestimation bias |
0.3 | 1 | 2017 | Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.3 | 1 | 2017 | Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017 |
Machine learning › Reinforcement learning
value function estimation |
0.3 | 1 | 2017 | Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning · NIPS 2017 |
Machine learning › Reinforcement learning › exploration
optimistic exploration |
0.2 | 1 | 2015 | Increasingly Cautious Optimism for Practical PAC-MDP Exploration · IJCAI 2015 |
Methods — techniques the papers use, named apart from their topics
explicit planning · 0.4augmented MDP · 0.4log-normal distribution analysis · 0.3PAC-MDP · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A large language model-driven reward design framework via dynamic feedback for reinforcement learning
Shengjie Sun 0002, Runze Liu 0002, Jiafei Lyu, Liangpeng Zhang, Xiu Li 0001 |
Knowl. Based Syst. | 5 |
| 2019 | Explicit Planning for Efficient Exploration in Reinforcement LearningabstractEfficient exploration is crucial to achieving good performance in reinforcement learning. Existing systematic exploration strategies (R-MAX, MBIE, UCRL, etc.), despite being promising theoretically, are essentially greedy strategies that follow some predefined heuristics. When the heuristics do not match the dynamics of Markov decision processes (MDPs) well, an excessive amount of time can be wasted in travelling through already-explored states, lowering the overall efficiency. We argue that explicit planning for exploration can help alleviate such a problem, and propose a Value Iteration for Exploration Cost (VIEC) algorithm which computes the optimal exploration scheme by solving an augmented MDP. We then present a detailed analysis of the exploration behaviour of some popular strategies, showing how these strategies can fail and spend O(n^2 md) or O(n^2 m + nmd) steps to collect sufficient data in some tower-shaped MDPs, while the optimal exploration scheme, which can be obtained by VIEC, only needs O(nmd), where n, m are the numbers of states and actions and d is the data demand. The analysis not only points out the weakness of existing heuristic-based strategies, but also suggests a remarkable potential in explicit planning for exploration. Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001 |
NeurIPS | 1 |
| 2017 | Relief R-CNN: Utilizing Convolutional Features for Fast Object Detection
Guiying Li 0002, Junlong Liu, Chunhui Jiang, Liangpeng Zhang, Minlong Lin, Ke Tang 0001 |
ISNN (1) | 4 |
| 2017 | Log-normality and Skewness of Estimated State/Action Values in Reinforcement LearningabstractUnder/overestimation of state/action values are harmful for reinforcement learning agents. In this paper, we show that a state/action value estimated using the Bellman equation can be decomposed to a weighted sum of path-wise values that follow log-normal distributions. Since log-normal distributions are skewed, the distribution of estimated state/action values can also be skewed, leading to an imbalanced likelihood of under/overestimation. The degree of such imbalance can vary greatly among actions and policies within a single problem instance, making the agent prone to select actions/policies that have inferior expected return and higher likelihood of overestimation. We present a comprehensive analysis to such skewness, examine its factors and impacts through both theoretical and empirical results, and discuss the possible ways to reduce its undesirable effects. Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001 |
NIPS | 1 |
| 2015 | Increasingly Cautious Optimism for Practical PAC-MDP Exploration
Liangpeng Zhang, Ke Tang 0001, Xin Yao 0001 |
IJCAI | 1 |