VLDB 2026 Research / reviewers in the wild / expert
Zhaolong Shen
dblp:57/8133
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Motion planning and robot control · 67% Reinforcement learning · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control › robot control
learning control |
0.9 | 1 | 2025 | DOPT: D-Learning with Off-Policy Target toward Sample Efficiency and Fast Convergence Control · ICRA 2025 |
Robotics › Motion planning and robot control › robot control
lyapunov stability |
0.9 | 1 | 2025 | DOPT: D-Learning with Off-Policy Target toward Sample Efficiency and Fast Convergence Control · ICRA 2025 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.9 | 1 | 2025 | DOPT: D-Learning with Off-Policy Target toward Sample Efficiency and Fast Convergence Control · ICRA 2025 |
Methods — techniques the papers use, named apart from their topics
off-policy learning · 0.9neural network control · 0.9lyapunov theory · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DOPT: D-Learning with Off-Policy Target toward Sample Efficiency and Fast Convergence ControlabstractIn recent times, Lyapunov theory has been in-corporated into learning-based control methods to provide a stability guarantee. However, merely satisfying the Lyapunov conditions does not fully leverage the capabilities of the Neural Network (NN) controller. Furthermore, training an effective Lyapunov candidate requires substantial data, which inherently results in sample inefficiency. To address these limitations, we propose an off-policy variant of the vanilla D-learning method that uses current and historical data to iteratively enhance the NN controller within the framework of Lyapunov theory. Our method outperforms the Deep Deterministic Policy Gradient (DDPG) and D-learning in terms of stability, sample efficiency, and the quality of the trained controllers and Lyapunov candidates. Link to code: github.com/Shenzhaolong1330/DOPT Zhaolong Shen, Quan Quan |
ICRA | 1 |
| 2025 | DL-Clip: Online D-Learning with Clipping Operation for Fast Model-Free Stabilizing ControlabstractIn this paper, we present DL-Clip, an innovative online learning approach for nonlinear stabilizing control that operates without prior knowledge of system dynamics or reward signals, while significantly improving training efficiency. DL-Clip introduces a novel integration of stabilizing control with efficient Reinforcement Learning (RL) training mechanisms. The algorithm uses Lyapunov functions to ensure system stability and employs clipping operations to optimize policy updates, achieving faster convergence. We evaluate the effectiveness of DL-Clip through experiments, including simulations of the inverted pendulum and the Image-Based Visual Servoing (IBVS) for multicopter position stabilization. In addition, we validate the approach through a real flight experiment based on the IBVS problem, demonstrating its practical applicability. Zhaolong Shen, Quan Quan |
IROS | 3 |