Botao Dong

dblp:343/4091 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-2026-6856ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 93% Optimization for machine learning · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › offline reinforcement learning
conservative q-learning
0.812024
Efficient Offline Reinforcement Learning With Relaxed Conservatism · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Efficient Offline Reinforcement Learning With Relaxed Conservatism · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning › value-based reinforcement learning
q-value estimation
0.812024
Efficient Offline Reinforcement Learning With Relaxed Conservatism · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
Efficient Offline Reinforcement Learning With Relaxed Conservatism · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Optimization for machine learning
convergence analysis
0.212024
Efficient Offline Reinforcement Learning With Relaxed Conservatism · IEEE Trans. Pattern Anal. Mach. Intell. 2024

Methods — techniques the papers use, named apart from their topics

variance reduction · 0.8conservative policy improvement · 0.8
YearPublicationVenuePosition
2026 Horizon-Greedy Q-Ensembles Regularized Decision Transformer: An Offline Reinforcement Learning Approach for Robotic Tasks
abstract
Offline reinforcement learning (RL) offers a promising paradigm for learning policies from precollected datasets. Nonetheless, applying it to robotic control poses significant challenges, including non-Markovian dynamics and the scarcity of high-quality demonstrations, both of which can undermine the performance of existing methods. To address these issues, this work introduces the horizon-GreedyQ-ensemblesRegularized decisionTransformer (GQRT), an offline RL algorithm tailored for robotic tasks. GQRT leverages a transformer-based policy for trajectory modeling, thereby enabling effective long-horizon decision-making in non-Markovian settings. To alleviate the lack of expert demonstrations, we develop a multistep horizon-greedy policy evaluation mechanism that stitches suboptimal sequences into improved trajectories. Furthermore, to cope with mixed-quality demonstrations collected from diverse sources, we incorporate aQ-ensemble with lower confidence bound regularization, which ensures more stable and reliable value estimation. Extensive experiments on robotic locomotion and manipulation benchmarks demonstrate that GQRT achieves state-of-the-art performance, validating its robustness in complex robotic scenarios.
Botao Dong, Xin Dong 0021, Mingxuan Wang, Hongtian Chen
IEEE Trans. Ind. Informatics1
2025 Multilevel Distributed Fuzzy Optimum Policy Iteration Pareto-Nash Equilibrium Seeking of Multiagent Multiobjective General Sum Games
abstract
Seeking the Pareto-Nash equilibrium in multi-agent, multi-objective general-sum games (MMGSG) poses a significant challenge, particularly in accurately capturing individual preferences and adhering to the fairness principle of the solution. To address this issue, this paper introduces, for the first time, a multi-level distributed fuzzy optimum policy iteration (MDFOPI) method for identifying the Pareto-Nash equilibrium point in MMGSG. This approach is grounded in fuzzy optimal membership degrees, and employs fuzzy measures and$\lambda$-mean classification to construct the coupled multi-objective optimum matrix, utilizing the strategy space as the foundation. The Pareto-Nash equilibrium point is sought through the MDFOPI method, with the multi-objective optimal membership degree matrix used to organize the sampled data and integrate the results of multi-objective evaluations. This work rigorously proves the existence of Nash equilibria in MMGSG and establishes the convergence of the MDFOPI method to a fixed point, specifically a Pareto-Nash equilibrium point. The accuracy and practical applicability of the research findings are verified through simulation experiments.
Xiwen Ma, Wei Xie 0009, Botao Dong, Jingsong Yang, Hongtian Chen, Weidong Zhang 0004
IEEE Trans. Fuzzy Syst.3
2025 Historical Decision-Making Regularized Maximum Entropy Reinforcement Learning
abstract
The challenge of the exploration-exploitation dilemma persists in off-policy reinforcement learning (RL) algorithms, impeding the improvement of policy performance and sample efficiency. To tackle this challenge, a novel historical decision-making regularized maximum entropy (HDMRME) RL algorithm is developed to strike the balance between exploration and exploitation. Built upon the maximum entropy RL framework, the historical decision-making regularization method is proposed to enhance the exploitation capability of RL policies. The theoretical analysis involves proving the convergence of HDMRME, investigating the tradeoff between exploration and exploitation of HDMRME, examining the disparity between the Q-function learned through HDMRME and the classic one, and analyzing the suboptimality of the trained policy. The performance of HDMRME is evaluated across various continuous-action control tasks from Mujoco and OpenAI Gym platforms. Comparative experiments demonstrate that HDMRME exhibits superior sample efficiency and achieves more competitive performance compared with other state-of-the-art RL algorithms.
Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen, Weidong Zhang 0004
IEEE Trans. Neural Networks Learn. Syst.1
2025 Visionary Policy Iteration for Continuous Control
abstract
In this article, a novel visionary policy iteration (VPI) framework is proposed to address the continuous-action reinforcement learning (RL) tasks. In VPI, a visionary Q-function is constructed by incorporating the successor state into the standard Q-function. Due to the introduction of the successor state, the proposed visionary Q-function captures information about state transitions within the Markov decision process (MDP), thereby providing a forward-looking perspective that enables a more accurate and foresighted evaluation of potential action outcomes. The relationship between the visionary Q-function and the standard Q-function is analyzed. Subsequently, both the policy evaluation and policy improvement rules in VPI are designed based on the proposed visionary Q-function. The convergence proof for VPI is provided, ensuring that the iterative policy sequence in VPI will converge to the optimal policy. By combining the VPI framework with the twin delayed deep deterministic policy gradient (TD3) algorithm, a visionary TD3 (VTD3) algorithm is developed. The evaluation of VTD3 is performed on multiple continuous-action control tasks from Mujoco and OpenAI Gym platforms. The results of comparative experiments demonstrate that VTD3 can achieve more competitive performance than other state-of-the-art (SOTA) RL approaches. Additionally, the experimental results indicate that VPI enhances decision-making capability, reduces Q-function estimation bias, and improves sample efficiency, thereby boosting the performance of existing RL algorithms.
Botao Dong, Longyang Huang, Xiwen Ma, Hongtian Chen, Weidong Zhang 0004
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Efficient Offline Reinforcement Learning With Relaxed Conservatism
abstract
Offline reinforcement learning (RL) aims at learning an optimal policy from a static offline data set, without interacting with the environment. However, the theoretical understanding of the existing offline RL methods needs further studies, among which the conservatism of the learned Q-function and the learned policy is a major issue. In this article, we propose a simple and efficient offline RL with relaxed conservatism (ORL-RC) framework for addressing this concern by learning a Q-function that is close to the true Q-function under the learned policy. The conservatism of learned Q-functions and policies of offline RL methods is analyzed. The analysis results support that the conservatism can lead to policy performance degradation. We establish the convergence results of the proposed ORL-RC, and the bounds of learned Q-functions with and without sampling errors, respectively, suggesting that the gap between the learned Q-function and the true Q-function can be reduced by executing the conservative policy improvement. A practical implementation of ORL-RC is presented and the experimental results on the D4RL benchmark suggest that ORL-RC exhibits superior performance and substantially outperforms existing state-of-the-art offline RL methods.
Longyang Huang, Botao Dong, Weidong Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Offline Reinforcement Learning With Behavior Value Regularization
abstract
Offline reinforcement learning (offline RL) aims to find task-solving policies from prerecorded datasets without online environment interaction. It is unfortunate that extrapolation errors can cause over-optimistic Q-value estimates when learning with a fixed dataset, limiting the performance of the learned policy. To tackle this issue, this article proposes an offline actor-critic with behavior value regularization (OAC-BVR) method. In the policy evaluation stage, the difference between the Q-function and the value of the behavior policy is considered as the regularization term, driving the learned value function to approach the value of the behavior policy. The convergence of the proposed policy evaluation with behavior value regularization (PE-BVR) and the value function difference are analyzed, respectively. Compared with existing offline actor-critic methods, the proposed OAC-BVR method integrates the value of the behavior policy, thereby simultaneously alleviating over-optimistic Q-value estimates and reducing Q-function bias. Experimental results on the D4RL MuJoCo and Maze2d datasets demonstrate the validity of the proposed PE-BVR and the performance advantage of OAC-BVR over the state-of-the-art offline RL algorithms. The code of OAC-BVR is available at https://github.com/LongyangHuang/OAC-BVR.
Longyang Huang, Botao Dong, Wei Xie 0009, Weidong Zhang 0004
IEEE Trans. Cybern.2
2024 Mild Policy Evaluation for Offline Actor-Critic
abstract
In offline actor-critic (AC) algorithms, the distributional shift between the training data and target policy causes optimistic value estimates for out-of-distribution (OOD) actions. This leads to learned policies skewed toward OOD actions with falsely high values. The existing value-regularized offline AC algorithms address this issue by learning a conservative value function, leading to a performance drop. In this article, we propose a mild policy evaluation (MPE) by constraining the difference between the values of actions supported by the target policy and those of actions contained within the offline dataset. The convergence of the proposed MPE, the gap between the learned value function and the true one, and the suboptimality of the offline AC with MPE are analyzed, respectively. A mild offline AC (MOAC) algorithm is developed by integrating MPE into off-policy AC. Compared with existing offline AC algorithms, the value function gap of MOAC is bounded by the existence of sampling errors. Moreover, in the absence of sampling errors, the true state value function can be obtained. Experimental results on the D4RL benchmark dataset demonstrate the effectiveness of MPE and the performance superiority of MOAC compared to the state-of-the-art offline reinforcement learning (RL) algorithms.
Longyang Huang, Botao Dong, Jinhui Lu, Weidong Zhang 0004
IEEE Trans. Neural Networks Learn. Syst.2