Jie Li 0042

dblp:17/2703-42 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-3718-5593ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Robust guaranteed neural learning-based output tracking control for uncertain nonlinear systems: An uncertainty feedback compensation method
Chengbo Dai, Jie Li 0042, Zhenlong Wu, Donghai Li
Eng. Appl. Artif. Intell.2
2026 VLM-driven causal auditing: A counterfactual framework for revealing causal confusion in end-to-end driving
Guofa Li, Yuhao Wei, Qi Lan, Jie Li 0042, Xiangyun Ren
Expert Syst. Appl.5
2026 Reinforcement Learning for Driving Policy Generalization via Risk-Aware Expert Policy and Conditional Diffusion
abstract
In real-world autonomous driving applications, reinforcement learning (RL) faces several persistent challenges including limited critical training data, low exploration efficiency, and poor policy generalization. These limitations often result in slow convergence or stagnation in local optima, undermining the safety and generalization of intelligent vehicles. In order to address these issues, we propose a generation-enhanced RL framework for challenging decision-making scenarios, which comprises three stages: data acquisition, data utilization, and data generation. Firstly, a risk-aware expert policy is developed to guide exploration on regions proximal to the boundary between safe and unsafe actions during the early phases of training. Secondly, we propose a hierarchical dynamic prioritized replay mechanism to enhance the utilization of collected experiences. The replay buffer dynamically ranks experience samples based on their relevance to current policy updates. By weighting essential transitions more heavily during replay, the agent gains enhanced capability to learn from critical scenarios. Thirdly, to mitigate the scarcity of high-risk transitions in the training data, we propose a conditional generation method based on diffusion model. This model synthesizes diverse and structurally relevant transitions with statistical realism, which is concentrated near the policy decision boundary. By employing the high-risk metric as conditional input, the generation model supplements the replay buffer by covering distributional gaps, leading to policy generalization in rare but safety-critical scenarios. Experimental results demonstrate that our method significantly accelerates policy learning and improves driving performance under complex scenarios. The learned policies also outperform existing baselines, indicating their advantages in terms of safety and generalization.
Guofa Li, Yingchen Wang, Delin Ouyang, Jie Li 0042, Xiangyun Ren
IEEE Trans. Intell. Transp. Syst.4
2026 Learning Optimal Robust Control for Nonlinear Mixed Traffic Under External Disturbance
abstract
The integration of connected and automated vehicles (CAVs) into traffic systems holds potential to mitigate undesired disturbances. Nevertheless, coexisting human-driven vehicles (HDVs) introduce complex behavioral disturbances, which has imposed critical challenges for control robustness. This study develops a computational framework based on policy iteration to derive robust control policies with optimized attenuation performance for nonlinear mixed traffic flow. Specifically, robust$H_{\infty }$control problem is solved by applying the framework of zero-sum game, whose solution at the Nash equilibrium is transformed into a Hamilton–Jacobi (HJ) inequality with a Hamiltonian constraint. For achieving desired attenuation performance, the value function is updated by gradient descent based on counterexamples violating Hamiltonian and monotonicity constraints, where the positive definiteness of the value function is ensured by convex neural networks, facilitating the analysis of control stability via Lyapunov methods. By utilizing constraint gaps, the attenuation level is optimized through the analytical formulae derived from the HJ inequality. The stability and algorithm convergence are proved. Experimental results demonstrate the capability of the learned controller to effectively attenuate disturbance propagation and stabilize mixed traffic flow.
Jie Li 0042, Jiawei Wang 0001, Yangang Ren, Shen Li 0001, Guofa Li, Shengbo Eben Li
IEEE Trans. Intell. Transp. Syst.1
2025 Alternating interaction fusion of Image-Point cloud for Multi-Modal 3D object detection
Guofa Li, Haifeng Lu, Jie Li 0042, Zhenning Li 0001, Qingkun Li, Xiangyun Ren
Adv. Eng. Informatics3
2025 Lightweight Strategies for Decision-Making of Autonomous Vehicles in Lane Change Scenarios Based on Deep Reinforcement Learning
abstract
High-performance vision-based decision-making networks are often limited by hardware capabilities in practical applications. To address this challenge, this study proposes lightweight optimization strategies for decision-making models from the aspects of parameter size, training memory usage, and inference speed. Specifically, an innovative solution is proposed to achieve lightweight parameters. The Video Swin Transformer is employed to simultaneously extract temporal and spatial features, with the network trained using a Prioritized Replay Deep Q-Network (PRDQN) that incorporates risk assessment. To further reduce training memory usage, the Q-target network in PRDQN is removed, and the mellowmax operator is integrated to enhance the training process, resulting in the PRDeepMellow Swin Transformer. After analyzing the inference speed problems encountered by the algorithm in practical applications, the vanilla self-attention is replaced by a linear self-attention based on double softmax, namely Double Softmax Linear Video Swin Transformer (DSLVS Transformer) which improves the inference speed for long sequences. The proposed methods were evaluated across three high-speed lane change scenarios (a static scenario, a dynamic scenario, and a randomly changing scenario). Experimental results demonstrate that the proposed methods can still maintain excellent decision performance after the corresponding lightweight optimizations.
Guofa Li, Yifan Qiu, Qingkun Li, Jie Li 0042, Shengbo Eben Li
IEEE Trans. Intell. Transp. Syst.5
2025 Decision-Making for Autonomous Vehicles in Multi-Scenarios With Global Map Model and Dynamic Safe Topological Structure
abstract
Autonomous vehicles are expected to navigate safely and efficiently in dynamic environments, which requires the seamless integration of efficient global route planning with stable and safe local motion planning. To enhance global route planning, we construct a state-value function to represent the global map and then iterate an inheritable state-transition matrix that can be used to rapidly search globally optimal routes. To improve the safety, stability and adaptability of local motion planning, we propose an adaptive dynamic safe topological structure that combines hierarchical computational layers to decompose decision-making processes and adaptively regulate decision outputs in diverse scenarios. These methods are integrated into a unified framework for cooperative decision-making, which is validated in various scenarios in CARLA. The experimental results demonstrate that the global route planning method efficiently searches optimal routes, while the local motion planning method ensures safe and robust decisions across various driving scenarios and styles. Additionally, the integrated framework achieves effective cooperation between efficient global navigation and effective local decision-making.
Delin Ouyang, Jie Li 0042, Guofa Li, Xiangyun Ren, Qianlei Peng
IEEE Trans. Intell. Transp. Syst.2
2025 Robust Approximate Dynamic Programming for Nonlinear Systems With Both Model Error and External Disturbance
abstract
Model error and external disturbance have been separately addressed by optimizing the definite performance in standard linear control problems. However, the concurrent handling of both introduces uncertainty and nonconvexity into the performance, posing a huge challenge for solving nonlinear problems. This article introduces an additional cost function in the augmented Hamilton-Jacobi-Isaacs (HJI) equation of zero-sum games to simultaneously manage the model error and external disturbance in nonlinear robust performance problems. For satisfying the Hamilton-Jacobi inequality in nonlinear robust control theory under all considered model errors, the relationship between the additional cost function and model uncertainty is revealed. A critic online learning algorithm, applying Lyapunov stabilizing terms and historical states to reinforce training stability and achieve persistent learning, is proposed to approximate the solution of the augmented HJI equation. By constructing a joint Lyapunov candidate about the critic weight and system state, both stability and convergence are proved by the second method of Lyapunov. Theoretical results also show that introducing historical data reduces the ultimate bounds of system state and critic error. Three numerical examples are conducted to demonstrate the effectiveness of the proposed method.
Jie Li 0042, Ryozo Nagamune, Yuhang Zhang 0018, Shengbo Eben Li
IEEE Trans. Neural Networks Learn. Syst.1
2024 Feasible Reachable Policy Iteration
abstract
The goal-reaching tasks with safety constraints are common control problems in real world, such as intelligent driving and robot manipulation. The difficulty of this kind of problem comes from the exploration termination caused by safety constraints and the sparse rewards caused by goals. The existing safe RL avoids unsafe exploration by restricting the search space to a feasible region, the essence of which is the pruning of the search space. However, there are still many ineffective explorations in the feasible region because of the ignorance of the goals. Our approach considers both safety and goals; the policy space pruning is achieved by a function called feasible reachable function, which describes whether there is a policy to make the agent safely reach the goals in the finite time domain. This function naturally satisfies the self-consistent condition and the risky Bellman equation, which can be solved by the fixed point iteration method. On this basis, we propose feasible reachable policy iteration (FRPI), which is divided into three steps: policy evaluation, region expansion, and policy improvement. In the region expansion step, by using the information of agent to reach the goals, the convergence of the feasible region is accelerated, and simultaneously a smaller feasible reachable region is identified. The experimental results verify the effectiveness of the proposed FR function in both improving the convergence speed of better or comparable performance without sacrificing safety and identifying a smaller policy space with higher sample efficiency.
Shentao Qin, Yao Mu 0001, Jie Li 0042, Wenjun Zou, Jingliang Duan, Shengbo Eben Li
ICML4
2024 Optimization Landscape of Policy Gradient Methods for Discrete-Time Static Output Feedback
abstract
In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control, output-feedback control is more prevalent since the underlying state of the system may not be fully observed in many practical settings. This article analyzes the optimization landscape inherent to policy gradient methods when applied to static output feedback (SOF) control in discrete-time LTI systems subject to quadratic cost. We begin by establishing crucial properties of the SOF cost, encompassing coercivity, L -smoothness, and M -Lipschitz continuous Hessian. Despite the absence of convexity, we leverage these properties to derive novel findings regarding convergence (and nearly dimension-free rate) to stationary points for three policy gradient methods, including the vanilla policy gradient method, the natural policy gradient method, and the Gauss-Newton method. Moreover, we provide proof that the vanilla policy gradient method exhibits linear convergence toward local minima when initialized near such minima. This article concludes by presenting numerical examples that validate our theoretical findings. These results not only characterize the performance of gradient descent for optimizing the SOF problem but also provide insights into the effectiveness of general policy gradient methods within the realm of reinforcement learning.
Jingliang Duan, Jie Li 0042, Kai Zhao 0004, Shengbo Eben Li, Lin Zhao 0009
IEEE Trans. Cybern.2
2023 Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control
abstract
The Hamilton-Jacobi-Bellman (HJB) equation serves as the necessary and sufficient condition for the optimal solution to the continuous-time (CT) optimal control problem (OCP). Compared with the infinite-horizon HJB equation, the solving of the finite-horizon (FH) HJB equation has been a long-standing challenge, because the partial time derivative of the value function is involved as an additional unknown term. To address this problem, this study first-time bridges the link between the partial time derivative and the terminal-time utility function, and thus it facilitates the use of the policy iteration (PI) technique to solve the CT FH OCPs. Based on this key finding, the FH approximate dynamic programming (ADP) algorithm is proposed leveraging an actor-critic framework. It is shown that the algorithm exhibits important properties in terms of convergence and optimality. Rather importantly, with the use of multilayer neural networks (NNs) in the actor-critic architecture, the algorithm is suitable for CT FH OCPs toward more general nonlinear and complex systems. Finally, the effectiveness of the proposed algorithm is demonstrated by conducting a series of simulations on both a linear quadratic regulator (LQR) problem and a nonlinear vehicle tracking problem.
Ziyu Lin, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Jie Li 0042, Jianyu Chen 0002, Bo Cheng 0003, Jun Ma 0008
IEEE Trans. Neural Networks Learn. Syst.5
2021 Cover: International Journal of Intelligent Systems, Volume 36 Issue 8 August 2021
abstract
Cover Caption: The cover image is based on the Research Article Direct and indirect reinforcement learning by Yang Guan et al., https://doi.org/10.1002/int.22466.
Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003
Int. J. Intell. Syst.4
2021 Direct and indirect reinforcement learning
abstract
Reinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision-making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal policy by directly maximizing an objective function using gradient descent methods, in which the objective function is usually the expectation of accumulative future rewards. The latter indirectly finds the optimal policy by solving the Bellman equation, which is the sufficient and necessary condition from Bellman's principle of optimality. We study policy gradient (PG) forms of direct and indirect RL and show that both of them can derive the actor–critic architecture and can be unified into a PG with the approximate value function and the stationary state distribution, revealing the equivalence of direct and indirect RL. We employ a Gridworld task to verify the influence of different forms of PG, suggesting their differences and relationships experimentally. Finally, we classify current mainstream RL algorithms using the direct and indirect taxonomy, together with other ones, including value-based and policy-based, model-based and model-free.
Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003
Int. J. Intell. Syst.4
2020 Accelerated Convergence of Time-Splitting Algorithm for MPC using Cross-Node Consensus
abstract
The splitting strategy over prediction horizon of model predictive control (MPC) has the potential to compute optimal action in a parallel way. However, such time-splitting algorithms often lead to very slow convergence speed because the state consensus only happens in each pair of adjacent nodes, i.e. a point-to-point topology. This paper proposes a generic cross-node consensus method to extend the shortcoming of limiting to point-to-point topology for the purpose of accelerating the convergence of time-splitting MPC. The cross-node consensus is realized by predicting the state transition from one node to another using plant prediction model, which can increase the information exchange efficiency in the prediction horizon. The time-splitting optimization algorithm is implemented by combing with alternating directions method of multipliers (ADMM). Simulations with autonomous driving show that this new algorithm significantly reduces the number of iterations in time-splitting MPC, averagely about 81% compared with classic time-splitting technique.
Maierdanjiang Maihemuti, Shengbo Eben Li, Jie Li 0042, Jiaxin Gao 0002, Bo Cheng 0003
IV3
2020 Robust Distributed Consensus Control of Uncertain Multiagents Interacted by Eigenvalue-Bounded Topologies
abstract
The uncertainties arising from the plant model and topologies have been a major challenge in multiagent consensus control. This article presents a distributed robust control method for an uncertain multiagent system with eigenvalue-bounded topologies. The heterogeneity of node dynamics is described as the uncertainties of a linear model with a common certain part. The linear transformation method is adopted to decompose topologically coupled controllers. Then, the linear matrix inequalities (LMIs) technique is used to numerically solve the distributed robust controller problem. It is proved that such a controller is robust stable under the condition that the topology is eigenvalue-bounded. The effectiveness of this method is validated by the simulation of a group of unmanned ground vehicles compared with the LQR controller.
Keqiang Li 0002, Shengbo Eben Li, Feng Gao 0007, Ziyu Lin, Jie Li 0042, Qi Sun 0004
IEEE Internet Things J.5