Jingliang Duan

dblp:208/9091 · DBLP profile ↗
← Back
37ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0002-3697-1576ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Transformer-based explicit model predictive control with variable prediction horizon
Sichao Wu, Xingyu Cao, Fawang Zhang, Guangyuan Yu, Jingliang Duan
Eng. Appl. Artif. Intell.9
2026 A smooth reinforcement learning method for trajectory tracking and collision avoidance of wheeled vehicle
Liangfa Chen, Xujie Song, Wenxuan Wang 0004, Liming Xiao, Shengbo Eben Li, Jingliang Duan
Expert Syst. Appl.9
2026 Multi-agent reinforcement learning with beta distribution for thrust sampling in stochastic orbital pursuit-evasion games
Yuqiao Zhao, Jingliang Duan, Shengbo Eben Li, Chang Liu 0002
Neurocomputing2
2025 ODE-based Smoothing Neural Network for Reinforcement Learning Tasks
abstract
The smoothness of control actions is a significant challenge faced by deep reinforcement learning (RL) techniques in solving optimal control problems. Existing RL-trained policies tend to produce non-smooth actions due to high-frequency input noise and unconstrained Lipschitz constants in neural networks. This article presents a Smooth ODE (SmODE) network capable of simultaneously addressing both causes of unsmooth control actions, thereby enhancing policy performance and robustness under noise condition. We first design a smooth ODE neuron with first-order low-pass filtering expression, which can dynamically filter out high frequency noises of hidden state by a learnable state-based system time constant. Additionally, we construct a state-based mapping function, $g$, and theoretically demonstrate its capacity to control the ODE neuron's Lipschitz constant. Then, based on the above neuronal structure design, we further advanced the SmODE network serving as RL policy approximators. This network is compatible with most existing RL algorithms, offering improved adaptability compared to prior approaches. Various experiments show that our SmODE network demonstrates superior anti-interference capabilities and smoother action outputs than the multi-layer perception and smooth network architectures like LipsNet.
Wenxuan Wang 0004, Xujie Song, Yuming Yin, Liangfa Chen, Jingliang Duan, Shengbo Eben Li
ICLR8
2025 LipsNet++: Unifying Filter and Controller into a Policy Network
abstract
Deep reinforcement learning (RL) is effective for decision-making and control tasks like autonomous driving and embodied AI. However, RL policies often suffer from the action fluctuation problem in real-world applications, resulting in severe actuator wear, safety risk, and performance degradation. This paper identifies the two fundamental causes of action fluctuation: observation noise and policy non-smoothness. We propose LipsNet++, a novel policy network with Fourier filter layer and Lipschitz controller layer to separately address both causes. The filter layer incorporates a trainable filter matrix that automatically extracts important frequencies while suppressing noise frequencies in the observations. The controller layer introduces a Jacobian regularization technique to achieve a low Lipschitz constant, ensuring smooth fitting of a policy function. These two layers function analogously to the filter and controller in classical control theory, suggesting that filtering and control capabilities can be seamlessly integrated into a single policy network. Both simulated and real-world experiments demonstrate that LipsNet++ achieves the state-of-the-art noise robustness and action smoothness. The code and videos are publicly available at https://xjsong99.github.io/LipsNet_v2.
Xujie Song, Liangfa Chen, Wenxuan Wang 0004, Shentao Qin, Yinsong Ma, Jingliang Duan, Shengbo Eben Li
ICML8
2025 Enhanced DACER Algorithm with Multimodal Q-value Distribution for Risk-Sensitive Stochastic Vehicle Environments
abstract
Reinforcement learning demonstrates strong capabilities in handling complex control tasks, especially in the field of autonomous driving where vehicles cope with uncertain environments. Existing reinforcement learning methods attempt to model the value distribution using unimodal, but in this modeling process, a significant amount of the complete distribution information is lost. In response to this problem, we propose the DACER++, an online multimodal distributional RL algorithm. DACER++ characterize the value distribution as multimodal will enhance the accuracy of characterizing the value distribution and improve algorithm performance. We construct the quantiles value network and use quantile regression to approximate the full quantile function of the state-action return distribution. This method allows for the precise modeling of multi-modal distributions, and formulates risk-sensitive policies adaptable to different environment. Then, We integrate quantiles value network with the actor-critic architecture algorithm DACER. Experiments on multi-goal tasks and MuJoCo benchmarks show that DACER++ not only has multimodal policy representation capability, but also achieves state-of-the-art performance. In stochastic vehicle meeting environments, DACER++ can learn different multimodal value distributions and multimodal trajectories according to various risk preferences, including the conservative and aggressive driving style.
Xujie Song, Wenjun Zou, Bin Shuai, Weixian He, Jingliang Duan, Shengbo Eben Li
IV8
2025 Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-Lane Scenarios
abstract
Reinforcement learning (RL), known for its self-evolution capability, offers a promising approach to training high-level autonomous driving systems. However, handling constraints remains a significant challenge for existing RL algorithms, particularly in real-world applications. In this paper, we propose a new safety-oriented training technique called harmonic policy iteration (HPI). At each RL iteration, it first calculates two policy gradients associated with efficient driving and safety constraints, respectively. Then, a harmonic gradient is derived for policy updating, minimizing conflicts between the two gradients and consequently enabling a more balanced and stable training process. Furthermore, we adopt the state-of-the-art DSAC algorithm as the backbone and integrate it with our HPI to develop a new safe RL algorithm, DSAC-H. Extensive simulations in multi-lane scenarios demonstrate that DSAC-H achieves efficient driving performance with near-zero safety constraint violations.
Feihong Zhang, Guojian Zhan, Bin Shuai, Jingliang Duan, Shengbo Eben Li
IV5
2025 Off-policy Reinforcement Learning with Model-based Exploration Augmentation
abstract
Exploration is crucial in Reinforcement Learning (RL) as it enables the agent to understand the environment for better decision-making. Existing exploration methods fall into two paradigms: active exploration, which injects stochasticity into the policy but struggles in high-dimensional environments, and passive exploration, which manages the replay buffer to prioritize under-explored regions but lacks sample diversity. To address the limitation in passive exploration, we propose Modelic Generative Exploration (MoGE), which augments exploration through the generation of under-explored critical states and synthesis of dynamics-consistent experiences. MoGE consists of two components: (1) a diffusion generator for critical states under the guidance of entropy and TD error, and (2) a one-step imagination world model for constructing critical transitions for agent learning. Our method is simple to implement and seamlessly integrates with mainstream off-policy RL algorithms without structural modifications. Experiments on OpenAI Gym and DeepMind Control Suite demonstrate that MoGE, as an exploration augmentation, significantly enhances efficiency and performance in complex tasks.
Xiangteng Zhang, Guojian Zhan, Wenxuan Wang 0004, Jingliang Duan, Shengbo Eben Li
NeurIPS7
2025 Smooth policy iteration for zero-sum Markov Games
Yangang Ren, Yao Lyu, Wenxuan Wang 0004, Shengbo Eben Li, Zeyang Li 0001, Jingliang Duan
Neurocomputing6
2025 Distributional Soft Actor-Critic With Three Refinements
abstract
Reinforcement learning (RL) has shown remarkable success in solving complex decision-making and control tasks. However, many model-free RL algorithms experience performance degradation due to inaccurate value estimation, particularly the overestimation of Q-values, which can lead to suboptimal policies. To address this issue, we previously proposed the Distributional Soft Actor-Critic (DSAC or DSACv1), an off-policy RL algorithm that enhances value estimation accuracy by learning a continuous Gaussian value distribution. Despite its effectiveness, DSACv1 faces challenges such as training instability and sensitivity to reward scaling, caused by high variance in critic gradients due to return randomness. In this paper, we introduce three key refinements to DSACv1 to overcome these limitations and further improve Q-value estimation accuracy: expected value substitution, twin value distribution learning, and variance-based critic gradient adjustment. The enhanced algorithm, termed DSAC with Three refinements (DSAC-T or DSACv2), is systematically evaluated across a diverse set of benchmark tasks. Without the need for task-specific hyperparameter tuning, DSAC-T consistently matches or outperforms leading model-free RL algorithms, including SAC, TD3, DDPG, TRPO, and PPO, in all tested environments. Additionally, DSAC-T ensures a stable learning process and maintains robust performance across varying reward scales. Its effectiveness is further demonstrated through real-world application in controlling a wheeled robot, highlighting its potential for deployment in practical robotic tasks.
Jingliang Duan, Wenxuan Wang 0004, Liming Xiao, Jiaxin Gao 0002, Shengbo Eben Li, Chang Liu 0002, Ya-Qin Zhang, Bo Cheng 0003, Keqiang Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Tractor Semi-Trailer Off-Tracking and Stability Approximate Bi-Level Policy Optimization
abstract
Trajectory tracking control of tractor semi-trailer vehicles poses significant challenges due to inherent off-tracking behavior and roll instability risks. While existing approaches have demonstrated effectiveness, they often rely on computationally intensive numerical solvers and require time-consuming manual tuning of cost function weights. This paper presents an approximate bi-level policy optimization (ABPO) framework that simultaneously optimizes the cost function and synthesizes an explicit control policy to minimize off-tracking while reducing computational complexity. The proposed framework employs a hierarchical structure: the upper level updates cost weights based on the trailer’s stability trajectory, while the lower level derives an approximate optimal policy by solving the tractor’s control problem. By leveraging Pontryagin’s Maximum Principle (PMP), we have developed a novel method to analytically compute cost weight gradients through differentiation of the PMP conditions. This enables the formulation of a related optimal control problem (OCP) whose solutions directly yield gradients for cost parameter updates. The ABPO framework achieves automatic weight coefficient adjustment, enhances trajectory tracking accuracy for both tractor and trailer units, and significantly reduces computational burden. Simulation and experimental validation across 4 classical scenarios demonstrates that the learned policy reduces rearward amplification by 17.82%, lateral tracking errors by 84.15%, and rollover by 64.19%, respectively. Notably, the control policy computation requires less than 10 ms, making it suitable for real-time applications. The source code for the algorithms described in this paper is publicly available at https://github.com/TroyResearch/ABPO.git.
Fawang Zhang, Jingliang Duan, Hui Liu 0001, Xingyu Cao, Shida Nie, Congshuai Guo, Yujia Xie, Jun Ma 0008, Shangli Wang
IEEE Trans Autom. Sci. Eng.2
2025 Multimodal Reinforcement Learning With Score-Based Policy
abstract
Learning multimodal policies is crucial for enhancing exploration in online reinforcement learning (RL), especially in tasks with continuous action spaces and non-convex reward landscapes. While recent diffusion policies show promise, they often suffer from low computational efficiency in online settings. A more training-efficient paradigm involves modeling the policy as a Boltzmann distribution and guiding the action sampling directly with the gradient of the Q-value with respect to the action (proportional to the score function of the policy), such as via Langevin dynamics. However, analysis in this paper reveals that this gradient-guided approach suffers from two critical challenges: sampling instability caused by the widely varying magnitude of action gradients; and mode imbalance, where the sampling process inaccurately represents the weights of different high-value action modes. To address these challenges, this paper introduces three targeted techniques: score normalization and reshaping to stabilize the sampling process, and value-based resampling to correct mode imbalance. These techniques are then integrated into an actor-critic framework, resulting in the Score-Enhanced Actor-Critic (SEAC) algorithm. Simulation and real-world experiments demonstrate that SEAC not only effectively learns multimodal behaviors but also achieves state-of-the-art performance and high computational efficiency compared to prior multimodal RL methods. The code of this paper is available at https://github.com/THUzouwenjun/SEAC.
Wenjun Zou, Bin Shuai, Liming Xiao, Yinsong Ma, Jingliang Duan, Shengbo Eben Li
IEEE Trans Autom. Sci. Eng.7
2025 Feasible Policy Iteration With Guaranteed Safe Exploration
abstract
Safety guarantee is an important topic when training real-world tasks with reinforcement learning (RL). During online environmental exploration, any constraint violation can lead to significant property damage and risks to personnel. Existing safe RL methods either exclusively address safety concerns after reaching optimality or incorporate a certain degree of tolerance for constraint violations during training. This article proposes a feasible policy iteration framework that can guarantee absolute safety during online exploration, i.e., constraint violations never happen in real-world interactions. The key to maintaining absolute safety lies in confining the environmental exploration at each step always within the feasible region of the current policy. This feasible region is described by a newly defined constraint decay function with uncertainty, ensuring the forward invariance of the feasible region under the worst case. Within the proposed framework, the feasible region maintains its monotonic expanding property and converges to its maximum extent, even though only local samples are available, i.e., the agent only has access to samples within the feasible region. Meanwhile, the trained policy also improves monotonically within its corresponding feasible region if one can use different updating rules inside and outside the feasible region. Finally, practical algorithms are designed with the actor-critic-scenery architecture, consisting of three modules: 1) safe exploration; 2) model error estimation; and 3) network update. Experimental results indicate that our algorithms achieve performance comparable to baselines while maintaining zero constraint violation throughout the entire training process. In contrast, the baseline algorithm typically requires thousands of constraint violations to achieve the same performance. These findings suggest a substantial potential for applying feasible policy iteration in real-world tasks, enabling the online evolution of intricate systems.
Yuhang Zhang 0018, Shengbo Eben Li, Yao Lyu, Jingliang Duan, Zhilong Zheng, Dezhao Zhang
IEEE Trans. Cybern.5
2025 Real-Time Resilient Tracking Control for Autonomous Vehicles Through Triple Iterative Approximate Dynamic Programming
abstract
Enhancing control precision, mitigating external disturbances, and ensuring real-time responsiveness stand as the cornerstone of autonomous vehicle tracking endeavors, each of which intricately interwoven to uphold operational safety. In pursuit of addressing these issues, this paper presents a triple iterative control method inspired by approximate dynamic programming (ADP) tailored for real-time disturbance avoidance. The control framework orchestrates simultaneous iterations of value function, control policy, and disturbance policy, engineered to optimize tracking control amidst external disturbances cast as a zero-sum differential game, tackled adeptly through deep neural networks. Rigorous mathematical proof underpins its triple iteration, coupled with assurances of residual error convergence, solidifying its safety guarantee ability and algorithmic resilience. To validate its effectiveness, both numerical simulations and experiments on a real micro-vehicle platform were conducted. Results underscore the feasibility of this new method, showcasing its energy-saving capability and a four-times acceleration compared to conventional model predictive control (MPC) approaches when confronted with lateral disturbances. Notably, the single-step calculation time of this method on the Raspberry Pi is only 1.44ms, affirming its practical viability and real-world applicability.
Jiale Geng, Yunqi Cheng, Liye Tang, Jingliang Duan, Feng Duan 0006, Shengbo Eben Li
IEEE Trans. Intell. Transp. Syst.5
2025 Diffeomorphism-Transformed Iterative Linear Quadratic Regulator for Constrained Motion Planning in Autonomous Driving
abstract
Ensuring safe driving and real-time execution is a crucial requirement in the motion planning process for autonomous vehicles. Hence, there is a compelling demand for advanced motion planning algorithms that exhibit effective management of inequality constraints and exceptional computational performance. This paper investigates a diffeomorphism-transformed iterative linear quadratic regulator (DTiLQR) algorithm for addressing constrained motion planning problems in autonomous vehicles with nonlinear dynamics and multiple inequality constraints. With regard to the state and input constraints, a novel state-and-input diffeomorphism is proposed to transform the constrained state/input space into an unconstrained one. Subsequently, these inequality constraints are systematically incorporated into the vehicle dynamics, thereby leading to the newly constructed system in this context. Then, we reformulate and incorporate the obstacle avoidance constraint into the objective function using state diffeomorphism and logarithmic barrier function. With this, the original optimization problem is converted to the unconstrained counterpart, adhering only to the constructed system dynamics. In this sense, featuring a streamlined single-loop architecture (which is essentially different from the dual-loop algorithmic design of existing constrained iLQR algorithms), DTiLQR is used to solve the optimization problem effectively while maintaining motion performance and constraint satisfaction for the resulting optimal trajectory. Ultimately, case studies across various driving situations showcase the effectiveness and exceptional computational efficiency of the proposed DTiLQR algorithm.
Zicheng Zhu, Haichao Liu 0003, Jingliang Duan, Han Zhao 0007, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.4
2025 Conformal Symplectic Optimization for Stable Reinforcement Learning
abstract
Training deep reinforcement learning (RL) agents necessitates overcoming the highly unstable nonconvex stochastic optimization inherent in the trial-and-error mechanism. To tackle this challenge, we propose a physics-inspired optimization algorithm called relativistic adaptive gradient descent (RAD), which enhances long-term training stability. By conceptualizing neural network (NN) training as the evolution of a conformal Hamiltonian system, we present a universal framework for transferring long-term stability from conformal symplectic integrators to iterative NN updating rules, where the choice of kinetic energy governs the dynamical properties of resulting optimization algorithms. By utilizing relativistic kinetic energy, RAD incorporates principles from special relativity and limits parameter updates below a finite speed, effectively mitigating abnormal gradient influences. In addition, RAD models NN optimization as the evolution of a multiparticle system where each trainable parameter acts as an independent particle with an individual adaptive learning rate. We prove RAD's sublinear convergence under general nonconvex settings, where smaller gradient variance and larger batch sizes contribute to tighter convergence. Notably, RAD degrades to the well-known adaptive moment estimation (ADAM) algorithm when its speed coefficient is chosen as one and symplectic factor as a small positive value. Experimental results show RAD outperforming nine baseline optimizers with five RL algorithms across twelve environments, including standard benchmarks and challenging scenarios. Notably, RAD achieves up to a 155.1% performance improvement over ADAM in Atari games, showcasing its efficacy in stabilizing and accelerating RL training.
Yao Lyu, Xiangteng Zhang, Shengbo Eben Li, Jingliang Duan, Letian Tao, Qing Xu 0010, Keqiang Li 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 Feasible Reachable Policy Iteration
abstract
The goal-reaching tasks with safety constraints are common control problems in real world, such as intelligent driving and robot manipulation. The difficulty of this kind of problem comes from the exploration termination caused by safety constraints and the sparse rewards caused by goals. The existing safe RL avoids unsafe exploration by restricting the search space to a feasible region, the essence of which is the pruning of the search space. However, there are still many ineffective explorations in the feasible region because of the ignorance of the goals. Our approach considers both safety and goals; the policy space pruning is achieved by a function called feasible reachable function, which describes whether there is a policy to make the agent safely reach the goals in the finite time domain. This function naturally satisfies the self-consistent condition and the risky Bellman equation, which can be solved by the fixed point iteration method. On this basis, we propose feasible reachable policy iteration (FRPI), which is divided into three steps: policy evaluation, region expansion, and policy improvement. In the region expansion step, by using the information of agent to reach the goals, the convergence of the feasible region is accelerated, and simultaneously a smaller feasible reachable region is identified. The experimental results verify the effectiveness of the proposed FR function in both improving the convergence speed of better or comparable performance without sacrificing safety and identifying a smaller policy space with higher sample efficiency.
Shentao Qin, Yao Mu 0001, Jie Li 0042, Wenjun Zou, Jingliang Duan, Shengbo Eben Li
ICML6
2024 Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator
Jingliang Duan, Lin Zhao 0009
IJCAI2
2024 Guiding Reinforcement Learning with Incomplete System Dynamics
abstract
Model-free reinforcement learning (RL) is inherently a reactive method, operating under the assumption that it starts with no prior knowledge of the system and entirely depends on trial-and-error for learning. This approach faces several challenges, such as poor sample efficiency, generalization, and the need for well-designed reward functions to guide learning effectively. On the other hand, controllers based on complete system dynamics do not require data. This paper addresses the intermediate situation where there is not enough model information for complete controller design, but there is enough to suggest that a model-free approach is not the best approach either. By carefully decoupling known and unknown information about the system dynamics, we obtain an embedded controller guided by our partial model and thus improve the learning efficiency of an RL-enhanced approach. A modular design allows us to deploy mainstream RL algorithms to refine the policy. Simulation results show that our method significantly improves sample efficiency compared with standard RL methods on continuous control tasks, and also offers enhanced performance over traditional control approaches. Experiments on a real ground vehicle also validate the performance of our method, including generalization and robustness.
Shuyuan Wang, Jingliang Duan, Nathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni, Lixian Zhang 0001
IROS2
2024 Diffusion Actor-Critic with Entropy Regulator
abstract
Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to acquire complex policies. In response to this problem, we propose an online RL algorithm termed diffusion actor-critic with entropy regulator (DACER). This algorithm conceptualizes the reverse process of the diffusion model as a novel policy function and leverages the capability of the diffusion model to fit multimodal distributions, thereby enhancing the representational capacity of the policy. Since the distribution of the diffusion policy lacks an analytical expression, its entropy cannot be determined analytically. To mitigate this, we propose a method to estimate the entropy of the diffusion policy utilizing Gaussian mixture model. Building on the estimated entropy, we can learn a parameter $\alpha$ that modulates the degree of exploration and exploitation. Parameter $\alpha$ will be employed to adaptively regulate the variance of the added noise, which is applied to the action output by the diffusion model. Experimental trials on MuJoCo benchmarks and a multimodal task demonstrate that the DACER algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting a stronger representational capacity of the diffusion policy.
Yuxuan Jiang 0011, Wenjun Zou, Xujie Song, Wenxuan Wang 0004, Liming Xiao, Jingliang Duan, Shengbo Eben Li
NeurIPS10
2024 Safe Model-Based Reinforcement Learning With an Uncertainty-Aware Reachability Certificate
abstract
Safe reinforcement learning (RL) that solves constraint-satisfactory policies provides a promising way to the broader safety-critical applications of RL in real-world problems such as robotics. Among all safe RL approaches, model-based methods reduce training time violations further due to their high sample efficiency. However, lacking safety robustness against the model uncertainties remains an issue in safe model-based RL, especially in training time safety. In this paper, we propose a distributional reachability certificate (DRC) and its Bellman equation to address model uncertainties and characterize robust persistently safe states. Furthermore, we build a safe RL framework to resolve constraints required by the DRC and its corresponding shield policy. We also devise a line search method to maintain safety and reach higher returns simultaneously while leveraging the shield policy. Comprehensive experiments on classical benchmarks such as constrained tracking and navigation indicate that the proposed algorithm achieves comparable returns with much fewer constraint violations during training. Our code is available at https://github.com/ManUtdMoon/Distributional-Reachability-Policy-Optimization.Note to Practitioners—Although it has been proven that RL can be applied in complex robotics control tasks, the training process of an RL control policy induces frequent failures because the agent needs to learn safety through constraint violations. This issue hinders the promotion of RL because a large amount of failure of robots is too expensive to afford. This paper aims to reduce the training-time violations of RL-based control methods, enabling RL to be leveraged in a broader application area. To achieve the goal, we first introduce a safety quantity describing the distribution of potential constraint violations in the long term. By imposing constraints on the quantile of the safety distribution, we can realize safety robust to the model uncertainty, which is necessary for real-world robot learning with environment uncertainty. Second, we further devise a shield policy aiming to minimize the constraint violation. The policy will intervene when the agent is about to violate state constraints, further enhancing exploration safety. Third, we implement a line search method to find an action pursuing near-optimal performance when fulfilling safety requirements strictly. Our experimental results indicate that the proposed algorithm reduces training-time violations significantly while maintaining competitive task performance. We make a step towards applying RL safely in real-world tasks. Our future work includes conducting physical verification on real robots to evaluate the algorithm and improving safety further by starting from an initially safe control policy that comes from domain knowledge.
Dongjie Yu, Wenjun Zou, Haitong Ma, Shengbo Eben Li, Yuming Yin, Jianyu Chen 0002, Jingliang Duan
IEEE Trans Autom. Sci. Eng.8
2024 Optimization Landscape of Policy Gradient Methods for Discrete-Time Static Output Feedback
abstract
In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control, output-feedback control is more prevalent since the underlying state of the system may not be fully observed in many practical settings. This article analyzes the optimization landscape inherent to policy gradient methods when applied to static output feedback (SOF) control in discrete-time LTI systems subject to quadratic cost. We begin by establishing crucial properties of the SOF cost, encompassing coercivity, L -smoothness, and M -Lipschitz continuous Hessian. Despite the absence of convexity, we leverage these properties to derive novel findings regarding convergence (and nearly dimension-free rate) to stationary points for three policy gradient methods, including the vanilla policy gradient method, the natural policy gradient method, and the Gauss-Newton method. Moreover, we provide proof that the vanilla policy gradient method exhibits linear convergence toward local minima when initialized near such minima. This article concludes by presenting numerical examples that validate our theoretical findings. These results not only characterize the performance of gradient descent for optimizing the SOF problem but also provide insights into the effectiveness of general policy gradient methods within the realm of reinforcement learning.
Jingliang Duan, Jie Li 0042, Kai Zhao 0004, Shengbo Eben Li, Lin Zhao 0009
IEEE Trans. Cybern.1
2024 Inverse Model Predictive Control: Learning Optimal Control Cost Functions for MPC
abstract
Inverse optimal control (IOC) seeks to infer a control cost function that captures the underlying goals and preferences of expert demonstrations. While significant progress has been made in finite-horizon IOC, which focuses on learning control cost functions based on rollout trajectories rather than actual trajectories, the application of IOC to receding horizon control, also known as model predictive control (MPC), has been overlooked. MPC is more prevalent in practical settings and poses additional challenges for IOC learning since it is complicated to calculate the gradient of actual trajectories with respect to cost parameters. In light of this, we propose the inverse MPC (IMPC) method to identify the optimal cost function that effectively minimizes the discrepancy between the actual trajectory and its associated demonstration. To compute the gradient of actual trajectories with respect to cost parameters, we first establish two differential Pontryagin's maximum principle (PMP) conditions by differentiating the traditional PMP conditions with respect to cost parameters and initial states, respectively. We then formulate two auxiliary optimal control problems based on the derived differentiated PMP conditions, whose solutions can be directly used to determine the gradient for updating cost parameters. We validate the efficacy of the proposed method through experiments involving five simulation tasks and two real-world mobile robot control tasks. The results consistently demonstrate that IMPC outperforms existing finite-horizon IOC methods across all experiments.
Fawang Zhang, Jingliang Duan, Hao Chen 0108, Hui Liu 0001, Shida Nie, Shengbo Eben Li
IEEE Trans. Ind. Informatics2
2024 Enhance Sample Efficiency and Robustness of End-to-End Urban Autonomous Driving via Semantic Masked World Model
abstract
End-to-end autonomous driving provides a feasible way to automatically maximize overall driving system performance by directly mapping the raw pixels from a front-facing camera to control signals. Recent advanced methods construct a latent world model to map the high dimensional observations into compact latent space. However, the latent states embedded by the world model proposed in previous works may contain a large amount of task-irrelevant information, resulting in low sampling efficiency and poor robustness to input perturbations. Meanwhile, the training data distribution is usually unbalanced, and the learned policy is challenging to cope with the corner cases during the driving process. To solve the above challenges, we present aSEManticMasked recurrent world model (SEM2), which introduces a semantic filter to extract key driving-relevant features and make decisions via the filtered features, and is trained with a multi-source data sampler, which aggregates common data and multiple corner case data in a single batch, to balance the data distribution. Extensive experiments on CARLA show our method outperforms the state-of-the-art approaches in terms of sample efficiency and robustness to input permutations.
Yao Mu 0001, Chen Chen 0068, Jingliang Duan, Ping Luo 0002, Shengbo Eben Li
IEEE Trans. Intell. Transp. Syst.4
2024 Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral Lagrangian
abstract
Safety is essential for reinforcement learning (RL) applied in the real world. Adding chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under uncertainty. Existing chance-constrained RL methods, such as the penalty methods and the Lagrangian methods, either exhibit periodic oscillations or learn an overconservative or unsafe policy. In this article, we address these shortcomings by proposing a separated proportional-integral Lagrangian (SPIL) algorithm. We first review the constrained policy optimization process from a feedback control perspective, which regards the penalty weight as the control input and the safe probability as the control output. Based on this, the penalty method is formulated as a proportional controller, and the Lagrangian method is formulated as an integral controller. We then unify them and present a proportional-integral Lagrangian method to get both their merits with an integral separation technique to limit the integral value to a reasonable range. To accelerate training, the gradient of safe probability is computed in a model-based manner. The convergence of the overall algorithm is analyzed. We demonstrate that our method can reduce the oscillations and conservatism of RL policy in a car-following simulation. To prove its practicality, we also apply our method to a real-world mobile robot navigation task, where our robot successfully avoids a moving obstacle with highly uncertain or even aggressive behaviors.
Baiyu Peng, Jingliang Duan, Jianyu Chen 0002, Shengbo Eben Li, Genjin Xie, Congsheng Zhang, Yang Guan, Yao Mu 0001, Enxin Sun
IEEE Trans. Neural Networks Learn. Syst.2
2023 Global Convergence of Two-Timescale Actor-Critic for Solving Linear Quadratic Regulator
abstract
The actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly consider the uncommon double-loop variant or basic models with finite state and action space. We investigate the more practical single-sample two-timescale AC for solving the canonical linear quadratic regulator (LQR) problem, where the actor and the critic update only once with a single sample in each iteration on an unbounded continuous state and action space. Existing analysis cannot conclude the convergence for such a challenging case. We develop a new analysis framework that allows establishing the global convergence to an epsilon-optimal solution with at most an order of epsilon to -2.5 sample complexity. To our knowledge, this is the first finite-time convergence analysis for the single sample two-timescale AC for solving LQR with global optimality. The sample complexity improves those of other variants by orders, which sheds light on the practical wisdom of single sample algorithms. We also further validate our theoretical findings via comprehensive simulation comparisons.
Jingliang Duan, Yingbin Liang, Lin Zhao 0009
AAAI2
2023 LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal Control
abstract
Deep reinforcement learning (RL) is a powerful approach for solving optimal control problems. However, RL-trained policies often suffer from the action fluctuation problem, where the consecutive actions significantly differ despite only slight state variations. This problem results in mechanical components' wear and tear and poses safety hazards. The action fluctuation is caused by the high Lipschitz constant of actor networks. To address this problem, we propose a neural network named LipsNet. We propose the Multi-dimensional Gradient Normalization (MGN) method, to constrain the Lipschitz constant of networks with multi-dimensional input and output. Benefiting from MGN, LipsNet achieves Lipschitz continuity, allowing smooth actions while preserving control performance by adjusting Lipschitz constant. LipsNet addresses the action fluctuation problem at network level rather than algorithm level, which can serve as actor networks in most RL algorithms, making it more flexible and user-friendly than previous works. Experiments demonstrate that LipsNet has good landscape smoothness and noise robustness, resulting in significantly smoother action compared to the Multilayer Perceptron.
Xujie Song, Jingliang Duan, Wenxuan Wang 0004, Shengbo Eben Li, Chen Chen 0068, Bo Cheng 0003, Junqing Wei, Xiaoming Simon Wang
ICML2
2023 Integrated Decision and Control: Toward Interpretable and Computationally Efficient Driving Intelligence
abstract
Decision and control are core functionalities of high-level automated vehicles. Current mainstream methods, such as functional decomposition and end-to-end reinforcement learning (RL), suffer high time complexity or poor interpretability and adaptability on real-world autonomous driving tasks. In this article, we present an interpretable and computationally efficient framework called integrated decision and control (IDC) for automated vehicles, which decomposes the driving task into static path planning and dynamic optimal tracking that are structured hierarchically. First, the static path planning generates several candidate paths only considering static traffic elements. Then, the dynamic optimal tracking is designed to track the optimal path while considering the dynamic obstacles. To that end, we formulate a constrained optimal control problem (OCP) for each candidate path, optimize them separately, and follow the one with the best tracking performance. To unload the heavy online computation, we propose a model-based RL algorithm that can be served as an approximate-constrained OCP solver. Specifically, the OCPs for all paths are considered together to construct a single complete RL problem and then solved offline in the form of value and policy networks for real-time online path selecting and tracking, respectively. We verify our framework in both simulations and the real world. Results show that compared with baseline methods, IDC has an order of magnitude higher online computing efficiency, as well as better driving performance, including traffic efficiency and safety. In addition, it yields great interpretability and adaptability among different driving scenarios and tasks.
Yang Guan, Yangang Ren, Qi Sun 0004, Shengbo Eben Li, Haitong Ma, Jingliang Duan, Bo Cheng 0003
IEEE Trans. Cybern.6
2023 Policy Iteration Based Approximate Dynamic Programming Toward Autonomous Driving in Constrained Dynamic Environment
abstract
In the area of autonomous driving, it typically brings great difficulty in solving the motion planning problem since the vehicle model is nonlinear and the driving scenarios are complex. Particularly, most of the existing methods cannot be generalized to dynamically changing scenarios with varying surrounding vehicles. To address this problem, this development here investigates the framework of integrated decision and control. As part of the modules, static path planning determines the reference candidates ahead, and then the optimal path-tracking controller realizes the specific autonomous driving task. An innovative and effective constrained finite-horizon approximate dynamic programming (ADP) algorithm is herein presented to generate the desired control policy for effective path tracking. With the generalized policy neural network that maps from the state to the control input, the proposed algorithm preserves the high effectiveness for the motion planning problem towards changing driving environments with varying surrounding vehicles. Moreover, the algorithm attains the noteworthy advantage of alleviating the typically heavy computational loads with the mode of offline training and online execution. As a result of the utilization of multi-layer neural networks in conjunction with the actor-critic framework, the constrained ADP method is capable of handling complex and multidimensional scenarios. Finally, various simulations have been carried out to show that the constrained ADP algorithm is effective.
Ziyu Lin, Jun Ma 0008, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Bo Cheng 0003, Tong Heng Lee
IEEE Trans. Intell. Transp. Syst.3
2023 Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control
abstract
The Hamilton-Jacobi-Bellman (HJB) equation serves as the necessary and sufficient condition for the optimal solution to the continuous-time (CT) optimal control problem (OCP). Compared with the infinite-horizon HJB equation, the solving of the finite-horizon (FH) HJB equation has been a long-standing challenge, because the partial time derivative of the value function is involved as an additional unknown term. To address this problem, this study first-time bridges the link between the partial time derivative and the terminal-time utility function, and thus it facilitates the use of the policy iteration (PI) technique to solve the CT FH OCPs. Based on this key finding, the FH approximate dynamic programming (ADP) algorithm is proposed leveraging an actor-critic framework. It is shown that the algorithm exhibits important properties in terms of convergence and optimality. Rather importantly, with the use of multilayer neural networks (NNs) in the actor-critic architecture, the algorithm is suitable for CT FH OCPs toward more general nonlinear and complex systems. Finally, the effectiveness of the proposed algorithm is demonstrated by conducting a series of simulations on both a linear quadratic regulator (LQR) problem and a nonlinear vehicle tracking problem.
Ziyu Lin, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Jie Li 0042, Jianyu Chen 0002, Bo Cheng 0003, Jun Ma 0008
IEEE Trans. Neural Networks Learn. Syst.2
2022 Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints
Jingliang Duan, Shengbo Eben Li, Qi Sun 0004, Zhenzhong Jia, Bo Cheng 0003
Neurocomputing1
2022 Fixed-Dimensional and Permutation Invariant State Representation of Autonomous Driving
abstract
In this paper, we propose a new state representation method, called encoding sum and concatenation (ESC), to describe the environment observation for decision-making in autonomous driving. Unlike existing state representation methods, ESC is applicable to the situation where the number of surrounding vehicles is variable and eliminates the need for manually pre-designed sorting rules, leading to higher representation ability and generality. The proposed ESC method introduces a feature neural network (NN) to encode the real-valued feature of each surrounding vehicle into an encoding vector, and then adds these vectors up to obtain the representation vector of the set of surrounding vehicles. Then, a fixed-dimensional and permutation-invariance state representation can be obtained by concatenating the set representation with other variables, such as indicators of the ego vehicle and road. By introducing the sum-of-power mapping, this paper has further proved that the injectivity of the ESC state representation can be guaranteed if the output dimension of the feature NN is greater than the number of variables of all surrounding vehicles. This means that the ESC representation can be used to describe the environment and taken as the inputs of learning-based policy functions. Experiments demonstrate that compared with the fixed-permutation representation method, the policy learning accuracy based on ESC representation is improved by 62.2%.
Jingliang Duan, Dongjie Yu, Shengbo Eben Li, Wenxuan Wang 0004, Yangang Ren, Ziyu Lin, Bo Cheng 0003
IEEE Trans. Intell. Transp. Syst.1
2022 Self-Learned Intelligence for Integrated Decision and Control of Automated Vehicles at Signalized Intersections
abstract
Intersection is one of the most accident-prone urban scenarios for autonomous driving wherein making safe and computationally efficient decisions is non-trivial. Current research mainly focuses on the simplified traffic conditions while ignoring the existence of mixed traffic flows, i.e., vehicles, cyclists and pedestrians. For urban roads, different participants lead to a quite dynamic and complex interaction, posing great difficulty to learn an intelligent policy. This paper develops the dynamic permutation state representation in the framework of integrated decision and control (IDC) to handle signalized intersections with mixed traffic flows. Specially, this representation introduces an encoding function and summation operator to construct driving states from environmental observation, capable of dealing with different types and variant number of traffic participants. A constrained optimal control problem is built wherein the objective involves tracking performance and the constraints for different participants, roads and signal lights are designed respectively to assure safety. We solve this problem by gradient-based optimization, wherein the reasonable state will be given by the encoding function and then served as the input of policy and value function. An off-policy training is designed to reuse observations from driving environment and backpropagation through time is utilized to update the policy function and encoding function jointly. Verification result shows that the dynamic permutation state representation can enhance the driving performance of IDC, including comfort, decision compliance and safety with a large margin. The trained driving policy can realize efficient and smooth passing in the complex intersection, guaranteeing driving intelligence and safety simultaneously.
Yangang Ren, Jianhua Jiang, Guojian Zhan, Shengbo Eben Li, Chen Chen 0068, Keqiang Li 0002, Jingliang Duan
IEEE Trans. Intell. Transp. Syst.7
2022 Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors
abstract
In reinforcement learning (RL), function approximation errors are known to easily lead to the Q -value overestimations, thus greatly reducing policy performance. This article presents a distributional soft actor-critic (DSAC) algorithm, which is an off-policy RL method for continuous control setting, to improve the policy performance by mitigating Q -value overestimations. We first discover in theory that learning a distribution function of state-action returns can effectively mitigate Q -value overestimations because it is capable of adaptively adjusting the update step size of the Q -value function. Then, a distributional soft policy iteration (DSPI) framework is developed by embedding the return distribution function into maximum entropy RL. Finally, we present a deep off-policy actor-critic variant of DSPI, called DSAC, which directly learns a continuous return distribution by keeping the variance of the state-action returns within a reasonable range to address exploding and vanishing gradient problems. We evaluate DSAC on the suite of MuJoCo continuous control tasks, achieving the state-of-the-art performance.
Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun 0004, Bo Cheng 0003
IEEE Trans. Neural Networks Learn. Syst.1
2021 Separated Proportional-Integral Lagrangian for Chance Constrained Reinforcement Learning
abstract
Safety is essential for reinforcement learning (RL) applied in real-world tasks like autonomous driving. Imposing chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under model uncertainty. Existing chance constrained RL methods like the penalty methods and the Lagrangian methods either exhibit periodic oscillations or learn an over-conservative or unsafe policy. In this paper, we address these shortcomings by elegantly combining these two methods and propose a separated proportional-integral Lagrangian (SPIL) algorithm. We first rewrite penalty methods as optimizing safe probability according to the proportional value of constraint violation, and Lagrangian methods as optimizing according to the integral value of the violation. Then we propose to add up both the integral and proportion values to optimize the policy, with an integral separation technique to limit the integral value within a reasonable range. Besides, the gradient of policy is computed in a model-based paradigm to accelerate training. The proposed method is proved to reduce oscillations and conservatism while ensuring safety by a car-following experiment.
Baiyu Peng, Yao Mu 0001, Jingliang Duan, Yang Guan, Shengbo Eben Li, Jianyu Chen 0002
IV3
2021 Cover: International Journal of Intelligent Systems, Volume 36 Issue 8 August 2021
abstract
Cover Caption: The cover image is based on the Research Article Direct and indirect reinforcement learning by Yang Guan et al., https://doi.org/10.1002/int.22466.
Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003
Int. J. Intell. Syst.3
2021 Direct and indirect reinforcement learning
abstract
Reinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision-making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal policy by directly maximizing an objective function using gradient descent methods, in which the objective function is usually the expectation of accumulative future rewards. The latter indirectly finds the optimal policy by solving the Bellman equation, which is the sufficient and necessary condition from Bellman's principle of optimality. We study policy gradient (PG) forms of direct and indirect RL and show that both of them can derive the actor–critic architecture and can be unified into a PG with the approximate value function and the stationary state distribution, revealing the equivalence of direct and indirect RL. We employ a Gridworld task to verify the influence of different forms of PG, suggesting their differences and relationships experimentally. Finally, we classify current mainstream RL algorithms using the direct and indirect taxonomy, together with other ones, including value-based and policy-based, model-based and model-free.
Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003
Int. J. Intell. Syst.3