VLDB 2026 Research / reviewers in the wild / expert
Shengbo Eben Li
dblp:148/6962 · also Sheng-bo Li 0001, Shengbo Li 0001
· DBLP profile ↗
110ranked-venue papers
5as first author
84since 2021 · last 2026
0000-0003-4923-3633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 3 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 2 first-author · 32 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Synthesis of Formally Verified Neural Value Function for Hamilton-Jacobi Reachability Analysis (Abstract Reprint)abstractHamilton-Jacobi (HJ) reachability analysis provides a formal method for guaranteeing safety in constrained control problems. It synthesizes a value function to represent a long-term safe set called feasible region. Early synthesis methods based on state space discretization cannot scale to high-dimensional problems, while recent methods that use neural networks to approximate value functions result in unverifiable feasible regions. To achieve both scalability and verifiability, we propose a framework for synthesizing verified neural value functions for HJ reachability analysis. Our framework consists of three stages: pre-training, adversarial training, and verification-guided training. We design three techniques to address three challenges to improve scalability respectively: boundary-guided backtracking (BGB) to improve counterexample search efficiency, entering state regularization (ESR) to enlarge feasible region, and activation pattern alignment (APA) to accelerate neural network verification. We also provide a neural safety certificate synthesis and verification benchmark called Cersyve-9, which includes nine commonly used safe control tasks and supplements existing neural network verification benchmarks. Our framework successfully synthesizes verified neural value functions on all tasks, and our proposed three techniques exhibit superior scalability and efficiency compared with existing methods. Hanjiang Hu, Tianhao Wei, Shengbo Eben Li, Changliu Liu |
AAAI | 4 |
| 2026 | A smooth reinforcement learning method for trajectory tracking and collision avoidance of wheeled vehicle
Liangfa Chen, Xujie Song, Wenxuan Wang 0004, Liming Xiao, Shengbo Eben Li, Jingliang Duan |
Expert Syst. Appl. | 7 |
| 2026 | Multi-agent reinforcement learning with beta distribution for thrust sampling in stochastic orbital pursuit-evasion games
Yuqiao Zhao, Jingliang Duan, Shengbo Eben Li, Chang Liu 0002 |
Neurocomputing | 3 |
| 2026 | Nonlinear Bayesian Filtering With Natural Gradient Gaussian ApproximationabstractPractical Bayes filters often assume the state distribution of each time step to be Gaussian for computational tractability, resulting in the so-called Gaussian filters. When facing nonlinear systems, Gaussian filters such as extended Kalman filter (EKF) or unscented Kalman filter (UKF) typically rely on certain linearization techniques, which can introduce large estimation errors. To address this issue, this paper reconstructs the prediction and update steps of Gaussian filtering as solutions to two distinct optimization problems, whose optimal conditions are found to have analytical forms from Stein's lemma. It is observed that the stationary point for the prediction step requires calculating the first two moments of the prior distribution, which is equivalent to that step in existing moment-matching filters. In the update step, instead of linearizing the model to approximate the stationary points, we propose an iterative approach to directly minimize the update step's objective to avoid linearization errors. For the purpose of performing the steepest descent on the Gaussian manifold, we derive its natural gradient that leverages Fisher information matrix to adjust the gradient direction, accounting for the curvature of the parameter space. Combining this update step with moment matching in the prediction step, we introduce a new iterative filter for nonlinear systems called Natural Gradient Gaussian Approximation filter, or NANO filter for short. We prove that NANO filter locally converges to the optimal Gaussian approximation at each time step. Furthermore, the estimation error is proven exponentially bounded for nearly linear measurement equation and low noise levels through constructing a supermartingale-like property across consecutive time steps. Real-world experiments demonstrate that, compared to popular Gaussian filters such as EKF, UKF, iterated EKF, and posterior linearization filter, NANO filter reduces the average root mean square error by approximately 45% while maintaining a comparable computational burden. Wenhan Cao, Zeju Sun, Chang Liu 0002, Stephen S.-T. Yau, Shengbo Eben Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | On the Equilibrium Between Feasible Zone and Uncertain Model in Safe ExplorationabstractEnsuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a way to ensure safety, key questions remain unresolved: what is the maximum feasible zone achievable through exploration, and how can it be identified? This paper, for the first time, answers these questions by revealing that the goal of safe exploration is to find the equilibrium between the feasible zone and the environment model. This conclusion is based on the understanding that these two components are interdependent: a larger feasible zone leads to a more accurate environment model, and a more accurate model, in turn, enables exploring a larger zone. We propose the first equilibrium-oriented safe exploration framework called safe equilibrium exploration (SEE), which alternates between finding the maximum feasible zone and the least uncertain model. Using a graph formulation of the uncertain model, we prove that the uncertain model obtained by SEE is monotonically refined, the feasible zones monotonically expand, and both converge to the equilibrium of safe exploration. Experiments on classic control tasks show that our algorithm successfully expands the feasible zones with zero constraint violation, and achieves the equilibrium of safe exploration within a few iterations. Zhilong Zheng, Shengbo Eben Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Enhanced Integrated Decision and Control for High-Level Automated Vehicles and Its Experiment VerificationabstractLearning through experience is essential for high-level autonomous driving systems, as it has the potential to enhance driving performance in corner cases. However, current decision and control modules adopt an empirical design paradigm for engineering efficiency, relying heavily on expert rules or real-vehicle data, failing to fully cover and optimize edge scenarios. To address this gap, we propose an enhanced integrated decision and control method that leverages reinforcement learning as the optimal control problem solver, endowing high-level automated vehicles with experience data usage. Specifically, a constrained mixed policy gradient algorithm is developed, which combines and dynamically adjusts the application ratio of experience data and the environmental model during training. This approach achieves fast convergence while maintaining high performance even with inaccurate analytic models. Furthermore, an attention based encoding network is designed to accommodate diverse driving states in urban traffic, integrating an embedding network for feature extraction and a weighting network for feature fusion, realizing order-insensitive encoding and importance differentiation of road users. The trained policy is deployed on a fully functional autonomous vehicle. Experiments at a signalized intersection show that the proposed method can accurately identify critical surrounding obstacles and execute safe, efficient, and intelligent driving behaviors across 32 scenarios. Yang Guan, Liye Tang, Yao Lyu, Shengbo Eben Li, Kehua Sheng, Keqiang Li 0002 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Learning Optimal Robust Control for Nonlinear Mixed Traffic Under External DisturbanceabstractThe integration of connected and automated vehicles (CAVs) into traffic systems holds potential to mitigate undesired disturbances. Nevertheless, coexisting human-driven vehicles (HDVs) introduce complex behavioral disturbances, which has imposed critical challenges for control robustness. This study develops a computational framework based on policy iteration to derive robust control policies with optimized attenuation performance for nonlinear mixed traffic flow. Specifically, robust$H_{\infty }$control problem is solved by applying the framework of zero-sum game, whose solution at the Nash equilibrium is transformed into a Hamilton–Jacobi (HJ) inequality with a Hamiltonian constraint. For achieving desired attenuation performance, the value function is updated by gradient descent based on counterexamples violating Hamiltonian and monotonicity constraints, where the positive definiteness of the value function is ensured by convex neural networks, facilitating the analysis of control stability via Lyapunov methods. By utilizing constraint gaps, the attenuation level is optimized through the analytical formulae derived from the HJ inequality. The stability and algorithm convergence are proved. Experimental results demonstrate the capability of the learned controller to effectively attenuate disturbance propagation and stabilize mixed traffic flow. Jie Li 0042, Jiawei Wang 0001, Yangang Ren, Shen Li 0001, Guofa Li, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Generation of High-Coverage Traffic Scenarios for Efficient Simulation Testing of Automated Driving SystemsabstractVirtual simulation testing is crucial for ensuring automated vehicles safety, which offers low cost and good repeatability. The key is to test in various virtual driving scenarios, but often fails to strike a balance between scenario coverage and test efficiency. To address this issue, we propose a scenario generation method based on a Genetic Algorithm optimized Hamiltonian Monte Carlo sampling approach. Specifically, a Markov chain is constructed converging to the joint probability density distribution function of scenario parameters. By defining a Hamiltonian function with a potential energy term related to the posterior distribution and a kinetic energy term, the sampling process moves efficiently towards high probability regions and achieves faster convergence to the target distribution. Moreover, Jensen-Shannon divergence between generated samples and raw data is proposed to evaluate the scenario coverage, and used as the objective function in Genetic Algorithm to optimize the algorithm parameters. The proposed method is validated by lead vehicle deceleration scenarios generation, where an S-shaped deceleration model is proposed to parameterize the scenario and 22,343 segments extracted from a naturalistic driving dataset are used to calibrate the scenario parameters. Subsequently, the joint probability distribution density function of scenario parameters fitted by a gaussian mixture model is imported into our proposed method. As the results, 1,649 concrete lead vehicle deceleration scenarios are generated with a coverage of 99.14% for the dataset. Compared with the previous Markov Chain Monte Carlo sampling method, our method achieved 9 times higher coverage with 28 times fewer scenarios. Henan Yuan, Qianru Dong, Ruidong Yan, Chunjiao Dong, Chen Chen 0068, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Bicriteria Policy Optimization for High-Accuracy Reinforcement LearningabstractIn essence, reinforcement learning (RL) solves optimal control problem (OCP) by employing a neural network (NN) to fit the optimal policy from state to action. The accuracy of policy approximation is often very low in complex control tasks, leading to unsatisfactory control performance compared with online optimal controllers. A primary reason is that the landscape of value function is always not only rugged in most areas but also flat on the bottom, which damages the convergence to the minimum point. To address this issue, we develop a bicriteria policy optimization (BPO) algorithm, which leverages a few optimal demonstration trajectories to guide the policy search at the gradient level. Different from conventional problem definition, BPO seeks to solve a bicriteria OCP, which has two homomorphic objectives: one is from the standard reward signals and the other is to align the demonstration trajectories. We introduce two co-state variables, one for each objectives, and formulate two Hamiltonians for this bicriteria OCP. The resulting new optimality condition preserves the minimum values of both Hamiltonians. Furthermore, we find that gradient conflict is a key obstacle to simultaneously descending both Hamiltonians, and its impact is negatively proportional to the inner product between the ideal and actual gradients. A minimax optimization problem is built at each RL iteration to minimize conflicts between two homomorphic objectives, whose solution for policy updating is referred to as harmonic gradient. By converting its inner optimization loop into a linear programming with convex trust region constraint, we simplify this problem into a single-loop maximization problem with much increased computational efficiency. Experiment tests on both linear and nonlinear control tasks validate the effectiveness of our BPO algorithm on the accuracy improvement of policy network. Guojian Zhan, Xiangteng Zhang, Feihong Zhang, Letian Tao, Shengbo Eben Li |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | ODE-based Smoothing Neural Network for Reinforcement Learning TasksabstractThe smoothness of control actions is a significant challenge faced by deep reinforcement learning (RL) techniques in solving optimal control problems. Existing RL-trained policies tend to produce non-smooth actions due to high-frequency input noise and unconstrained Lipschitz constants in neural networks. This article presents a Smooth ODE (SmODE) network capable of simultaneously addressing both causes of unsmooth control actions, thereby enhancing policy performance and robustness under noise condition. We first design a smooth ODE neuron with first-order low-pass filtering expression, which can dynamically filter out high frequency noises of hidden state by a learnable state-based system time constant. Additionally, we construct a state-based mapping function, $g$, and theoretically demonstrate its capacity to control the ODE neuron's Lipschitz constant. Then, based on the above neuronal structure design, we further advanced the SmODE network serving as RL policy approximators. This network is compatible with most existing RL algorithms, offering improved adaptability compared to prior approaches. Various experiments show that our SmODE network demonstrates superior anti-interference capabilities and smoother action outputs than the multi-layer perception and smooth network architectures like LipsNet. Wenxuan Wang 0004, Xujie Song, Yuming Yin, Liangfa Chen, Jingliang Duan, Shengbo Eben Li |
ICLR | 9 |
| 2025 | Diffusion-Based Planning for Autonomous Driving with Flexible GuidanceabstractAchieving human-like driving behaviors in complex open-world environments is a critical challenge in autonomous driving. Contemporary learning-based planning approaches such as imitation learning methods often struggle to balance competing objectives and lack of safety assurance,due to limited adaptability and inadequacy in learning complex multi-modal behaviors commonly exhibited in human planning, not to mention their strong reliance on the fallback strategy with predefined rules. We propose a novel transformer-based Diffusion Planner for closed-loop planning, which can effectively model multi-modal driving behavior and ensure trajectory quality without any rule-based refinement. Our model supports joint modeling of both prediction and planning tasks under the same architecture, enabling cooperative behaviors between vehicles. Moreover, by learning the gradient of the trajectory score function and employing a flexible classifier guidance mechanism, Diffusion Planner effectively achieves safe and adaptable planning behaviors. Evaluations on the large-scale real-world autonomous planning benchmark nuPlan and our newly collected 200-hour delivery-vehicle driving dataset demonstrate that Diffusion Planner achieves state-of-the-art closed-loop performance with robust transferability in diverse driving styles. Yinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng, Liyuan Mao, Weihao Gu, Rui Ai 0001, Shengbo Eben Li, Xianyuan Zhan |
ICLR | 9 |
| 2025 | LipsNet++: Unifying Filter and Controller into a Policy NetworkabstractDeep reinforcement learning (RL) is effective for decision-making and control tasks like autonomous driving and embodied AI. However, RL policies often suffer from the action fluctuation problem in real-world applications, resulting in severe actuator wear, safety risk, and performance degradation. This paper identifies the two fundamental causes of action fluctuation: observation noise and policy non-smoothness. We propose LipsNet++, a novel policy network with Fourier filter layer and Lipschitz controller layer to separately address both causes. The filter layer incorporates a trainable filter matrix that automatically extracts important frequencies while suppressing noise frequencies in the observations. The controller layer introduces a Jacobian regularization technique to achieve a low Lipschitz constant, ensuring smooth fitting of a policy function. These two layers function analogously to the filter and controller in classical control theory, suggesting that filtering and control capabilities can be seamlessly integrated into a single policy network. Both simulated and real-world experiments demonstrate that LipsNet++ achieves the state-of-the-art noise robustness and action smoothness. The code and videos are publicly available at https://xjsong99.github.io/LipsNet_v2. Xujie Song, Liangfa Chen, Wenxuan Wang 0004, Shentao Qin, Yinsong Ma, Jingliang Duan, Shengbo Eben Li |
ICML | 9 |
| 2025 | Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement LearningabstractEnd-to-end autonomous driving offers a stream-lined alternative to the traditional modular pipeline, integrating perception, prediction, and planning within a single framework. While Deep Reinforcement Learning (DRL) has recently gained traction in this domain, existing approaches often overlook the critical connection between feature extraction of DRL and perception. In this paper, we bridge this gap by mapping the DRL feature extraction network directly to the perception phase, en-abling clearer interpretation through semantic segmentation. By leveraging Bird's-Eye- View (BEV) representations, we propose a novel DRL-based end-to-end driving framework that utilizes multi-sensor inputs to construct a unified three-dimensional understanding of the environment. This BEV-based system extracts and translates critical environmental features into high-level abstract states for DRL, facilitating more informed control. Extensive experimental evaluations demonstrate that our approach not only enhances interpretability but also significantly outperforms state-of-the-art methods in autonomous driving control tasks, reducing the collision rate by 20 %. Siyi Lu, Shengbo Eben Li, Yugong Luo, Jianqiang Wang 0003, Keqiang Li 0002 |
ICRA | 3 |
| 2025 | Vision-Driven 2D Supervised Fine-Tuning Framework for Bird's Eye View PerceptionabstractVisual bird’s eye view (BEV) perception, dute to its excellent perceptual capabilities, is progressively replacing costly LiDAR-based perception systems, especially in the realm of urban intelligent driving. However, this type of perception still relies on LiDAR data to construct ground truth databases, a process that is both cumbersome and time-consuming. Additionally, most mass-produced autonomous driving systems are equipped solely with surround camera sensors and lack the LiDAR data necessary for precise annotation. To tackle this challenge, we propose a fine-tuning method for BEV perception network based on visual 2D semantic perception, aimed at enhancing the model’s generalization capabilities in new scene data. Leveraging the maturity of 2D perception technologies, our method utilizes only 2D semantic segmentation labels and monocular depth estimations, thereby significantly reducing the dependence on expensive BEV ground truths and offering strong potential for industrial deployment. Extensive experiments and comparative analyses on the nuScenes and Waymo datasets demonstrate the effectiveness of our method. Specifically, it improves mAP and NDS by 2.51% and 1.93% on nuScenes, and by 1.21% and 0.78% on Waymo, respectively, validating its practical utility and robustness across diverse domains. Qiaoyi Wang, Honglin Sun, Qing Xu 0010, Bolin Gao, Shengbo Eben Li, Jianqiang Wang 0003, Keqiang Li 0002 |
IROS | 6 |
| 2025 | Transferable Latent-To-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged RobotsabstractReinforcement learning (RL) has demonstrated remarkable capability in acquiring robot skills, but learning each new skill still requires substantial data collection for training. The pretrain-and-finetune paradigm offers a promising approach for efficiently adapting to new robot entities and tasks. Inspired by the idea that acquired knowledge can accelerate learning new tasks with the same robot and help a new robot master a trained task, we propose a latent training framework where a transferable latent-to-latent locomotion policy is pretrained alongside diverse task-specific observation encoders and action decoders. This policy in latent space processes encoded latent observations to generate latent actions to be decoded, with the potential to learn general abstract motion skills. To retain essential information for decision-making and control, we introduce a diffusion recovery module that minimizes information reconstruction loss during pretrain stage. During fine-tune stage, the pretrained latent-to-latent locomotion policy remains fixed, while only the lightweight task-specific encoder and decoder are optimized for efficient adaptation. Our method allows a robot to leverage its own prior experience across different tasks as well as the experience of other morphologically diverse robots to accelerate adaptation. We validate our approach through extensive simulations and real-world experiments, demonstrating that the pretrained latent-to-latent locomotion policy effectively generalizes to new robot entities and tasks with improved efficiency. Ziang Zheng, Guojian Zhan, Bin Shuai, Shentao Qin, Shengbo Eben Li |
IROS | 7 |
| 2025 | Explicit Nonlinear Control for Optimal Trajectory Tracking of Autonomous VehiclesabstractMotion control of autonomous vehicles (AVs) that considers nonlinear dynamics and multidimensional motion coupling characteristics represents a critical research direction, particularly for vehicles operating under extreme conditions. However, the nonlinearity of model and the time-varying characteristics of reference states make control law design challenging, often resorting to computationally expensive approaches such as model predictive control (MPC). This paper proposes an explicit nonlinear control design framework for optimal trajectory tracking of AVs. Specifically, taking the three-degree-of-freedom vehicle dynamics model as an example, we first augment the system state to yield an affine representation, followed by input-output feedback linearization. The discrete-time linearized system is then augmented with reference states from the preview horizon, and a linear quadratic regulator is designed to provide an analytical solution for the virtual control inputs. Finally, the original control inputs are computed using the feedback linearization control law. We performed simulations to evaluate the effectiveness of our proposed approach and compared it against MPC. Results indicate that our proposed approach achieves tracking accuracy, smoothness, and robustness comparable to MPC, while significantly reducing computational requirements, with a computing time of nearly 0 ms. Weixian He, Bin Shuai, Chen Chen 0068, Chang Liu 0002, Shengbo Eben Li |
IV | 8 |
| 2025 | Enhanced DACER Algorithm with Multimodal Q-value Distribution for Risk-Sensitive Stochastic Vehicle EnvironmentsabstractReinforcement learning demonstrates strong capabilities in handling complex control tasks, especially in the field of autonomous driving where vehicles cope with uncertain environments. Existing reinforcement learning methods attempt to model the value distribution using unimodal, but in this modeling process, a significant amount of the complete distribution information is lost. In response to this problem, we propose the DACER++, an online multimodal distributional RL algorithm. DACER++ characterize the value distribution as multimodal will enhance the accuracy of characterizing the value distribution and improve algorithm performance. We construct the quantiles value network and use quantile regression to approximate the full quantile function of the state-action return distribution. This method allows for the precise modeling of multi-modal distributions, and formulates risk-sensitive policies adaptable to different environment. Then, We integrate quantiles value network with the actor-critic architecture algorithm DACER. Experiments on multi-goal tasks and MuJoCo benchmarks show that DACER++ not only has multimodal policy representation capability, but also achieves state-of-the-art performance. In stochastic vehicle meeting environments, DACER++ can learn different multimodal value distributions and multimodal trajectories according to various risk preferences, including the conservative and aggressive driving style. Xujie Song, Wenjun Zou, Bin Shuai, Weixian He, Jingliang Duan, Shengbo Eben Li |
IV | 9 |
| 2025 | A Path-Driven Probabilistic Framework for Simulating Abnormal Behavior in Autonomous Driving ScenariosabstractAchieving high-level autonomous driving poses significant challenges, primarily due to the extensive and time-consuming testing required for low-probability abnormal behavior scenarios. Existing simulation platforms are limited by microscopic traffic flow models based on car-following and lane-changing theories (focused on the behavior of individual vehi-cles), which limits the implementation of complex and diverse abnormal behaviors such as road deviation and cutting in. This paper proposes an Abnormal behavior Generation model based on Event Probability triggers and Static paths (AGEPS model), which is capable of continuously and automatically generating abnormal behaviors without disrupting normal traffic flow. The AGEPS model is computationally efficient and scalable, as it decouples path optimization and trajectory tracking from obstacle avoidance, using explicit control law for both tracking and avoidance. The paper demonstrates three typical abnormal behaviors: Overtaking on the Right (OOR), Driving on the Lane Line (DOL), and Sudden Braking (SUB), indicating that the AGEPS model effectively generates these behaviors by selecting and tracking target static paths while adjusting lateral offsets and desired speeds. Experimental results validate the AGEPS model's effectiveness and its ability to generate abnormal behaviors continuously. Simulation results indicate that the AGEPS model reduces single-step computation time (including decision-making and control) by over 92.8 % compared to the MPC controller, with a single-step execution time of just 4 ms. Chen Chen 0068, Teh Jing Lin, Zi'ang Zheng, Bo Cheng 0003, Shengbo Eben Li |
IV | 7 |
| 2025 | Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-Lane ScenariosabstractReinforcement learning (RL), known for its self-evolution capability, offers a promising approach to training high-level autonomous driving systems. However, handling constraints remains a significant challenge for existing RL algorithms, particularly in real-world applications. In this paper, we propose a new safety-oriented training technique called harmonic policy iteration (HPI). At each RL iteration, it first calculates two policy gradients associated with efficient driving and safety constraints, respectively. Then, a harmonic gradient is derived for policy updating, minimizing conflicts between the two gradients and consequently enabling a more balanced and stable training process. Furthermore, we adopt the state-of-the-art DSAC algorithm as the backbone and integrate it with our HPI to develop a new safe RL algorithm, DSAC-H. Extensive simulations in multi-lane scenarios demonstrate that DSAC-H achieves efficient driving performance with near-zero safety constraint violations. Feihong Zhang, Guojian Zhan, Bin Shuai, Jingliang Duan, Shengbo Eben Li |
IV | 6 |
| 2025 | One Filters All: A Generalist Filter For State EstimationabstractEstimating hidden states in dynamical systems, also known as optimal filtering, is a long-standing problem in various fields of science and engineering. In this paper, we introduce a general filtering framework, $\textbf{LLM-Filter}$, which leverages large language models (LLMs) for state estimation by embedding noisy observations with text prototypes. In a number of experiments for classical dynamical systems, we find that first, state estimation can significantly benefit from the knowledge embedded in pre-trained LLMs. By achieving proper modality alignment with the frozen LLM, LLM-Filter outperforms the state-of-the-art learning-based approaches. Second, we carefully design the prompt structure, System-as-Prompt (SaP), incorporating task instructions that enable LLMs to understand tasks and adapt to specific systems. Guided by these prompts, LLM-Filter exhibits exceptional generalization, capable of performing filtering tasks accurately in changed or even unseen environments. We further observe a scaling-law behavior in LLM-Filter, where accuracy improves with larger model sizes and longer training times. These findings make LLM-Filter a promising foundation model of filtering. Wenhan Cao, Chang Liu 0002, Shengbo Eben Li |
NeurIPS | 7 |
| 2025 | Off-policy Reinforcement Learning with Model-based Exploration AugmentationabstractExploration is crucial in Reinforcement Learning (RL) as it enables the agent to understand the environment for better decision-making. Existing exploration methods fall into two paradigms: active exploration, which injects stochasticity into the policy but struggles in high-dimensional environments, and passive exploration, which manages the replay buffer to prioritize under-explored regions but lacks sample diversity. To address the limitation in passive exploration, we propose Modelic Generative Exploration (MoGE), which augments exploration through the generation of under-explored critical states and synthesis of dynamics-consistent experiences. MoGE consists of two components: (1) a diffusion generator for critical states under the guidance of entropy and TD error, and (2) a one-step imagination world model for constructing critical transitions for agent learning. Our method is simple to implement and seamlessly integrates with mainstream off-policy RL algorithms without structural modifications. Experiments on OpenAI Gym and DeepMind Control Suite demonstrate that MoGE, as an exploration augmentation, significantly enhances efficiency and performance in complex tasks. Xiangteng Zhang, Guojian Zhan, Wenxuan Wang 0004, Jingliang Duan, Shengbo Eben Li |
NeurIPS | 8 |
| 2025 | Bootstrap Off-policy with World ModelabstractOnline planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learning and policy improvement. To address this, we propose BOOM (Bootstrap Off-policy with WOrld Model), a framework that tightly integrates planning and off-policy learning through a bootstrap loop: the policy initializes the planner, and the planner refines actions to bootstrap the policy through behavior alignment. This loop is supported by a jointly learned world model, which enables the planner to simulate future trajectories and provides value targets to facilitate policy improvement. The core of BOOM is a likelihood-free alignment loss that bootstraps the policy using the planner’s non-parametric action distribution, combined with a soft value-weighted mechanism that prioritizes high-return behaviors and mitigates variability in the planner’s action quality within the replay buffer. Experiments on the high-dimensional DeepMind Control Suite and Humanoid-Bench show that BOOM achieves state-of-the-art results in both training stability and final performance. The code is accessible at \url{https://github.com/molumitu/BOOM_MBRL}. Guojian Zhan, Xiangteng Zhang, Jiaxin Gao 0002, Masayoshi Tomizuka, Shengbo Eben Li |
NeurIPS | 6 |
| 2025 | Smooth policy iteration for zero-sum Markov Games
Yangang Ren, Yao Lyu, Wenxuan Wang 0004, Shengbo Eben Li, Zeyang Li 0001, Jingliang Duan |
Neurocomputing | 4 |
| 2025 | Scalable Synthesis of Formally Verified Neural Value Function for Hamilton-Jacobi Reachability AnalysisabstractHamilton-Jacobi (HJ) reachability analysis provides a formal method for guaranteeing safety in constrained control problems. It synthesizes a value function to represent a long-term safe set called feasible region. Early synthesis methods based on state space discretization cannot scale to high-dimensional problems, while recent methods that use neural networks to approximate value functions result in unverifiable feasible regions. To achieve both scalability and verifiability, we propose a framework for synthesizing verified neural value functions for HJ reachability analysis. Our framework consists of three stages: pre-training, adversarial training, and verification-guided training. We design three techniques to address three challenges to improve scalability respectively: boundary-guided backtracking (BGB) to improve counterexample search efficiency, entering state regularization (ESR) to enlarge feasible region, and activation pattern alignment (APA) to accelerate neural network verification. We also provide a neural safety certificate synthesis and verification benchmark called Cersyve-9, which includes nine commonly used safe control tasks and supplements existing neural network verification benchmarks. Our framework successfully synthesizes verified neural value functions on all tasks, and our proposed three techniques exhibit superior scalability and efficiency compared with existing methods. Hanjiang Hu, Tianhao Wei, Shengbo Eben Li, Changliu Liu |
J. Artif. Intell. Res. | 4 |
| 2025 | Distributional Soft Actor-Critic With Three RefinementsabstractReinforcement learning (RL) has shown remarkable success in solving complex decision-making and control tasks. However, many model-free RL algorithms experience performance degradation due to inaccurate value estimation, particularly the overestimation of Q-values, which can lead to suboptimal policies. To address this issue, we previously proposed the Distributional Soft Actor-Critic (DSAC or DSACv1), an off-policy RL algorithm that enhances value estimation accuracy by learning a continuous Gaussian value distribution. Despite its effectiveness, DSACv1 faces challenges such as training instability and sensitivity to reward scaling, caused by high variance in critic gradients due to return randomness. In this paper, we introduce three key refinements to DSACv1 to overcome these limitations and further improve Q-value estimation accuracy: expected value substitution, twin value distribution learning, and variance-based critic gradient adjustment. The enhanced algorithm, termed DSAC with Three refinements (DSAC-T or DSACv2), is systematically evaluated across a diverse set of benchmark tasks. Without the need for task-specific hyperparameter tuning, DSAC-T consistently matches or outperforms leading model-free RL algorithms, including SAC, TD3, DDPG, TRPO, and PPO, in all tested environments. Additionally, DSAC-T ensures a stable learning process and maintains robust performance across varying reward scales. Its effectiveness is further demonstrated through real-world application in controlling a wheeled robot, highlighting its potential for deployment in practical robotic tasks. Jingliang Duan, Wenxuan Wang 0004, Liming Xiao, Jiaxin Gao 0002, Shengbo Eben Li, Chang Liu 0002, Ya-Qin Zhang, Bo Cheng 0003, Keqiang Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | RINO: Accurate, Robust Radar-Inertial Odometry With Non-Iterative EstimationabstractOdometry in adverse weather conditions, such as fog, rain, and snow, presents significant challenges, as traditional vision- and LiDAR-based methods often suffer from degraded performance. Radar-Inertial Odometry (RIO) has emerged as a promising solution due to its resilience in such environments. In this paper, we present RINO, a non-iterative RIO framework implemented in an adaptively loosely coupled manner. Building upon ORORA as the baseline for radar odometry, RINO introduces several key advancements, including improvements in keypoint extraction, motion distortion compensation, and pose estimation via an adaptive voting mechanism. This voting strategy facilitates efficient polynomial-time optimization while simultaneously quantifying the uncertainty in the radar module’s pose estimation. The estimated uncertainty is subsequently integrated into the maximum a posteriori (MAP) estimation within a Kalman filter framework. Unlike prior loosely coupled odometry systems, RINO not only retains the global and robust registration capabilities of the radar component but also dynamically accounts for the real-time operational state of each sensor during fusion. Experimental results conducted on publicly available datasets demonstrate that RINO reduces translation and rotation errors by 1.06% and 0.09°/100m, respectively, when compared to the baseline method, thus significantly enhancing its accuracy. Furthermore, RINO achieves performance comparable to state-of-the-art methods. Our code is available at https://github.com/yangsc4063/rino. Shuocheng Yang, Yueming Cao, Shengbo Eben Li, Jianqiang Wang 0003, Shaobing Xu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Multimodal Reinforcement Learning With Score-Based PolicyabstractLearning multimodal policies is crucial for enhancing exploration in online reinforcement learning (RL), especially in tasks with continuous action spaces and non-convex reward landscapes. While recent diffusion policies show promise, they often suffer from low computational efficiency in online settings. A more training-efficient paradigm involves modeling the policy as a Boltzmann distribution and guiding the action sampling directly with the gradient of the Q-value with respect to the action (proportional to the score function of the policy), such as via Langevin dynamics. However, analysis in this paper reveals that this gradient-guided approach suffers from two critical challenges: sampling instability caused by the widely varying magnitude of action gradients; and mode imbalance, where the sampling process inaccurately represents the weights of different high-value action modes. To address these challenges, this paper introduces three targeted techniques: score normalization and reshaping to stabilize the sampling process, and value-based resampling to correct mode imbalance. These techniques are then integrated into an actor-critic framework, resulting in the Score-Enhanced Actor-Critic (SEAC) algorithm. Simulation and real-world experiments demonstrate that SEAC not only effectively learns multimodal behaviors but also achieves state-of-the-art performance and high computational efficiency compared to prior multimodal RL methods. The code of this paper is available at https://github.com/THUzouwenjun/SEAC. Wenjun Zou, Bin Shuai, Liming Xiao, Yinsong Ma, Jingliang Duan, Shengbo Eben Li |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Feasible Policy Iteration With Guaranteed Safe ExplorationabstractSafety guarantee is an important topic when training real-world tasks with reinforcement learning (RL). During online environmental exploration, any constraint violation can lead to significant property damage and risks to personnel. Existing safe RL methods either exclusively address safety concerns after reaching optimality or incorporate a certain degree of tolerance for constraint violations during training. This article proposes a feasible policy iteration framework that can guarantee absolute safety during online exploration, i.e., constraint violations never happen in real-world interactions. The key to maintaining absolute safety lies in confining the environmental exploration at each step always within the feasible region of the current policy. This feasible region is described by a newly defined constraint decay function with uncertainty, ensuring the forward invariance of the feasible region under the worst case. Within the proposed framework, the feasible region maintains its monotonic expanding property and converges to its maximum extent, even though only local samples are available, i.e., the agent only has access to samples within the feasible region. Meanwhile, the trained policy also improves monotonically within its corresponding feasible region if one can use different updating rules inside and outside the feasible region. Finally, practical algorithms are designed with the actor-critic-scenery architecture, consisting of three modules: 1) safe exploration; 2) model error estimation; and 3) network update. Experimental results indicate that our algorithms achieve performance comparable to baselines while maintaining zero constraint violation throughout the entire training process. In contrast, the baseline algorithm typically requires thousands of constraint violations to achieve the same performance. These findings suggest a substantial potential for applying feasible policy iteration in real-world tasks, enabling the online evolution of intricate systems. Yuhang Zhang 0018, Shengbo Eben Li, Yao Lyu, Jingliang Duan, Zhilong Zheng, Dezhao Zhang |
IEEE Trans. Cybern. | 3 |
| 2025 | Zeroth-Order Actor-Critic: An Evolutionary Framework for Sequential Decision ProblemsabstractEvolutionary algorithms (EAs) have shown promise in solving sequential decision problems (SDPs) by simplifying them to static optimization problems and searching for the optimal policy parameters in a zeroth-order way. While these methods are highly versatile, they often suffer from high sample complexity due to their ignorance of the underlying temporal structures. In contrast, reinforcement learning (RL) methods typically formulate SDPs as Markov Decision Process (MDP). Although more sample efficient than EAs, RL methods are restricted to differentiable policies and prone to getting stuck in local optima. To address these issues, we propose a novel evolutionary framework Zeroth-Order Actor-Critic (ZOAC). We propose to use step-wise exploration in parameter space and theoretically derive the zeroth-order policy gradient. We further utilize the actor-critic architecture to effectively leverage the Markov property of SDPs and reduce the variance of gradient estimators. In each iteration, ZOAC employs samplers to collect trajectories with parameter space exploration, and alternates between first-order policy evaluation (PEV) and zeroth-order policy improvement (PIM). To evaluate the effectiveness of ZOAC, we apply it to a challenging multi-lane driving task, optimizing the parameters in a rule-based, non-differentiable driving policy that consists of three sub-modules: behavior selection, path planning, and trajectory tracking. We also compare it with gradient-based RL methods on three Gymnasium tasks, optimizing neural network policies with thousands of parameters. Experimental results demonstrate the strong capability of ZOAC in solving SDPs. ZOAC significantly outperforms EAs that treat the problem as static optimization and matches the performance of gradient-based RL methods even without first-order information, in terms of total average return across all tasks. Yuheng Lei, Yao Lyu, Guojian Zhan, Jianyu Chen 0002, Shengbo Eben Li, Sifa Zheng |
IEEE Trans. Evol. Comput. | 7 |
| 2025 | Real-Time Resilient Tracking Control for Autonomous Vehicles Through Triple Iterative Approximate Dynamic ProgrammingabstractEnhancing control precision, mitigating external disturbances, and ensuring real-time responsiveness stand as the cornerstone of autonomous vehicle tracking endeavors, each of which intricately interwoven to uphold operational safety. In pursuit of addressing these issues, this paper presents a triple iterative control method inspired by approximate dynamic programming (ADP) tailored for real-time disturbance avoidance. The control framework orchestrates simultaneous iterations of value function, control policy, and disturbance policy, engineered to optimize tracking control amidst external disturbances cast as a zero-sum differential game, tackled adeptly through deep neural networks. Rigorous mathematical proof underpins its triple iteration, coupled with assurances of residual error convergence, solidifying its safety guarantee ability and algorithmic resilience. To validate its effectiveness, both numerical simulations and experiments on a real micro-vehicle platform were conducted. Results underscore the feasibility of this new method, showcasing its energy-saving capability and a four-times acceleration compared to conventional model predictive control (MPC) approaches when confronted with lateral disturbances. Notably, the single-step calculation time of this method on the Raspberry Pi is only 1.44ms, affirming its practical viability and real-world applicability. Jiale Geng, Yunqi Cheng, Liye Tang, Jingliang Duan, Feng Duan 0006, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Lightweight Strategies for Decision-Making of Autonomous Vehicles in Lane Change Scenarios Based on Deep Reinforcement LearningabstractHigh-performance vision-based decision-making networks are often limited by hardware capabilities in practical applications. To address this challenge, this study proposes lightweight optimization strategies for decision-making models from the aspects of parameter size, training memory usage, and inference speed. Specifically, an innovative solution is proposed to achieve lightweight parameters. The Video Swin Transformer is employed to simultaneously extract temporal and spatial features, with the network trained using a Prioritized Replay Deep Q-Network (PRDQN) that incorporates risk assessment. To further reduce training memory usage, the Q-target network in PRDQN is removed, and the mellowmax operator is integrated to enhance the training process, resulting in the PRDeepMellow Swin Transformer. After analyzing the inference speed problems encountered by the algorithm in practical applications, the vanilla self-attention is replaced by a linear self-attention based on double softmax, namely Double Softmax Linear Video Swin Transformer (DSLVS Transformer) which improves the inference speed for long sequences. The proposed methods were evaluated across three high-speed lane change scenarios (a static scenario, a dynamic scenario, and a randomly changing scenario). Experimental results demonstrate that the proposed methods can still maintain excellent decision performance after the corresponding lightweight optimizations. Guofa Li, Yifan Qiu, Qingkun Li, Jie Li 0042, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Cross-Driver Domain Generalization for Improved Drowsiness Recognition Based on EEG SignalsabstractDesigning brain-computer interface systems for electroencephalogram (EEG)-based driver drowsiness recognition remains a significant challenge due to the significant variation in EEG signals across subjects and recording sessions. To address this problem, this paper develops a novel two-stage information transfer strategy framework for domain generalization. The framework has two domain mappers to reduce the distribution differences of EEG features from different individuals, a mapper mix block for generating hybrid mapping features, and a domain adversarial neural network (DANN) for drowsiness recognition based on hybrid EEG features. In the process of DANN to capture common features, we additionally employ two models based on self-attention mechanism to capture domain-invariant attention relationships between electrode channels and between frequency bands. Experimental results show that the proposed framework achieves an average accuracy of 81.34% in the leave-one-out cross validation for driver drowsiness recognition, which is higher than the state-of-the-art model with the number of 79.37%. In addition, we explore the impact of EEG features from different frequency bands and brain regions on this cross-subject task. The results show that EEG features from delta, theta and alpha bands can achieve much better performance than the other two bands, and features from the frontal lobe region perform better than the other regions. These findings reveal domain-invariant features and their relationships with brain regions and frequency bands, enhancing our understanding of the underlying messages of EEG signals. Guofa Li, Delin Ouyang, Qingkun Li, Zhenning Li 0001, Shengbo Eben Li, Cristina Olaverri-Monreal |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Chain-of-Thought Guided Multimodal Large Language Models for Scene-Aware Accident Anticipation in Autonomous DrivingabstractAccurately anticipating traffic accidents is a fundamental task for the safe and effective deployment of autonomous vehicles (AVs). However, existing models primarily rely on dashcam footage and often fail to generalize across varied driving scenarios due to their dependence on visual data and the rarity of high-risk events in datasets. These limitations undermine their robustness and reduce practical applicability in dynamic, unpredictable environments. To address these challenges, this study proposes a novel approach, termed MLTA, which integrates multimodal learning with the hypergraph attention network to hierarchically extract and capture cross-modal interaction. It leverages LLava-next, a multimodal large language model (MLLM) guided by the Chain-of-Thought (CoT) prompting paradigm, to produce context-aware interpretations of traffic scenes. This is further enhanced by a human-inspired attention mechanism that mimics the decision-making priorities of experienced human drivers. This combination enables more accurate identification of critical elements in a scene, improving both prediction precision and timeliness. Extensive experiments on four real-world datasets—DAD, A3D, CCD, and DADA-2000—show that our approach consistently outperforms state-of-the-art (SOTA) methods, demonstrating strong adaptability and robustness in complex driving environments. Haicheng Liao, Bin Rao 0003, Chengyue Wang 0001, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Minds on the Move: Decoding Trajectory Prediction in Autonomous Driving With Cognitive InsightsabstractIn mixed autonomous driving environments, accurately predicting the future trajectories of surrounding vehicles is crucial for the safe operation of autonomous vehicles (AVs). In driving scenarios, a vehicle’s trajectory is determined by the decision-making process of human drivers. However, existing models primarily focus on the inherent statistical patterns in the data, often neglecting the critical aspect of understanding the decision-making processes of human drivers. This oversight results in models that fail to capture the true intentions of human drivers, leading to suboptimal performance in long-term trajectory prediction. To address this limitation, we introduce a Cognitive-Informed Transformer (CITF) that incorporates a cognitive concept, Perceived Safety, to interpret drivers’ decision-making mechanisms. Perceived Safety encapsulates the varying risk tolerances across drivers with different driving behaviors. Specifically, we develop a Perceived Safety-aware Module that includes a Quantitative Safety Assessment for measuring the subject risk levels within scenarios, and Driver Behavior Profiling for characterizing driver behaviors. Furthermore, we present a novel module, Leanformer, designed to capture social interactions among vehicles. CITF demonstrates significant performance improvements on three well-established datasets. In terms of long-term prediction, it surpasses existing benchmarks by 12.0% on the NGSIM, 28.2% on the HighD, and 20.8% on the MoCAD dataset. Additionally, its robustness in scenarios with limited or missing data is evident, surpassing most state-of-the-art (SOTA) baselines, and paving the way for real-world applications. Haicheng Liao, Chengyue Wang 0001, Kaiqun Zhu, Yilong Ren, Bolin Gao, Shengbo Eben Li, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | OE-BevSeg: An Object Informed and Environment Aware Multimodal Framework for Bird's-Eye-View Vehicle Semantic SegmentationabstractBird’s-eye-view (BEV) semantic segmentation is becoming crucial in autonomous driving systems. It realizes ego-vehicle surrounding environment perception by projecting 2D multi-view images into 3D world space. Recently, BEV segmentation has made notable progress, attributed to better view transformation modules, larger image encoders, or more temporal information. However, there are still two issues: 1) a lack of effective understanding and enhancement of BEV space features, particularly in accurately capturing long-distance environmental features and 2) recognizing fine details of target objects. To address these issues, we propose OE-BevSeg, an end-to-end multimodal framework that enhances BEV segmentation performance through global environment-aware perception and local target object enhancement. OE-BevSeg employs an environment-aware BEV compressor. Based on prior knowledge about the main composition of the BEV surrounding environment varying with the increase of distance intervals, long-sequence global modeling is utilized to improve the model’s understanding and perception of the environment. From the perspective of enriching target object information in segmentation results, we introduce the center-informed object enhancement module, using centerness information to supervise and guide the segmentation head, thereby enhancing segmentation performance from a local enhancement perspective. Additionally, we designed a multimodal fusion branch that integrates multi-view RGB image features with radar/LiDAR features, achieving significant performance improvements. Extensive experiments show that, whether in camera-only or multimodal fusion BEV segmentation tasks, our approach achieves state-of-the-art results by a large margin on the nuScenes dataset for vehicle segmentation, demonstrating superior applicability in the field of autonomous driving. Our code will be released at https://github.com/SunJ1025/OE-BevSeghttps://github.com/SunJ1025/OE-BevSeg. Jian Sun 0038, Yuqi Dai, Chi-Man Vong, Qing Xu 0010, Shengbo Eben Li, Jianqiang Wang 0003, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Risk-Aware Vehicle Trajectory Prediction Under Safety-Critical ScenariosabstractTrajectory prediction is significant for intelligent vehicles to achieve high-level autonomous driving, and a lot of relevant research achievements have been made recently. Despite the rapid development, most existing studies solely focus on normal and safe scenarios while largely neglecting safety-critical scenarios, particularly those involving imminent collisions. This oversight may result in autonomous vehicles lacking the essential predictive ability in such situations, posing a significant threat to safety. To tackle these, this paper proposes a risk-aware trajectory prediction framework tailored to safety-critical scenarios. Leveraging distinctive hazardous features, we develop three core risk-aware components. First, we introduce a risk-incorporated scene encoder, which augments conventional encoders with quantitative risk information to achieve risk-aware encoding of hazardous scene contexts. Next, we incorporate endpoint-risk-combined intention queries as prediction priors in the decoder to ensure that the predicted multimodal trajectories cover both various spatial intentions and risk levels. Lastly, an auxiliary risk prediction task is implemented for the ultimate risk-aware prediction. Furthermore, to support model training and performance evaluation, we introduce a safety-critical trajectory prediction dataset and tailored evaluation metrics. We conduct comprehensive evaluations and compare our model with several SOTA models. Results demonstrate the superior performance of our model, with a significant improvement in most metrics. This prediction advancement enables autonomous vehicles to execute correct collision avoidance maneuvers under safety-critical scenarios, eventually enhancing road traffic safety. Qingfan Wang, Gaoyuan Kuang, Chen Lv 0001, Shengbo Eben Li, Bingbing Nie |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Multi-Scale Reinforcement Learning of Dynamic Energy Controller for Connected Electrified VehiclesabstractThe synergy of reinforcement learning (RL)-based energy management and vehicle-to-everything communication has been proved effective in boosting the fuel economy of connected plug-in hybrid electric vehicles (PHEVs). However, the intricate coupling of mechanical, electrical, thermal states and driving cycle results in a high-dimensional complex energy control problem for PHEVs, which is challenging to optimally solve within the same time scale. To this end, this study designs a multi-horizon reinforcement learning (MHRL)-based energy management of PHEVs, aware of the traffic preview from intelligent transportation systems to optimize the energy flow and thermal states as well as the transient dynamics of the powertrain. The proposed strategy features a novel state space representation, and solves the coordinated training among multiple sub-networks belonging to different control tasks in various time scales. Simulation and hardware-in-the-loop experiments are carried out based on a standard driving cycle and a real-world driving cycle with real-time traffic data demonstrate that the MHRL strategy improves fuel economy by 3.0%~7.9% compared to conventional RL-based energy management under various coolant temperature conditions and dynamic driving scenarios. Hao Zhang 0131, Shengbo Eben Li, Junzhi Zhang, Zhi Wang 0024 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Robust Approximate Dynamic Programming for Nonlinear Systems With Both Model Error and External DisturbanceabstractModel error and external disturbance have been separately addressed by optimizing the definite performance in standard linear control problems. However, the concurrent handling of both introduces uncertainty and nonconvexity into the performance, posing a huge challenge for solving nonlinear problems. This article introduces an additional cost function in the augmented Hamilton-Jacobi-Isaacs (HJI) equation of zero-sum games to simultaneously manage the model error and external disturbance in nonlinear robust performance problems. For satisfying the Hamilton-Jacobi inequality in nonlinear robust control theory under all considered model errors, the relationship between the additional cost function and model uncertainty is revealed. A critic online learning algorithm, applying Lyapunov stabilizing terms and historical states to reinforce training stability and achieve persistent learning, is proposed to approximate the solution of the augmented HJI equation. By constructing a joint Lyapunov candidate about the critic weight and system state, both stability and convergence are proved by the second method of Lyapunov. Theoretical results also show that introducing historical data reduces the ultimate bounds of system state and critic error. Three numerical examples are conducted to demonstrate the effectiveness of the proposed method. Jie Li 0042, Ryozo Nagamune, Yuhang Zhang 0018, Shengbo Eben Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Conformal Symplectic Optimization for Stable Reinforcement LearningabstractTraining deep reinforcement learning (RL) agents necessitates overcoming the highly unstable nonconvex stochastic optimization inherent in the trial-and-error mechanism. To tackle this challenge, we propose a physics-inspired optimization algorithm called relativistic adaptive gradient descent (RAD), which enhances long-term training stability. By conceptualizing neural network (NN) training as the evolution of a conformal Hamiltonian system, we present a universal framework for transferring long-term stability from conformal symplectic integrators to iterative NN updating rules, where the choice of kinetic energy governs the dynamical properties of resulting optimization algorithms. By utilizing relativistic kinetic energy, RAD incorporates principles from special relativity and limits parameter updates below a finite speed, effectively mitigating abnormal gradient influences. In addition, RAD models NN optimization as the evolution of a multiparticle system where each trainable parameter acts as an independent particle with an individual adaptive learning rate. We prove RAD's sublinear convergence under general nonconvex settings, where smaller gradient variance and larger batch sizes contribute to tighter convergence. Notably, RAD degrades to the well-known adaptive moment estimation (ADAM) algorithm when its speed coefficient is chosen as one and symplectic factor as a small positive value. Experimental results show RAD outperforming nine baseline optimizers with five RL algorithms across twelve environments, including standard benchmarks and challenging scenarios. Notably, RAD achieves up to a 155.1% performance improvement over ADAM in Atari games, showcasing its efficacy in stabilizing and accelerating RL training. Yao Lyu, Xiangteng Zhang, Shengbo Eben Li, Jingliang Duan, Letian Tao, Qing Xu 0010, Keqiang Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Learn Zero-Constraint-Violation Safe Policy in Model-Free Constrained Reinforcement LearningabstractWe focus on learning the zero-constraint-violation safe policy in model-free reinforcement learning (RL). Existing model-free RL studies mostly use the posterior penalty to penalize dangerous actions, which means they must experience the danger to learn from the danger. Therefore, they cannot learn a zero-violation safe policy even after convergence. To handle this problem, we leverage the safety-oriented energy functions to learn zero-constraint-violation safe policies and propose the safe set actor-critic (SSAC) algorithm. The energy function is designed to increase rapidly for potentially dangerous actions, locating the safe set on the action space. Therefore, we can identify the dangerous actions prior to taking them and achieve zero-constraint violation. Our major contributions are twofold. First, we use the data-driven methods to learn the energy function, which releases the requirement of known dynamics. Second, we formulate a constrained RL problem to solve the zero-violation policies. We prove that our Lagrangian-based constrained RL solutions converge to the constrained optimal zero-violation policies theoretically. The proposed algorithm is evaluated on the complex simulation environments and a hardware-in-loop (HIL) experiment with a real autonomous vehicle controller. Experimental results suggest that the converged policies in all environments achieve zero-constraint violation and comparable performance with model-based baseline. Haitong Ma, Changliu Liu, Shengbo Eben Li, Sifa Zheng, Jianyu Chen 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SEPT: Towards Efficient Scene Representation Learning for Motion PredictionabstractMotion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper presents SEPT, a modeling framework that leverages self-supervised learning to develop powerful spatiotemporal understanding for complex traffic scenes. Specifically, our approach involves three masking-reconstruction modeling tasks on scene inputs including agents' trajectories and road network, pretraining the scene encoder to capture kinematics within trajectory, spatial structure of road network, and interactions among roads and agents. The pretrained encoder is then finetuned on the downstream forecasting task. Extensive experiments demonstrate that SEPT, without elaborate architectural design or manual feature engineering, achieves state-of-the-art performance on the Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming previous methods on all main metrics by a large margin. Zhiqian Lan, Yuxuan Jiang 0011, Yao Mu 0001, Chen Chen 0068, Shengbo Eben Li |
ICLR | 5 |
| 2024 | Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelabstractSafe offline reinforcement learning is a promising way to bypass risky online interactions towards safe policy learning. Most existing methods only enforce soft constraints, i.e., constraining safety violations in expectation below thresholds predetermined. This can lead to potentially unsafe outcomes, thus unacceptable in safety-critical scenarios. An alternative is to enforce the hard constraint of zero violation. However, this can be challenging in offline setting, as it needs to strike the right balance among three highly intricate and correlated aspects: safety constraint satisfaction, reward maximization, and behavior regularization imposed by offline datasets. Interestingly, we discover that via reachability analysis of safe-control theory, the hard safety constraint can be equivalently translated to identifying the largest feasible region given the offline dataset. This seamlessly converts the original trilogy problem to a feasibility-dependent objective, i.e., maximizing reward value within the feasible region while minimizing safety risks in the infeasible region. Inspired by these, we propose FISOR (FeasIbility-guided Safe Offline RL), which allows safety constraint adherence, reward maximization, and offline policy learning to be realized via three decoupled processes, while offering strong safety performance and stability. In FISOR, the optimal policy for the translated optimization problem can be derived in a special form of weighted behavior cloning, which can be effectively extracted with a guided diffusion model thanks to its expressiveness. We compare FISOR against baselines on DSRL benchmark for safe offline RL. Evaluation results show that FISOR is the only method that can guarantee safety satisfaction in all tasks, while achieving top returns in most tasks. Code: https://github.com/ZhengYinan-AIR/FISOR. Yinan Zheng, Dongjie Yu, Shengbo Eben Li, Xianyuan Zhan |
ICLR | 5 |
| 2024 | Feasible Reachable Policy IterationabstractThe goal-reaching tasks with safety constraints are common control problems in real world, such as intelligent driving and robot manipulation. The difficulty of this kind of problem comes from the exploration termination caused by safety constraints and the sparse rewards caused by goals. The existing safe RL avoids unsafe exploration by restricting the search space to a feasible region, the essence of which is the pruning of the search space. However, there are still many ineffective explorations in the feasible region because of the ignorance of the goals. Our approach considers both safety and goals; the policy space pruning is achieved by a function called feasible reachable function, which describes whether there is a policy to make the agent safely reach the goals in the finite time domain. This function naturally satisfies the self-consistent condition and the risky Bellman equation, which can be solved by the fixed point iteration method. On this basis, we propose feasible reachable policy iteration (FRPI), which is divided into three steps: policy evaluation, region expansion, and policy improvement. In the region expansion step, by using the information of agent to reach the goals, the convergence of the feasible region is accelerated, and simultaneously a smaller feasible reachable region is identified. The experimental results verify the effectiveness of the proposed FR function in both improving the convergence speed of better or comparable performance without sacrificing safety and identifying a smaller policy space with higher sample efficiency. Shentao Qin, Yao Mu 0001, Jie Li 0042, Wenjun Zou, Jingliang Duan, Shengbo Eben Li |
ICML | 7 |
| 2024 | Human Observation-Inspired Trajectory Prediction for Autonomous Driving in Mixed-Autonomy Traffic EnvironmentsabstractIn the burgeoning field of autonomous vehicles (AVs), trajectory prediction remains a formidable challenge, especially in mixed autonomy environments. Traditional approaches often rely on computational methods such as time-series analysis. Our research diverges significantly by adopting an interdisciplinary approach that integrates principles of human cognition and observational behavior into trajectory prediction models for AVs. We introduce a novel “adaptive visual sector” mechanism that mimics the dynamic allocation of attention human drivers exhibit based on factors like spatial orientation, proximity, and driving speed. Additionally, we develop a “dynamic traffic graph” using Convolutional Neural Networks (CNN) and Graph Attention Networks (GAT) to capture spatio-temporal dependencies among agents. Benchmark tests on the NGSIM, HighD, and MoCAD datasets reveal that our model (GAVA) outperforms state-of-the-art baselines by at least 15.2%, 19.4%, and 12.0%, respectively. Our findings underscore the potential of leveraging human cognition principles to enhance the proficiency and adaptability of trajectory prediction algorithms in AVs. Haicheng Liao, Shangqian Liu, Yongkang Li 0003, Zhenning Li 0001, Chengyue Wang 0001, Yunjian Li, Shengbo Eben Li, Cheng-Zhong Xu 0001 |
ICRA | 7 |
| 2024 | Synthesize Efficient Safety Certificates for Learning-Based Safe Control using Magnitude RegularizationabstractSafety certificates based on energy functions can provide demonstrable safety for complex robotic systems. However, all recent studies on learning-based energy function synthesis only consider the feasibility of the control policy, which might cause over-conservativeness and even fail to achieve the control goal. To solve the problem of over-conservative controllers, we proposed the magnitude regularization technique to improve the controller performance of safe controllers by reducing the conservativeness inside the energy function, while keeping the promising provable safety guarantees. Specifically, we quantify the conservativeness by the magnitude of the energy function, and we reduce the conservativeness by adding a magnitude regularization term to the synthesis loss. We propose an algorithm using reinforcement learning (RL) for synthesis to unify the learning process of safe controllers and energy functions. We conducted simulation experiments on Safety Gym and real-robot experiments using small quadrotors. Simulation results show that the proposed algorithm does reduce the conservativeness of the energy function and outperforms baselines in terms of controller performance while maintaining safety. Real-robot experiments have shown that the proposed algorithm indeed reduce conservativeness on the small quadrotors. Haitong Ma, Sifa Zheng, Shengbo Eben Li, Jianqiang Wang 0003 |
ICRA | 4 |
| 2024 | Rocket Landing Control with Random Annealing Jump Start Reinforcement LearningabstractRocket recycling is a crucial pursuit in aerospace technology, aimed at reducing costs and environmental impact in space exploration. The primary focus centers on rocket landing control, involving the guidance of a nonlinear under-actuated rocket with limited fuel in real-time. This challenging task prompts the application of reinforcement learning (RL), yet goal-oriented nature of the problem poses difficulties for standard RL algorithms due to the absence of intermediate reward signals. This paper, for the first time, significantly elevates the success rate of rocket landing control from 8% with a baseline controller to 97% on a high-fidelity rocket model using RL. Our approach, called Random Annealing Jump Start (RAJS), is tailored for real-world goal-oriented problems by leveraging prior feedback controllers as guide policy to facilitate environmental exploration and policy learning in RL. In each episode, the guide policy navigates the environment for the guide horizon, followed by the exploration policy taking charge to complete remaining steps. This jump-start strategy prunes exploration space, rendering the problem more tractable to RL algorithms. The guide horizon is sampled from a uniform distribution, with its upper bound annealing to zero based on performance metrics, mitigating distribution shift and mismatch issues in existing methods. Additional enhancements, including cascading jump start, refined reward and terminal condition, and action smoothness regulation, further improve policy performance and practical applicability. The proposed method is validated through extensive evaluation and Hardware-in-the-Loop testing, affirming the effectiveness, real-time feasibility, and smoothness of the proposed controller. Yuxuan Jiang 0011, Zhiqian Lan, Guojian Zhan, Shengbo Eben Li, Qi Sun 0004, Tianwen Yu, Changwu Zhang |
IROS | 5 |
| 2024 | Control Safety Function for Explicit Safety-Critical Control of Autonomous VehiclesabstractReal-time safety-critical control is essential for high-level autonomous driving. Existing methods usually formulate safety-critical control as a constrained optimal control problem (COCP), and suffer from high computational complexity of the underlying iterative optimization processes. To address the issue of complexity, this paper presents an explicit safety-critical control method called the Control Safety Function (CSF) approach, which can replace online optimization with an analytical control law, dramatically enhancing real-time control capabilities. The CSF is formulated as the weighted sum of Control Lyapunov Function (CLF) and Control Barrier Functions (CBFs), with the value of CSF increasing to infinity as the state approaches the boundary of a safe set. The explicit control law is then derived from the gradient of CSF and system dynamics. Different from existing explicit controllers that can only apply to systems of relative degree one, the CSF method provides an approach to enforce safety constraints to systems with high relative degree, making CSF especially suitable for autonomous driving. The CSF approach is evaluated in a vehicle path-tracking scenario with multiple obstacles, accompanied by a comparative analysis against the Model Predictive Control (MPC) method. Simulation results indicate that CSF achieves control accuracy comparable to MPC, with significant reduction in computation time - approximately 3.23 ms per step, which is about 94.0% faster. These results suggest that CSF is a promising approach for real-time safety-critical control of high-level autonomous driving. Dongyoon Kim, Sen Yang 0023, Wenjun Zou, Bin Shuai, Dezhao Zhang, Fang Zhang 0002, Chang Liu 0002, Shengbo Eben Li |
IV | 8 |
| 2024 | Diffusion Actor-Critic with Entropy RegulatorabstractReinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to acquire complex policies. In response to this problem, we propose an online RL algorithm termed diffusion actor-critic with entropy regulator (DACER). This algorithm conceptualizes the reverse process of the diffusion model as a novel policy function and leverages the capability of the diffusion model to fit multimodal distributions, thereby enhancing the representational capacity of the policy. Since the distribution of the diffusion policy lacks an analytical expression, its entropy cannot be determined analytically. To mitigate this, we propose a method to estimate the entropy of the diffusion policy utilizing Gaussian mixture model. Building on the estimated entropy, we can learn a parameter $\alpha$ that modulates the degree of exploration and exploitation. Parameter $\alpha$ will be employed to adaptively regulate the variance of the added noise, which is applied to the action output by the diffusion model. Experimental trials on MuJoCo benchmarks and a multimodal task demonstrate that the DACER algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting a stronger representational capacity of the diffusion policy. Yuxuan Jiang 0011, Wenjun Zou, Xujie Song, Wenxuan Wang 0004, Liming Xiao, Jingliang Duan, Shengbo Eben Li |
NeurIPS | 11 |
| 2024 | Safe Reinforcement Learning With Dual RobustnessabstractReinforcement learning (RL) agents are vulnerable to adversarial disturbances, which can deteriorate task performance or break down safety specifications. Existing methods either address safety requirements under the assumption of no adversary (e.g., safe RL) or only focus on robustness against performance adversaries (e.g., robust RL). Learning one policy that is both safe and robust under any adversaries remains a challenging open problem. The difficulty is how to tackle two intertwined aspects in the worst cases: feasibility and optimality. The optimality is only valid inside a feasible region (i.e., robust invariant set), while the identification of maximal feasible region must rely on how to learn the optimal policy. To address this issue, we propose a systematic framework to unify safe RL and robust RL, including the problem formulation, iteration scheme, convergence analysis and practical algorithm design. The unification is built upon constrained two-player zero-sum Markov games, in which the objective for protagonist is twofold. For states inside the maximal robust invariant set, the goal is to pursue rewards under the condition of guaranteed safety; for states outside the maximal robust invariant set, the goal is to reduce the extent of constraint violation. A dual policy iteration scheme is proposed, which simultaneously optimizes a task policy and a safety policy. We prove that the iteration scheme converges to the optimal task policy which maximizes the twofold objective in the worst cases, and the optimal safety policy which stays as far away from the safety boundary. The convergence of safety policy is established by exploiting the monotone contraction property of safety self-consistency operators, and that of task policy depends on the transformation of safety constraints into state-dependent action spaces. By adding two adversarial networks (one is for safety guarantee and the other is for task performance), we propose a practical deep RL algorithm for constrained zero-sum Markov games, called dually robust actor-critic (DRAC). The evaluations with safety-critical benchmarks demonstrate that DRAC achieves high performance and persistent safety under all scenarios (no adversary, safety adversary, performance adversary), outperforming all baselines by a large margin. Zeyang Li 0001, Chuxiong Hu, Yunan Wang, Shengbo Eben Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Safe Model-Based Reinforcement Learning With an Uncertainty-Aware Reachability CertificateabstractSafe reinforcement learning (RL) that solves constraint-satisfactory policies provides a promising way to the broader safety-critical applications of RL in real-world problems such as robotics. Among all safe RL approaches, model-based methods reduce training time violations further due to their high sample efficiency. However, lacking safety robustness against the model uncertainties remains an issue in safe model-based RL, especially in training time safety. In this paper, we propose a distributional reachability certificate (DRC) and its Bellman equation to address model uncertainties and characterize robust persistently safe states. Furthermore, we build a safe RL framework to resolve constraints required by the DRC and its corresponding shield policy. We also devise a line search method to maintain safety and reach higher returns simultaneously while leveraging the shield policy. Comprehensive experiments on classical benchmarks such as constrained tracking and navigation indicate that the proposed algorithm achieves comparable returns with much fewer constraint violations during training. Our code is available at https://github.com/ManUtdMoon/Distributional-Reachability-Policy-Optimization.Note to Practitioners—Although it has been proven that RL can be applied in complex robotics control tasks, the training process of an RL control policy induces frequent failures because the agent needs to learn safety through constraint violations. This issue hinders the promotion of RL because a large amount of failure of robots is too expensive to afford. This paper aims to reduce the training-time violations of RL-based control methods, enabling RL to be leveraged in a broader application area. To achieve the goal, we first introduce a safety quantity describing the distribution of potential constraint violations in the long term. By imposing constraints on the quantile of the safety distribution, we can realize safety robust to the model uncertainty, which is necessary for real-world robot learning with environment uncertainty. Second, we further devise a shield policy aiming to minimize the constraint violation. The policy will intervene when the agent is about to violate state constraints, further enhancing exploration safety. Third, we implement a line search method to find an action pursuing near-optimal performance when fulfilling safety requirements strictly. Our experimental results indicate that the proposed algorithm reduces training-time violations significantly while maintaining competitive task performance. We make a step towards applying RL safely in real-world tasks. Our future work includes conducting physical verification on real robots to evaluate the algorithm and improving safety further by starting from an initially safe control policy that comes from domain knowledge. Dongjie Yu, Wenjun Zou, Haitong Ma, Shengbo Eben Li, Yuming Yin, Jianyu Chen 0002, Jingliang Duan |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Optimization Landscape of Policy Gradient Methods for Discrete-Time Static Output FeedbackabstractIn recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control, output-feedback control is more prevalent since the underlying state of the system may not be fully observed in many practical settings. This article analyzes the optimization landscape inherent to policy gradient methods when applied to static output feedback (SOF) control in discrete-time LTI systems subject to quadratic cost. We begin by establishing crucial properties of the SOF cost, encompassing coercivity, L -smoothness, and M -Lipschitz continuous Hessian. Despite the absence of convexity, we leverage these properties to derive novel findings regarding convergence (and nearly dimension-free rate) to stationary points for three policy gradient methods, including the vanilla policy gradient method, the natural policy gradient method, and the Gauss-Newton method. Moreover, we provide proof that the vanilla policy gradient method exhibits linear convergence toward local minima when initialized near such minima. This article concludes by presenting numerical examples that validate our theoretical findings. These results not only characterize the performance of gradient descent for optimizing the SOF problem but also provide insights into the effectiveness of general policy gradient methods within the realm of reinforcement learning. Jingliang Duan, Jie Li 0042, Kai Zhao 0004, Shengbo Eben Li, Lin Zhao 0009 |
IEEE Trans. Cybern. | 5 |
| 2024 | Inverse Model Predictive Control: Learning Optimal Control Cost Functions for MPCabstractInverse optimal control (IOC) seeks to infer a control cost function that captures the underlying goals and preferences of expert demonstrations. While significant progress has been made in finite-horizon IOC, which focuses on learning control cost functions based on rollout trajectories rather than actual trajectories, the application of IOC to receding horizon control, also known as model predictive control (MPC), has been overlooked. MPC is more prevalent in practical settings and poses additional challenges for IOC learning since it is complicated to calculate the gradient of actual trajectories with respect to cost parameters. In light of this, we propose the inverse MPC (IMPC) method to identify the optimal cost function that effectively minimizes the discrepancy between the actual trajectory and its associated demonstration. To compute the gradient of actual trajectories with respect to cost parameters, we first establish two differential Pontryagin's maximum principle (PMP) conditions by differentiating the traditional PMP conditions with respect to cost parameters and initial states, respectively. We then formulate two auxiliary optimal control problems based on the derived differentiated PMP conditions, whose solutions can be directly used to determine the gradient for updating cost parameters. We validate the efficacy of the proposed method through experiments involving five simulation tasks and two real-world mobile robot control tasks. The results consistently demonstrate that IMPC outperforms existing finite-horizon IOC methods across all experiments. Fawang Zhang, Jingliang Duan, Hao Chen 0108, Hui Liu 0001, Shida Nie, Shengbo Eben Li |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | Quantifying the Individual Differences of Drivers' Risk Perception via Potential Damage Risk ModelabstractThere will be a time when automated vehicles coexist with human-driven ones. Understanding how drivers assess driving risks and modeling their differences is crucial for developing human-like and personalized behaviors in automated vehicles, gaining people’s trust and acceptance. However, existing driving risk models are usually developed at a statistical level, and no single model can accurately describe and explain the variations in risk perception among drivers. We propose a concise yet effective model known as the Potential Damage Risk (PODAR) model, which provides a universal and physically meaningful structure for estimating driving risk and explaining the reasons for differences in risk perception. Leveraging an open-access dataset collected from an obstacle avoidance experiment, this paper establishes individual risk perception models for drivers with high fitness performances. We conclude that the variations in risk perception among drivers stem from their assessments of potential damage, accounting for the uncertainty in both temporal and spatial dimensions. Our findings offer an explanation for human risk perceptions and present a promising risk model for autonomous vehicles to develop human-like behaviors and personalized services. Chen Chen 0068, Zhiqian Lan, Guojian Zhan, Yao Lyu, Bingbing Nie, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Enhance Sample Efficiency and Robustness of End-to-End Urban Autonomous Driving via Semantic Masked World ModelabstractEnd-to-end autonomous driving provides a feasible way to automatically maximize overall driving system performance by directly mapping the raw pixels from a front-facing camera to control signals. Recent advanced methods construct a latent world model to map the high dimensional observations into compact latent space. However, the latent states embedded by the world model proposed in previous works may contain a large amount of task-irrelevant information, resulting in low sampling efficiency and poor robustness to input perturbations. Meanwhile, the training data distribution is usually unbalanced, and the learned policy is challenging to cope with the corner cases during the driving process. To solve the above challenges, we present aSEManticMasked recurrent world model (SEM2), which introduces a semantic filter to extract key driving-relevant features and make decisions via the filtered features, and is trained with a multi-source data sampler, which aggregates common data and multiple corner case data in a single batch, to balance the data distribution. Extensive experiments on CARLA show our method outperforms the state-of-the-art approaches in terms of sample efficiency and robustness to input permutations. Yao Mu 0001, Chen Chen 0068, Jingliang Duan, Ping Luo 0002, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | A Reinforcement Learning Benchmark for Autonomous Driving in General Urban ScenariosabstractReinforcement learning (RL) has gained significant interest for its potential to improve decision and control in autonomous driving. However, current approaches have yet to demonstrate sufficient scenario generality and observation generality, hindering their wider utilization. To address these limitations, we propose a unified benchmark simulator for RL algorithms (called IDSim) to facilitate decision and control for high-level autonomous driving, with emphasis on diverse scenarios and a unified observation interface. IDSim is composed of a scenario library and a simulation engine, and is designed with execution efficiency and determinism in mind. The scenario library covers common urban scenarios, with automated random generation of road structure and traffic flow, and the simulation engine operates on the generated scenarios with dynamic interaction support. We conduct four groups of benchmark experiments with five common RL algorithms and focus on challenging signalized intersection scenarios with varying conditions. The results showcase the reliability of the simulator and reveal its potential to improve the generality of RL algorithms. Our analysis suggests that multi-task learning and observation design are potential areas for further algorithm improvement. Yuxuan Jiang 0011, Guojian Zhan, Zhiqian Lan, Chang Liu 0002, Bo Cheng 0003, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | A RGB-Thermal Image Segmentation Method Based on Parameter Sharing and Attention Fusion for Safe Autonomous DrivingabstractIn this paper, we propose a new RGB-thermal image segmentation method based on parameter sharing and attention fusion for safe autonomous driving. An encoder-decoder network structure is adopted. The encoder, which has shared convolution layer parameters and private batch normalization layer parameters (parameter sharing scheme), is used to extract features from RGB and thermal images. The extracted features are then fused by spatial and channel attention. The output of each residual block is fused, and the self-learning weight is used to integrate the fusion information of all residual blocks of the same levels. Subsequently, the fused features are integrated through a feature integration (FI) module in the decoder. Cross-entropy supervision of segmentation and edge is performed on the outputs of the decoders. Our proposed method is evaluated and compared with 17 state-of-the-art image segmentation methods, both qualitatively and quantitatively on the MFNet dataset which includes various objects in urban scenes. The results show that the proposed method outperforms previous methods by at least 0.3% and 1.8% in MRecall and MIoU, respectively, providing foundations for the development of autonomous driving technologies for safety enhancement. Guofa Li, Yongjie Lin, Delin Ouyang, Shen Li 0001, Xingda Qu, Dawei Pi, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | A Transformation-Aggregation Framework for State Representation of Autonomous Driving SystemsabstractThe effective design of state representation plays a pivotal role in bridging the gap between perception and downstream decision-making and control in autonomous driving systems based on reinforcement learning. Among perception observations, accurately representing the set of surrounding participants poses significant challenges due to its variable size and unordered nature. These challenges result in dimension sensitivity and permutation sensitivity issues, respectively. To systematically tackle these complexities, we introduce a general framework for learning a state representation module, which can be regarded as a combination of any transformation and aggregation modules. Specifically, we employ a transformer encoder as the transformation module to attract features individually and design a novel aggregation module called Feature-wise Sorted Query (FSQ) to effectively aggregate the feature set into a state representation vector while adaptively attending to essential participant features. In FSQ, the feature-wise sort and latent query attention mechanisms can address permutation sensitivity and dimension sensitivity, respectively. Additionally, we propose an offline training paradigm to decouple state representation learning from downstream reinforcement learning tasks, enhancing stability and avoiding overfitting for particular scenarios. Simulation and experimental results demonstrate the superior performance and a more stable training process of our method across six varying-size set regression tasks and downstream autonomous driving tasks, particularly when leveraging the offline training paradigm. Guojian Zhan, Yuxuan Jiang 0011, Shengbo Eben Li, Yao Lyu, Xiangteng Zhang, Yuming Yin |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral LagrangianabstractSafety is essential for reinforcement learning (RL) applied in the real world. Adding chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under uncertainty. Existing chance-constrained RL methods, such as the penalty methods and the Lagrangian methods, either exhibit periodic oscillations or learn an overconservative or unsafe policy. In this article, we address these shortcomings by proposing a separated proportional-integral Lagrangian (SPIL) algorithm. We first review the constrained policy optimization process from a feedback control perspective, which regards the penalty weight as the control input and the safe probability as the control output. Based on this, the penalty method is formulated as a proportional controller, and the Lagrangian method is formulated as an integral controller. We then unify them and present a proportional-integral Lagrangian method to get both their merits with an integral separation technique to limit the integral value to a reasonable range. To accelerate training, the gradient of safe probability is computed in a model-based manner. The convergence of the overall algorithm is analyzed. We demonstrate that our method can reduce the oscillations and conservatism of RL policy in a car-following simulation. To prove its practicality, we also apply our method to a real-world mobile robot navigation task, where our robot successfully avoids a moving obstacle with highly uncertain or even aggressive behaviors. Baiyu Peng, Jingliang Duan, Jianyu Chen 0002, Shengbo Eben Li, Genjin Xie, Congsheng Zhang, Yang Guan, Yao Mu 0001, Enxin Sun |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlabstractDeep reinforcement learning (RL) is a powerful approach for solving optimal control problems. However, RL-trained policies often suffer from the action fluctuation problem, where the consecutive actions significantly differ despite only slight state variations. This problem results in mechanical components' wear and tear and poses safety hazards. The action fluctuation is caused by the high Lipschitz constant of actor networks. To address this problem, we propose a neural network named LipsNet. We propose the Multi-dimensional Gradient Normalization (MGN) method, to constrain the Lipschitz constant of networks with multi-dimensional input and output. Benefiting from MGN, LipsNet achieves Lipschitz continuity, allowing smooth actions while preserving control performance by adjusting Lipschitz constant. LipsNet addresses the action fluctuation problem at network level rather than algorithm level, which can serve as actor networks in most RL algorithms, making it more flexible and user-friendly than previous works. Experiments demonstrate that LipsNet has good landscape smoothness and noise robustness, resulting in significantly smoother action compared to the Multilayer Perceptron. Xujie Song, Jingliang Duan, Wenxuan Wang 0004, Shengbo Eben Li, Chen Chen 0068, Bo Cheng 0003, Junqing Wei, Xiaoming Simon Wang |
ICML | 4 |
| 2023 | Integrated Decision and Control: Toward Interpretable and Computationally Efficient Driving IntelligenceabstractDecision and control are core functionalities of high-level automated vehicles. Current mainstream methods, such as functional decomposition and end-to-end reinforcement learning (RL), suffer high time complexity or poor interpretability and adaptability on real-world autonomous driving tasks. In this article, we present an interpretable and computationally efficient framework called integrated decision and control (IDC) for automated vehicles, which decomposes the driving task into static path planning and dynamic optimal tracking that are structured hierarchically. First, the static path planning generates several candidate paths only considering static traffic elements. Then, the dynamic optimal tracking is designed to track the optimal path while considering the dynamic obstacles. To that end, we formulate a constrained optimal control problem (OCP) for each candidate path, optimize them separately, and follow the one with the best tracking performance. To unload the heavy online computation, we propose a model-based RL algorithm that can be served as an approximate-constrained OCP solver. Specifically, the OCPs for all paths are considered together to construct a single complete RL problem and then solved offline in the form of value and policy networks for real-time online path selecting and tracking, respectively. We verify our framework in both simulations and the real world. Results show that compared with baseline methods, IDC has an order of magnitude higher online computing efficiency, as well as better driving performance, including traffic efficiency and safety. In addition, it yields great interpretability and adaptability among different driving scenarios and tasks. Yang Guan, Yangang Ren, Qi Sun 0004, Shengbo Eben Li, Haitong Ma, Jingliang Duan, Bo Cheng 0003 |
IEEE Trans. Cybern. | 4 |
| 2023 | Safe-State Enhancement Method for Autonomous Driving via Direct Hierarchical Reinforcement LearningabstractReinforcement learning (RL) has shown excellent performance in the sequential decision-making problem, where safety in the form of state constraints is of great significance in the design and application of RL. Simple constrained end-to-end RL methods might lead to significant failure in a complex system like autonomous vehicles. In contrast, some hierarchical RL (HRL) methods generate driving goals directly, which could be closely combined with motion planning. With safety requirements, some safe-enhanced RL methods add post-processing modules to avoid unsafe goals or achieve expectation-based safety, which accepts the existence of unsafe states and allows some violations of safe constraints. However, ensuring state safety is vital for autonomous vehicles. Therefore, this paper proposes a state-based safety enhancement method for autonomous driving via direct hierarchical reinforcement learning. Finally, we design a constrained reinforcement learner based on the State-based Constrained Markov Decision Process (SCMDP), where a learnable safety module could adjust the constraint strength adaptively. We integrate a dynamic module in the policy training and generate future goals considering safety, temporal-spatial continuity, and dynamic feasibility, which could eliminate dependence on the prior model. Simulations in the typical highway scenes with uncertainties show that the proposed method has better training performance, higher driving safety in interactive scenes, more decision intelligence in traffic congestions, and better economic driving ability on roads with changing slopes. Ziqing Gu, Lingping Gao, Haitong Ma, Shengbo Eben Li, Sifa Zheng, Junbo Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Policy Iteration Based Approximate Dynamic Programming Toward Autonomous Driving in Constrained Dynamic EnvironmentabstractIn the area of autonomous driving, it typically brings great difficulty in solving the motion planning problem since the vehicle model is nonlinear and the driving scenarios are complex. Particularly, most of the existing methods cannot be generalized to dynamically changing scenarios with varying surrounding vehicles. To address this problem, this development here investigates the framework of integrated decision and control. As part of the modules, static path planning determines the reference candidates ahead, and then the optimal path-tracking controller realizes the specific autonomous driving task. An innovative and effective constrained finite-horizon approximate dynamic programming (ADP) algorithm is herein presented to generate the desired control policy for effective path tracking. With the generalized policy neural network that maps from the state to the control input, the proposed algorithm preserves the high effectiveness for the motion planning problem towards changing driving environments with varying surrounding vehicles. Moreover, the algorithm attains the noteworthy advantage of alleviating the typically heavy computational loads with the mode of offline training and online execution. As a result of the utilization of multi-layer neural networks in conjunction with the actor-critic framework, the constrained ADP method is capable of handling complex and multidimensional scenarios. Finally, various simulations have been carried out to show that the constrained ADP algorithm is effective. Ziyu Lin, Jun Ma 0008, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Bo Cheng 0003, Tong Heng Lee |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal ControlabstractThe Hamilton-Jacobi-Bellman (HJB) equation serves as the necessary and sufficient condition for the optimal solution to the continuous-time (CT) optimal control problem (OCP). Compared with the infinite-horizon HJB equation, the solving of the finite-horizon (FH) HJB equation has been a long-standing challenge, because the partial time derivative of the value function is involved as an additional unknown term. To address this problem, this study first-time bridges the link between the partial time derivative and the terminal-time utility function, and thus it facilitates the use of the policy iteration (PI) technique to solve the CT FH OCPs. Based on this key finding, the FH approximate dynamic programming (ADP) algorithm is proposed leveraging an actor-critic framework. It is shown that the algorithm exhibits important properties in terms of convergence and optimality. Rather importantly, with the use of multilayer neural networks (NNs) in the actor-critic architecture, the algorithm is suitable for CT FH OCPs toward more general nonlinear and complex systems. Finally, the effectiveness of the proposed algorithm is demonstrated by conducting a series of simulations on both a linear quadratic regulator (LQR) problem and a nonlinear vehicle tracking problem. Ziyu Lin, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Jie Li 0042, Jianyu Chen 0002, Bo Cheng 0003, Jun Ma 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Flow-based Recurrent Belief State Learning for POMDPsabstractPartially Observable Markov Decision Process (POMDP) provides a principled and generic framework to model real world sequential decision making processes but yet remains unsolved, especially for high dimensional continuous space and unknown models. The main challenge lies in how to accurately obtain the belief state, which is the probability distribution over the unobservable environment states given historical information. Accurately calculating this belief state is a precondition for obtaining an optimal policy of POMDPs. Recent advances in deep learning techniques show great potential to learn good belief states. However, existing methods can only learn approximated distribution with limited flexibility. In this paper, we introduce the \textbf{F}l\textbf{O}w-based \textbf{R}ecurrent \textbf{BE}lief \textbf{S}tate model (FORBES), which incorporates normalizing flows into the variational inference to learn general continuous belief states for POMDPs. Furthermore, we show that the learned belief states can be plugged into downstream RL algorithms to improve performance. In experiments, we show that our methods successfully capture the complex belief states that enable multi-modal predictions as well as high quality reconstructions, and results on challenging visual-motor control tasks show that our method achieves superior performance and sample efficiency. Yao Mu 0001, Ping Luo 0002, Shengbo Eben Li, Jianyu Chen 0002 |
ICML | 4 |
| 2022 | Reachability Constrained Reinforcement LearningabstractConstrained reinforcement learning (CRL) has gained significant interest recently, since safety constraints satisfaction is critical for real-world problems. However, existing CRL methods constraining discounted cumulative costs generally lack rigorous definition and guarantee of safety. In contrast, in the safe control research, safety is defined as persistently satisfying certain state constraints. Such persistent safety is possible only on a subset of the state space, called feasible set, where an optimal largest feasible set exists for a given environment. Recent studies incorporate feasible sets into CRL with energy-based methods such as control barrier function (CBF), safety index (SI), and leverage prior conservative estimations of feasible sets, which harms the performance of the learned policy. To deal with this problem, this paper proposes the reachability CRL (RCRL) method by using reachability analysis to establish the novel self-consistency condition and characterize the feasible sets. The feasible sets are represented by the safety value function, which is used as the constraint in CRL. We use the multi-time scale stochastic approximation theory to prove that the proposed algorithm converges to a local optimum, where the largest feasible set can be guaranteed. Empirical results on different benchmarks validate the learned feasible set, the policy performance, and constraint satisfaction of RCRL, compared to CRL and safe control baselines. Dongjie Yu, Haitong Ma, Shengbo Eben Li, Jianyu Chen 0002 |
ICML | 3 |
| 2022 | Cola-HRL: Continuous-Lattice Hierarchical Reinforcement Learning for Autonomous DrivingabstractReinforcement learning (RL) has shown promising performance in autonomous driving applications in recent years. The early end-to-end RL method is usually unexplainable and fails to generate stable actions, while the hierarchical RL (HRL) method can tackle the above issues by dividing complex problems into multiple sub-tasks. Prior HRL works either select discrete driving behaviors with continuous control commands, or generate expected goals for the low-level controller. However, they typically have strong scenario dependence or fail to generate goals with good quality. To address the above challenges, we propose a Continuous-Lattice Hierarchical RL (Cola-HRL) method for autonomous driving tasks to make high-quality decisions in various scenarios. We utilize the continuous-lattice module to generate reasonable goals, ensuring temporal and spatial reachability. Then, we train and evaluate our method under different traffic scenarios based on real-world High Definition maps. Experimental results show our method can handle multiple scenarios. In addition, our method also demonstrates better performance and driving behaviors compared to existing RL methods. Lingping Gao, Ziqing Gu, Cong Qiu, Lanxin Lei, Shengbo Eben Li, Sifa Zheng, Junbo Chen |
IROS | 5 |
| 2022 | Beyond backpropagate through time: Efficient model-based training through time-splittingabstractModel-based policy gradient (MBPG) has been employed to seek an approximate solution to the optimal control problem. However, there is coupling between adjacent states due to temporal dependencies, making the training time grow linearly with the time horizon. This paper reshapes the training process of MBPG with the time-splitting technique to establish a time-independent algorithm called Training Through Time-Splitting (T3S). First, copy the coupled variables to obtain two independent variables. Meanwhile, an extra variable together with an equivalence constraint is introduced for problem consistency. Then, the transformed problem divides into subproblems with carefully derived loss functions. Subproblems own decoupled variables and shared policy networks, which means they can be optimized concurrently. Guided by the algorithm design, this paper further proposes an asynchronous parallel training scheme to accelerate training efficiency. Numerical simulation shows that the T3S algorithm outperforms the MBPG algorithm by 83.6% in wall-clock time with a trajectory tracking task. Jiaxin Gao 0002, Yang Guan, Shengbo Eben Li, Junqing Wei, Keqiang Li 0002 |
Int. J. Intell. Syst. | 4 |
| 2022 | Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints
Jingliang Duan, Shengbo Eben Li, Qi Sun 0004, Zhenzhong Jia, Bo Cheng 0003 |
Neurocomputing | 3 |
| 2022 | Fuel Economy Optimization for Platooning Vehicle Swarms via Distributed Economic Model Predictive ControlabstractCooperation among multiple connected vehicles (CVs) brings swarm intelligence to our transportation systems and helps improve their performance. This study proposes a fuel economy optimization approach for a platooning vehicle swarm via distributed economic model predictive control (DEMPC). Within the DEMPC framework, each CV shares its assumed trajectory in the predictive horizon with its neighboring CVs at each control loop. With its neighbors’ and its own assumed trajectories, each CV first solves an open-loop control optimization problem for platoon formation, and then solves an open-loop economic optimization problem for direct fuel economy improvement. In particular, the optimal cost of the former optimization problem is used in the latter one to build an upperbound constraint for stability guarantee. The asymptotic convergence of assumed terminal states is given in the analysis, and the recursive feasibility of the two optimization problems are proved with an explicit constraint on the weight matrices in the open-loop control optimization problem. Based on these analyses, the asymptotic stability of the closed-loop system is finally proved through Lyapunov analysis. Numerical simulation results validate the effectiveness of the proposed approach in terms of closed-loop stability and fuel economy improvement. Note to Practitioners—This work aims to optimize the fuel economy for a swarm of vehicles running in a platoon. Existing approaches on platoon control mainly focus on the accurate tracking of the following vehicles, but this may cause aggressive control and hence lead to poor fuel economy. Therefore, this paper proposes a distributed economic model predictive control approach to explicitly optimize the fuel consumption rate of each vehicle. The proposed method can improve fuel economy and reduce communication requirements and burdens for platooning vehicle swarms. The paper assumes no communication time delay and packet drops, so the impact of unreliable communication will be studied in the future work. Yougang Bian, Changkun Du, Manjiang Hu, Shengbo Eben Li, Haikuo Liu, Chongkang Li |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Exploring Behavioral Patterns of Lane Change Maneuvers for Human-Like Autonomous DrivingabstractDue to the growing interest in automated driving, a deep understanding on the characteristics of human driving behavior is critical for human-like autonomous vehicles. Among various driving behaviors, lane change is the most important one for vehicle lateral driving safety. This study proposes an unsupervised method to extract and discover the behavioral patterns of lane change maneuvers for the purpose of exploring the composed behavioral patterns during lane change. This method involves two phases: Firstly, the lane change sequences will be segmented into blocks using time-series segmentation algorithms. Three segmentation algorithms were utilized in this study. In the second phase, the segments will be clustered to find the corresponding behavioral pattern of each segment. Two extended latent Dirichlet allocation (LDA) models were adopted to cluster the segments. The combination of different segmentation and clustering algorithms were evaluated and compared by employing entropy and perplexity as the evaluation criteria. Collected lane change data from naturalistic driving were applied to examine its effectiveness. The results show that this method could effectively mine descriptive behavioral patterns from lane change data. This study provides a promising data mining solution to facilitating deep and comprehensive understanding on driver lane change behaviors, which will promote the development of human-like autonomous vehicles. Yaoyu Chen, Guofa Li, Shen Li 0001, Wenjun Wang 0005, Shengbo Eben Li, Bo Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Interpretable End-to-End Urban Autonomous Driving With Latent Deep Reinforcement LearningabstractUnlike popular modularized framework, end-to-end autonomous driving seeks to solve the perception, decision and control problems in an integrated way, which can be more adapting to new scenarios and easier to generalize at scale. However, existing end-to-end approaches are often lack of interpretability, and can only deal with simple driving tasks like lane keeping. In this article, we propose an interpretable deep reinforcement learning method for end-to-end autonomous driving, which is able to handle complex urban scenarios. A sequential latent environment model is introduced and learned jointly with the reinforcement learning process. With this latent model, a semantic birdeye mask can be generated, which is enforced to connect with certain intermediate properties in today’s modularized framework for the purpose of explaining the behaviors of learned policy. The latent space also significantly reduces the sample complexity of reinforcement learning. Comparison tests in a realistic driving simulator show that the performance of our method in urban scenarios with crowded surrounding vehicles dominates many baselines including DQN, DDPG, TD3 and SAC. Moreover, through masked outputs, the learned model is able to provide a better explanation of how the car reasons about the driving environment. Jianyu Chen 0002, Shengbo Eben Li, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Fixed-Dimensional and Permutation Invariant State Representation of Autonomous DrivingabstractIn this paper, we propose a new state representation method, called encoding sum and concatenation (ESC), to describe the environment observation for decision-making in autonomous driving. Unlike existing state representation methods, ESC is applicable to the situation where the number of surrounding vehicles is variable and eliminates the need for manually pre-designed sorting rules, leading to higher representation ability and generality. The proposed ESC method introduces a feature neural network (NN) to encode the real-valued feature of each surrounding vehicle into an encoding vector, and then adds these vectors up to obtain the representation vector of the set of surrounding vehicles. Then, a fixed-dimensional and permutation-invariance state representation can be obtained by concatenating the set representation with other variables, such as indicators of the ego vehicle and road. By introducing the sum-of-power mapping, this paper has further proved that the injectivity of the ESC state representation can be guaranteed if the output dimension of the feature NN is greater than the number of variables of all surrounding vehicles. This means that the ESC representation can be used to describe the environment and taken as the inputs of learning-based policy functions. Experiments demonstrate that compared with the fixed-permutation representation method, the policy learning accuracy based on ESC representation is improved by 62.2%. Jingliang Duan, Dongjie Yu, Shengbo Eben Li, Wenxuan Wang 0004, Yangang Ren, Ziyu Lin, Bo Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Multisource Adaption for Driver Attention Prediction in Arbitrary Driving ScenesabstractDriver attention cues contribute to the following intended maneuver prediction and provide a risk indicator for the advanced driver-assistance systems in complex driving scenarios. The diverse traffic scenes result in a challenging task to predict human visual attention with high generalization capability. The data heterogeneity caused by multiple sources with different data characteristics, such as video sources, was investigated in the current research and mitigated using the domain adaption modules (i.e., domain-specific batch normalization, Gaussian priors, and smoothing filter) and the domain-specific focal loss. Inspired by human attention mechanism, generic coders and task-driven attention modules were incorporated into a lightweight network to replicate human-like perceptual patterns, such as the perception of latent risk. Integrating these adaptive modules, we proposed the adaptive driver attention (ADA) model to predict salient regions in different traffic scenes. Consequently, the ADA model trained jointly on four driver attention datasets achieves the best performance against the state-of-the-art methods across seven metrics. Retrospective visualizations of the network and cross-validation results further explain the merit of the proposed method. Since these approaches are generic, adaptive modules are universally applicable for standard deep neural network architectures to alleviate the heterogeneity across datasets and easily extended for arbitrary traffic scenes in the real world. Shun Gan, Xizhe Pei, Yulong Ge, Qingfan Wang, Shi Shang, Shengbo Eben Li, Bingbing Nie |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Structural Transformer Improves Speed-Accuracy Trade-Off in Interactive Trajectory Prediction of Multiple Surrounding VehiclesabstractFast and accurate long-term trajectory prediction of surrounding vehicles (SVs) is critical to autonomous driving systems. In high-density traffic flows, strongly correlated vehicle behaviors require considering the interactions among multiple SVs when predicting their future trajectories. However, existing interactive prediction methods, most based on Long Short-Term Memory (LSTM), are suffering from slow prediction because they analyze SVs one by one and analyze trajectory sequence node by node. This paper presents a fast interactive trajectory prediction method called Structural Transformer which learns both spatial and temporal dependencies among multiple SVs in parallel. Specifically, our model first removes the internal states and loops of LSTM and replaces with a weighted self-reference mapping to realize parallel computation. Then, it embeds the relative spatial information of multiple SVs into trajectory states and reorganizes the self-reference mapping with neighbor-only interaction masks to achieve interactive prediction. Results on the NGSIM dataset show satisfyingly speed and accuracy performance on long-term trajectory prediction of multiple SVs. The longitudinal and lateral errors are reduced to 2.67m and 0.25m over 5s time horizon. The computational time of each step is only 12ms on a 2080ti GPU, which is over 4 times faster than the Structural LSTM. Lian Hou, Shengbo Eben Li, Bo Yang 0044, Zheng Wang 0039, Kimihiko Nakano |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Indirect Shared Control Through Non-Zero Sum Differential Game for Cooperative Automated DrivingabstractCooperative driving of human driver and automated system can effectively reduce the necessity of extremely accurate environment perception of highly automated vehicles, and enhance the robustness of decision-making and motion control. However, due to the two players’ different intentions, severe conflicts may exist during the cooperation, which often result in negative consequences on driving safety and maneuverability. This paper presents an indirect shared control method to model the situation and improve the driving performance, which focus on the affine input nonlinear vehicle dynamic system for shared controller design under the framework of non-zero sum differential game. The Nash equilibria strategy indicates the best response for the automated system, which can guide the automated controller to act more safely and comfortably. Aimed to obtain fast solution for practical application, approximate dynamic programming is utilized to find the Nash equilibria, which is represented by deep neural networks and solved iteratively. Driver-in-the-loop tests on a driving simulator were conducted to verify the performance of the proposed method under highway driving scenarios. The results show that the designed controller is able to reduce the driving workload and ensure the driving safety. Qingkun Li, Shengbo Eben Li, Renjie Li 0004, Yangang Ren, Wenjun Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Sensitivity of Electrodermal Activity Features for Driver Arousal Measurement in Cognitive Load: The Application in Automated Driving SystemsabstractDriver’s under-arousal occurred in automated driving systems (ADS) impairs takeover safety. This study aims to determine electrodermal activity (EDA) features’ importance for driver’s arousal quantification. A car-following simulator study was conducted with participants concurrently executing four levels of cognitive tasks, triggering four levels of arousal. Participants’ skin conductance (SC) data were collected and decomposed into tonic (skin conductance level, SCL) and phasic (skin conductance response, SCR) components. Seventeen features extracted from SC, SCL and SCR were compared. As a result, SCR-relevant features showed higher significance and larger effect size than SC and SCL features in response to cognitive load, which suggests the phasic component dominates changes in EDA under varying cognitive load. Moreover, the SCR rate TTP.nSCRs, identified by$0.03 ~\mu \text{S}$thresholds, attained the largest effect size among all features for driver’s arousal measurement. A varying time windows (TW) analysis showed that TTP.nSCRs was the most suggested arousal metric when TW was over 20 s, whereas the sum of SCRs amplitudes TTP.AmpSum was preferred when TW was less than 20 s. For driver’s arousal quantification with multi-features, the top five suggested features were TTP.nSCRs, SC_Rate5, CDA.SCR (or CDA.ISCR), CDA.AmpSum, and TTP.AmpSum. Although male drivers showed higher values of EDA features than female drivers, the sensitivity of the proposed EDA features stands across gender and individuals. This study promotes an improved understanding of EDA changes in human cognitive process. The sensitive EDA features proposed could be used from uni- or multi-modalities in driver state management and takeover-safety prediction for ADS. Changxu Wu, Bingbing Nie, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Self-Learned Intelligence for Integrated Decision and Control of Automated Vehicles at Signalized IntersectionsabstractIntersection is one of the most accident-prone urban scenarios for autonomous driving wherein making safe and computationally efficient decisions is non-trivial. Current research mainly focuses on the simplified traffic conditions while ignoring the existence of mixed traffic flows, i.e., vehicles, cyclists and pedestrians. For urban roads, different participants lead to a quite dynamic and complex interaction, posing great difficulty to learn an intelligent policy. This paper develops the dynamic permutation state representation in the framework of integrated decision and control (IDC) to handle signalized intersections with mixed traffic flows. Specially, this representation introduces an encoding function and summation operator to construct driving states from environmental observation, capable of dealing with different types and variant number of traffic participants. A constrained optimal control problem is built wherein the objective involves tracking performance and the constraints for different participants, roads and signal lights are designed respectively to assure safety. We solve this problem by gradient-based optimization, wherein the reasonable state will be given by the encoding function and then served as the input of policy and value function. An off-policy training is designed to reuse observations from driving environment and backpropagation through time is utilized to update the policy function and encoding function jointly. Verification result shows that the dynamic permutation state representation can enhance the driving performance of IDC, including comfort, decision compliance and safety with a large margin. The trained driving policy can realize efficient and smooth passing in the complex intersection, guaranteeing driving intelligence and safety simultaneously. Yangang Ren, Jianhua Jiang, Guojian Zhan, Shengbo Eben Li, Chen Chen 0068, Keqiang Li 0002, Jingliang Duan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation ErrorsabstractIn reinforcement learning (RL), function approximation errors are known to easily lead to the Q -value overestimations, thus greatly reducing policy performance. This article presents a distributional soft actor-critic (DSAC) algorithm, which is an off-policy RL method for continuous control setting, to improve the policy performance by mitigating Q -value overestimations. We first discover in theory that learning a distribution function of state-action returns can effectively mitigate Q -value overestimations because it is capable of adaptively adjusting the update step size of the Q -value function. Then, a distributional soft policy iteration (DSPI) framework is developed by embedding the return distribution function into maximum entropy RL. Finally, we present a deep off-policy actor-critic variant of DSPI, called DSAC, which directly learns a continuous return distribution by keeping the variance of the state-action returns within a reasonable range to address exploding and vanishing gradient problems. We evaluate DSAC on the suite of MuJoCo continuous control tasks, achieving the state-of-the-art performance. Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Model-based Constrained Reinforcement Learning using Generalized Control Barrier FunctionabstractModel information can be used to predict future trajectories, so it has huge potential to avoid dangerous regions when applying reinforcement learning (RL) on real-world tasks, like autonomous driving. However, existing studies mostly use model-free constrained RL, which causes inevitable constraint violations. This paper proposes a model-based feasibility enhancement technique of constrained RL, which enhances the feasibility of policy using generalized control barrier function (GCBF) defined on the distance to constraint boundary. By using the model information, the policy can be optimized safely without violating actual safety constraints, and the sample efficiency is increased. The infeasibility in solving the constrained policy gradient is handled by an adaptive coefficient mechanism. We evaluate the proposed method in both simulations and real vehicle experiments in a complex autonomous driving collision avoidance task. The proposed method achieves up to four times fewer constraint violations and converges 3.36 times faster than baseline constrained RL approaches. Haitong Ma, Jianyu Chen 0002, Shengbo Eben Li, Ziyu Lin, Yang Guan, Yangang Ren, Sifa Zheng |
IROS | 3 |
| 2021 | Separated Proportional-Integral Lagrangian for Chance Constrained Reinforcement LearningabstractSafety is essential for reinforcement learning (RL) applied in real-world tasks like autonomous driving. Imposing chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under model uncertainty. Existing chance constrained RL methods like the penalty methods and the Lagrangian methods either exhibit periodic oscillations or learn an over-conservative or unsafe policy. In this paper, we address these shortcomings by elegantly combining these two methods and propose a separated proportional-integral Lagrangian (SPIL) algorithm. We first rewrite penalty methods as optimizing safe probability according to the proportional value of constraint violation, and Lagrangian methods as optimizing according to the integral value of the violation. Then we propose to add up both the integral and proportion values to optimize the policy, with an integral separation technique to limit the integral value within a reasonable range. Besides, the gradient of policy is computed in a model-based paradigm to accelerate training. The proposed method is proved to reduce oscillations and conservatism while ensuring safety by a car-following experiment. Baiyu Peng, Yao Mu 0001, Jingliang Duan, Yang Guan, Shengbo Eben Li, Jianyu Chen 0002 |
IV | 5 |
| 2021 | Model-Based Reinforcement Learning via Imagination with Derived MemoryabstractModel-based reinforcement learning aims to improve the sample efficiency of policy learning by modeling the dynamics of the environment. Recently, the latent dynamics model is further developed to enable fast planning in a compact space. It summarizes the high-dimensional experiences of an agent, which mimics the memory function of humans. Learning policies via imagination with the latent model shows great potential for solving complex tasks. However, only considering memories from the true experiences in the process of imagination could limit its advantages. Inspired by the memory prosthesis proposed by neuroscientists, we present a novel model-based reinforcement learning framework called Imagining with Derived Memory (IDM). It enables the agent to learn policy from enriched diverse imagination with prediction-reliability weight, thus improving sample efficiency and policy robustness. Experiments on various high-dimensional visual control tasks in the DMControl benchmark demonstrate that IDM outperforms previous state-of-the-art methods in terms of policy robustness and further improves the sample efficiency of the model-based method. Yao Mu 0001, Yuzheng Zhuang, Bin Wang 0034, Guangxiang Zhu, Wulong Liu, Jianyu Chen 0002, Ping Luo 0002, Shengbo Eben Li, Chongjie Zhang, Jianye Hao |
NeurIPS | 8 |
| 2021 | Cover: International Journal of Intelligent Systems, Volume 36 Issue 8 August 2021abstractCover Caption: The cover image is based on the Research Article Direct and indirect reinforcement learning by Yang Guan et al., https://doi.org/10.1002/int.22466. Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
Int. J. Intell. Syst. | 2 |
| 2021 | Direct and indirect reinforcement learningabstractReinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision-making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal policy by directly maximizing an objective function using gradient descent methods, in which the objective function is usually the expectation of accumulative future rewards. The latter indirectly finds the optimal policy by solving the Bellman equation, which is the sufficient and necessary condition from Bellman's principle of optimality. We study policy gradient (PG) forms of direct and indirect RL and show that both of them can derive the actor–critic architecture and can be unified into a PG with the approximate value function and the stationary state distribution, revealing the equivalence of direct and indirect RL. We employ a Gridworld task to verify the influence of different forms of PG, suggesting their differences and relationships experimentally. Finally, we classify current mainstream RL algorithms using the direct and indirect taxonomy, together with other ones, including value-based and policy-based, model-based and model-free. Yang Guan, Shengbo Eben Li, Jingliang Duan, Jie Li 0042, Yangang Ren, Qi Sun 0004, Bo Cheng 0003 |
Int. J. Intell. Syst. | 2 |
| 2021 | Indirect Shared Control for Cooperative Driving Between Driver and Automation in Steer-by-Wire VehiclesabstractIt is widely acknowledged that drivers should remain in the control loop before automated vehicles completely meet real-world operational conditions. This paper presents an “indirect shared control” framework for steer-by-wire vehicles, which allows the control authority to be continuously shared between the driver and automation through an weighted-input-summation method. A “best-response” driver steering model based on model predictive control (MPC) for indirect shared control is proposed. Unlike any conventional driver model for manual driving, this model assumes that drivers can learn and incorporate the controller strategy into their internal model for predictive path following. The analytic solution to the driver model is provided to enable off-line simulations. A driving-simulator experiment was conducted to demonstrate the advantages of the indirect shared control system in a highway lane-keeping task. The result showed that the proposed indirect shared control method was effective to improve the subjects’ lane-keeping performance and reduce steering control effort. The proposed driver steering model was also validated by the experiment data, which produced a smaller prediction error than the conventional MPC driver model. Renjie Li 0004, Yanan Li 0001, Shengbo Eben Li, Chaofei Zhang, Etienne Burdet, Bo Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Accelerated Convergence of Time-Splitting Algorithm for MPC using Cross-Node ConsensusabstractThe splitting strategy over prediction horizon of model predictive control (MPC) has the potential to compute optimal action in a parallel way. However, such time-splitting algorithms often lead to very slow convergence speed because the state consensus only happens in each pair of adjacent nodes, i.e. a point-to-point topology. This paper proposes a generic cross-node consensus method to extend the shortcoming of limiting to point-to-point topology for the purpose of accelerating the convergence of time-splitting MPC. The cross-node consensus is realized by predicting the state transition from one node to another using plant prediction model, which can increase the information exchange efficiency in the prediction horizon. The time-splitting optimization algorithm is implemented by combing with alternating directions method of multipliers (ADMM). Simulations with autonomous driving show that this new algorithm significantly reduces the number of iterations in time-splitting MPC, averagely about 81% compared with classic time-splitting technique. Maierdanjiang Maihemuti, Shengbo Eben Li, Jie Li 0042, Jiaxin Gao 0002, Bo Cheng 0003 |
IV | 2 |
| 2020 | Robust Distributed Consensus Control of Uncertain Multiagents Interacted by Eigenvalue-Bounded TopologiesabstractThe uncertainties arising from the plant model and topologies have been a major challenge in multiagent consensus control. This article presents a distributed robust control method for an uncertain multiagent system with eigenvalue-bounded topologies. The heterogeneity of node dynamics is described as the uncertainties of a linear model with a common certain part. The linear transformation method is adopted to decompose topologically coupled controllers. Then, the linear matrix inequalities (LMIs) technique is used to numerically solve the distributed robust controller problem. It is proved that such a controller is robust stable under the condition that the topology is eigenvalue-bounded. The effectiveness of this method is validated by the simulation of a group of unmanned ground vehicles compared with the LQR controller. Keqiang Li 0002, Shengbo Eben Li, Feng Gao 0007, Ziyu Lin, Jie Li 0042, Qi Sun 0004 |
IEEE Internet Things J. | 2 |
| 2020 | Interactive Trajectory Prediction of Surrounding Road Users for Autonomous Driving Using Structural-LSTM NetworkabstractAccurate trajectory prediction of surrounding road users is critical to autonomous driving systems. In mixed traffic flows, road users with different kinds of behaviors and styles bring complexity to the environment, which requires considering interactions among road users when anticipating their future trajectories. This paper presents a long-term interactive trajectory prediction method for surrounding vehicles using a hierarchical multi-sequence learning network. In contrast to non-interactive method which assumes that road users are independent of each other, this method can automatically learn high-level dependencies among multiple interacting vehicles through the proposed structural-LSTM (long short-term memory) network. Specifically, structural-LSTM first assigns one LSTM for each interacting vehicle. Then these LSTMs share their cell states and hidden states with their spatial-neighboring LSTMs by a radial connection, and recurrently analyze the output state of itself as well as the other LSTMs in a deeper layer. Finally based on all output states, the network predicts trajectories for surrounding vehicles. The proposed method is evaluated on the NGSIM dataset, and its results show that satisfyingly accurate prediction performance of long-term trajectories of surrounding vehicles is accessible, e.g., longitudinal and lateral RMS error can be reduced to less than 1.93m and 0.31m over 5s time horizon, respectively. Lian Hou, Long Xin, Shengbo Eben Li, Bo Cheng 0003, Wenjun Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Pedestrian Trajectory Prediction with Learning-based Approaches: A Comparative StudyabstractTo enable safe and efficient navigations through the urban environment, autonomous vehicles need to anticipate the future motions of the walking pedestrians who might collide with them. The dynamic and stochastic behavior characteristics of the pedestrians make the trajectory prediction challengeable for most kinematics-based approaches. This paper presents a comparative study of six state-of-the-art learning-based methods for pedestrian trajectory prediction, including Gaussian Process (GP), LSTM, GP-LSTM, Character-based LSTM, Sequence-to-Sequence (Seq2Seq), and attention-based Seq2Seq. The trajectory prediction is formulated as the regression task or sequence generation problem that predicts future trajectories based on observed trajectories. We evaluate the performance of the learning-based methods on a public real-world pedestrian dataset. To address the concern of data scarcity, we employ three forms of data augmentation (i.e., translation, rotation, and stretch) to enlarge the dataset, which produce the transformed trajectories from the original trajectories. By comparison, those learning-based approaches are ranked based on prediction accuracy from high to low as Seq2Seq, attention-based Seq2Seq, C-LSTM, LSTM, GP, and GP-LSTM. Particularly, Seq2Seq model outperforms all baseline approaches, with the mean and final point errors less than 15cm in normal scenarios when predicting 1s ahead. Yang Li 0093, Long Xin, Dameng Yu, Pengwen Dai, Jianqiang Wang 0003, Shengbo Eben Li |
IV | 6 |
| 2019 | Synthesis of Robust Lane Keeping Systems: Impact of Controller and Design Parameters on System PerformanceabstractThe lane keeping system (LKS), a promising driver assistance system, is essential for autonomous vehicles. In realworld road conditions, it can be quite challenging because LKS must stay within the lane without causing passenger discomfort while both disturbances (e.g., road curvature, wind gusts, and hydroplaning) and model uncertainties in parameters [e.g., vehicle mass, center of gravity (CG), and tire cornering stiffness] are present. In this paper, the performance limits and tradeoffs between three performance criteria (lane tracking, stability robustness, and passenger comfort) are first investigated by exploring the entire design space of three prominent controllers, i.e., proportional-integral-derivative, linear-quadratic- Gaussian, and H-infinity (H∞). Then, a sensitivity study on the vehicle parameters is conducted in order to investigate the impact of the parameters on the three performance metrics. Based on the aforementioned studies, this paper concludes that a robust controller can provide the maximum performance limit with respect to the lane tracking and stability robustness, when properly designed. However, it is observed that the robust controller is still sensitive to a few design and model parameters, such as look-ahead distance and CG. Therefore, the sensitivity study suggests that for vehicles with excessive mass and CG changes, such as SUVs and trucks, the adaptation of controller and look-ahead distance may be necessary to maximize both tracking performance and passenger comfort over a wide range of vehicle speeds. Kibeom Lee, Shengbo Eben Li, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Cooperative Method of Traffic Signal Optimization and Speed Control of Connected Vehicles at Isolated IntersectionsabstractSignalized intersections play an important role in transportation efficiency and vehicle fuel economy in urban areas. This paper proposes a cooperative method of traffic signal control and vehicle speed optimization for connected automated vehicles, which optimizes the traffic signal timing and vehicles' speed trajectories at the same time. The method consists of two levels, i.e., roadside traffic signal optimization and onboard vehicle speed control. The former calculates the optimal traffic signal timing and vehicles' arrival time to minimize the total travel time of all vehicles; the latter optimizes the engine power and brake force to minimize the fuel consumption of individual vehicles. The enumeration method and the pseudospectral method are applied in roadside and onboard optimization, respectively. Simulation studies are conducted to compare the proposed method with benchmark methods. The results show significant improvement of transportation efficiency and fuel economy by the cooperation method. Xuegang Ban, Yougang Bian, Jianqiang Wang 0003, Shengbo Eben Li, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2018 | Continuous Decision Making for On-road Autonomous Driving under Uncertain and Interactive EnvironmentsabstractAlthough autonomous driving techniques have achieved great improvements, challenges still exist in decision making for variety of different scenarios under uncertain and interactive environments. A good decision maker must satisfy the following requirements: (1) Be in a generic and unified form to cover as more scenarios as possible. (2) Be able to interact properly with other moving obstacles under the uncertainty of their motions. In this paper, the continuous decision making (CDM) framework is proposed to formulate different driving scenarios in a unified way, which encodes the high level decision making information into a continuous reference trajectory that can be naturally combined with a lower level trajectory planner. Within the framework, a maximum interaction defensive policy (MIDP) is proposed, which calculates the best action to interact with stochastic moving obstacles while guaranteeing safety. The method is applied to a ramp merging scenario and the stochastic behavior models of the surrounding vehicles are learned from the NGSIM dataset. Simulations are shown to visualize and analyze the results. Jianyu Chen 0002, Chen Tang 0001, Long Xin, Shengbo Eben Li, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 4 |
| 2018 | Platooning of Connected Vehicles With Undirected Topologies: Robustness Analysis and Distributed H-infinity Controller SynthesisabstractThis paper considers the robustness analysis and distributed 'I-1∞(H-infinity) controller synthesis for a platoon of connected vehicles with undirected topologies. We first formulate a unified model to describe the collective behavior of homogeneous platoons with external disturbances using graph theory. By exploiting the spectral decomposition of a symmetric matrix, the collective dynamics of a platoon is equivalently decomposed into a set of subsystems sharing the same size with one single vehicle. Then, we provide an explicit scaling trend of robustness measure γ-gain, and introduce a scalable multistep procedure to synthesize a distributed 'I-1∞controller for large-scale platoons. It is shown that communication topology, especially the leader's information, exerts great influence on both robustness performance and controller synthesis. Furthermore, an intuitive optimization problem is formulated to optimize an undirected topology for a platoon system, and the upper and lower bounds of the objective are explicitly analyzed, which hints us that coordination of multiple mini-platoons is one reasonable architecture to control large-scale platoons. Numerical simulations are conducted to illustrate our findings. Yang Zheng 0001, Shengbo Eben Li, Keqiang Li 0002, Wei Ren 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Robust Longitudinal Control of Multi-Vehicle Systems - A Distributed H-Infinity MethodabstractThe platooning of automated vehicles has the potential to significantly benefit road traffic. This paper presents a distributed$\text{H}_{\mathrm {\infty }}$control method for multi-vehicle systems with identical dynamic controllers and rigid formation geometry. After compensating for the powertrain nonlinearity, the node dynamics in a platoon is mathematically described by a multiplicative uncertainty model. The platoon control system is then decomposed into an uncertain part and a diagonal nominal system through linear transformation and eigenvalue decomposition of the information-exchange-topology matrix. Robust stability, string stability, and distance tracking performance of the designed platoons are analyzed theoretically under the decoupled$\text{H}_{\mathrm {\infty }}$framework. A comparative simulation with non-robust controllers is used to demonstrate the effectiveness of this method. Shengbo Eben Li, Feng Gao 0007, Keqiang Li 0002, Le Yi Wang, Keyou You, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Driver-automation indirect shared control of highly automated vehicles with intention-aware authority transitionabstractShared control is an important approach to avoid the driver-out-of-the-loop problems brought by imperfect autonomous driving. Steer-by-wire technology allows the mechanical decoupling between the steering wheel and the road wheels. On steer-by-wire vehicles, the automation can join the control loop by correcting the driver steering input, which forms a new paradigm of shared control. The new framework, under which the driver indirectly controls the vehicle through the automation's input transformation, is called indirect shared control. This paper presents an indirect shared control system, which realizes the dynamic control authority allocation with respect to the driver's authority intention. The simulation results demonstrate the effectiveness and benefits of the proposed control authority adaptation method. Renjie Li 0004, Yanan Li 0001, Shengbo Eben Li, Etienne Burdet, Bo Cheng 0003 |
Intelligent Vehicles Symposium | 3 |
| 2017 | Instantaneous Feedback Control for a Fuel-Prioritized Vehicle Cruising System on Highways With a Varying SlopeabstractThis paper presents two fuel-prioritized feedback controllers, which are called the estimated minimum principle (EMP) and kinetic energy conversion (KEC), to realize eco-cruising on varying slopes for vehicles with conventional powertrains. The former is derived from the minimum principle with an estimated Hamiltonian, and the latter is designed based on the equivalent conversion between the kinetic-energy change of vehicle body and the fuel consumption of the engine. They are implemented with analytical control laws and rely on current road slope information only without look-ahead prediction. This feature results in a very light computing load, with the average computing time of each step less than one millisecond. Their fuel-saving performances are quantitatively studied and compared with a model predictive control and a constant speed control. As an expansion, the control rule for avoiding rear-end collision is also designed by using a safety-guaranteed car-following model to constrain the high-risk behaviors. Shaobing Xu, Shengbo Eben Li, Bo Cheng 0003, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Detection of driver cognitive distraction: An SVM based real-time algorithm and its comparison study in typical driving scenariosabstractDetection of driver cognitive distraction is critical for active safety systems of road vehicles. Compared with visual distraction, cognitive distraction is more challenging for detection due to the lack of apparent exterior features. This paper presents a novel real-time detection algorithm for driver cognitive distraction by using support vector machine (SVM). Data are collected from 26 subjects, driving in typical urban and highway scenarios in a simulator. The chosen urban scenario is the stop-controlled intersection and the highway scenario is the speed-limited highway. Driver cognitive distraction while driving is induced by clock tasks which compete with the main driving tasks for visuospatial short working memory. For each subject, distracted driving instances and the equal number of non-distracted driving instances were collected (24 for urban scenario and 20 for highway scenario in total). Features concerning both driving performance and eye movement are used for training and validation. The proposed algorithm have correct rate of 93.0% and 98.5% for highway and urban scenarios respectively. Results also show that driver distraction can be recognized 6.5 s to 9.0 s after its happening, indicating good performance of the detection algorithm. Yuan Liao 0002, Shengbo Eben Li, Guofa Li, Wenjun Wang 0005, Bo Cheng 0003, Fang Chen 0006 |
Intelligent Vehicles Symposium | 2 |
| 2016 | Dynamical tracking of surrounding objects for road vehicles using linearly-arrayed ultrasonic sensorsabstractAccurate detection and tracking of traffic participants are crucial to advanced driver assistance systems. This paper presents a centralized object tracking approach for surrounding objects in road traffic environments by using multiple linearly arrayed ultrasonic sensors. An ultrasonic sensor model is specifically developed for traffic environment, which consists of detection scope, chance of detection and ranging error, incorporating factors of object shapes, materials, distances and orientations. A centralized filter is designed to selectively fuse new measurements that are obtained using the Extended Kalman Filter (EKF) from certain sensors to conduct object tracking at each step. The effectiveness of proposed method is validated by simulations, which is found to have superior tracking performance compared to traditional triangle localization method, with more stable and smaller tracking error, especially when the object is entering or leaving the detection area. Jiaying Yu, Shengbo Eben Li, Chang Liu 0002, Bo Cheng 0003 |
Intelligent Vehicles Symposium | 2 |
| 2016 | Efficient and accurate computation of model predictive control using pseudospectral discretization
Shengbo Eben Li, Shaobing Xu, Dongsuk Kum |
Neurocomputing | 1 |
| 2016 | Detection of Driver Cognitive Distraction: A Comparison Study of Stop-Controlled Intersection and Speed-Limited HighwayabstractDriver distraction has been identified as one major cause of unsafe driving. The existing studies on cognitive distraction detection mainly focused on high-speed driving situations, but less on low-speed traffic in urban driving. This paper presents a method for the detection of driver cognitive distraction at stop-controlled intersections and compares its feature subsets and classification accuracy with that on a speed-limited highway. In the simulator study, 27 subjects were recruited to participate. Driver cognitive distraction is induced by the clock task that taxes visuospatial working memory. The support vector machine (SVM) recursive feature elimination algorithm is used to extract an optimal feature subset out of features constructed from driving performance and eye movement. After feature extraction, the SVM classifier is trained and cross-validated within subjects. On average, the classifier based on the fusion of driving performance and eye movement yields the best correct rate and F-measure (correctrate = 95.8 ± 4.4%; for stop-controlled intersections and correct rate = 93.7 ± 5.0%; for a speed-limited highway) among four types of the SVM model based on different candidate features. The comparisons of extracted optimal feature subsets and the SVM performance between two typical driving scenarios are presented. Yuan Liao 0002, Shengbo Eben Li, Wenjun Wang 0005, Guofa Li, Bo Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | A Forward Collision Warning Algorithm With Adaptation to Driver BehaviorsabstractSignificant effort has been made on designing user-acceptable driver assistance systems. To adapt to driver characteristics, this paper proposes a forward collision warning (FCW) algorithm that can adjust its warning thresholds in a real-time manner according to driver behavior changes, including both behavioral fluctuation and individual difference. This adaptive FCW algorithm overcomes the limit of traditional FCW with fixed risk evaluation models and fixed triggering thresholds by continuously monitoring driver braking behaviors in multiple lanes. A real-time identification algorithm for the warning thresholds is designed by using the recursive least squares method. Based on naturalistic experimental data, offline simulations show that this algorithm can match driver behavioral fluctuation and individual difference in long-time driving condition, and as time goes on, the adaptability to driver behavior is gradually improved, thus decreasing the false-alarm rate of FCW. Jianqiang Wang 0003, Chenfei Yu, Shengbo Eben Li |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Stability and Scalability of Homogeneous Vehicular Platoon: Study on the Influence of Information Flow TopologiesabstractIn addition to decentralized controllers, the information flow among vehicles can significantly affect the dynamics of a platoon. This paper studies the influence of information flow topology on the internal stability and scalability of homogeneous vehicular platoons moving in a rigid formation. A linearized vehicle longitudinal dynamic model is derived using the exact feedback linearization technique, which accommodates the inertial delay of powertrain dynamics. Directed graphs are adopted to describe different types of allowable information flow interconnecting vehicles, including both radar-based sensors and vehicle-to-vehicle (V2V) communications. Under linear feedback controllers, a unified internal stability theorem is proved by using the algebraic graph theory and Routh-Hurwitz stability criterion. The theorem explicitly establishes the stabilizing thresholds of linear controller gains for platoons, under a large class of different information flow topologies. Using matrix eigenvalue analysis, the scalability is investigated for platoons under two typical information flow topologies, i.e., 1) the stability margin of platoon decays to zero as 0(1/N2) for bidirectional topology; and 2) the stability margin is always bounded and independent of the platoon size for bidirectional-leader topology. Numerical simulations are used to illustrate the results. Yang Zheng 0001, Shengbo Eben Li, Jianqiang Wang 0003, Dongpu Cao, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Lane change maneuver recognition via vehicle state and driver operation signals - Results from naturalistic driving dataabstractLane change maneuver recognition is critical in driver characteristics analysis and driver behavior modeling for active safety systems. This paper presents an enhanced classification method to recognize lane change maneuver by using optimized features exclusively extracted from vehicle state and driver operation signals. The sequential forward floating selection (SFFS) algorithm was adopted to select the optimized feature set to maximize the k-nearest-neighbor classifier performance. The hidden Markov models (HMMs), based on the optimized feature set, were developed to classify driver lane change and lane keeping maneuvers. Fifteen drivers participated in the road test for validation with an accumulation of 2,200 km naturalistic driving data, from which 372 lane changes were extracted. Results show that the recognition rate of lane change maneuver achieves 88.2%. The numbers are 87.6% and 88.8% for left and right lane change maneuvers, respectively, superior to the results from conventional classifiers. Guofa Li, Shengbo Eben Li, Yuan Liao 0002, Wenjun Wang 0005, Bo Cheng 0003, Fang Chen 0006 |
Intelligent Vehicles Symposium | 2 |
| 2015 | An overview of vehicular platoon control under the four-component frameworkabstractThe platooning of autonomous ground vehicles has potential to largely benefit the road traffic, including enhancing highway safety, improving traffic utility and reducing fuel consumption. The main goal of platoon control is to ensure all the vehicles in the same group to move at consensual speed while maintaining desired spaces between adjacent vehicles. This paper presents an overview of vehicular platoon control techniques from networked control perspective, which naturally decomposes a platoon into four interrelated components, i.e., 1) node dynamics (ND), 2) information flow topology (IFT), 3) distributed controller (DC) and, 4) geometry formation (GF). Under the four-component framework, existing literature are categorized and analyzed according to their technical features. Three main performance metrics, i.e. string stability, stability margin and coherence behavior, are also discussed. Shengbo Eben Li, Yang Zheng 0001, Keqiang Li 0002, Jianqiang Wang 0003 |
Intelligent Vehicles Symposium | 1 |
| 2015 | The impact of driver cognitive distraction on vehicle performance at stop-controlled intersectionsabstractDriver distraction has been identified as an important driving safety issue. However, existed studies focused less on low-speed condition, especially at intersections. This paper aims to find the impact of driver cognitive distraction on vehicle performance at stop-controlled intersections. Eight subjects (young adult: 4, older adult: 4) participated in this study and each of them drove through 40 stop-controlled intersections. The intersections were presented randomly at two levels of FOV (field of view). Driver cognitive distraction was induced by a one-back task and a clock task. Results showed that the cognitive tasks led to more abrupt steering in both age groups while significant influence on lane-keeping capability was only observed in the young group. Steering smoothness was mainly influenced by the cognitive tasks at brake on-restart phase in the young group while at after-restart phase in the older group. Impaired longitudinal control (stop for watching) was observed in the older adult group. These findings can be applied to automatically recognize driver distraction at stop-controlled intersections in future. Yuan Liao 0002, Shengbo Eben Li, Wenjun Wang 0005, Guofa Li, Bo Cheng 0003 |
Intelligent Vehicles Symposium | 2 |
| 2015 | Synthesis of multiple model switching controllers using H∞ theory for systems with large uncertainties
Feng Gao 0007, Shengbo Eben Li, Dongsuk Kum, Hui Zhang 0019 |
Neurocomputing | 2 |
| 2015 | Coordinated Adaptive Cruise Control System With Lane-Change AssistanceabstractTo address the problem caused by a conventional adaptive cruise control (ACC) system, which hinders drivers from changing lanes, in this study we propose a novel coordinated ACC system with a lane-change assistance function, which enables dual-target tracking, safe lane change, and longitudinal ride comfort. We first analyze lane-change risk by calculating minimum safety spacing between the host vehicle and surrounding vehicles and then develop a coordinated control algorithm using model predictive control theory. Tracking performance is designed on the basis of tracking errors of the host car and two leading vehicles, safety performance is realized by considering the safe distance between the host car and surrounding vehicles, and ride comfort performance is realized by limiting the vehicle's longitudinal acceleration. Driver-in-the-loop tests performed on a driving simulator confirm that the proposed ACC system can overcome the disadvantages of conventional ACC and achieves multiobjective coordination in the lane-change process. Ruina Dang, Jianqiang Wang 0003, Shengbo Eben Li, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Fast Online Computation of a Model Predictive Controller and Its Application to Fuel Economy-Oriented Adaptive Cruise ControlabstractThe recent progress of advanced vehicle control systems presents a great opportunity for the application of model predictive control (MPC) in the automotive industry. However, high computational complexity inherently associated with the receding horizon optimization must be addressed to achieve real-time implementation. This paper presents a generic scale reduction framework to reduce the online computational burden of MPC controllers. A lower dimensional MPC algorithm is formulated by combining an existing “move blocking ” strategy with a “constraint-set compression” strategy, which is proposed to further reduce the problem scale by partially relaxing inequality constraints in the prediction horizon. The closed-loop stability is guaranteed by adding terminal zero-state constraint. The tradeoff between control optimality and computational intensity is achieved by proper design of the blocking and compression matrices. The fast algorithm has been applied on intelligent vehicular longitudinal automation, implemented as a fuel economy-oriented adaptive cruise controller and experimentally evaluated by a series of real-time simulations and field tests. These results indicate that the proposed method significantly improves the computational speed while maintaining satisfactory control optimality without sacrificing the desired performance. Shengbo Eben Li, Zhenzhong Jia, Keqiang Li 0002, Bo Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Fuel-Optimal Cruising Strategy for Road Vehicles With Step-Gear Mechanical TransmissionabstractThis paper studies the principles and mechanism of a fuel-optimal strategy in cruising scenarios, i.e., the pulse and glide (PnG) operation, for road vehicles equipped with a step-gear transmission. In the PnG strategy, the control of the engine and the transmission determines the fuel-saving performance, and it is obtained by solving an optimal control problem (OCP). Due to a discrete gear ratio, strong nonlinear engine fuel characteristics, and different dynamics in the pulse/glide mode, the OCP is a switching nonlinear mixed-integer problem. This challenging problem is converted by a knotting technique and the Legendre pseudospectral method to a nonlinear programming problem, which then solves the optimal engine torque and transmission gear position. The optimization results show the significant fuel saving of the PnG operation as compared with the constant-speed cruising strategy. The underlying fuel-saving mechanism of the PnG strategy is explained graphically. For a real-time implementation, a near-optimal practical rule that enables a driver and/or an automatic control system to fast select gear positions and engine torque profile is proposed with only slightly deteriorated fuel saving. Shaobing Xu, Shengbo Eben Li, Bo Cheng 0003, Huei Peng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Periodicity based cruising control of passenger cars for optimized fuel consumptionabstractEco-driving technologies are able to largely reduce the fuel consumption of ground vehicles. This paper presents how to determine the fuel-optimized operating strategies of passenger cars under cruising process. The design naturally casts into an optimal control problem with the S-shaped engine fueling rate as the integrand of cost function. The solutions are numerically solved by the Legendre pseudospectral method, of which many are found to demonstrate periodic behaviors. In the periodic operation, the engine switches between the minimum brake specific fuel consumption (BSFC) point and the idling point, while the vehicle speed oscillates between its upper and lower bounds. The formation of periodic operation are analyzed and explained by the π-test theory and steady state analysis method. Shengbo Eben Li, Shaobing Xu, Guofa Li, Bo Cheng 0003 |
Intelligent Vehicles Symposium | 1 |
| 2014 | Legendre pseudospectral computation of optimal speed profiles for vehicle eco-driving systemabstractThis paper presents a computational framework to solve optimal control problems (OCPs) using Legendre Pseudospectral (PS) method and its application to obtain eco-driving strategies for ground vehicles. Both control and state variables of OCPs are approximated by Lagrange interpolating polynomials at the Legendre-Gauss-Lobatto (LGL) collocation points. The OCP is converted into a nonlinear programming (NLP) problem, and numerically solved by matured optimization algorithms. To implement the PS method, we developed a computational package, called Pseudospectral Optimal control Problem Solver (POPS) in Matlab environment. Further, the POPS is applied to obtain fuel-optimized driving strategies for automated vehicles in hilly road conditions. Shaobing Xu, Shengbo Eben Li, Bo Cheng 0003 |
Intelligent Vehicles Symposium | 3 |