EDBT 2026 Demo / reviewers in the wild / expert
Jingda Wu
dblp:205/4902
· DBLP profile ↗
24ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-7336-4492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CyGen-SAC: A data-driven framework for representativeness metric and policy optimization in generation of multidimensional driving cycles
Julin Hu, Hongwen He, Jingda Wu |
Adv. Eng. Informatics | 3 |
| 2026 | Energy-Efficient integrated thermal management for electric vehicles using evolutionary deep reinforcement learning
Jiankun Peng, Jingda Wu, Chunye Ma |
Expert Syst. Appl. | 3 |
| 2026 | Global context alignment and separable fusion for generalizable multi-modal 3D object detection
Yingjuan Tang, Hongwen He, Jingda Wu, Yongpeng Shen, Yong Wang 0044, Yifan Wu 0026 |
Expert Syst. Appl. | 3 |
| 2026 | Q-Advantage Integrated Human-Guided Reinforcement Learning for Safe End-to-End Autonomous DrivingabstractReinforcement learning (RL) is a promising approach for end-to-end autonomous driving, but its practical deployment remains challenging due to low sample efficiency and sensitivity to reward design. To address these challenges, this study presents a novel Q-advantage integrated human-guided reinforcement learning (QIHG-RL) framework that effectively combines the strengths of machine learning and human expertise. The QIHG-RL framework features: 1) an ensemble Q-advantage function that aggregates multiple value networks to enhance value estimation, and 2) an integration mechanism that embeds the Q-advantage into both the actor-critic network and the prioritized experience replay. This design allows the agent to leverage sparse and sub-optimal human demonstrations, accelerating policy learning in the early training phase while gradually enhancing exploration as training progresses. The framework is evaluated across three safety-critical driving tasks. Experimental results show a 167% improvement in sample efficiency compared to standard RL methods and a 14% performance gain over a state-of-the-art human-guided RL baseline. Furthermore, a Sim2Real pipeline combining domain randomization and semantic denoised remapping facilitates successful deployment on a real-world autonomous vehicle. Yong Wang 0044, Hongwen He, Jingda Wu, Yingjuan Tang, Zirui Kuang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Trust-Calibrated Human-in-the-Loop Reinforcement Learning for Safe and Efficient Autonomous NavigationabstractAutonomous navigation technology for autonomous ground vehicles (AGVs) is currently a highly active research area. With the advancement of internet of things (IoT) technologies, AGVs increasingly leverage interconnected systems, such as onboard sensors, vehicle-to-everything (V2X) communication, and cloud data sharing, to enhance navigation capabilities. Human-in-the-loop (HIL) guidance has been shown to be effective in improving the performance of reinforcement learning (RL) algorithms. Existing methods in this field often assume that human guidance is always beneficial. However, incorrect guidance can cause oscillations or even divergences in RL training. In this study, we propose an innovative trust-calibrated HIL-RL approach to address these gaps. First, human guidance is introduced into the RL framework to enhance learning performance through intervention and demonstration. This process includes adding behavior cloning (BC) objectives to the RL policy and an adaptive experience replay mechanism. Second, a trust evaluation mechanism is incorporated within the HIL-RL framework to calculate a belief value, which not only ensures that human guidance is trustworthy but also dynamically optimizes the BC weight. This improvement enhances the training efficiency of RL agents and supports steady performance improvement, even when exposed to potentially detrimental external intervention. The results in simulation show that the proposed method achieves an improvement in success rate of 22–26% over vanilla RL. In real-world experiments, the proposed method achieved a 100% success rate and demonstrated outstanding navigation efficiency, validating the effectiveness of the trust evaluation mechanism. Guangzhong Zhou, Jingda Wu, Chao Huang 0006 |
IEEE Internet Things J. | 3 |
| 2025 | Toward Multi-Task Generalization in Autonomous Navigation: A Human-in-the-Loop Adversarial Reinforcement Learning With Diffusion PolicyabstractDue to the complexity and variability of real-world environments, data-driven autonomous navigation strategies for autonomous ground vehicles have significant potential to improve performance and adaptability in diverse scenarios. Reinforcement learning (RL) has emerged as a promising approach for autonomous navigation. However, existing RL methods often struggle with low sample efficiency, limited adaptability, and poor generalization in dynamic multi-task scenarios. To address these issues, we propose a novel framework: human-in-the-loop adversarial RL with diffusion policy, designed for scalable and robust policy learning. This framework leverages a diffusion model as policy network, effectively exploring and learning high-dimensional, multi-modal behavior distributions. It also integrates human feedback to improve data efficiency and stabilize policy training. On top of this, adversarial training is employed to improve robustness and adaptability to change in tasks and distributions. The proposed method is trained in simulation, and then the well-trained policy is transferred to the real-world. Experimental results demonstrate that this approach significantly outperforms existing methods in terms of efficiency, stability, generalization, and multi-task adaptability, offering a promising solution for the next generation of autonomous navigation systems. The supplementary video is available at https://youtu.be/JH3knw0I5lU Chao Huang 0006, Jingda Wu, Xin Yuan 0008 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Enhancing Fuel Cell Electric Vehicle Efficiency With an Information-Bridged Hierarchical Reinforcement Learning MethodabstractThe intelligent transportation system furnishes electrified vehicles with multi-source traffic information, thereby enhancing the potential for greater energy efficiency. Eco-driving and internal energy management represent dual pathways to achieving these efficiencies. Departing from existing studies that typically investigate these pathways independently, this paper introduces a collaborative hierarchical reinforcement learning (RL) method that synchronizes the optimization of both eco-driving and energy management in an integrated solution. To advance RL performance, we propose a novel information bridge scheme at the methodological level. This scheme optimizes the actor-critic RL algorithm’s learning mechanism, facilitating improved information exchange between the dual agents: eco-driving (upper layer) and energy management (lower layer). The critic value of the lower-level agent as a conduit for transmitting condensed state information into the upper, improving the holistic performance. Furthermore, convolutional networks are employed to enhance traffic information extraction. The effectiveness of our method is demonstrated through SUMO simulations, showing a 30.28% improvement in energy efficiency with only a 7.53% compromise in timeliness compared to prevailing Krauss eco-driving. Results also indicate improved training stability and adaptability of our method. This research not only contributes to the optimization of eco-driving for electrified vehicles but also has values in other multi-agent collaborative optimization fields. Xiangqi Wan, Jingda Wu, Mei Yan, Hongwen He |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Human-Guided Continual Learning for Personalized Decision-Making of Autonomous DrivingabstractLearning-based techniques hold considerable promise in achieving human-like autonomous driving. However, one deployed policy encounters difficulties in satisfying the drivers’ diverse decision-making preferences simultaneously. Meanwhile, training personalized policies for each driver from scratch is time-consuming and resource-intensive. To address these challenges, this paper proposes a human-guided continual learning framework, wherein the human drivers could real-time take over a deployed policy when it performs unsatisfactorily, and the autonomous vehicle (AV) agent would automatically acquire human demonstrations and dynamically alter itself in accordance with personalized decision-making preference. Furthermore, a priority experience memory-enabled elastic weight consolidation (PEM-EWC) mechanism is developed to prevent the AV agent from overfitting to a limited number of human demonstrations and catastrophically forgetting its acquired fundamental driving abilities. Driver-in-the-loop simulations and real-world experiments are conducted in representative autonomous driving decision-making scenarios, and experimental results demonstrate the superior equilibrium of our proposed approach in terms of driving safety, human likeness, and training efficiency, compared to other baselines, which suggests that it provides a promising solution for personalized decision-making in autonomous driving. The supplementary video is available athttps://youtu.be/HKF0ayxMycc. Haohan Yang, Yanxin Zhou, Jingda Wu, Lie Yang, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Active Scene Recognition for Domestic Robots: Observing, Moving, and Recognizing
Chao Huang 0006, Hailong Huang 0001, Jingda Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Learning Based Model Predictive Path Tracking Control for Autonomous BusesabstractIn addressing the trade-off between prediction model accuracy and computational cost in the context of path tracking control, this paper proposes a learning-based model predictive control (LB-MPC) strategy for autonomous buses. A three-degree-of-freedom (DOF) single-track vehicle dynamic model is established, and an in-depth analysis is conducted on its step response error with respect to variations in vehicle speed, pedal position, and front wheel steering angle compared to the IPG TruckMaker model. Methods for constructing error datasets and receding horizon updates are designed, and a Gaussian process regression (GPR) is employed to establish an error fitting model for real-time error compensation and correction of the nominal single-track model. The error correction model is utilized as the prediction model, and a path tracking cost function is designed to formulate a quadratic programming (QP) optimization problem, proposing an LB-MPC path tracking control architecture. Through joint simulations using the IPG TruckMaker & Simulink platform and real bus experiment, the real-time performance and effectiveness of the proposed GPR error correction model and LB-MPC path tracking control strategy are verified. Results demonstrate that compared to traditional MPC path tracking control strategy, the proposed LB-MPC strategy reduces the average path tracking error by 79.00%. Mo Han, Hongwen He, Jianfei Cao, Jingda Wu, Wei Liu 0058, Man Shi |
IV | 4 |
| 2024 | Towards efficient multi-modal 3D object detection: Homogeneous sparse fuse network
Yingjuan Tang, Hongwen He, Yong Wang 0044, Jingda Wu |
Expert Syst. Appl. | 4 |
| 2024 | Fear-Neuro-Inspired Reinforcement Learning for Safe Autonomous DrivingabstractEnsuring safety and achieving human-level driving performance remain challenges for autonomous vehicles, especially in safety-critical situations. As a key component of artificial intelligence, reinforcement learning is promising and has shown great potential in many complex tasks; however, its lack of safety guarantees limits its real-world applicability. Hence, further advancing reinforcement learning, especially from the safety perspective, is of great importance for autonomous driving. As revealed by cognitive neuroscientists, the amygdala of the brain can elicit defensive responses against threats or hazards, which is crucial for survival in and adaptation to risky environments. Drawing inspiration from this scientific discovery, we present a fear-neuro-inspired reinforcement learning framework to realize safe autonomous driving through modeling the amygdala functionality. This new technique facilitates an agent to learn defensive behaviors and achieve safe decision making with fewer safety violations. Through experimental tests, we show that the proposed approach enables the autonomous driving agent to attain state-of-the-art performance compared to the baseline agents and perform comparably to 30 certified human drivers, across various safety-critical scenarios. The results demonstrate the feasibility and effectiveness of our framework while also shedding light on the crucial role of simulating the amygdala function in the application of reinforcement learning to safety-critical autonomous driving domains. Xiangkun He, Jingda Wu, Zhiyu Huang, Zhongxu Hu, Jun Wang 0012, Alberto L. Sangiovanni-Vincentelli, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Ensembled Traffic-Aware Transformer-Based Predictive Energy Management for Electrified VehiclesabstractThe predictive energy management strategy (PEMS) offers potential advantages in enhancing the driving economy of electrified vehicles using vehicle speed prediction. However, realizing accurate predictions in practical contexts remains a challenge. Departing from conventional PEMS that rely on historical speed or static traffic data, we introduce a real-time traffic-aware PEMS for improved performance. To better understand the interplay between the host vehicle and its surrounding traffic, we use a Transformer network as the predictor that employs the speeds and relative distances of the surrounding six vehicles to forecast future speed sequences for the host vehicle. To augment this data-driven approach, we develop a dual-predictor strategy based on the deep ensemble technique. This strategy measures the Transformer’s output uncertainty to gauge prediction reliability and introduce an automated threshold mechanism. Based on this threshold and real-time uncertainties, the strategy chooses between the Transformer and an exponential predictor to achieve improved prediction outcomes. A reinforcement learning method is integrated as the PEMS optimizer. For validation, we generate training data with traffic information based on the next generation simulation (NGSIM) dataset and create a test scenario in the SUMO simulator. The results confirm that speed predictions based on real-time traffic data surpass traditional PEMS, either directly inputting traffic data or excluding it. The Transformer predictor significantly outperforms the state-of-the-art predictor. Importantly, our dual-predictor design amplifies prediction accuracy by 27.2% against the standard single-network predictor under non-training conditions. Overall, our PEMS enhances driving economy by 11.1% relative to traffic-unaware models and 8.0% over non-Transformer schemes. Jingda Wu, Zhongbao Wei, Hongwen He, Henglai Wei, Shuangqi Li, Fei Gao 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Differentiable Integrated Motion Prediction and Planning With Learnable Cost Function for Autonomous DrivingabstractPredicting the future states of surrounding traffic participants and planning a safe, smooth, and socially compliant trajectory accordingly are crucial for autonomous vehicles (AVs). There are two major issues with the current autonomous driving system: the prediction module is often separated from the planning module, and the cost function for planning is hard to specify and tune. To tackle these issues, we propose a differentiable integrated prediction and planning (DIPP) framework that can also learn the cost function from data. Specifically, our framework uses a differentiable nonlinear optimizer as the motion planner, which takes as input the predicted trajectories of surrounding agents given by the neural network and optimizes the trajectory for the AV, enabling all operations to be differentiable, including the cost function weights. The proposed framework is trained on a large-scale real-world driving dataset to imitate human driving trajectories in the entire driving scene and validated in both open-loop and closed-loop manners. The open-loop testing results reveal that the proposed method outperforms the baseline methods across a variety of metrics and delivers planning-centric prediction results, allowing the planning module to output trajectories close to those of human drivers. In closed-loop testing, the proposed method outperforms various baseline methods, showing the ability to handle complex urban driving scenarios and robustness against the distributional shift. Importantly, we find that joint training of planning and prediction modules achieves better performance than planning with a separate trained prediction module in both open-loop and closed-loop tests. Moreover, the ablation study indicates that the learnable components in the framework are essential to ensure planning stability and performance. Code and Supplementary Videos are available at https://mczhi.github.io/DIPP/. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Sampling Efficient Deep Reinforcement Learning Through Preference-Guided Stochastic ExplorationabstractStochastic exploration is the key to the success of the deep -network (DQN) algorithm. However, most existing stochastic exploration approaches either explore actions heuristically regardless of their values or couple the sampling with values, which inevitably introduce bias into the learning process. In this article, we propose a novel preference-guided -greedy exploration algorithm that can efficiently facilitate exploration for DQN without introducing additional bias. Specifically, we design a dual architecture consisting of two branches, one of which is a copy of DQN, namely, the branch. The other branch, which we call the preference branch, learns the action preference that the DQN implicitly follows. We theoretically prove that the policy improvement theorem holds for the preference-guided -greedy policy and experimentally show that the inferred action preference distribution aligns with the landscape of corresponding values. Intuitively, the preference-guided -greedy exploration motivates the DQN agent to take diverse actions, so that actions with larger values can be sampled more frequently, and those with smaller values still have a chance to be explored, thus encouraging the exploration. We comprehensively evaluate the proposed method by benchmarking it with well-known DQN variants in nine different environments. Extensive results confirm the superiority of our proposed method in terms of performance and convergence speed. Wenhui Huang 0001, Jingda Wu, Xiangkun He, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Prioritized Experience-Based Reinforcement Learning With Human Guidance for Autonomous DrivingabstractReinforcement learning (RL) requires skillful definition and remarkable computational efforts to solve optimization and control problems, which could impair its prospect. Introducing human guidance into RL is a promising way to improve learning performance. In this article, a comprehensive human guidance-based RL framework is established. A novel prioritized experience replay mechanism that adapts to human guidance in the RL process is proposed to boost the efficiency and performance of the RL algorithm. To relieve the heavy workload on human participants, a behavior model is established based on an incremental online learning method to mimic human actions. We design two challenging autonomous driving tasks for evaluating the proposed algorithm. Experiments are conducted to access the training and testing performance and learning mechanism of the proposed algorithm. Comparative results against the state-of-the-art methods suggest the advantages of our algorithm in terms of learning efficiency, performance, and robustness. Jingda Wu, Zhiyu Huang, Wenhui Huang 0001, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Human-Guided Deep Reinforcement Learning for Optimal Decision Making of Autonomous VehiclesabstractAlthough deep reinforcement learning (DRL) methods are promising for making behavioral decisions in autonomous vehicles (AVs), their low training efficiency and difficulty to adapt to untrained cases hinder their applications. Introducing a human role in the DRL paradigm could improve training efficiency by using human prior knowledge and overcome untrained cases in deployment by online human takeover. In this study, a novel value-based DRL algorithm that leverages human guidance to improve its performance is proposed for addressing high-level decision-making problems in autonomous driving. We develop a new learning objective for DRL to increase the value of the human policy over the undertrained DRL policy so that the DRL agent can be encouraged to mimic human behaviors and thereby utilizing human guidance more efficiently. Our method can autonomously evaluate the importance of different human guidance, which makes it more robust for variation of human performance. The proposed DRL algorithm was used to address a challenging multiobjective lane-change decision-making problem. We collected human guidance from a human-in-the-loop driving experiment and evaluated our method in a high-fidelity simulator. Results validated the advantages of the proposed algorithm in terms of training efficiency and optimality in the decision-making problem compared to the baselines of state-of-the-art existing methods. Results also revealed the favorable fine-tuning ability of the proposed algorithm, which is promising for addressing the long-tail issue in DRL-based autonomous driving. Our methodology does not introduce additional domain knowledge so that it can be seamlessly applied to other similar issues. The supplementary video is available at https://youtu.be/Ec7WkqeLsB8. Jingda Wu, Haohan Yang, Lie Yang, Yi Huang 0038, Xiangkun He, Chen Lv 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Human-Guided Reinforcement Learning With Sim-to-Real Transfer for Autonomous NavigationabstractReinforcement learning (RL) is a promising approach in unmanned ground vehicles (UGVs) applications, but limited computing resource makes it challenging to deploy a well-behaved RL strategy with sophisticated neural networks. Meanwhile, the training of RL on navigation tasks is difficult, which requires a carefully-designed reward function and a large number of interactions, yet RL navigation can still fail due to many corner cases. This shows the limited intelligence of current RL methods, thereby prompting us to rethink combining RL with human intelligence. In this paper, a human-guided RL framework is proposed to improve RL performance both during learning in the simulator and deployment in the real world. The framework allows humans to intervene in RL's control progress and provide demonstrations as needed, thereby improving RL's capabilities. An innovative human-guided RL algorithm is proposed that utilizes a series of mechanisms to improve the effectiveness of human guidance, including human-guided learning objective, prioritized human experience replay, and human intervention-based reward shaping. Our RL method is trained in simulation and then transferred to the real world, and we develop a denoised representation for domain adaptation to mitigate the simulation-to-real gap. Our method is validated through simulations and real-world experiments to navigate UGVs in diverse and dynamic environments based only on tiny neural networks and image inputs. Our method performs better in goal-reaching and safety than existing learning- and model-based navigation approaches and is robust to changes in input features and ego kinetics. Furthermore, our method allows small-scale human demonstrations to be used to improve the trained RL agent and learn expected behaviors online. Jingda Wu, Yanxin Zhou, Haohan Yang, Zhiyu Huang, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Conditional Predictive Behavior Planning With Inverse Reinforcement Learning for Human-Like Autonomous DrivingabstractMaking safe and human-like decisions is an essential capability of autonomous driving systems, and learning-based behavior planning presents a promising pathway toward achieving this objective. Distinguished from existing learning-based methods that directly output decisions, this work introduces a predictive behavior planning framework that learns to predict and evaluate from human driving data. This framework consists of three components: a behavior generation module that produces a diverse set of candidate behaviors in the form of trajectory proposals, a conditional motion prediction network that predicts future trajectories of other agents based on each proposal, and a scoring module that evaluates the candidate plans using maximum entropy inverse reinforcement learning (IRL). We validate the proposed framework on a large-scale real-world urban driving dataset through comprehensive experiments. The results show that the conditional prediction model can predict distinct and reasonable future trajectories given different trajectory proposals and the IRL-based scoring module can select plans that are close to human driving. The proposed framework outperforms other baseline methods in terms of similarity to human driving trajectories. Additionally, we find that the conditional prediction model improves both prediction and planning performance compared to the non-conditional model. Lastly, we note that the learning of the scoring module is crucial for aligning the evaluations with human drivers. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Efficient Deep Reinforcement Learning With Imitative Expert Priors for Autonomous DrivingabstractDeep reinforcement learning (DRL) is a promising way to achieve human-like autonomous driving. However, the low sample efficiency and difficulty of designing reward functions for DRL would hinder its applications in practice. In light of this, this article proposes a novel framework to incorporate human prior knowledge in DRL, in order to improve the sample efficiency and save the effort of designing sophisticated reward functions. Our framework consists of three ingredients, namely, expert demonstration, policy derivation, and RL. In the expert demonstration step, a human expert demonstrates their execution of the task, and their behaviors are stored as state-action pairs. In the policy derivation step, the imitative expert policy is derived using behavioral cloning and uncertainty estimation relying on the demonstration data. In the RL step, the imitative expert policy is utilized to guide the learning of the DRL agent by regularizing the KL divergence between the DRL agent's policy and the imitative expert policy. To validate the proposed method in autonomous driving applications, two simulated urban driving scenarios (unprotected left turn and roundabout) are designed. The strengths of our proposed method are manifested by the training results as our method can not only achieve the best performance but also significantly improve the sample efficiency in comparison with the baseline algorithms (particularly 60% improvement compared with soft actor-critic). In testing conditions, the agent trained by our method obtains the highest success rate and shows diverse and human-like driving behaviors as demonstrated by the human expert. We also find that using the imitative expert policy trained with the ensemble method that estimates both policy and model uncertainties, as well as increasing the training sample size, can result in better training and testing performance, especially for more difficult tasks. As a result, the proposed method has shown its potential to facilitate the applications of DRL-enabled human-like autonomous driving systems in practice. The code and supplementary videos are also provided. [https://mczhi.github.io/Expert-Prior-RL/]. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Improved Deep Reinforcement Learning with Expert Demonstrations for Urban Autonomous DrivingabstractLearning-based approaches, such as reinforcement learning (RL) and imitation learning (IL), have indicated superiority over rule-based approaches in complex urban autonomous driving environments, showing great potential to make intelligent decisions. However, current RL and IL approaches still have their own drawbacks, such as low data efficiency for RL and poor generalization capability for IL. In light of this, this paper proposes a novel learning-based method that combines deep reinforcement learning and imitation learning from expert demonstrations, which is applied to longitudinal vehicle motion control in autonomous driving scenarios. Our proposed method employs the soft actor-critic structure and modifies the learning process of the policy network to incorporate both the goals of maximizing reward and imitating the expert. Moreover, an adaptive prioritized experience replay is designed to sample experience from both the agent’s self-exploration and expert demonstration, in order to improve sample efficiency. The proposed method is validated in a simulated urban roundabout scenario and compared with various prevailing RL and IL baseline approaches. The results manifest that the proposed method has a faster training speed, as well as better performance in navigating safely and time-efficiently. Zhiyu Huang, Jingda Wu |
IV | 3 |
| 2022 | Driving Behavior Modeling Using Naturalistic Human Driving Data With Inverse Reinforcement LearningabstractDriving behavior modeling is of great importance for designing safe, smart, and personalized autonomous driving systems. In this paper, an internal reward function-based driving model that emulates the human’s decision-making mechanism is utilized. To infer the reward function parameters from naturalistic human driving data, we propose a structural assumption about human driving behavior that focuses on discrete latent driving intentions. It converts the continuous behavior modeling problem to a discrete setting and thus makes maximum entropy inverse reinforcement learning (IRL) tractable to learn reward functions. Specifically, a polynomial trajectory sampler is adopted to generate candidate trajectories considering high-level intentions and approximate the partition function in the maximum entropy IRL framework. An environment model considering interactive behaviors among the ego and surrounding vehicles is built to better estimate the generated trajectories. The proposed method is applied to learn personalized reward functions for individual human drivers from the NGSIM highway driving dataset. The qualitative results demonstrate that the learned reward functions are able to explicitly express the preferences of different drivers and interpret their decisions. The quantitative results reveal that the learned reward functions are robust, which is manifested by only a marginal decline in proximity to the human driving trajectories when applying the reward function in the testing conditions. For the testing performance, the personalized modeling method outperforms the general modeling approach, significantly reducing the modeling errors in human likeness (a custom metric to gauge accuracy), and these two methods deliver better results compared to other baseline methods. Moreover, it is found that predicting the response actions of surrounding vehicles and incorporating their potential decelerations caused by the ego vehicle are critical in estimating the generated trajectories, and the accuracy of personalized planning using the learned reward functions relies on the accuracy of the forecasting model. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Battery Thermal- and Health-Constrained Energy Management for Hybrid Electric Bus Based on Soft Actor-Critic DRL AlgorithmabstractEnergy management is critical to reducing the size and operating cost of hybrid energy systems, so as to expedite on-the-move electric energy technologies. This article proposes a novel knowledge-based, multiphysics-constrained energy management strategy for hybrid electric buses, with an emphasized consciousness of both thermal safety and degradation of onboard lithium-ion battery (LIB) system. Particularly, a multiconstrained least costly formulation is proposed by augmenting the overtemperature penalty and multistress-driven degradation cost of LIB into the existing indicators. Further, a soft actor-critic deep reinforcement learning strategy is innovatively exploited to make an intelligent balance over conflicting objectives and virtually optimize the power allocation with accelerated iterative convergence. The proposed strategy is tested under different road missions to validate its superiority over existing methods in terms of the converging effort, as well as the enforcement of LIB thermal safety and the reduction of overall driving cost. Jingda Wu, Zhongbao Wei, Yu Wang 0071, Yunwei Li 0001, Dirk Uwe Sauer |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Reference-Free Human-Automation Shared Control for Obstacle Avoidance of Automated VehiclesabstractIn this paper, a novel reference-free shared control system is designed for obstacle avoidance for automated vehicles. Rather than using a reference path to guide the driver, the proposed framework constrains the vehicle's status to guarantee the safety without scarifying the driver's freedom. The constrained Delaunay triangle method is introduced to identify the vehicle's position constraints and the constraints of obstacle avoidance, vehicle stability and physical limitations are investigated and unified. A nonlinear predictive control problem, which is constructed accounting nonlinear vehicle dynamics and given driver actions, is designed to optimize the steering and braking actions needed to keep the vehicle safe. The automation is supposed to correct the driver's steering or braking actions to prevent constraint violation and losing the control of vehicle. The simulation results show that the automation can assist the driver to avoid obstacles and guarantee the vehicle's stability with minimal control intervention. Chao Huang 0006, Peng Hang, Jingda Wu, Anh-Tu Nguyen, Chen Lv 0001 |
SMC | 3 |