VLDB 2026 Research / reviewers in the wild / expert
Zhong Cao 0003
dblp:27/8404-3
· DBLP profile ↗
15ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-2243-5705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving PolicyabstractWith the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement data (often occurring as a single instance) restricts training effectiveness. Furthermore, some disengagement data should be excluded since the disengagement may not always come from the failure of driving policies, e.g. the driver may casually intervene for a while. To this end, this work proposes disengagement-reason-augmented reinforcement learning (DRARL), which enhances driving policy improvement process according to the reason of disengagement cases. Specifically, the reason of disengagement is identified by an out-of-distribution (OOD) state estimation model. When the reason doesn’t exist, the case will be identified as a casual disengagement case, which doesn’t require additional policy adjustment. Otherwise, the policy can be updated under a reason-augmented imagination environment, improving the policy performance of disengagement cases with similar reasons. The method is evaluated using real-world disengagement cases collected by autonomous driving robotaxi. Experimental results demonstrate that the method accurately identifies policy-related disengagement reasons, allowing the agent to handle both original and semantically similar cases through reason-augmented training. Furthermore, the approach prevents the agent from becoming overly conservative after policy adjustments. Overall, this work provides an efficient way to improve driving policy performance with disengagement cases. Weitao Zhou, Bo Zhang 0106, Zhong Cao 0003, Xiang Li 0001, Diange Yang |
IROS | 3 |
| 2025 | From Prediction to Planning: Comprehensive Uncertainty Management in Autonomous Driving
Wenbo Shao, Zhong Cao 0003, Hong Wang 0014, Jun Li 0082 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Stochastic pedestrian avoidance for autonomous vehicles using hybrid reinforcement learningabstractEnsuring the safety of pedestrians is essential and challenging when autonomous vehicles are involved. Classical pedestrian avoidance strategies cannot handle uncertainty, and learning-based methods lack performance guarantees. In this paper we propose a hybrid reinforcement learning (HRL) approach for autonomous vehicles to safely interact with pedestrians behaving uncertainly. The method integrates the rule-based strategy and reinforcement learning strategy. The confidence of both strategies is evaluated using the data recorded in the training process. Then we design an activation function to select the final policy with higher confidence. In this way, we can guarantee that the final policy performance is not worse than that of the rule-based policy. To demonstrate the effectiveness of the proposed method, we validate it in simulation using an accelerated testing technique to generate stochastic pedestrians. The results indicate that it increases the success rate for pedestrian avoidance to 98.8%, compared with 94.4% of the baseline method. Huiqian Li, Jin Huang 0002, Zhong Cao 0003, Diange Yang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Reliable Autonomous Driving Environment Model With Unified State-Extended BoundaryabstractFrom the early stage of robotic applications to current autonomous driving technologies, environment modeling has been acting as the middleware for connecting perception and decision layers. In robotic applications, space-oriented models (e.g., grid map, drivable area) are widely applied to faithfully reflect the space occupation. With the development of autonomous driving, highly dynamic and complex road environment brings rising need to understand the type and motion status of objects, thus element list has became the mainstream environment model. However, along comes the reliablity problem caused by missed detection and irregular objects, which is still inevitable despite the detection accuracy improvement. In view of this, a new view of driving environment is proposed as the unified state-extended boundary (USEB), aiming to improve the reliablity of element-oriented model. For driving decision requirements, different types of elements are consistently converted into driving constraints. Semantics and dynamics are expressed as the status of drivable area boundary, making it possible to merge space occupation to improve reliability against missed detection and irregular objects. Evaluation of USEB is carried out on the nuScenes dataset. Comparative results show that the proposed USEB could cover the required information for driving decision, whereas achieving higher reliability than the commonly applied element-oriented model. Xinyu Jiao, Kun Jiang 0002, Yunlong Wang 0009, Zhong Cao 0003, Mengmeng Yang 0001, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Semantic Traffic Law Adaptive Decision-Making for Self-Driving VehiclesabstractFacts proved that obeying traffic laws keeps the promise to promote the safety of self-driving vehicles. Current self-driving vehicles usually have fixed algorithms during autonomous driving, however the traffic laws may differ or change in different regions or times, e.g., tidal lanes. It raises a crucial requirement to make self-driving vehicles adapt to the newly received traffic laws. The challenges are that traffic laws are usually semantic and manually designed, but the original algorithms may not always contain the pre-designed interface to adapt to emerging laws. To this end, this work proposes a traffic law adaptive decision-making platform, which uses the linear temporal logic (LTL) formula to consistently describe the semantic traffic laws. Then, an LTL-based reinforcement learning framework is designed to estimate the probability of illegal behavior under different traffic laws. Finally, a law-specific backup policy is designed to maintain the performance threshold by monitoring the probability of illegal behavior. This work takes three typical scenarios where the traffic laws differ for instance to prove the effectiveness of the proposed approach, i.e., law amendment presented by the government, law difference between different regions, and temporary traffic control. The results show that the proposed method can help the original decision-making algorithms adapt to the traffic laws well without pre-defined interfaces. This method provides a way to administer on-road driving self-driving vehicles. Hong Wang 0014, Zhong Cao 0003, Wenhao Yu 0006, Chengxiang Zhao, Ding Zhao, Diange Yang, Jun Li 0082 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Identify, Estimate and Bound the Uncertainty of Reinforcement Learning for Autonomous DrivingabstractDeep reinforcement learning (DRL) has emerged as a promising approach for developing more intelligent autonomous vehicles (AVs). A typical DRL application on AVs is to train a neural network-based driving policy. However, the black-box nature of neural networks can result in unpredictable decision failures, making such AVs unreliable. To this end, this work proposes a method to identify and protect unreliable decisions of a DRL driving policy. The basic idea is to estimate and constrain the policy’s performance uncertainty, which quantifies potential performance drop due to insufficient training data or network fitting errors. By constraining the uncertainty, the DRL model’s performance is always greater than that of a baseline policy. The uncertainty caused by insufficient data is estimated by the bootstrapped method. Then, the uncertainty caused by the network fitting error is estimated using an ensemble network. Finally, a baseline policy is added as the performance lower bound to avoid potential decision failures. The overall framework is called uncertainty-bound reinforcement learning (UBRL). The proposed UBRL is evaluated on DRL policies with different amounts of training data, taking an unprotected left-turn driving case as an example. The result shows that the UBRL method can identify potentially unreliable decisions of DRL policy. The UBRL guarantees to outperform baseline policy even when the DRL policy is not well-trained and has high uncertainty. Meanwhile, the performance of UBRL improves with more training data. Such a method is valuable for the DRL application on real-road driving and provides a metric to evaluate a DRL policy. Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Dynamically Conservative Self-Driving Planner for Long-Tail CasesabstractSelf-driving vehicles (SDVs) are becoming reality but still suffer from “long-tail” challenges during natural driving: the SDVs will continually encounter rare, safety-critical cases that may not be included in the dataset they were trained. Some safety-assurance planners solve this problem by being conservative in all possible cases, which may significantly affect driving mobility. To this end, this work proposes a method to automatically adjust the conservative level according to each case’s “long-tail” rate, named dynamically conservative planner (DCP). We first define the “long-tail” rate as an SDV’s confidence to pass a driving case. The rate indicates the probability of safe-critical events and is estimated using the statistics bootstrapped method with historical data. Then, a reinforcement learning-based planner is designed to contain candidate policies with different conservative levels. The final policy is optimized based on the estimated “long-tail” rate. In this way, the DCP is designed to automatically adjust to be more conservative in low-confidence “long-tail” cases while keeping efficient otherwise. The DCP is evaluated in the CARLA simulator using driving cases with “long-tail” distributed training data. The results show that the DCP can accurately estimate the “long-tail” rate to identify potential risks. Based on the rate, the DCP automatically avoids potential collisions in “long-tail” cases using conservative decisions while not affecting the average velocity in other typical cases. Thus, the DCP is safer and more efficient than the baselines with fixed conservative levels, e.g., an always conservative planner. This work provides a technique to guarantee SDV’s performance in unexpected driving cases without resorting to a global conservative setting, which contributes to solving the “long-tail” problem practically. Weitao Zhou, Zhong Cao 0003, Nanshan Deng, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Confidence-Aware Reinforcement Learning for Self-Driving CarsabstractReinforcement learning (RL) can be used to design smart driving policies in complex situations where traditional methods cannot. However, they are frequently black-box in nature, and the resulting policy may perform poorly, including in scenarios where few training cases are available. In this paper, we propose a method to use RL under two conditions: (i) RL works together with a baseline rule-based driving policy; and (ii) the RL intervenes only when the rule-based method seems to have difficulty handling and when the confidence of the RL policy is high. Our motivation is to use a not-well trained RL policy to reliably improve AV performance. The confidence of the policy is evaluated by Lindeberg-Levy Theorem using the recorded data distribution in the training process. The overall framework is named “confidence-aware reinforcement learning” (CARL). The condition to switch between the RL policy and the baseline policy is analyzed and presented. Driving in a two-lane roundabout scenario is used as the application case study. Simulation results show the proposed method outperforms the pure RL policy and the baseline rule-based policy. Zhong Cao 0003, Shaobing Xu, Huei Peng, Diange Yang, Robert Zidek |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | A General Autonomous Driving Planner Adaptive to Scenario CharacteristicsabstractAutonomous vehicle requires a general planner for all possible scenarios. Existing researches design such a planner by a unified scenario description. However, it may significantly increase the planner complexity even in some simple tasks, e.g., car following, further resulting in unsatisfactory driving performance. This work aims to design a general planner which can 1) drive in all possible scenarios and 2) have lower complexity in some common scenarios. To this end, this work proposes a pertinent boundary for multi-scenario driving planning. The total approach is named as Pertinent Boundary-based Unified Decision system. Based on the original drivable area, the pertinent boundary can further support motion status and semantics of the traffic elements, which provides the potential of pertinent performance for given scenarios. The pertinent boundary can support unified driving with the drivable area, in the meantime, can be pertinently modified to support the pertinent driving decisions for identified driving scenarios (e.g., car-following, junction left turning). It will further avoid the bump between the connections of the scenarios due to the continuity of space boundary. Thus, the planner is suitable for the fully autonomous driving. The proposed method is validated in different classical driving decision scenarios. Results show that the proposed method can support pertinent driving decisions in identified scenarios, in the meantime, assure generalized cross-scenario planning when no scenario information is available. Such a method shed light on fully autonomous driving by pertinence improvement of multi-scenario decision in the complex real world. Xinyu Jiao, Zhong Cao 0003, Kun Jiang 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | PNNUAD: Perception Neural Networks Uncertainty Aware Decision-Making for Autonomous VehicleabstractMost environment perception methods in autonomous vehicles rely on deep neural networks because of their impressive performance. However, neural networks have black-box characteristics in nature, which may lead to perception uncertainty and untrustworthy autonomous vehicles. Thus, this work proposes a decision-making method to adapt the potential perception uncertainty due to the sensor noises, fuzzy features, and unfamiliar inputs. The whole method is named as Perception Neural Networks Uncertainty Aware Decision-Making (PNNUAD) method. PNNUAD first uses the Monte Carlo dropout method to estimate the perception neural network uncertainty into a distribution around the original output. Then, the perception uncertainty will be considered in a designed reinforcement learning-based planner using a distributed value function. Finally, a backup policy will maintain the vehicle’s performance to avoid disastrous perception uncertainty. The evaluation section uses an augmented reality urban driving scenario; namely, the scenario builds in the CARLA simulator while the perception uncertainty comes from the real dataset. This case study focuses on the object class uncertainty of a widely used neural network, i.e., YOLO-V3. The results indicate that the proposed method can maintain AV safety even with poor perception performance. Meanwhile, the AV has not become too conservative by defending the perception uncertainty. This work is necessary for applying the statistics neural networks to safety-critical autonomous vehicles, and the source code will be open-source in this work. Hong Wang 0014, Zhong Cao 0003, Diange Yang, Jun Li 0082 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | System and Experiments of Model-Driven Motion Planning and Control for Autonomous VehiclesabstractThis article presents a model-based motion planning and control system for autonomous vehicles and its experimental validation. The system consists of four modules: 1) global routing; 2) behavior planner; 3) local trajectory generation; and 4) trajectory tracking. The algorithm and software of each module are detailed, including a behavior planner with unified models to handle typical scenarios in both highway and urban driving, a deterministic sampling algorithm for robust responsive trajectory generation, and a dynamics-and-delay-aware preview algorithm to achieve accurate trajectory tracking. The developed system is implemented and tested at the Mcity test facility with a full-size automated car and a dozen of challenging traffic scenarios. Shaobing Xu, Robert Zidek, Zhong Cao 0003, Pingping Lu, Xinpeng Wang 0002, Boqi Li 0001, Huei Peng |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | LiDAR-based Object Detection Failure Tolerated Autonomous Driving Planning SystemabstractA typical autonomous driving system usually relies on the detected objects from an environment perception module. Current research still cannot guarantee a perfect perception, and failure detections may cause collisions, leading to untrustworthy autonomous vehicles. This work proposes a trajectory planner to tolerate the detection failure of the LiDAR sensors. This method will plan the path relying on the detected objects as well as the raw sensor data. The overlapping and contradiction of both perception routes will be carefully addressed for safe and efficient driving. The object detector in this work uses a deep learning-based method, i.e., CNN-Segmentation neural network. The designed trajectory planner has multi-layers to handle the multi-resolution environment formed by different perception routes. The final system will dynamically adjust its attention to the detected objects or the point cloud to avoid collision due to detection failures. This method is implemented on a real autonomous vehicle to drive in an open urban area. The results show that when the autonomous vehicle fails to detect a surrounding object, e.g., vehicles or some undefined objects, the autonomous vehicles still can plan an efficient and safe trajectory. In the meantime, when the perception system works well, the A V will not be affected by the point clouds. This technology can make the autonomous vehicle trustworthy even with the black-box neural networks. The codes are open-source with our autonomous driving platform to help other researchers for A V development. Zhong Cao 0003, Weitao Zhou, Xinyu Jiao, Diange Yang |
IV | 1 |
| 2021 | Highway Exiting Planner for Automated Vehicles Using Reinforcement LearningabstractExiting from highways in crowded dynamic traffic is an important path planning task for autonomous vehicles (AVs). This task can be challenging because of the uncertain motion of surrounding vehicles and limited sensing/observing window. Conventional path planning methods usually compute a mandatory lane change (MLC) command, but the lane change behavior (e.g., vehicle speed and gap acceptance) should also adapt to traffic conditions and the urgency for exiting. In this paper, we propose a reinforcement learning-enhanced highway-exit planner. The learning-based strategy learns from past failures and adjusts the vehicle motion when the AV fails to exit. The reinforcement learning is based on the Monte Carlo tree search (MCTS) approach. The proposed learning-enhanced highway-exit planner is tested 6000 times in stochastic simulations. The results indicate that the proposed planner achieves a higher probability of successful highway exiting than a benchmark MLC planner. Zhong Cao 0003, Diange Yang, Shaobing Xu, Huei Peng, Boqi Li 0001, Shuo Feng 0002, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Monocular Depth Prediction through Continuous 3D LossabstractThis paper reports a new continuous 3D loss function for learning depth from monocular images. The dense depth prediction from a monocular image is supervised using sparse LIDAR points, which enables us to leverage available open source datasets with camera-LIDAR sensor suites during training. Currently, accurate and affordable range sensor is not readily available. Stereo cameras and LIDARs measure depth either inaccurately or sparsely/costly. In contrast to the current point-to-point loss evaluation approach, the proposed 3D loss treats point clouds as continuous objects; therefore, it compensates for the lack of dense ground truth depth due to LIDAR's sparsity measurements. We applied the proposed loss in three state-of-the-art monocular depth prediction approaches DORN, BTS, and Monodepth2. Experimental evaluation shows that the proposed loss improves the depth prediction accuracy and produces point-clouds with more consistent 3D geometric structures compared with all tested baselines, implying the benefit of the proposed loss on general depth prediction networks. A video demo of this work is available at https://youtu.be/5HL8BjSAY4Y. Minghan Zhu, Maani Ghaffari Jadidi, Yuanxin Zhong, Pingping Lu, Zhong Cao 0003, Ryan M. Eustice, Huei Peng |
IROS | 5 |
| 2020 | CLAP: Cloud-and-Learning-compatible Autonomous driving PlatformabstractAutonomous driving (AV) has been intensively researched over the last decade. In this paper, we introduce an open autonomous driving software stack to enable faster design, development, and testing of algorithms on simulated or experimental vehicles, which we hope will become a useful tool for AV researchers. Yuanxin Zhong, Zhong Cao 0003, Minghan Zhu, Xinpeng Wang 0002, Diange Yang, Huei Peng |
IV | 2 |