VLDB 2026 Research / reviewers in the wild / expert
Shi-tao Chen
dblp:195/4429 · also Shitao Chen
· DBLP profile ↗
41ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0002-9422-3058ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-performance multi-agent path finding in high-obstacle-density and large-size maps
Shiguang Sun, Chang Tang, Shi-tao Chen, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan |
Neurocomputing | 3 |
| 2026 | Multiscale Spatio-Temporal Graph Convolutional Network for UAV Anomaly Detection
Gang Hu 0018, Zhongliang Zhou, Shi-tao Chen |
IEEE Internet Things J. | 4 |
| 2026 | MonoA2: Adaptive depth with augmented head for monocular 3D object detection
Jinpeng Dong, Sanping Zhou, Jingjing Jiang, Weiliang Zuo, Shi-tao Chen, Nanning Zheng 0001 |
Pattern Recognit. | 7 |
| 2026 | CNDM: Customized Noise Diffusion Model for Trajectory PredictionabstractPrecise and reliable multi-agent trajectory prediction is fundamental to enabling safe autonomous navigation. While denoising diffusion models have demonstrated remarkable capability in capturing the multimodal uncertainty inherent in this task, their reliance on a generic, isotropic Gaussian noise prior poses a critical limitation. This uniform prior is fundamentally misaligned with the highly structured, heterogeneous uncertainty of agent motion—shaped by dynamics, type, and environmental constraints—forcing the denoising network to implicitly relearn complex motion priors from scratch. To bridge this gap, we propose the Customized Noise Diffusion Model (CNDM), a novel framework that introduces a learned, agent-specific noise prior. At the core of CNDM is a Prior-Guidance Network (PGN) that distills an agent’s history, type, and scene context into a parametric, anisotropic Gaussian distribution. This customized prior provides a physically-grounded starting point for the diffusion process. To enable efficient training on such anisotropic noise, we leverage a Mahalanobis whitening transformation to standardize the denoising task. Extensive experiments on the Waymo Open Motion and Argoverse 2 datasets show that CNDM achieves competitive performance against state-of-the-art methods, excelling particularly in capturing diverse motion patterns and improving probabilistic calibration. Ablation studies confirm that the performance gains are directly attributable to our customized noise design, underscoring the importance of integrating structured domain knowledge into the generative foundation of diffusion models. Entao Chang, Jiawei Fu 0001, Wenjie Gao 0001, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | FlowCalib: Targetless Infrastructure LiDAR-Camera Extrinsic Calibration Based on Optical Flow and Scene FlowabstractRecently, multi-sensor fusion-based vehicle infrastructure cooperative perception has aroused extensive attention due to the demands for the safety of autonomous driving and traffic monitoring. An accurate calibration between different sensors is a critical foundation for most sensor fusion systems. For LiDAR-camera calibration, high accuracy can be achieved with the help of artificial calibration targets, such as a checkerboard. However, unlike autonomous vehicles, roadside sensors monitor traffic scenes with continuous traffic flow from a fixed viewpoint, posing challenges for conventional calibration methods. There, a calibration method suitable for roadside scenes is required for infrastructure sensors. In this paper, we propose FlowCalib, a novel targetless infrastructure LiDAR-camera spatial calibration method through alignment of scene flow and optical flow. The main idea is to leverage the inherent consistency of moving objects in traffic flow across two types of sensor data. Firstly, the moving objects are extracted by optical flow and scene flow. Then, the extrinsic parameters are obtained in two steps: rough calibration and calibration refinement. In rough calibration, the center and motion flow of each moving instance are calculated by clustering methods separately in the point cloud and image. Based on this, the possible initial value set of extrinsic parameters is estimated by two-step parameter sampling. The initial parameters are obtained by distance of center and motion flow in point cloud and image based scoring. Subsequently, the extrinsic parameters are refined by optimization of instance alignment loss and flow alignment loss of moving objects. In the end, quantitative and qualitative experiments are conducted to validate the effectiveness of the algorithm across both simulated datasets and real-world datasets. Renwei Hai, Yanqing Shen, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and LightingabstractHigh-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies (domain gaps) of input images. To address this issue, we propose SAMap, a novel map learning framework that enhances domain generalization capabilities of existing models by reducing domain discrepancies in input images. SAMap innovatively introduces a Semantic Aligner, an image-to-image transformation module that aligns images from different domains into a unified domain space while preserving semantic consistency. To train this aligner, we leverage Vision-Language Models (VLMs) that have acquired image-text alignment capabilities. Specifically, we first train a Prompt Learner that combines handcrafted and learnable prompts to capture domain-invariant semantic information. We then train Semantic Aligner through dual supervision mechanisms: a content preservation loss that maintains feature consistency across transformations and a semantic alignment loss that leverages VLM’s encoders to align transformed images with domain-invariant textual representations. Adequate experiments on the NuScenes dataset demonstrate that when integrated with three existing HD map prediction methods, SAMap achieves a performance improvement of up to 11.6% on unseen domains (rain or night conditions), effectively validating its generalization capabilities across domains. Wenjie Gao 0001, Haodong Jing, Jiawei Fu 0001, Shi-tao Chen, Nanning Zheng 0001 |
IROS | 5 |
| 2025 | Modeling Human-like Driving Behavior Based on Maximum Entropy Deep Inverse Reinforcement LearningabstractModeling expert driving behavior is crucial for the successful implementation of human-like autonomous driving. In this paper, we propose a new sampling-based Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) framework. It leverages naturalistic human driving data to train the reward model and thus evaluates driving behaviors from the reward of sampled candidate trajectories. The proposed framework utilizes deep neural networks to learn the feature-reward mapping, which offers superior fitting capabilities compared to traditional linear reward functions. A polynomial trajectory sampler for long-term decision making and a dynamic window trajectory sampler for short-term planning are adopted to simplify the calculation of partition function in the MEDIRL algorithm. In addition, the proposed framework offers a solution to the probability estimation of driving behaviors by calculating the likelihood of sampled candidate trajectories based on their reward values. Comparative experiments are conducted on the NGSIM US-101 Highway dataset, and the experimental results demonstrate the superiority of the proposed model in personalizing reward functions, as well as the applicability of the proposed method in modeling driving behaviors across various time horizons. Jiamin Shi, Tangyike Zhang, Shi-tao Chen, Nanning Zheng 0001, Jingmin Xin |
IROS | 3 |
| 2025 | StructVPR++: Distill Structural and Semantic Knowledge With Weighting Samples for Visual Place RecognitionabstractVisual place recognition is a challenging task for autonomous driving and robotics, which is usually considered as an image retrieval problem. A commonly used two-stage strategy involves global retrieval followed by re-ranking using patch-level descriptors. Most deep learning-based methods in an end-to-end manner cannot extract global features with sufficient semantic information from RGB images. In contrast, re-ranking can utilize more explicit structural and semantic information in one-to-one matching process, but it is time-consuming. To bridge the gap between global retrieval and re-ranking and achieve a good trade-off between accuracy and efficiency, we propose StructVPR++, a framework that embeds structural and semantic knowledge into RGB global representations via segmentation-guided distillation. Our key innovation lies in decoupling label-specific features from global descriptors, enabling explicit semantic alignment between image pairs without requiring segmentation during deployment. Furthermore, we introduce a sample-wise weighted distillation strategy that prioritizes reliable training pairs while suppressing noisy ones. Experiments on four benchmarks demonstrate that StructVPR++ surpasses state-of-the-art global methods by 5-23% in Recall@1 and even outperforms many two-stage approaches, achieving real-time efficiency with a single RGB input. Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Leveraging Anchor-Based LiDAR 3D Object Detection via Point Assisted Sample Selectionabstract3D object detection based on LiDAR point cloud and prior anchor boxes is a critical technology for autonomous driving environment perception and understanding. Nevertheless, an overlooked practical issue in existing methods is the ambiguity in training sample allocation based on box Intersection over Union (IoUbox). This problem impedes further enhancements in the performance of anchor-based LiDAR 3D object detectors. To tackle this challenge, this paper introduces a new training sample selection method that utilizes point cloud distribution for anchor sample quality measurement, named Point Assisted Sample Selection (PASS). This method has undergone rigorous evaluation on four widely utilized datasets. Experimental results demonstrate that the application of PASS elevates the average precision of anchor-based LiDAR 3D object detectors to a novel state-of-the-art, thereby proving the effectiveness of the proposed approach. The codes will be made available at https://github.com/ XJTU-Haolin/Point_Assisted_Sample_Selection. Shi-tao Chen, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map ConstructionabstractHigh-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion by nearby vehicles, the performance of these methods significantly declines in complex scenarios and long-range detection tasks. In this paper, we explore a new perspective that boosts HD map construction through the use of satellite maps to complement onboard sensors. We initially generate the satellite map tiles for each sample in nuScenes and release a complementary dataset for further research. To enable better integration of satellite maps with existing methods, we propose a hierarchical fusion module, which includes feature-level fusion and BEV-level fusion. The feature-level fusion, composed of a mask generator and a masked cross-attention mechanism, is used to refine the features from onboard sensors. The BEV-level fusion mitigates the coordinate differences between features obtained from onboard sensors and satellite maps through an alignment module. The experimental results on the augmented nuScenes showcase the seamless integration of our module into three existing HD map construction methods. The satellite maps and our proposed module notably enhance their performance in both HD map semantic segmentation and instance detection tasks. Our code will be available at https://github.com/xjtu-csgao/SatforHDMap. Wenjie Gao 0001, Jiawei Fu 0001, Yanqing Shen, Haodong Jing, Shi-tao Chen, Nanning Zheng 0001 |
ICRA | 5 |
| 2024 | POAQL: A Partially Observable Altruistic Q-Learning Method for Cooperative Multi-Agent Reinforcement LearningabstractMulti-Agent Path Finding (MAPF) is an important issue in multi-agent cooperation. Many studies apply MultiAgent Reinforcement Learning (MARL) to solve MAPF in partially observable settings. The objective of cooperative MARL is to maximize the cumulative team reward. Nevertheless, in partially observable settings, the team reward is misleading due to unpredictable factors from the behavior and state of unobserved agents. To address this issue, we propose a Partially Observable Altruistic Q-learning (POAQL) method. POAQL considers the cumulative reward of the observed subteam instead of the whole team, where Altruistic Q-learning plays an important role in learning the subteam action value. In addition, we design a new conflict resolution without additional guidance to emphasize the cooperative nature of MARL frameworks. Experimental results show that POAQL outperforms existing reinforcement learning methods in terms of efficiency and performance. Lesong Tao, Miao Kang, Jinpeng Dong, Songyi Zhang, Ke Ye, Shi-tao Chen, Nanning Zheng 0001 |
ICRA | 6 |
| 2024 | Fast Multi-Class Vehicle Cooperative Path Optimization in Complex Urban V2X Transportation: A Novel Parallel Multi-Agent Reinforcement Learning ApproachabstractUrban road traffic systems are advancing into sophisticated networks, underscoring the importance of real-time collaborative decision-making. This study tackles the intricate challenge of cooperative path planning under complex urban conditions, taking into account a variety of vehicle types and their respective priorities. While conventional path planning techniques struggle with such intricate coordination, reinforcement learning, though theoretically capable, is hindered by its limited model reusability and protracted training times. To address these issues, we present a novel parallel multi-agent reinforcement learning strategy for path planning that is adaptable to various vehicle types. The problem is initially cast as a multi-agent Markov Decision Process (MDP), followed by the introduction of a parallel training approach within the Q-learning framework. This approach leverages tensor computation to transform the Q-table, state, and reward, thereby markedly accelerating the training process. Empirical simulations demonstrate the approach’s efficacy, achieving a 0.84% reduction in training time (from approximately 771.611 seconds to 0.654 seconds), achieving a 93.94% lower probability of path overlap though the total distance increased by 7.69%. Shi-tao Chen, Shuyang Cai, Ziheng Tang, Donghe Li, Nanning Zheng 0001 |
IV | 1 |
| 2024 | The Optimal Horizon Model Predictive Control Planning for Autonomous Vehicles in Dynamic EnvironmentsabstractA primary challenge in autonomous driving is achieving safe and efficient trajectory planning in complex dynamic environments. This task requires adherence to traffic laws and vehicle dynamics models as well as an understanding of the spatial distributions and behavior of various traffic participants in densely populated areas. Model Predictive Control (MPC) and its variants typically employ a fixed prediction horizon, which results in limited adaptability in dynamic environments. A long prediction horizon escalates computational costs, while a short prediction horizon may impact real-time performance adversely. To tackle this challenge, our study introduces an optimal horizon MPC planning approach. This method incorporates a sliding horizon window founded on reinforcement learning and interactive MPC planning, making it versatile for a variety of driving scenarios. Additionally, our approach implicitly models the spatio-temporal interactions among traffic participants, thereby enriching the information pool for effective planning. Rigorous tests and validations conducted using the real-world dataset nuPlan affirm that our proposed method delivers robust planning performance, facilitating safe and efficient trajectory planning for autonomous vehicles. Shi-tao Chen, Jiamin Shi, Nanning Zheng 0001 |
IV | 2 |
| 2024 | Human-Like Reverse Parking using Deep Reinforcement Learning with Attention MechanismabstractThis study explores efficient and safe Automated Valet Parking (AVP) strategies in unstructured and dynamic environments. Existing approaches utilizing reinforcement learning neglected the interaction between dynamic agents and ego vehicle, and disregarded human driving patterns, leading to their ineffectiveness in unstructured dynamic environments. We propose a novel hybrid attention mechanism that comprehends the mixed interactions between static and dynamic elements, aiding autonomous vehicles in advanced planning. We implemented a guidance system based on human preferences, eliminating the need for expert data and expediting the training process via intermediate planning stages, thereby facilitating parking maneuvers akin to human drivers. The model was trained and validated in a range of parking situations. The experimental outcomes indicate that our method possesses robust adaptability and navigation skills in static and dynamic environments. Zhuo Qiu, Shi-tao Chen, Jiamin Shi, Nanning Zheng 0001 |
IV | 2 |
| 2023 | StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionabstractVisual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial structural information for better performance. In this paper, we propose StructVPR, a novel training architecture for VPR, to enhance structural knowledge in RGB global features and thus improve feature stability in a constantly changing environment. Specifically, StructVPR uses segmentation images as a more definitive source of structural knowledge input into a CNN network and applies knowledge distillation to avoid online segmentation and inference of seg-branch in testing. Considering that not all samples contain high-quality and helpful knowledge, and some even hurt the performance of distillation, we partition samples and weigh each sample's distillation loss to enhance the expected knowledge precisely. Finally, StructVPR achieves impressive performance on several benchmarks using only global retrieval and even outperforms many two-stage approaches by a large margin. After adding additional re-ranking, ours achieves state-of-the-art performance while maintaining a low computational cost. Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001 |
CVPR | 5 |
| 2023 | MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection
Zewei Lin, Yanqing Shen, Sanping Zhou, Shi-tao Chen, Nanning Zheng 0001 |
ICANN (7) | 4 |
| 2023 | InteractionNet: Joint Planning and Prediction for Autonomous Driving with TransformersabstractPlanning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interaction and dynamic changes of traffic scenarios. To address this challenge, we propose InteractionNet, which leverages transformer to share global contextual reasoning among all traffic participants to capture interaction and interconnect planning and prediction to achieve joint. Besides, InteractionNet deploys another transformer to help the model pay extra attention to the perceived region containing critical or unseen vehicles. InteractionNet outperforms other baselines in several benchmarks, especially in terms of safety, which benefits from the joint consideration of planning and forecasting. The code will be available at https://github.com/fujiawei0724/InteractionNet. Jiawei Fu 0001, Yanqing Shen, Zhiqiang Jian, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IROS | 4 |
| 2023 | Efficient Lane-changing Behavior Planning via Reinforcement Learning with Imitation Learning InitializationabstractRobust lane-changing behavior planning is critical to ensuring the safety and comfort of autonomous vehicles. In this paper, we proposed an efficient and robust vehicle lane-changing behavior decision-making method based on reinforcement learning (RL) and imitation learning (IL) initialization which learns the potential lane-changing driving mechanisms from driving mechanism from the interactions between vehicle and environment, so as to simplify the manual driving modeling and have good adaptability to the dynamic changes of lane-changing scene. Our method further makes the following improvements on the basis of the Proximal Policy Optimization (PPO) algorithm: (1) A dynamic hybrid reward mechanism for lane-changing tasks is adopted; (2) A state space construction method based on fuzzy logic and deformation pose is presented to enable behavior planning to learn more refined tactical decision-making; (3) An RL initialization method based on imitation learning which only requires a small amount of scene data is introduced to solve the low efficiency of RL learning under sparse reward. Experiments on the SUMO show the effectiveness of the proposed method, and the test on the CARLA simulator also verifies the generalization ability of the method. Jiamin Shi, Tangyike Zhang, Junxiang Zhan, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IV | 4 |
| 2023 | Design Hybrid Computing Architecture for Accelerating Point Cloud RegistrationabstractHigh-precision simultaneous localization and mapping (SLAM) is one of the core technologies of unmanned driving. LiDAR-based SLAM algorithms are often complex and computationally intensive, and usually are deployed on high performance CPU or GPU computing architecture with high power consumption and low energy efficiency ratio, which is not conducive to vehicle-level applications. In this paper, we design and implement a low power CPU and FPGA hybrid computing architecture for accelerating the key algorithm of LiDAR-based localization scheme. More specifically, we propose a software and hardware co-design strategy: (1) we first propose chain representation as a new type of map representation, which uses the depth discontinuity region as the segmentation location to segment the point cloud data. Our method not only reduces noise issues for down-sampling operation in point cloud representation, but also has the same computational and storage overhead as point cloud representation. (2) We further exploit the inherent parallelism in the algorithms to design a pipeline hardware architecture, which can effectively improve the speed of the algorithm in the embedded platform. Deployed on the Xilinx ZCU102 platform, our system achieves 24.4x and 3.2x speedups compared to the ARM Cortex A53 processor and the Intel i7-10700 processor, respectively, at 4.204W power consumption without severely degrading the final output quality. Xiao Wang 0002, Xiaodong Deng, Yingxiang Li, Shi-tao Chen, Longjun Liu, Nanning Zheng 0001 |
IV | 4 |
| 2023 | Onboard Sensors-Based Self-Localization for Autonomous Vehicle With Hierarchical MapabstractLocalization is a fundamental and crucial module for autonomous vehicles. Most of the existing localization methodologies, such as signal-dependent methods (RTK-GPS and Bluetooth), simultaneous localization and mapping (SLAM), and map-based methods, have been utilized in outdoor autonomous driving vehicles and indoor robot positioning. However, they suffer from severe limitations, such as signal-blocked scenes of GPS, computing resource occupation explosion in large-scale scenarios, intolerable time delay, and registration divergence of SLAM/map-based methods. In this article, a self-localization framework, without relying on GPS or any other wireless signals, is proposed. We demonstrate that the proposed homogeneous normal distribution transform algorithm and two-way information interaction mechanism could achieve centimeter-level localization accuracy, which reaches the requirement of autonomous vehicle localization for instantaneity and robustness. In addition, benefitting from hardware and software co-design, the proposed localization approach is extremely light-weighted enough to be operated on an embedded computing system, which is different from other LiDAR localization methods relying on high-performance CPU/GPU. Experiments on a public dataset (Baidu Apollo SouthBay dataset) and real-world verified the effectiveness and advantages of our approach compared with other similar algorithms. Yanqing Shen, Yuedong Yang, Xiaodong Deng, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | Construct Effective Geometry Aware Feature Pyramid Network for Multi-Scale Object DetectionabstractFeature Pyramid Network (FPN) has been widely adopted to exploit multi-scale features for scale variation in object detection. However, intrinsic defects in most of the current methods with FPN make it difficult to adapt to the feature of different geometric objects. To address this issue, we introduce geometric prior into FPN to obtain more discriminative features. In this paper, we propose Geometry-aware Feature Pyramid Network (GaFPN), which mainly consists of the novel Geometry-aware Mapping Module and Geometry-aware Predictor Head.The Geometry-aware Mapping Module is proposed to make full use of all pyramid features to obtain better proposal features by the weight-generation subnetwork. The weights generation subnetwork generates fusion weight for each layer proposal features by using the geometric information of the proposal. The Geometry-aware Predictor Head introduces geometric prior into predictor head by the embedding generation network to strengthen feature representation for classification and regression. Our GaFPN can be easily extended to other two-stage object detectors with feature pyramid and applied to instance segmentation task. The proposed GaFPN significantly improves detection performance compared to baseline detectors with ResNet-50-FPN: +1.9, +2.0, +1.7, +1.3, +0.8 points Average Precision (AP) on Faster-RCNN, Cascade R-CNN, Dynamic R-CNN, SABL, and AugFPN respectively on MS COCO dataset. Jinpeng Dong, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001 |
AAAI | 4 |
| 2022 | Parametric Path Optimization for Wheeled Robots NavigationabstractCollision risk and smoothness are the most important factors in global path planning. Currently, planning methods that reduce global path collision risk and improve its smoothness through numerical optimization have achieved good results. However, these methods cannot always optimize the path. The reason is all points on the path are considered as decision variables, which leads to the high dimensionality of the defined optimization problem. Therefore, we propose a novel global path optimization method. The method characterizes the path as a parametric curve and then optimizes the curve's parameters with a defined objective function, which successfully reduces the dimension of optimization problem. The proposed method is compared with baseline and state-of-the-art methods. Experimental results show the path optimized by our method is not only optimal in collision risk, but also in efficiency and smoothness. Furthermore, the proposed method is also implemented and tested in both simulation and real robots. Zhiqiang Jian, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001 |
ICRA | 4 |
| 2022 | Multi-Model-Based Local Path Planning Methodology for Autonomous Driving: An Integrated FrameworkabstractAutonomous driving systems (ADSs) need to be able to respond quickly to changes in the dynamic traffic scenario. However, regardless of the changes occurring in traffic scenes, the current local path planning frameworks of ADSs are based on the fixed frequency re-planning path (i.e., running their planning algorithms repeatedly). This planning method makes it difficult to provide a reasonable traveling path, agility, and comfort for driverless vehicles in changing traffic scenarios. Therefore, this article performs an in-depth analysis of the problems of traditional planning frameworks which use a fixed frequency to replan the path and proposes a novel path planning framework that is universal based on multiple-models. The proposed framework divides the planning process into several layers, each of which has different functions. With this framework, the ADS can adaptively adjust the planning process according to the changes in traffic scenes and then provide different path planning algorithms to ensure its safety and flexibility in the process of driving. Moreover, the problems caused by the traditional planning framework can be solved. This framework has been applied to the autonomous vehicle “Pioneer”, which won first place in the 2019 China Intelligent Vehicle Future Challenge (IVFC). The effectiveness and rationality of the integrated framework of local path planning proposed in this article were verified by a large number of tests in real-world traffic scenarios. Zhiqiang Jian, Shi-tao Chen, Songyi Zhang, Yu Chen 0040, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | SceGene: Bio-Inspired Traffic Scenario Generation for Autonomous Driving TestingabstractThe core value of simulation-based autonomy tests is to create densely extreme traffic scenarios to test the performance and robustness of the algorithms and systems. Test scenarios are usually designed or extracted manually from the real-world data, which is inefficient with a remarkable domain gap compared with testing in real scenarios. Therefore, it is crucial to automatically generate realistic and diverse dynamic traffic scenarios making autonomy tests efficient. Moreover, scenario generation is expected to be interpretable, controllable, and diversified, which can be hard to achieve simultaneously by methods based on rules or deep networks. In this paper, we propose a dynamic traffic scenario generation method called SceGene, inspired by genetic inheritance and mutation processes in biological intelligence. SceGene applies biological processes, such as crossover and mutation, to exchange and mutate the content of scenarios, and involves the natural selection process to control generation direction. SceGene has three main parts: 1) a new representation method for describing the traffic scenarios’ feature; 2) a new scenario generation algorithm based on crossover, mutation, and selection; and 3) an abnormal scenario information repair method based on the microscopic driving model. Evaluation on the public traffic scenario dataset shows that SceGene can ensure highly realistic and diversified scenario generation in an interpretable and controllable way, significantly improving the efficiency of the simulation-based autonomy tests. Ao Li 0006, Shi-tao Chen, Nanning Zheng 0001, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | From Human Driving to Automated Driving: What Do We Know About Drivers?abstractHumanlike automated driving (AD) strategies which are inspired by drivers’ cognition ways may show advantages in dealing with complicated scenarios. However, many humanlike AD strategies just mimic drivers’ behaviors or some specific characteristic. Learning algorithms are powerful technics to realize these strategies, but the architectures in learning-based strategies are too simple or with no detailed foundations. Therefore, we mean to summarize drivers’ cognition characteristics and design a comprehensive and well-founded architecture for humanlike AD solutions. We review the massive studies about drivers with human driving or AD and summarize the characteristics from three perspectives, cognition foundation, cognition process, and cognition strategies. As for cognition foundation, we propose a simple analogy to show the working mechanisms of biological neural networks; as for cognition foundation, the important role of previous experience is highlighted; as for cognition strategies, we discuss drivers’ cognition compensation strategies under the influences of environment, vehicle automation, and personal states systematically. After the above review of drivers’ characteristics, we classify the methods to model drivers. We find that models based on cognition processes can maintain more cognition details, and thus we design a driving-dedicated cognitive architecture. This architecture works by the cooperation of several modules including long-term memory, management module, and so on. It has solid theoretical and factual foundations and can reflect drivers’ cognition characteristics comprehensively. Finally, we discuss what needs to be done in the near future for us to improve humanlike AD solutions gradually. Shi-tao Chen, Jingyue Zheng, Masayoshi Tomizuka, Nanning Zheng 0001, Jianqiang Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Hierarchical Motion Planning for Autonomous Driving in Large-Scale Complex ScenariosabstractMotion planning algorithms, an essential part of the autonomous driving system, have been extensively studied. However, in large-scale complex scenarios, how to develop an optimal path to comply with the requirements of smoothness and safety remains a vital issue. In this study, a hierarchical search spacial scales-based hybrid A* (termed as HHA*) motion planning method is proposed, capable of efficiently generating smooth and safe paths. The proposed HHA* method covers two stages. First, the search space is divided on a coarse scale to generate local goals. Subsequently, the novel heuristic function and exploration strategies are adopted in the fine-scale search space to generate paths like that with a human driver guided by the local goals. Moreover, with the usage of the clothoid, the smoothness of the generated path is improved to be G2–continuous (i.e., curvature continuous), which fits the vehicle’s kinematic constraints without the need for later smoothing. Numerous experimental results from the simulation and on-road tests indicate that the proposed method can effectively perform motion planning that meets smoothness and safety in large-scale complex scenarios. Songyi Zhang, Zhiqiang Jian, Xiaodong Deng, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | DSP-Net: Dense-to-Sparse Proposal Generation Approach for 3D Object Detection on Point CloudabstractObject proposals generated based on sparse points from the raw point cloud have been widely used in 3D object detection. However, following the above scheme, most existing proposal generators have two problems, one is that the features for proposal generation constrain the detection performance by containing insufficient information; the other is that the sparse points obtained from the raw point cloud are misaligned with their corresponding objects in location and feature aspects. In this paper, we propose a dense-to-sparse proposal generation approach for 3D object detection, which can deal with the two problems simultaneously. Our approach utilizes the 3D CNN backbone to output dense features as a supplement to the original sparse point features for proposal generation. Besides, an object-aware feature pooling module is designed to address the misalignment between sparse points and corresponding objects. Experiments on the KITTI dataset show that our method outperforms the existing sparse-style methods and other published state-of-the-art methods. Xinrui Yan, Shi-tao Chen, Zhixiong Nan, Jingmin Xin, Nanning Zheng 0001 |
IJCNN | 3 |
| 2021 | MPR-Net: Multi-Scale Key Points Regression for Lane DetectionabstractLane detection is usually regarded as a semantic segmentation task, however, segmentation-based methods require expensive computational costs, and it is difficult to segment the lanes with heavy noises. Therefore, this paper proposes a lane detection method based on multi-scale key point regression, which directly uses the position of the key points of the lane instead of predicting the pixel-wise outputs. First, we design a lightweight backbone to extract a series of feature maps with a forward view image as input, and then apply a multi-scale fusion network on these feature maps to obtain the location and confidence information of the key points of the lane. Finally, a clustering and curve fitting mechanism with quadratic inverse proportion are used to obtain the final lane detection. Our proposed model can recognize dashed lane markings and handle many extreme scenarios where lanes are completely occluded or heavily noised. In addition, our model uses a relatively explicit framework, which contributes to ensuring real-time performance at 30Hz. In order to prove our method's performance, we conduct experiments on the TuSimple benchmark and RVD dataset, and results demonstrate that our method achieves competitive results compared with other methods. Dantong Zhu, Shengqi Wang, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001 |
IV | 4 |
| 2021 | Predicting short-term next-active-object through visual attention and hand position
Jingjing Jiang, Zhixiong Nan, Hui Chen 0036, Shi-tao Chen, Nanning Zheng 0001 |
Neurocomputing | 4 |
| 2020 | HeLPS: Heterogeneous LiDAR-based Positioning System for Autonomous VehicleabstractLiDAR-based positioning systems are widely used in unmanned systems. However, affected by the high computational complexity of high-precision positioning algorithms, the current positioning system is supported by hardware with low power efficiency and thus hard to integrate into many platforms. In this paper, we analyze features of the positioning system in autonomous driving application, design, and apply Heterogeneous LiDAR-based Positioning System, HeLPS, with software-hardware co-design methodology to achieve better efficiency. Our contributions can be concluded in three aspects. Firstly, we design the CPU-FPGA heterogeneous positioning system accelerating Iterative Closest Point (ICP) algorithm and achieves improvements on both speed and power-efficiency. Secondly, we exploit the spatial locality in the point cloud and design a new compressed data structure for fast neighbor accessing. The experiment reports a significant speedup comparing with other data structures. Lastly, we explore the data access pattern in positioning application and develop a specific cache system and out-of-order execution system reducing memory burden. Our system is deployed on a small and cheap Xilinx Zynq7000 ARM+FPGA platform, which achieves 983.3x speedup compared with Cortex-A9 CPU, and 31.8x speedup compared with i7-7820 CPU, with only 2.37W power consumption. Yuedong Yang, Xiaodong Deng, Yanqing Shen, Shi-tao Chen, Nanning Zheng 0001 |
IECON | 5 |
| 2020 | Mixed Test Environment-based Vehicle-in-the-loop Validation - A New Testing Approach for Autonomous VehiclesabstractThe current test of autonomous driving technology requires extensive experimental verification, whether in simulation or on real roads. Netherthless, how to test autonomous vehicles thoroughly in a safe and comprehensive manner remains a major challenge. To achieve safer and more effective autonomous driving testing, this paper proposes a novel mixed test environment-based validation method with vehicle in the loop(ViL) : (1) our method supports more realistic drive safety tests in mixed scenarios which integrate the synthetic and the real-world scenarios. Synthetic scenarios offer complex traffic simulation with diverse road conditions. The real-world scenarios introduce the real autonomous driving vehicle, the real sensor suite as well as the test field to the test loop, having further bridged the gap between Hardware-in-the-Loop(HiL) testing and real road tests than ViL; (2) virtual perceptional results are simulated directly and delivered to the real vehicle in Unified Fusion Data Format(UFDF), without rendering virtual detection data for reduced resource consumption; (3) diverse test scenarios are configurable and reproducible with OSM-based High Definition(HD) map, enabling the simulation to be decoupled from a specific test filed or traffic facilities. A series of experiments on the application of our method have been demonstrated, and our approach is proved to be a promising drive safety testing technique before actual road testing. Yu Chen 0040, Shi-tao Chen, Tong Xiao 0006, Songyi Zhang, Qian Hou, Nanning Zheng 0001 |
IV | 2 |
| 2020 | Modeling Methodology of Driver-Vehicle-Environment System Dynamics in Mixed Driving SituationabstractThe interactions between driver, vehicle and environment generate vehicle's behaviors, and the interactions of vehicle groups shape the traffic modes. Therefore, the conditionally or fully automated driving technologies should be developed and tested in the driver-vehicle-environment (DVE) system rather than being developed and tested individually. To build DVE system dynamics, firstly, we propose an architecture to cope with the complicated interactions in automated vehicles (AVs) and in mixed traffic situations. Then we summarize the driver behavior models and compare the differences of intelligence between human driver and automated vehicle. Finally, we summarize the feasible modeling approaches of DVE system into five categories. The primary distinctions are the modeling methods of human drivers' roles, which are realized by human driver per se, human driver's cognitive architecture, psychological motivation model, mechanism imitation, and specific mechanism transfer respectively. Taking the applications of human machine interface and AD strategy developments as examples, we analyze the benefits and drawbacks of these approaches. Shi-tao Chen, Nanning Zheng 0001, Jianqiang Wang 0003 |
IV | 2 |
| 2020 | MuRF-Net: Multi-Receptive Field Pillars for 3D Object Detection from Point CloudabstractIn this paper, we propose a point cloud based 3D object detection framework that accounts for both contextual and local information by leveraging multi-receptive field pillars, named as MuRF-Net. Recently, common pipelines can be divided into a voxel-based feature encoder and an object detector. During the feature encoding steps, contextual information is neglected, which is critical for the 3D object detection task. Thus, the encoded features are not suitable to input to the subsequent object detector. To address this challenge, we propose the MuRF-Net with a multi-receptive field voxelization mechanism to capture both contextual and local information. After the voxelization, the voxelized points (pillars) are processed by a feature encoder, and a channel-wise feature reconfiguration module is proposed to combine the features with different receptive fields using a lateral enhanced fusion network. In addition, to handle the increase of memory and computational cost brought by multi-receptive field voxelization, a dynamic voxel encoder is applied taking advantage of the sparseness of the point cloud. Experiments on the KITTI benchmark for both 3D object and Bird's Eye View (BEV) detection tasks on car class are conducted and MuRF-Net achieved the state-of-the-art results compared with other voxel-based methods. Besides, the MuRF-Net can achieve nearly real-time speed with 20Hz. Xinrui Yan, Shi-tao Chen, Jinpeng Dong, Ziyi Liu 0001, Zhixiong Nan, Nanning Zheng 0001 |
IV | 4 |
| 2020 | Traffic Agent Trajectory Prediction Using Social Convolution and Attention MechanismabstractThe trajectory prediction is significant for the decision-making of autonomous driving vehicles. In this paper, we propose a model to predict the trajectories of target agents around an autonomous vehicle. The main idea of our method is considering the history trajectories of the target agent and the influence of surrounding agents on the target agent. To this end, we encode the target agent history trajectories as an attention mask and construct a social map to encode the interactive relationship between the target agent and its surrounding agents. Given a trajectory sequence, the LSTM networks are firstly utilized to extract the features for all agents, based on which the attention mask and social map are formed. Then, the attention mask and social map are fused to get the fusion feature map, which is processed by the social convolution to obtain a fusion feature representation. Finally, this fusion feature is taken as the input of a variable-length LSTM to predict the trajectory of the target agent. We note that the variable-length LSTM enables our model to handle the case that the number of agents in the sensing scope is highly dynamic in traffic scenes. To verify the effectiveness of our method, we widely compare with several methods on a public dataset, achieving a 20% error decrease. In addition, the model satisfies the real-time requirement with the 32 fps. Tao Yang 0032, Zhixiong Nan, Shi-tao Chen, Nanning Zheng 0001 |
IV | 4 |
| 2020 | ATV Navigation in Complex and Unstructured Environment Containing StairsabstractSelf-driving and robotic technologies have been widely used in recent years. However, before L5 autonomy has been mature, it is more important and meaningful to apply these technologies to some wheeled robots on specific scenes. All-terrain vehicles (ATV) are widely used in urban search and rescue. An unstructured complex scene on which an ATV works usually contains many unusual obstacles, such as stairs. As a result, it is important for autonomous ATVs to detect, localize, and traverse stairs. In this paper, a real-time outdoor stair detection and localization method is proposed. A VLP-16 LIDAR is used to collect environment data, and the stairs are detected and localized using a single frame of LIDAR data by their geometric features, such as slope and parallel edges. A stair navigation strategy is also proposed in this paper. 3-axis attitudes measured by IMU are used during stair climbing on the basis of the coupling between roll and yaw on a slope. Experiments are carried out and the results prove the robustness and accuracy of detection and localization algorithm. The navigation strategy is proven to be safe and feasible. Kongtao Zhu, Junxiang Zhan, Shi-tao Chen, Zhixiong Nan, Tangyike Zhang, Dantong Zhu, Nanning Zheng 0001 |
IV | 3 |
| 2019 | High-Definition Map Combined Local Motion Planning and Obstacle Avoidance for Autonomous DrivingabstractLocal motion planning plays an important role in an autonomous driving system. And applying mature local motion planning methods to real traffic scenarios with regular constraints is one of the keys to the applications of autonomous vehicles. In this paper, we present a local motion planning method combined with High-Definition (HD) maps. Through the HD map defined by OpenStreetMap, the local motion planner can obtain the prior knowledge of traffic scenarios and achieve path planning and optimization accordingly. In order to improve the safety and comfort of the obstacle avoiding process, we also propose an inertia-like path selection algorithm based on this planning method. We evaluated the proposed method on our designed autonomous driving experimental platform `Pioneer' and participated in the 2018 Intelligent Vehicles Future Challenge. In the competition, the `Pioneer' successfully completed all the races and won the championship without any manual intervention. Zhiqiang Jian, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001 |
IV | 3 |
| 2019 | Robust Extrinsic Parameter Calibration of 3D LIDAR Using Lie AlgebrasabstractIn the field of autonomous driving, multi-beam light detection and ranging (3D LIDAR) system and global navigation satellite system/integrated inertial navigation system (GNSS/INS) are widely used in high-definition map construction, localization and obstacle detection. As 3D LIDAR system and INS have their own coordinate systems, the calibration of the two mentioned systems is required. In this paper, a novel algorithm for calibrating the coordinate system of 3D LIDAR and INS is proposed, which consists of three parts. The first procedure is to project two point clouds to the world coordinate system based on the initial transform matrix between 3D LIDAR and INS with the real-time data from INS. Then optimal point-to-point correspondences can be found between two frames of point cloud data through registration method. Finally, the loss function is constructed with the sum of the Euclidean distances of the corresponding points and optimized by using perturbation model of Lie algebras, so as to obtain the optimal transform matrix. With different given initial calibration parameters, test results of both simulation and real experiments validate the proposed algorithm and quantify its accuracy and robustness. Yanqing Shen, Tangyike Zhang, Songyi Zhang, Yongbo Huo, Shi-tao Chen, Nanning Zheng 0001 |
IV | 6 |
| 2019 | Autonomous driving: cognitive construction and situation understanding
Shi-tao Chen, Zhiqiang Jian, Yu Chen 0040, Zhuoli Zhou, Nanning Zheng 0001 |
Sci. China Inf. Sci. | 1 |
| 2018 | Autonomous Vehicle Testing and Validation Platform: Integrated Simulation System with Hardware in the LoopabstractWith the development of autonomous driving, offline testing remains an important process allowing low-cost and efficient validation of vehicle performance and vehicle control algorithms in multiple virtual scenarios. This paper aims to propose a novel simulation platform with hardware in the loop (HIL). This platform comprises of four layers: the vehicle simulation layer, the virtual sensors layer, the virtual environment layer and the Electronic Control Unit (ECU) layer for hardware control. Our platform has attained multiple capabilities: (1) it enables the construction and simulation of kinematic car models, various sensors and virtual testing fields; (2) it performs a closed-loop evaluation of scene perception, path planning, decision-making and vehicle control algorithms, whilst also having multi-agent interaction system; (3) it further enables rapid migrations of control and decision-making algorithms from the virtual environment to real self-driving cars. In order to verify the effectiveness of our simulation platform, several experiments have been performed with self-defined car models in virtual scenarios of a public road and an open parking lot and the results are substantial. Yu Chen 0040, Shi-tao Chen, Tangyike Zhang, Songyi Zhang, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 2 |
| 2018 | Model-Based Decision Making With Imagination for Autonomous ParkingabstractAutonomous parking technology is a key concept within autonomous driving research. This paper will propose an imaginative autonomous parking algorithm to solve issues concerned with parking. The proposed algorithm consists of three parts: an imaginative model for anticipating results before parking, an improved rapid-exploring random tree (RRT) for planning a feasible trajectory from a given start point to a parking lot, and a path smoothing module for optimizing the efficiency of parking tasks. Our algorithm is based on a real kinematic vehicle model; which makes it more suitable for algorithm application on real autonomous cars. Furthermore, due to the introduction of the imagination mechanism, the processing speed of our algorithm is ten times faster than that of traditional methods, permitting the realization of real-time planning simultaneously. In order to evaluate the algorithm's effectiveness, we have compared our algorithm with traditional RRT, within three different parking scenarios. Ultimately, results show that our algorithm is more stable than traditional RRT and performs better in terms of efficiency and quality. Ziyue Feng, Shi-tao Chen, Yu Chen 0040, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 2 |
| 2017 | Hybrid-augmented intelligence: collaboration and cognitionabstractThe long-term goal of artificial intelligence (AI) is to make machines learn and think like human beings. Due to the high levels of uncertainty and vulnerability in human life and the open-ended nature of problems that humans are facing, no matter how intelligent machines are, they are unable to completely replace humans. Therefore, it is necessary to introduce human cognitive capabilities or human-like cognitive models into AI systems to develop a new form of AI, that is, hybrid-augmented intelligence. This form of AI or machine intelligence is a feasible and important developing model. Hybrid-augmented intelligence can be divided into two basic models: one is human-in-the-loop augmented intelligence with human-computer collaboration, and the other is cognitive computing based augmented intelligence, in which a cognitive model is embedded in the machine learning system. This survey describes a basic framework for human-computer collaborative hybrid-augmented intelligence, and the basic elements of hybrid-augmented intelligence based on cognitive computing. These elements include intuitive reasoning, causal models, evolution of memory and knowledge, especially the role and basic principles of intuitive reasoning for complex problem solving, and the cognitive learning framework for visual scene understanding based on memory and reasoning. Several typical applications of hybrid-augmented intelligence in related fields are given. Nanning Zheng 0001, Ziyi Liu 0001, Pengju Ren, Shi-tao Chen, Si-yu Yu, Jianru Xue, Badong Chen, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |