Xinyu Jiao

dblp:303/6781 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-2462-1691ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion Prediction
abstract
LiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion
Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang
IEEE Trans. Intell. Transp. Syst.3
2025 PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field Priors
abstract
Reliable spatial and motion perception is essential for safe autonomous navigation. Recently, class-agnostic motion prediction on bird's-eye view (BEV) cell grids derived from LiDAR point clouds has gained significant attention. However, existing frameworks typically perform cell classification and motion prediction on a per-pixel basis, neglecting important motion field priors such as rigidity constraints, temporal consistency, and future interactions between agents. These limitations lead to degraded performance, particularly in sparse and distant regions. To address these challenges, we introduce \textbf{PriorMotion}, an innovative generative framework designed for class-agnostic motion prediction that integrates essential motion priors by modeling them as distributions within a structured latent space. Specifically, our method captures structured motion priors using raster-vector representations and employs a variational autoencoder with distinct dynamic and static components to learn future motion distributions in the latent space. Experiments on the nuScenes dataset demonstrate that \textbf{PriorMotion} outperforms state-of-the-art methods across both traditional metrics and our newly proposed evaluation criteria. Notably, we achieve improvements of approximately 15.24\% in accuracy for fast-moving objects, an 3.59\% increase in generalization, a reduction of 0.0163 in motion stability, and a 31.52\% reduction in prediction errors in distant regions. Further validation on FMCW LiDAR sensors confirms the robustness of our approach.
Kangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo, Zheng Fu, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Diange Yang
ICCV3
2023 Reliable Autonomous Driving Environment Model With Unified State-Extended Boundary
abstract
From the early stage of robotic applications to current autonomous driving technologies, environment modeling has been acting as the middleware for connecting perception and decision layers. In robotic applications, space-oriented models (e.g., grid map, drivable area) are widely applied to faithfully reflect the space occupation. With the development of autonomous driving, highly dynamic and complex road environment brings rising need to understand the type and motion status of objects, thus element list has became the mainstream environment model. However, along comes the reliablity problem caused by missed detection and irregular objects, which is still inevitable despite the detection accuracy improvement. In view of this, a new view of driving environment is proposed as the unified state-extended boundary (USEB), aiming to improve the reliablity of element-oriented model. For driving decision requirements, different types of elements are consistently converted into driving constraints. Semantics and dynamics are expressed as the status of drivable area boundary, making it possible to merge space occupation to improve reliability against missed detection and irregular objects. Evaluation of USEB is carried out on the nuScenes dataset. Comparative results show that the proposed USEB could cover the required information for driving decision, whereas achieving higher reliability than the commonly applied element-oriented model.
Xinyu Jiao, Kun Jiang 0002, Yunlong Wang 0009, Zhong Cao 0003, Mengmeng Yang 0001, Diange Yang
IEEE Trans. Intell. Transp. Syst.1
2022 A General Autonomous Driving Planner Adaptive to Scenario Characteristics
abstract
Autonomous vehicle requires a general planner for all possible scenarios. Existing researches design such a planner by a unified scenario description. However, it may significantly increase the planner complexity even in some simple tasks, e.g., car following, further resulting in unsatisfactory driving performance. This work aims to design a general planner which can 1) drive in all possible scenarios and 2) have lower complexity in some common scenarios. To this end, this work proposes a pertinent boundary for multi-scenario driving planning. The total approach is named as Pertinent Boundary-based Unified Decision system. Based on the original drivable area, the pertinent boundary can further support motion status and semantics of the traffic elements, which provides the potential of pertinent performance for given scenarios. The pertinent boundary can support unified driving with the drivable area, in the meantime, can be pertinently modified to support the pertinent driving decisions for identified driving scenarios (e.g., car-following, junction left turning). It will further avoid the bump between the connections of the scenarios due to the continuity of space boundary. Thus, the planner is suitable for the fully autonomous driving. The proposed method is validated in different classical driving decision scenarios. Results show that the proposed method can support pertinent driving decisions in identified scenarios, in the meantime, assure generalized cross-scenario planning when no scenario information is available. Such a method shed light on fully autonomous driving by pertinence improvement of multi-scenario decision in the complex real world.
Xinyu Jiao, Zhong Cao 0003, Kun Jiang 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.1
2022 Distributed Car-Following Control for Intelligent Connected Vehicle Using Improved Super-Twisting Compensator Subject to Sudden Velocity Changes of Leading Vehicle
abstract
The optimal velocity-based model has been successfully applied to distributed car-following systems. However, the car-following performance is inevitably affected by a series of disturbances, particularly, sudden velocity changes of leading vehicle. To improve accuracy and response rate of car-following control in the presence of such disturbances, an improved super-twisting compensator (ISTC) is proposed and a composite controller is designed by combining ISTC with a finite- time controller. A second-order nominal system is constructed by using a virtual measurement signal along with its integration to facilitate the design of ISTC. By introducing the feedback of high-order estimation error, the accuracy and response rate of ISTC are increased significantly as compared with the conventional one under same gains. Such improvement further enhances the disturbance rejection ability of the composite controller. Both Lyapunov approach and numerical simulations are carried out to verify the effectiveness of the proposed method.
Ruidong Yan, Diange Yang, Jin Huang 0002, Kun Jiang 0002, Xinyu Jiao
IEEE Trans. Intell. Transp. Syst.5
2021 LiDAR-based Object Detection Failure Tolerated Autonomous Driving Planning System
abstract
A typical autonomous driving system usually relies on the detected objects from an environment perception module. Current research still cannot guarantee a perfect perception, and failure detections may cause collisions, leading to untrustworthy autonomous vehicles. This work proposes a trajectory planner to tolerate the detection failure of the LiDAR sensors. This method will plan the path relying on the detected objects as well as the raw sensor data. The overlapping and contradiction of both perception routes will be carefully addressed for safe and efficient driving. The object detector in this work uses a deep learning-based method, i.e., CNN-Segmentation neural network. The designed trajectory planner has multi-layers to handle the multi-resolution environment formed by different perception routes. The final system will dynamically adjust its attention to the detected objects or the point cloud to avoid collision due to detection failures. This method is implemented on a real autonomous vehicle to drive in an open urban area. The results show that when the autonomous vehicle fails to detect a surrounding object, e.g., vehicles or some undefined objects, the autonomous vehicles still can plan an efficient and safe trajectory. In the meantime, when the perception system works well, the A V will not be affected by the point clouds. This technology can make the autonomous vehicle trustworthy even with the black-box neural networks. The codes are open-source with our autonomous driving platform to help other researchers for A V development.
Zhong Cao 0003, Weitao Zhou, Xinyu Jiao, Diange Yang
IV4
2021 Bridging the Gap of Lane Detection Performance Between Different Datasets: Unified Viewpoint Transformation
abstract
Convolutional neural networks (CNNs) have shown excellent performance for vision-based lane detection. However, maintaining the performance of the trained models under new test scenarios still remains challenging due to the dataset bias between the training and test datasets; In lane detection processes, the dataset bias can be categorized into lane position bias and lane pattern bias, with the former one particularly influences the lane detection performance. To tackle this dataset bias, this article proposes aunified viewpoint transformation (UVT)method that transforms the camera viewpoints of different datasets into a common virtual world coordinate system, such that the mismatched lane position distributions can be effectively aligned. Experiments are conducted on multiple datasets including the Caltech[1], Tusimple[2], and KITTI[3]dataset. The results demonstrate the effectiveness of the UVT algorithm in improving the lane detection performance on the test datasets. Moreover, by incorporating the UVT into other techniques that tackling the dataset bias, the lane position and pattern differences are disentangled and separately minimized. As a result, the performance gap between the training data and the test scenarios can be bridged. Specifically, the model trained on the KITTI dataset have achieved high performance in the Tusimple and the Caltech dataset (F1-score: 84.8 and 87.1%). With the proposed algorithm, a lane detection model trained on one dataset can be effectively applied to datasets with different camera settings in vastly different localities, and achieve better generalization ability compared to the state of the art methods.
Tuopu Wen, Diange Yang, Kun Jiang 0002, Chun-lei Yu, Benny Wijaya, Xinyu Jiao
IEEE Trans. Intell. Transp. Syst.7