Jian Zhou 0011

dblp:97/97-11 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0001-6707-6542ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A data-driven spatio-temporal driving risk field mechanism for path planning
Zhuoer Wang, Baohan Shi, Jian Zhou 0011, Bingrong Xu
Expert Syst. Appl.5
2026 Efficient Topology-Aware Motion Planning for AVP in Large-Scale Occupancy Map
abstract
With the development of autonomous driving technology, autonomous valet parking (AVP) has become a key technology to solve the problem of urban parking. Current commercial AVP systems generally adopt solutions based on semantic maps, which achieve high-precision parking in small-scale scenarios. However, when the parking environment is expanded to large underground parking lots, semantic and occupancy grid maps face bottleneck problems such as a sharp drop in path generation efficiency and delayed parking space retrieval response. In addition, traditional High-definition maps (HD maps) rely on manual annotation and complex post-processing. In response to the above challenges, this article proposes an efficient adaptive topology plan for AVP in large-scale occupancy map: first, a scale-adaptive index model based on the R-tree structure is constructed to achieve hierarchical storage and dynamic resolution selection of grid map data; secondly, a multi - scale feature fusion topology aware method is designed to generate the environment topology; finally, a multi-path parallel hybrid A${}^{\ast }$algorithm is proposed for efficient planning. A comparison of our framework with both traditional and state-of-the-art methods shows that the framework is capable of enhancing planning efficiency and reducing average path generation time in large parking lots. Through simulations and real-world tests, the method has been shown to reduce path search time whilst generating paths that are easier to track with less tracking error.
Jian Zhou 0011, Fuyu Nie, Haoran Li 0022, Jinsheng Xiao
IEEE Trans. Intell. Transp. Syst.2
2025 DeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud Analysis
abstract
Due to the irregular and disordered data structure in 3D point clouds, prior works have focused on designing more sophisticated local representation methods to capture these complex local patterns. However, the recognition performance has saturated over the past few years, indicating that increasingly complex and redundant designs no longer make improvements to local learning. This phenomenon prompts us to diverge from the trend in 3D vision and instead pursue an alternative and successful solution: deeper neural networks. In this paper, we propose DeepLA-Net, a series of very deep networks for point cloud analysis. The key insight of our approach is to exploit a small but mighty local learning block, which uses 10× fewer FLOPs, enabling the construction of very deep networks. Furthermore, we design a training supervision strategy to ensure smooth gradient backpropagation and optimization in very deep networks. We construct the DeepLA-Net family with a depth of up to 120 blocks — at least 5× deeper than recent methods — trained on a single RTX 3090. An ensemble of the DeepLA-Net achieves state-of-the-art performance on classification and segmentation tasks of S3DIS Area5 (+2.2% mIoU), ScanNet test set (+1.6% mIoU), ScanObjectNN (+2.1% OA), and ShapeNet-Part (+0.9% cls.mIoU). The code are released at https://github.com/zeng-ziyin/DeepLA-Net.
Ziyin Zeng, Mingyue Dong, Jian Zhou 0011, Huan Qiu, Zhen Dong 0005
CVPR3
2025 Lane-level map matching for vehicles using mask-based raster high-definition maps
Ziyue Tian, Jian Zhou 0011, Zhang-Cai Yin, Quanhua Dong, Shen Ying
Expert Syst. Appl.4
2025 HGCM: A hybrid radar feature and GNSS based continuous mapping method
abstract
Online mapping and localization are critical technologies in autonomous driving and robotics systems, currently being a prominent research topic. However, the majority of proposed methods have primarily focused on camera and LiDAR (Light Detection and Ranging) sensors, which are inadequate for addressing more challenging weather conditions. Radar-based simultaneous localization and mapping (SLAM) is essential for autonomous navigation in complex and dynamic environments where other sensors may fail. However, there is a gap in accuracy and stability compared with LiDAR-based SLAM, thus limiting its application for autonomous driving, drones and robotics. High noise levels, low feature stability and the lack of global constraints would result in the deterioration of pose estimation. This paper presents a Simultaneous Localization and Mapping (SLAM) system that combines hybrid features from a spinning radar and Global Navigation Satellite System (GNSS) measurements to achieve accurate mapping results in complex urban environments. To ensure intra-frame and inter-frame stability of pose estimation, we applied structural noise removal, feature ranking and motion buffer techniques. Focusing on improving trajectory consistency, we introduce a GNSS factor and a motion factor to provide stable global constraints and mitigate the impact of GNSS noise on the system, respectively. Experimental results on the Mulran dataset, which includes complex urban scenarios, demonstrate that our system achieves state-of-the-art performance, with an average ATE(Absolute Trajectory Error) of 2.71m and a standard deviation of 1.10m.
Youchen Tang, Jian Zhou 0011, Maosheng Yan, Yuanxian Huang, Jinsheng Xiao
Expert Syst. Appl.2
2025 TCCBRN: A Causal Convolutional Recurrent Network for Indirect Trajectory Prediction Without Its Position Information
Zhuoer Wang, Zuo Huang, Jian Zhou 0011, Haomiao Bian
IEEE Internet Things J.5
2025 BEVFix: Deep feature enhancement for robust 3D object detection
Jian Zhou 0011, Chi Chen 0002, Hongkai Yu, Bo Du 0001, Qin Zou 0001
Neural Networks2
2025 Revisiting the Learning Stage in Range View Representation for Autonomous Driving
abstract
LiDAR segmentation is crucial for autonomous driving perception. Range view methods have been widely adopted for these applications due to their intuitiveness and ease of implementation. However, the inherent shortcomings of the range view approach (e.g., assuming that point clouds within the same pixel of a range image have the same semantic class) make it difficult to perform accurate fine-grained segmentation tasks, thus limiting its potential in practical applications. To address these issues, we propose RangeFusion, an end-to-end framework that greatly improves the ability to learn and process LiDAR point clouds from range views by employing a multispatial learning model. A novel range-scan space (RSS) is proposed to address the inability of existing range view methods to accurately aggregate features of neighboring points. This space achieves accurate and efficient neighboring point feature aggregation with linear time complexity. In addition, a supervised label smoothing method called multilevel feature selection heads (MFSHs) is designed, which achieves more fine-grained semantic prediction by subdividing the full point cloud into multisemantic hierarchical subclouds and adaptively fusing the features with confidence filtering. The performance of the proposed method was evaluated on several benchmarks, including SemanticKITTI and nuScenes. On these two datasets, mean intersection over union (mIoU) scores of 67.9% and 80.2% were achieved, respectively. This demonstrates that the proposed approach outperforms existing range view- and multiview-based approaches while maintaining efficient performance at 26.5 FPS. In addition, real road data were collected for testing. The code is available athttps://github.com/Wansit99/RangeFusion.
Jinsheng Xiao, Siyuan Wang 0011, Jian Zhou 0011, Ziyin Zeng, Ruijia Chen
IEEE Trans. Geosci. Remote. Sens.3
2025 RoCalib: Large-Scale Autonomous Geo-Calibration for Roadside Lidar With High-Definition Map
abstract
To address the challenges of low overlap and viewing direction differences in large-scale roadside LiDAR calibration, this paper proposes RoCalib, a novel automatic roadside LiDAR calibration method based on the high-definition (HD) map. This method enables geographic registration of the LiDAR without any specific target and improves the efficiency and safety of the large-scale roadside facility calibration and maintenance. First, a novel virtual reprojection model is designed to construct a virtual mapping from the HD map to LiDAR, reducing representation differences. Based on this, a universal spatial context descriptor is introduced, applicable to various LiDAR systems, facilitating rapid retrieval of LiDAR positions within the HD map. Finally, based on the multi-feature optimization method considering the road structure, the fine registration and parameter calibration of the roadside LiDAR and the HD map are completed. The proposed framework is validated on simulated, public, and self-collected datasets, demonstrating that this method can automatically and accurately achieve multi-LiDAR geographic calibration, yielding superior performance.
Cong Duan, Jian Zhou 0011, Zhen Dong 0005, Youchen Tang, Jinsheng Xiao
IEEE Trans. Intell. Transp. Syst.2
2025 MIM: High-Definition Maps Incorporated Multi-View 3D Object Detection
abstract
3D object detection has aroused increasing interest as a crucial component of autonomous driving systems. While recent works have explored various multi-modal fusion methods to enhance accuracy and robustness, fusing multi-view images and high-definition (HD) maps remains uncharted. Inspired by our previous work, we endeavor to introduce HD maps to camera-based detection, prompting the design of a new framework. To address this, we first analyze the function of HD maps in object detection to understand their benefits and the rationale for their fusion. From this analysis, we identify key disparities in view, semantics, and scale, leading to the development of MIM, a framework for HD Maps Incorporated Multi-view 3D object detection. HD maps are enriched in semantics by sampling unlabeled areas and encoding them into map features. Simultaneously, multi-view images are transformed into features in bird’s-eye view (BEV) using the adopted baseline. These features are then fused using attention mechanisms to align scales. Experiments conducted on the nuScenes dataset demonstrate that MIM outperforms camera-based methods. Moreover, an in-depth analysis investigates how HD maps impact object detection regarding each semantic layer. The results underscore the operational intricacies of HD maps in perception, setting the stage for future research. Code is available athttps://github.com/WHU-xjs/MIM-3D-Det.
Jinsheng Xiao, Jian Zhou 0011, Ziyue Tian, Hongping Zhang, Yuan-Fang Wang
IEEE Trans. Intell. Transp. Syst.3
2024 A lane-level localization method via the lateral displacement estimation model on expressway
Jian Zhou 0011, Quanhua Dong, Yaoan Bian, Zhijiang Li, Jinsheng Xiao
Expert Syst. Appl.2
2024 GPR-Former: Detection and Parametric Reconstruction of Hyperbolas in GPR B-Scan Images With Transformers
abstract
Ground Penetrating Radar (GPR) enables the non-invasive detection of various subsurface objects such as pipes, stones, etc. The location and size of the object in the medium could be obtained by fitting the generated hyperbolic signatures within the GPR B-scan and analyzing its parameters. In this paper, GPR-Former is proposed for automatic target detection and hyperbola fitting on GPR B-scan images. We have designed a transformer-based neural network to extract features to directly regress the parameters of hyperbolic signatures in the GPR B-scan data to detect targets beneath the ground automatically. A symmetry-constrained analytical solution for the hyperbolic parameters is proposed to refine the parameters derived from the transformer network, serving the extraction and analysis of buried objects in underground opaque spaces. Experiments are conducted on three datasets for the qualitative and quantitative validation of the GPR-Former, including ground-penetrating radar detection of submarine pipelines and land pipelines. Results show that the proposed method is able to automatically and efficiently extract hyperbolas from GPR B-scan images. True hyperbola-point precision (TP_Pre) and true hyperbola-point recall (TP_Rec) metrics are introduced to evaluate performances in parametric hyperbola extraction and fitting. The results show that the TP_Pre and TP_Rec of the proposed method reach 0.867, 0.402, 0.744 and 0.762, 0.736, 0.723, with an improvement of 6%, 22%, 4% compared with the state-of-the-art methods (C3 algorithm and migration learning-based method proposed by Yang), respectively.
Ang Jin, Chi Chen 0002, Bisheng Yang, Qin Zou 0001, Zhiye Wang, Zhengfei Yan, Shaolong Wu, Jian Zhou 0011
IEEE Trans. Geosci. Remote. Sens.8
2024 PointNAT: Large-Scale Point Cloud Semantic Segmentation via Neighbor Aggregation With Transformer
abstract
Given the prominence of 3D sensors in recent years, 3D point clouds are worthy to be further investigated for environment perception and scene understanding. Learning accurate local and global contexts in point clouds is pivotal for semantic segmentation, and neighbor aggregation and Transformers have achieved notable success in local and global perception in point cloud analysis, respectively. Nevertheless, studying each independently is far from the optimal solution for comprehensive feature learning. To address this, we take a novel step towards investigating and integrating the structures of neighbor aggregation and Transformers. In this paper, we introduce Point Neighbor Aggregation with Transformer (PointNAT), a conceptually straightforward and effective approach aiming to enhance the performance of 3D point cloud semantic segmentation. PointNAT consists of a Neighbor Aggregation Block (NAB) for local perception, a Point Transformer Block (PTB) for global modeling, and a Hybrid Block to connect NABs and PTBs. NABs effectively learn complex local features at varying scales through an improved neighbor aggregation operation and a multi-head mechanism. PTBs efficiently perform global attention using a small set of learnable key points. Hybrid Blocks serve as high-and-low frequency signal hybridizers, merging the strengths of these two blocks by adaptively assigning hybrid weights to local and global contexts. We have evaluated the performance of PointNAT with state-of-the-art networks on several benchmarks, including S3DIS, Toronto3D, and SensatUrban. PointNAT achieves mIoU scores of 77.8%, 84.7%, and 65.2% in these three dataset, respectively. Furthermore, it outperforms the baseline approach PointNeXt by 3.0%, 1.3%, and 4.2%, respectively, while utilizing only 59.9% of the parameters and 15.2% of the FLOPs. The results demonstrate PointNAT’s superior ability in accurately segmenting large-scale 3D point cloud scenes, emphasizing its potential to advance environment perception and scene understanding. Our code is available at https://github.com/zeng-ziyin/PointNAT.
Ziyin Zeng, Huan Qiu, Jian Zhou 0011, Zhen Dong 0005, Jinsheng Xiao
IEEE Trans. Geosci. Remote. Sens.3
2024 SGSR-Net: Structure Semantics Guided LiDAR Super-Resolution Network for Indoor LiDAR SLAM
abstract
Multi-Beam LiDAR (MBL) sensors sample the real-world with discrete 3D point clouds (PC) and have become a major and essential 3D sensing capability for autonomous robots. To ensure an accurate point sampling on surfaces, high-resolution MBL sensors (e.g., Ouster OS0-128) are commonly used to collect dense point clouds for robot tasks, including object detection and tracking, simultaneous localization and mapping (SLAM), in applications such as autonomous driving vehicles (ADVs). However, the high cost and large volume/weight/energy consumption of such sensors limit their usage in broader applications such as UAV/UGV swarms with small-scale agents with limited payload. Existing studies on Super-Resolution (SR) upsampling of the PC from low-resolution MBL have not considered the geometry semantics of the scenes, thus resulting in less optimal SR points for downstream subtasks (e.g., SLAM). Thus, this article proposes SGSR-Net, a structure semantics-guided MBL Super-Resolution network. SGSR-Net takes the low-resolution range images of the MBL sensors as input and produces dense and structure-aware Super-Resolution point cloud from those sparse measurements through a vertical spatial and channel attention-enhanced CNN model coupling with guided Monte Carlo filtering, for indoor LiDAR-SLAM applications. The SGSR-Net is validated using datasets collected by a UGV equipped with multiple MBL sensors. The results demonstrate that the proposed CG-LSR (CASE Attention Guided Encoder-Decoder LiDAR Super-Resolution Network) reduces the MAE of the SR points by 12.4% down to 0.177 m when compared with the state-of-the-art (SOTA) method Shan et al. (2020), Ren et al. (2021), Kwon et al. (2022), Long and Wang (2022). The indoor SLAM results with SR-points produced by SGSR-Net show that the mean and RMSE of the absolute pose error (APE) are decreased by 27% and 30%, down to 0.849 m and 0.902 m, respectively, which significantly improve the indoor-SLAM performance and stability of SOTA LiDAR-SLAM systems (i.e. LeGO-LOAM Shan and Englot (2018), Dellenbach et al. (2022), Vizzo et al. (2023), Zhang and Singh (2014)).
Chi Chen 0002, Ang Jin, Zhiye Wang, Yongwei Zheng, Bisheng Yang, Jian Zhou 0011, Zhigang Tu 0001
IEEE Trans. Multim.6
2023 FinGuard: A Multimodal AIGC Guardrail in Financial Scenarios
abstract
Recently, the development of foundation models has led to significant advances in the ability of artificial intelligence (AI) to generate multimodal content such as text and images. However, specialized industrial scenarios such as finance, which require high levels of security and compliance, pose challenges for the application of generative AI due to its uncontrollability. To address this issue, we propose FinGuard, a multimodal AI-generated content (AIGC) guardrail specifically designed for financial scenarios. We provide detailed definitions of the general quality, financial compliance, and security dimensions of AIGC, and implement the evaluation and inspection of multimodal AIGC including text and images. Our proposed FinGuard has been applied to a financial marketing application serving hundreds of millions of users.
Wenlong Du, Qingquan Li 0003, Jian Zhou 0011, Zhongjun Zhou
MMAsia3
2023 Tiny object detection with context enhancement and feature purification
Jinsheng Xiao, Haowen Guo, Jian Zhou 0011, Qiuze Yu, Yunhua Chen, Zhongyuan Wang 0001
Expert Syst. Appl.3
2023 FDLR-Net: A feature decoupling and localization refinement network for object detection in remote sensing images
Jinsheng Xiao, Yuntao Yao, Jian Zhou 0011, Haowen Guo, Qiuze Yu, Yuan-Fang Wang
Expert Syst. Appl.3
2023 Lane Information Extraction for High Definition Maps Using Crowdsourced Data
abstract
Lane information plays an important role in high-definition (HD) maps because it provides invaluable road information for autonomous vehicles. The most widely used lane extraction method for HD maps is based on mobile mapping systems, and it is prohibitively costly and time-consuming. In this study, we propose a novel approach for lane information extraction based on crowdsourcing vehicles equipped with monocular cameras and global navigation satellite system devices for recording road images and position data. First, we propose a lane mask propagation network to detect the lane markings in images, which are then projected from a perspective space into a three-dimensional space in accordance with the position data. Second, we propose a data management method to store lane information in the cloud data center. An improved density-based spatial clustering of applications with noise clustering algorithm and a gradual fitting algorithm are used to remove the outliers and improve the lane data accuracy. The proposed method is quantitatively evaluated against a real-world HD map produced by a mobile mapping vehicle. The experimental results show that more than 80% of the extracted lane markings meet the accuracy requirements of HD maps. In conclusion, our method can be used as a low-cost and efficient approach for updating the lane information in HD maps for autonomous vehicles.
Jian Zhou 0011, Yaoan Bian, Yuanxian Huang
IEEE Trans. Intell. Transp. Syst.1
2018 Lane Detection and Road Surface Reconstruction Based on Multiple Vanishing Point & Symposia
abstract
Lane detection algorithm based on monocular camera is one of the most popular methods in recent years, which can meet the requirement of real-time and robust for autonomous vehicle. In this way, the position of lane markers can be transferred from perspective space to road space base on the planar road assumption. However, large numbers of road scenes, especially the up and down slope road environment, cannot meet this requirement.In this paper, we propose a multiple vanishing point detection method to reconstruct the road space in slope scenes. In order to improve the accuracy of vanishing point estimation, the road images are decomposed into near and far regions. We extract candidate lane markers in near region by using multiscale convolution kernel and Hough Transform at first. Then, the lane markers in far region can be detected based on the result of near region. At last, different vanishing points are extracted in near and far regions. With the help of a vanishing point based on camera model, we can project both of near and far regions into road space. The experiment is conducted on our self-driving car `TuLian' in campus environment.
Jian Zhou 0011, Jinsheng Xiao, Weicheng Zeng
Intelligent Vehicles Symposium3
2017 Image Noise Estimation Based on Principal Component Analysis and Variance-Stabilizing Transformation
Huying Zhang, Jinsheng Xiao, Jian Zhou 0011
ICIG (3)5
2014 An approach to speed up RRT
abstract
We present an improved algorithm to RRT∗in this paper. Our algorithm tries to increase the efficiency by replacing the Initialtree() of RRT∗with a RRT tree only contains one solution to speed up finding a feasible solution, and then applying a RRT∗_S algorithm to optimize the current solution. This RRT∗_S is similar with RRT∗in principle, but it does not insert extra nodes into the current tree and optimizes just in fixed nodes instead of all the nearby nodes. We demonstrate the performance of this improved algorithm on two benchmark scenarios to show the effectiveness compared with RRT∗and RRT.
Yun-xiao Shan, Jian Zhou 0011
Intelligent Vehicles Symposium3