Zhenzhong Cao

dblp:194/4259 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6038-4147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HTMNet: A hybrid transformer-mamba network for LiDAR-based 3D detection and semantic segmentation
Jinzheng Guang, Yongru Wang, Zhenzhong Cao, Jingtai Liu
Expert Syst. Appl.4
2026 SLVD-CC: Enhancing statement-level vulnerability detection via context clarification
Yinglong Chen, Zhenzhong Cao, Guangfa Lyu
Sci. Comput. Program.2
2026 PipeCLIP: Defect-Conditioned and Cross-Focus-Driven Vision-Language Model for Video-Based Sewer Defect Inspection
abstract
In recent years, vision-language models (VLMs), such as CLIP, have excelled in the visual domain due to the availability of vast paired image-text data. However, directly adapting CLIP to the sewer defect inspection achieves a poor performance, since the vocabularies of defect category are highly “unfamiliar” to CLIP, resulting in the weak representations among the defect categories in the text embedding space. Besides, the substantial differences of pipe characteristics are challenging to align the multiple pairs of visual and text features across multi-focus segments. We propose PipeCLIP for adapting the text information to the video-based multi-label sewer defect classification, which is the first method to integrate a CLIP-based model into the sewer defect inspection. First, expert prior descriptions (EPD) are introduced to differentiate the distinctions between defect category in the text embedding space. Second, defect-attribute coupling (DAC) prompt is proposed to strengthen the coupling relationship between the sewer defects and pipe multi-attributes. The two prompts proposed above together with the traditional category prompt, collectively constitute defect-conditioned text prompt (DecTP) for CLIP. Then, cross-focus temporal (CFT) module is designed to integrate feature information from different focal length, strengthening the visual-text alignment across multi-focus segments. Extensive experiments are conducted on the public benchmark, in which the superiority of PipeCLIP is demonstrated compared with the state-of-the-art methods. Code is available at: https://anonymous.4open.science/r/PipeCLIP-0925.
Chenyang Zhao 0009, Chuanfei Hu, Zhenzhong Cao, Yinuo Song, Jingtai Liu
IEEE Trans Autom. Sci. Eng.3
2025 ELPTNet: An Efficient LiDAR-based 3D Pedestrian Tracking Network for Autonomous Navigation Social Robots
abstract
Autonomous navigation social robots need to track pedestrian movements in real-time with high precision to optimize path planning and avoid collisions. However, the main challenge of pedestrian tracking lies in the significant variations in human posture, which differ from rigid-body structures like vehicles. In this paper, we propose an Efficient LiDAR-based 3D Pedestrian Tracking Network (ELPTNet). First, our ELPTNet employs a 3D object detector to extract directional 3D pedestrian bounding boxes from LiDAR point clouds. Then, our ELPTNet employs a Constant Acceleration (CA) model and prediction confidence for target trajectory prediction. During the data association process, it integrates geometric, appearance, and motion features to enhance the robustness and real-time performance of 3D MOT when targets are temporarily occluded. Experimental results demonstrate that our ELPTNet achieves the highest ranking on the large-scale JRDB dataset for the 3D tracking task, outperforming previous state-of-the-art (SOTA) methods with improvements of 8.4% in MOTA and 6.6% in HOTA. Additionally, our ELPTNet attains an inference speed of 61 frames per second (FPS) on a single CPU. Therefore, our method enables accurate and real-time tracking of multiple pedestrians. The code is publicly available at https://github.com/jinzhengguang/ELPTNet.
Jinzheng Guang, Zhenzhong Cao, Yinuo Song, Jingtai Liu
IROS2
2025 DYO-SLAM: Visual Localization and Object Mapping in Dynamic Scenes
abstract
Addressing the impact of dynamic factors on localization accuracy and constructing a long-term consistent map containing only static elements are two crucial tasks in visual simultaneous localization and mapping (SLAM) for dynamic scenes. The introduction of dynamic elements can compromise the geometric constraints essential for visual SLAM, leading to a decrease in localization accuracy. Existing related research faces challenges in simultaneously ensuring localization accuracy in both low-dynamic and high-dynamic scenarios, while also maintaining the system’s real-time performance. To address this issue, we propose a two-stage, coarse-to-fine static-probability-based localization scheme. The construction of object-level maps offers strong support for tasks involving higher-level intelligent agent manipulation as well as augmented reality (AR). However, current research is inadequate for dynamic scenes where the objects to be modeled are frequently and irregularly obscured by dynamic objects, and where there are significant challenges such as severe image and point cloud noise, semantic noise, and lack of observational perspectives. To overcome these challenges, we first propose an object parameter estimation algorithm that combines clustering, weighted Principal Component Analysis (PCA) based on an energy function, and a minimum bounding rectangle. Then, we design a multi-modal object data association strategy based on appearance, semantic, and spatial features. The proposed object parameter estimation algorithm and data association strategy demonstrate improved accuracy and robustness in dynamic scenes with the aforementioned challenges. Finally, based on the entire system, we further develop a dynamic object tracking algorithm and construct an AR system to demonstrate the system’s application prospects. A series of public datasets and real-world scene results have been used to evaluate the effectiveness of the proposed system.
Xinggang Hu, Yanmin Wu, Zhenzhong Cao, Xiangkui Zhang, Xiangyang Ji
IEEE Trans. Circuits Syst. Video Technol.4
2023 Object SLAM With Robust Quadric Initialization and Mapping for Dynamic Outdoors
abstract
Object SLAM is a popular approach for autonomous driving and robotics, but accurate object perception in outdoor environments remains a challenge. State-of-the-art object SLAM algorithms rely on assumptions and are sensitive to observation noise, limiting their application in real-world scenarios. To address these challenges, we propose a novel object SLAM system that utilizes a quadric initialization algorithm based on constrained quadric optimization, which does not rely on planar assumptions and is robust to partial observations. Additionally, we introduce an automatic object data association algorithm capable of detecting motion states while associating objects across frames. To further enhance the accuracy of the quadric mapping, an extra thread is used to refine the ellipsoid parameters within a local sliding window composed of keyframes. Our system utilizes a joint optimization framework that optimizes camera poses, object landmarks, and point clouds in the local mapping thread for further global optimization while maintaining a consistent map. Experimental results on the real-world KITTI dataset show that the proposed system is more robust and significantly outperforms current state-of-the-art methods in quadric initialization and mapping in outdoor scenarios. Moreover, our system achieves real-time performance, making it suitable for practical applications.
Rui Tian 0002, Yunzhou Zhang, Zhenzhong Cao, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.3
2022 SemLoc: Accurate and Robust Visual Localization with Semantic and Structural Constraints from Prior Maps
abstract
Semantic information and geometrical structures of a prior map can be leveraged in visual localization to bound drift errors and improve accuracy. In this paper, we propose SemLoc, a pure visual localization system, for accurate localization in a prior semantic map. To tightly couple semantic and structure information from prior maps, a hybrid constraint is presented by using the Dirichlet distribution. Then, with the local landmarks and their semantic states tracked in the frontend, the camera poses and data associations are jointly optimized through Expectation-Maximization (EM) algorithm. We validate the effectiveness of our approach in both monocular and stereo modes on the public KITTI dataset. Experimental results demonstrate that our system can greatly reduce drift errors with an satisfying real-time performance. Compared with several state-of-the-art visual localization systems, the proposed framework achieves a competitive localization performance.
Shiwen Liang, Yunzhou Zhang, Rui Tian 0002, Delong Zhu 0001, Linghao Yang, Zhenzhong Cao
ICRA6
2022 CFP-SLAM: A Real-time Visual SLAM Based on Coarse-to-Fine Probability in Dynamic Environments
abstract
The dynamic factors in the environment will lead to the decline of camera localization accuracy due to the violation of the static environment assumption of SLAM algorithm. Recently, some related works generally use the combination of semantic constraints and geometric constraints to deal with dynamic objects, but problems can still be raised, such as poor real-time performance, easy to treat people as rigid bodies, and poor performance in low dynamic scenes. In this paper, a dynamic scene-oriented visual SLAM algorithm based on object detection and coarse-to-fine static probability named CFP-SLAM is proposed. The algorithm combines semantic constraints and geometric constraints to calculate the static probability of objects, keypoints and map points, and takes them as weights to participate in camera pose estimation. Extensive evaluations show that our approach can achieve almost the best results in high dynamic and low dynamic scenarios compared to the state-of-the-art dynamic SLAM methods, and shows quite high real-time ability.
Xinggang Hu, Yunzhou Zhang, Zhenzhong Cao, Yanmin Wu, Zhiqiang Deng, Wenkai Sun
IROS3
2021 SELF: A method of searching for library functions in stripped binary code
Xueqian Liu, Shoufeng Cao, Zhenzhong Cao, Qu Gao
Comput. Secur.3
2017 Sink location privacy protection under direction attack in wireless sensor networks
Zhenzhong Cao, Fengbo Lin
Wirel. Networks3