EDBT 2026 Demo / reviewers in the wild / expert
Zhiheng Li 0003
dblp:89/6935-3
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-1477-2066ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards 3D Object-Centric Feature Learning for Semantic Scene CompletionabstractVision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively. Yubo Cui, Xiangru Lin, Zhiheng Li 0003, Zheng Fang 0001 |
AAAI | 4 |
| 2026 | Dynamic clustering transformer for LiDAR-based 3D object detection
Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001 |
Pattern Recognit. | 2 |
| 2025 | LOMA: Language-assisted Semantic Occupancy Network via Triplane MambaabstractVision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic occupancy prediction. However, these works usually face two main challenges: 1) Limited geometric information. Due to the lack of geometric information in the image itself, it is challenging to directly predict 3D space information, especially in large-scale outdoor scenes. 2) Local restricted interaction. Due to the quadratic complexity of the attention mechanism, they often use modified local attention to fuse features, resulting in a restricted fusion. To address these problems, in this paper, we propose a language-assisted 3D semantic occupancy prediction network, named LOMA. In the proposed vision-language framework, we first introduce a VL-aware Scene Generator (VSG) module to generate the 3D language feature of the scene. By leveraging the vision-language model, this module provides implicit geometric knowledge and explicit semantic information from the language. Furthermore, we present a Tri-plane Fusion Mamba (TFM) block to efficiently fuse the 3D language feature and 3D vision feature. The proposed module not only fuses the two features with global modeling but also avoids too much computation costs. Experiments on the SemanticKITTI and SSCBench-KITTI360 datasets show that our algorithm achieves new state-of-the-art performances in both geometric and semantic completion tasks. Our code will be open soon. Yubo Cui, Zhiheng Li 0003, Jiaqiang Wang, Zheng Fang 0001 |
AAAI | 2 |
| 2025 | CAO-RONet: A Robust 4D Radar Odometry with Exploring More Information from Low-Quality PointsabstractRecently, 4D millimetre-wave radar exhibits more stable perception ability than LiDAR and camera under adverse conditions (e.g. rain and fog). However, low-quality radar points hinder its application, especially the odometry task that requires a dense and accurate matching. To fully explore the potential of 4D radar, we introduce a learning-based odometry framework, enabling robust ego-motion estimation from finite and uncertain geometry information. First, for sparse radar points, we propose a local completion to supplement missing structures and provide denser guideline for aligning two frames. Then, a context-aware association with a hierarchical structure flexibly matches points of different scales aided by feature similarity, and improves local matching consistency through correlation balancing. Finally, we present a window-based optimizer that uses historical priors to establish a coupling state estimation and correct errors of inter-frame matching. The superiority of our algorithm is confirmed on View-of-Delft dataset, achieving around a 50% performance improvement over previous approaches and delivering accuracy on par with LiDAR odometry. The code will be released at https://github.com/NEU-REAL/CAO-RONet. Zhiheng Li 0003, Yubo Cui, Ningyuan Huang, Chenglin Pang, Zheng Fang 0001 |
ICRA | 1 |
| 2025 | Target-Aware Viewpoint Generation for Active Robotic Exploration in Unknown EnvironmentsabstractWhen entering an unfamiliar environment, animals usually sweep off their surroundings to identify points of interest. In search and rescue robotics, autonomous exploration requires both coarse mapping of unknown areas and detailed target detection, which poses a significant challenge in balancing these tasks. To that end, we propose a target-aware robotic exploration framework that prioritizes both exploration efficiency and search completeness through three components: First, considering the computational limitations of robotic platforms, a lightweight 3D target detection method with post-fusion is introduced to detect target positions in real time. Secondly, we propose a target-aware viewpoint generation approach that integrates information gain and inspection gain to identify promising viewpoints for thorough target searches. Lastly, since a detailed examination of the environment demands numerous viewpoints, we propose a heuristic-based active exploration framework that employs a hierarchical structure to optimize exploration gain, traveling distance, and path smoothness to maximize the utility function of viewpoint sequences and ultimately find the optimal path. Extensive simulations and real-world experiments demonstrate our framework significantly enhances target search capabilities, achieving a 13 % average improvement in exploration efficiency over existing methods. Pu Xu, Zhiheng Li 0003, Zhaoqiang Bai, Zheng Fang 0001 |
ICRA | 3 |
| 2025 | RDN: An Efficient Denoising Network for 4D Radar Point CloudsabstractAccurate point cloud information is important for robot perception and autonomous driving. Although advanced 4D radar can provide point cloud with higher resolution than 3D radar, its data still contains a significant amount of noise due to measurement principle. To solve this issue, we propose RDN (Radar Denoising Network), a denoising network specifically designed for 4D radar. RDN includes three innovative modules: First, to overcome the noisy nature of radar points, we design a feature similarity-based farthest point sampling module (FS-FPS), which can extract representative sampling points from the noisy point cloud. Secondly, to address feature propagation issues caused by the sparse and long-range characteristics of 4D radar points, we introduce a virtual feature point prediction (VFP) module and an iterative upsampling (IUS) module. The VFP module generates virtual feature points through the network to serve as bridges for information transmission, while the IUS module uses an iterative approach to gradually refine feature propagation. The experiments on MSC-RAD4D and NTU4DRadLM datasets demonstrate the effectiveness and generalization of our method. Besides, odometry experiments prove the practical value of point cloud denoising in improving robot perception. Ningyuan Huang, Zhiheng Li 0003, Chenglin Pang, Zheng Fang 0001 |
IROS | 2 |
| 2025 | Coupling and Decoupling: Towards Temporal Feedback for 3D Object Detectionabstract3D object detection has garnered significant attention within the academic community, primarily due to its broad utility in domains such as autonomous driving and robotics. Prior research efforts have predominantly concentrated on leveraging temporal contextual information embedded within sequential data to enhance the current feature representations. However, a notable limitation of these endeavors lies in their inadequate treatment of the inherent noise present within historical sequences, thereby constraining the efficiency of fusion methods. In this paper, we propose a new temporal feedback network, named TFNet, to model and correct the temporal noise by designing acoupling-decouplingmechanism. Central to our approach are two distinct modules: (i) Foreground Feature Enhancement, which amplifies sparse instance details across temporal frames, thereby furnishing essential local information priors for subsequent fusion; and (ii) Coupling-Decoupling Feature Interaction, designed to first aggregate temporal contextual information and then disentangle fusion features into frame-specific representations. Leveraging a feedback strategy, this module can adaptively enhance useful information and eliminate noise within individual frame features. Empirical evaluations conducted on the nuScenes benchmark demonstrate the effectiveness of TFNet, achieving the new state-of-the-art performance without any bells and whistles. Yubo Cui, Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Zhiheng Li 0003, Zheng Fang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Trackingabstract3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over a series of frames, which would cause performance degradation in the scenes with sparse points. To break through this limitation, we introduce "Sequence-to-Sequence" tracking paradigm and a tracker named SeqTrack3D to capture target motion across continuous frames. Unlike previous methods that primarily adopted three strategies: matching two consecutive point clouds, predicting relative motion, or utilizing sequential point clouds to address feature degradation, our SeqTrack3D combines both historical point clouds and bounding box sequences. This novel method ensures robust tracking by leveraging location priors from historical boxes, even in scenes with sparse points. Extensive experiments conducted on large-scale datasets show that SeqTrack3D achieves new state-of-the-art performances, improving by 6.00% on NuScenes and 14.13% on Waymo dataset. The code will be made public at https://github.com/aron-lin/seqtrack3d. Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001 |
ICRA | 2 |
| 2024 | FlowTrack: Point-level Flow Network for 3D Single Object Trackingabstract3D single object tracking (SOT) is a crucial task in fields of mobile robotics and autonomous driving. Traditional motion-based approaches achieve target tracking by estimating the relative movement of target between two consecutive frames. However, they usually overlook local motion information of the target and fail to exploit historical frame information effectively. To overcome the above limitations, we propose a point-level flow method with multi-frame information for 3D SOT task, called FlowTrack. Specifically, by estimating the flow for each point in the target, our method could capture the local motion details of target, thereby improving the tracking performance. Meanwhile, to handle scenes with sparse points, we present a learnable target feature as the bridge to efficiently integrate target information from past frames. Moreover, we design a Instance Flow Head to transform dense point-level flow into instance-level motion, effectively aggregating local motion information to obtain global target motion. Finally, our method achieves competitive performance with improvements of 5.9% on the KITTI and 2.9% on the NuScenes, compared to the next best method. Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001 |
IROS | 3 |
| 2024 | Intersection Is Also Needed: A Novel LiDAR-Based Road Intersection Dataset and Detection Methodabstract3D object detection is crucial for autonomous driving. However, most existing methods focus on the foreground objects, such as vehicles and pedestrians, while ignoring some important background objects for traffic scene understanding, especially road intersections. Moreover, existing datasets (e.g., KITTI, Waymo) do not provide the labels for intersections, and the evaluation metric is also unsuitable for intersection detection. To address the above issues, we first present a LiDAR-based intersection dataset on the basis of KITTI dataset, calledKITTI-Intersection Dataset. The new dataset includes 4718 frames with 5178 instances belonging to Forkroad and Crossroad, respectively. To weaken the impact of uncertain intersection size on the performance evaluation, we introduce CEIOU instead of IOU as a new evaluation metric. Then, we proposeMInsectDetandMMInsectDet, two LiDAR-based detection methods, to solve the intersection detection problem. We start with a lightweight BEV backbone to alleviate the influence of numerous dynamic foreground objects at the intersection and obtain discriminative features. After that, to obtain more abundant and complete intersection features, we propose a Multi-Representation Backbone that integrates the BEV and voxel features to achieve better detection performance. Furthermore, in order to better adapt to various appearances and sizes of intersection, we propose a Class-Aware MultiHead, which classifies and regresses different categories with specific head. Finally, we evaluate our MInsectDet and MMInsectDet methods on the proposed KITTI-Intersection Dataset with the state-of-the-art foreground 3D detection methods. The results show that MMInsectDet achieves the best performance, and MInsectDet ranks second but could run at 65.0 FPS. Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |