Yubo Cui

dblp:256/3141 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5302-0484ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
abstract
Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively.
Yubo Cui, Xiangru Lin, Zhiheng Li 0003, Zheng Fang 0001
AAAI2
2026 Dynamic clustering transformer for LiDAR-based 3D object detection
Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001
Pattern Recognit.1
2025 LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba
abstract
Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic occupancy prediction. However, these works usually face two main challenges: 1) Limited geometric information. Due to the lack of geometric information in the image itself, it is challenging to directly predict 3D space information, especially in large-scale outdoor scenes. 2) Local restricted interaction. Due to the quadratic complexity of the attention mechanism, they often use modified local attention to fuse features, resulting in a restricted fusion. To address these problems, in this paper, we propose a language-assisted 3D semantic occupancy prediction network, named LOMA. In the proposed vision-language framework, we first introduce a VL-aware Scene Generator (VSG) module to generate the 3D language feature of the scene. By leveraging the vision-language model, this module provides implicit geometric knowledge and explicit semantic information from the language. Furthermore, we present a Tri-plane Fusion Mamba (TFM) block to efficiently fuse the 3D language feature and 3D vision feature. The proposed module not only fuses the two features with global modeling but also avoids too much computation costs. Experiments on the SemanticKITTI and SSCBench-KITTI360 datasets show that our algorithm achieves new state-of-the-art performances in both geometric and semantic completion tasks. Our code will be open soon.
Yubo Cui, Zhiheng Li 0003, Jiaqiang Wang, Zheng Fang 0001
AAAI1
2025 CAO-RONet: A Robust 4D Radar Odometry with Exploring More Information from Low-Quality Points
abstract
Recently, 4D millimetre-wave radar exhibits more stable perception ability than LiDAR and camera under adverse conditions (e.g. rain and fog). However, low-quality radar points hinder its application, especially the odometry task that requires a dense and accurate matching. To fully explore the potential of 4D radar, we introduce a learning-based odometry framework, enabling robust ego-motion estimation from finite and uncertain geometry information. First, for sparse radar points, we propose a local completion to supplement missing structures and provide denser guideline for aligning two frames. Then, a context-aware association with a hierarchical structure flexibly matches points of different scales aided by feature similarity, and improves local matching consistency through correlation balancing. Finally, we present a window-based optimizer that uses historical priors to establish a coupling state estimation and correct errors of inter-frame matching. The superiority of our algorithm is confirmed on View-of-Delft dataset, achieving around a 50% performance improvement over previous approaches and delivering accuracy on par with LiDAR odometry. The code will be released at https://github.com/NEU-REAL/CAO-RONet.
Zhiheng Li 0003, Yubo Cui, Ningyuan Huang, Chenglin Pang, Zheng Fang 0001
ICRA2
2025 Coupling and Decoupling: Towards Temporal Feedback for 3D Object Detection
abstract
3D object detection has garnered significant attention within the academic community, primarily due to its broad utility in domains such as autonomous driving and robotics. Prior research efforts have predominantly concentrated on leveraging temporal contextual information embedded within sequential data to enhance the current feature representations. However, a notable limitation of these endeavors lies in their inadequate treatment of the inherent noise present within historical sequences, thereby constraining the efficiency of fusion methods. In this paper, we propose a new temporal feedback network, named TFNet, to model and correct the temporal noise by designing acoupling-decouplingmechanism. Central to our approach are two distinct modules: (i) Foreground Feature Enhancement, which amplifies sparse instance details across temporal frames, thereby furnishing essential local information priors for subsequent fusion; and (ii) Coupling-Decoupling Feature Interaction, designed to first aggregate temporal contextual information and then disentangle fusion features into frame-specific representations. Leveraging a feedback strategy, this module can adaptively enhance useful information and eliminate noise within individual frame features. Empirical evaluations conducted on the nuScenes benchmark demonstrate the effectiveness of TFNet, achieving the new state-of-the-art performance without any bells and whistles.
Yubo Cui, Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Zhiheng Li 0003, Zheng Fang 0001
IEEE Trans. Multim.1
2024 SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Tracking
abstract
3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over a series of frames, which would cause performance degradation in the scenes with sparse points. To break through this limitation, we introduce "Sequence-to-Sequence" tracking paradigm and a tracker named SeqTrack3D to capture target motion across continuous frames. Unlike previous methods that primarily adopted three strategies: matching two consecutive point clouds, predicting relative motion, or utilizing sequential point clouds to address feature degradation, our SeqTrack3D combines both historical point clouds and bounding box sequences. This novel method ensures robust tracking by leveraging location priors from historical boxes, even in scenes with sparse points. Extensive experiments conducted on large-scale datasets show that SeqTrack3D achieves new state-of-the-art performances, improving by 6.00% on NuScenes and 14.13% on Waymo dataset. The code will be made public at https://github.com/aron-lin/seqtrack3d.
Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001
ICRA3
2024 FlowTrack: Point-level Flow Network for 3D Single Object Tracking
abstract
3D single object tracking (SOT) is a crucial task in fields of mobile robotics and autonomous driving. Traditional motion-based approaches achieve target tracking by estimating the relative movement of target between two consecutive frames. However, they usually overlook local motion information of the target and fail to exploit historical frame information effectively. To overcome the above limitations, we propose a point-level flow method with multi-frame information for 3D SOT task, called FlowTrack. Specifically, by estimating the flow for each point in the target, our method could capture the local motion details of target, thereby improving the tracking performance. Meanwhile, to handle scenes with sparse points, we present a learnable target feature as the bridge to efficiently integrate target information from past frames. Moreover, we design a Instance Flow Head to transform dense point-level flow into instance-level motion, effectively aggregating local motion information to obtain global target motion. Finally, our method achieves competitive performance with improvements of 5.9% on the KITTI and 2.9% on the NuScenes, compared to the next best method.
Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001
IROS2
2024 Intersection Is Also Needed: A Novel LiDAR-Based Road Intersection Dataset and Detection Method
abstract
3D object detection is crucial for autonomous driving. However, most existing methods focus on the foreground objects, such as vehicles and pedestrians, while ignoring some important background objects for traffic scene understanding, especially road intersections. Moreover, existing datasets (e.g., KITTI, Waymo) do not provide the labels for intersections, and the evaluation metric is also unsuitable for intersection detection. To address the above issues, we first present a LiDAR-based intersection dataset on the basis of KITTI dataset, calledKITTI-Intersection Dataset. The new dataset includes 4718 frames with 5178 instances belonging to Forkroad and Crossroad, respectively. To weaken the impact of uncertain intersection size on the performance evaluation, we introduce CEIOU instead of IOU as a new evaluation metric. Then, we proposeMInsectDetandMMInsectDet, two LiDAR-based detection methods, to solve the intersection detection problem. We start with a lightweight BEV backbone to alleviate the influence of numerous dynamic foreground objects at the intersection and obtain discriminative features. After that, to obtain more abundant and complete intersection features, we propose a Multi-Representation Backbone that integrates the BEV and voxel features to achieve better detection performance. Furthermore, in order to better adapt to various appearances and sizes of intersection, we propose a Class-Aware MultiHead, which classifies and regresses different categories with specific head. Finally, we evaluate our MInsectDet and MMInsectDet methods on the proposed KITTI-Intersection Dataset with the state-of-the-art foreground 3D detection methods. The results show that MMInsectDet achieves the best performance, and MInsectDet ranks second but could run at 65.0 FPS.
Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001
IEEE Trans. Intell. Transp. Syst.2
2023 Real-Time 3D Single Object Tracking With Transformer
abstract
LiDAR-based 3D single object tracking is a challenging issue in robotics and autonomous driving. Currently, existing approaches usually suffer from the problem that objects at long distance often have very sparse or partially-occluded point clouds, which makes the features extracted by the model ambiguous. Ambiguous features will make it hard to locate the target object and finally lead to bad tracking results. To solve this problem, we utilize the powerful Transformer architecture and propose aPoint-Track-Transformer (PTT)module for point cloud-based 3D single object tracking task. Specifically, PTT module generates fine-tuned attention features by computing attention weights, which guides the tracker focusing on the important features of the target and improves the tracking ability in complex scenarios. To evaluate our PTT module, we embed PTT into the dominant method and construct a novel 3D SOT tracker named PTT-Net. In PTT-Net, we embed PTT into the voting stage and proposal generation stage, respectively. PTT module in the voting stage could model the interactions among point patches, which learns context-dependent features. Meanwhile, PTT module in the proposal generation stage could capture the contextual information between object and background. We evaluate our PTT-Net on KITTI and NuScenes datasets. Experimental results demonstrate the effectiveness of PTT module and the superiority of PTT-Net, which surpasses the baseline by a noticeable margin,$\sim$10% in the Car category. Meanwhile, our method also has a significant performance improvement in sparse scenarios. In general, the combination of transformer and tracking pipeline enables our PTT-Net to achieve state-of-the-art performance on both two datasets. Additionally, PTT-Net could run in real-time at 40FPS on NVIDIA 1080Ti GPU. Our code is open-sourced for the research community athttps://github.com/shanjiayao/PTT.
Jiayao Shan, Sifan Zhou, Yubo Cui, Zheng Fang 0001
IEEE Trans. Multim.3
2021 3D Object Tracking with Transformer
Yubo Cui, Zheng Fang 0001, Jiayao Shan, Zuoxu Gu, Sifan Zhou
BMVC1
2021 PTT: Point-Track-Transformer Module for 3D Single Object Tracking in Point Clouds
abstract
3D single object tracking is a key issue for robotics. In this paper, we propose a transformer module called Point-Track-Transformer (PTT) for point cloud-based 3D single object tracking. PTT module contains three blocks for feature embedding, position encoding, and self-attention feature computation. Feature embedding aims to place features closer in the embedding space if they have similar semantic information. Position encoding is used to encode coordinates of point clouds into high dimension distinguishable features. Self-attention generates refined attention features by computing attention weights. Besides, we embed the PTT module into the open-source state-of-the-art method P2B to construct PTT-Net. Experiments on the KITTI dataset reveal that our PTT-Net surpasses the state-of-the-art by a noticeable margin $\left( {\sim 10\% } \right)$. Additionally, PTT-Net could achieve real-time performance (~40FPS) on NVIDIA 1080Ti GPU. Our code is open-sourced for the robotics community at https://github.com/shanjiayao/PTT.
Jiayao Shan, Sifan Zhou, Zheng Fang 0001, Yubo Cui
IROS4