VLDB 2026 Research / reviewers in the wild / expert
Zhe Liu 0033
dblp:70/1220-33
· DBLP profile ↗
19ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0002-2513-2662ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object Retrieval
Yang Zhou 0007, Zhe Liu 0033, Rui Yu 0002, Song Bai 0001, Yulong Wang 0002, Xinwei He 0001, Xiang Bai |
ICCV | 3 |
| 2025 | Hybrid Transformer-Mamba Model for 3D Semantic SegmentationabstractTransformer-based methods have demonstrated remarkable capabilities in 3D semantic segmentation through their powerful attention mechanisms, but the quadratic complexity limits their modeling of long-range dependencies in large-scale point clouds. While recent Mamba-based approaches offer efficient processing with linear complexity, they struggle with feature representation when extracting 3D features. However, effectively combining these complementary strengths remains an open challenge in this field. In this paper, we propose HybridTM, the first hybrid architecture that integrates Transformer and Mamba for 3D semantic segmentation. In addition, we propose the Inner Layer Hybrid Strategy, which combines attention and Mamba at a finer granularity, enabling simultaneous capture of long-range dependencies and fine-grained local features. Extensive experiments demonstrate the effectiveness and generalization of our HybridTM on diverse indoor and outdoor datasets. Furthermore, our HybridTM achieves state-of-the-art performance on ScanNet, ScanNet200, and nuScenes benchmarks. The code will be made available at https://github.com/deepinact/HybridTM. Xinyu Wang 0024, Jinghua Hou, Zhe Liu 0033, Yingying Zhu 0005 |
IROS | 3 |
| 2025 | AVS-Net: Point sampling with adaptive voxel size for 3D scene understanding
Hongcheng Yang, Dingkang Liang, Dingyuan Zhang, Zhe Liu 0033, Zhikang Zou, Xingyu Jiang 0005, Yingying Zhu 0005 |
Neurocomputing | 4 |
| 2025 | An Empirical Study of Ground Segmentation for 3-D Object DetectionabstractThe ratio of foreground and background points directly impacts the accuracy and speed of the lidar-based 3D object detection methods. However, existing methods generally ignore the impact of ground points. Although some traditional ground segmentation algorithms are available to remove ground point clouds, they usually suffer from over-segmentation, which leads to a sub-optimal and even worse performance for the downstream 3D detection task. We conduct an in-depth analysis and attribute this phenomenon to the reason that some crucial foreground points attached to the ground (e.g., the wheels of Cars, or the feet of Pedestrians) are directly removed due to over-segmentation. To this end, we propose a new Attached Point Restoring (APR) module to recover these discarded foreground points. Experimental results demonstrate the effectiveness and generalization of APR by integrating it into various ground segmentation algorithms to boost the performance or the running time of 3D detection on KITTI and Waymo datasets. Finally, we hope this paper can serve as a new guide to inspire future research in this field. Code is available athttps://github.com/yhc2021/GPR. Hongcheng Yang, Dingkang Liang, Zhe Liu 0033, Zhikang Zou, Xiaoqing Ye, Xiang Bai |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | OPEN: Object-Wise Position Embedding for Multi-view 3D Object Detection
Jinghua Hou, Xiaoqing Ye, Zhe Liu 0033, Shi Gong, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai |
ECCV (26) | 4 |
| 2024 | SEED: A Simple and Effective 3D DETR in Point Clouds
Zhe Liu 0033, Jinghua Hou, Xiaoqing Ye, Jingdong Wang 0001, Xiang Bai |
ECCV (11) | 1 |
| 2024 | LION: Linear Group RNN for 3D Object Detection in Point CloudsabstractThe benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward this goal, we propose a simple and effective window-based framework built on Linear group RNN (i.e., perform linear RNN for grouped features) for accurate 3D object detection, called LION. The key property is to allow sufficient feature interaction in a much larger group than transformer-based methods. However, effectively applying linear group RNN to 3D object detection in highly sparse point clouds is not trivial due to its limitation in handling spatial modeling. To tackle this problem, we simply introduce a 3D spatial feature descriptor and integrate it into the linear group RNN operators to enhance their spatial features rather than blindly increasing the number of scanning orders for voxel features. To further address the challenge in highly sparse point clouds, we propose a 3D voxel generation strategy to densify foreground features thanks to linear group RNN as a natural property of auto-regressive models.
Extensive experiments verify the effectiveness of the proposed components and the generalization of our LION on different linear group RNN operators including Mamba, RWKV, and RetNet. Furthermore, it is worth mentioning that our LION-Mamba achieves state-of-the-art on Waymo, nuScenes, Argoverse V2, and ONCE datasets. Last but not least, our method supports kinds of advanced linear RNN operators (e.g., RetNet, RWKV, Mamba, xLSTM and TTT) on small but popular KITTI dataset for a quick experience with our linear RNN-based framework. Zhe Liu 0033, Jinghua Hou, Xinyu Wang 0024, Xiaoqing Ye, Jingdong Wang 0001, Hengshuang Zhao, Xiang Bai |
NeurIPS | 1 |
| 2024 | SAM3D: zero-shot 3D object detection via the segment anything model
Dingyuan Zhang, Dingkang Liang, Hongcheng Yang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiang Bai |
Sci. China Inf. Sci. | 6 |
| 2024 | Multi-Modal 3D Object Detection by Box MatchingabstractMulti-modal 3D object detection has received growing attention as the information from different sensors like LiDAR and cameras are complementary. Most fusion methods for 3D detection rely on an accurate alignment and calibration between 3D point clouds and RGB images. However, such an assumption is not reliable in a real-world self-driving system, as the alignment between different modalities is easily affected by asynchronous sensors and disturbed sensor placement. We propose a novel Fusion network by Box Matching (FBMNet) for multi-modal 3D detection, which provides an alternative way for cross-modal feature alignment by learning the correspondence at the bounding box level to free up the dependency of calibration during inference. With the learned assignments between 3D and 2D object proposals, the fusion for detection can be effectively performed by combining their ROI features. Extensive experiments on the nuScenes dataset demonstrate that our method is much more robust in dealing with challenging cases such as asynchronous sensors, misaligned sensor placement, and degenerated camera images than existing fusion methods. We hope that our FBMNet could provide an available solution to dealing with these challenging cases for safety in real autonomous driving scenarios. Zhe Liu 0033, Xiaoqing Ye, Zhikang Zou, Xinwei He 0001, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object DetectionabstractIn this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. The key designs of StereoDistill are: the X-component Guided Distillation~(XGD) for regression and the Cross-anchor Logit Distillation~(CLD) for classification. In XGD, instead of empirically adopting a threshold to select the high-quality teacher predictions as soft targets, we decompose the predicted 3D box into sub-components and retain the corresponding part for distillation if the teacher component pilot is consistent with ground truth to largely boost the number of positive predictions and alleviate the mimicking difficulty of the student model. For CLD, we aggregate the probability distribution of all anchors at the same position to encourage the highest probability anchor rather than individually distill the distribution at the anchor level. Finally, our StereoDistill achieves state-of-the-art results for stereo-based 3D detection on the KITTI test benchmark and extensive experiments on KITTI and Argoverse Dataset validate the effectiveness. Zhe Liu 0033, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Xiang Bai |
AAAI | 1 |
| 2023 | A Simple Vision Transformer for Weakly Semi-supervised 3D Object DetectionabstractAdvanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can effectively reduce the annotation cost, they suffer from two limitations: 1) their performance is far inferior to the fully-supervised counterparts; 2) they are difficult to adapt to different detectors or scenes (e.g, indoor or outdoor). In this paper, we study weakly semi-supervised 3D object detection (WSS3D) with point annotations, where the dataset comprises a small number of fully labeled and massive weakly labeled data with a single point annotated for each 3D object. To fully exploit the point annotations, we employ the plain and non-hierarchical vision transformer to form a point-to-box converter, termed ViT-WSS3D. By modeling global interactions between LiDAR points and corresponding weak labels, our ViT-WSS3D can generate high-quality pseudo-bounding boxes, which are then used to train any 3D detectors without exhaustive tuning. Extensive experiments on indoor and outdoor datasets (SUN RGBD and KITTI) show the effectiveness of our method. In particular, when only using 10% fully labeled and the rest as point labeled data, our ViT-WSS3D can enable most detectors to achieve similar performance with the oracle model using 100% fully labeled data. Dingyuan Zhang, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiao Tan 0001, Xiang Bai |
ICCV | 6 |
| 2023 | DDS3D: Dense Pseudo-Labels with Dynamic Threshold for Semi-Supervised 3D Object DetectionabstractIn this paper, we present a simple yet effective semi-supervised 3D object detector named DDS3D. Our main contributions have two-fold. On the one hand, different from previous works using Non-Maximal Suppression (NMS) or its variants for obtaining the sparse pseudo labels, we propose a dense pseudo-label generation strategy to get dense pseudo-labels, which can retain more potential supervision information for the student network. On the other hand, instead of traditional fixed thresholds, we propose a dynamic threshold manner to generate pseudo-labels, which can guarantee the quality and quantity of pseudo-labels during the whole training process. Benefiting from these two components, our DDS3D outperforms the state-of-the-art semi-supervised 3d object detection with mAP of 3.1% on the pedestrian and 2.1% on the cyclist under the same configuration of 1% samples. Extensive ablation studies on the KITTI dataset demonstrate the effectiveness of our DDS3D. The code and models will be made publicly available at https://github.com/hust-jy/DDS3D Zhe Liu 0033, Jinghua Hou, Dingkang Liang |
ICRA | 2 |
| 2023 | Query-based Temporal Fusion with Explicit Motion for 3D Object DetectionabstractEffectively utilizing temporal information to improve 3D detection performance is vital for autonomous driving vehicles. Existing methods either conduct temporal fusion based on the dense BEV features or sparse 3D proposal features. However, the former does not pay more attention to foreground objects, leading to more computation costs and sub-optimal performance. The latter implements time-consuming operations to generate sparse 3D proposal features, and the performance is limited by the quality of 3D proposals. In this paper, we propose a simple and effective Query-based Temporal Fusion Network (QTNet). The main idea is to exploit the object queries in previous frames to enhance the representation of current object queries by the proposed Motion-guided Temporal Modeling (MTM) module, which utilizes the spatial position information of object queries along the temporal dimension to construct their relevance between adjacent frames reliably. Experimental results show our proposed QTNet outperforms BEV-based or proposal-based manners on the nuScenes dataset. Besides, the MTM is a plug-and-play module, which can be integrated into some advanced LiDAR-only or multi-modality 3D detectors and even brings new SOTA performance with negligible computation cost and latency on the nuScenes dataset. These experiments powerfully illustrate the superiority and generalization of our method. The code is available at https://github.com/AlmoonYsl/QTNet. Jinghua Hou, Zhe Liu 0033, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Xiang Bai |
NeurIPS | 2 |
| 2023 | Diffusion-Based 3D Object Detection with Random Boxes
Xin Zhou 0013, Jinghua Hou, Dingkang Liang, Zhe Liu 0033, Zhikang Zou, Xiaoqing Ye, Jianwei Cheng, Xiang Bai |
PRCV (2) | 5 |
| 2023 | EPNet++: Cascade Bi-Directional Fusion for Multi-Modal 3D Object DetectionabstractRecently, fusing the LiDAR point cloud and camera image to improve the performance and robustness of 3D object detection has received more and more attention, as these two modalities naturally possess strong complementarity. In this paper, we propose EPNet++ for multi-modal 3D object detection by introducing a novel Cascade Bi-directional Fusion (CB-Fusion) module and a Multi-Modal Consistency (MC) loss. More concretely, the proposed CB-Fusion module enhances point features with plentiful semantic information absorbed from the image features in a cascade bi-directional interaction fusion manner, leading to more powerful and discriminative feature representations. The MC loss explicitly guarantees the consistency between predicted scores from two modalities to obtain more comprehensive and reliable confidence scores. The experimental results on the KITTI, JRDB and SUN-RGBD datasets demonstrate the superiority of EPNet++ over the state-of-the-art methods. Besides, we emphasize a critical but easily overlooked problem, which is to explore the performance and robustness of a 3D detector in a sparser scene. Extensive experiments present that EPNet++ outperforms the existing SOTA methods with remarkable margins in highly sparse point cloud cases, which might be an available direction to reduce the expensive cost of LiDAR sensors. Zhe Liu 0033, Tengteng Huang, Bingling Li, Xiwu Chen, Xi Wang 0044, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | PRA-Net: Point Relation-Aware Network for 3D Point Cloud AnalysisabstractLearning intra-region contexts and inter-region relations are two effective strategies to strengthen feature representations for point cloud analysis. However, unifying the two strategies for point cloud representation is not fully emphasized in existing methods. To this end, we propose a novel framework named Point Relation-Aware Network (PRA-Net), which is composed of an Intra-region Structure Learning (ISL) module and an Inter-region Relation Learning (IRL) module. The ISL module can dynamically integrate the local structural information into the point features, while the IRL module captures inter-region relations adaptively and efficiently via a differentiable region partition scheme and a representative point-based strategy. Extensive experiments on several 3D benchmarks covering shape classification, keypoint estimation, and part segmentation have verified the effectiveness and the generalization ability of PRA-Net. Code will be available at https://github.com/XiwuChen/PRA-Net. Silin Cheng 0001, Xiwu Chen, Xinwei He 0001, Zhe Liu 0033, Xiang Bai |
IEEE Trans. Image Process. | 4 |
| 2020 | TANet: Robust 3D Object Detection from Point Clouds with Triple AttentionabstractIn this paper, we focus on exploring the robustness of the 3D object detection in point clouds, which has been rarely discussed in existing approaches. We observe two crucial phenomena: 1) the detection accuracy of the hard objects, e.g., Pedestrians, is unsatisfactory, 2) when adding additional noise points, the performance of existing approaches decreases rapidly. To alleviate these problems, a novel TANet is introduced in this paper, which mainly contains a Triple Attention (TA) module, and a Coarse-to-Fine Regression (CFR) module. By considering the channel-wise, point-wise and voxel-wise attention jointly, the TA module enhances the crucial information of the target while suppresses the unstable cloud points. Besides, the novel stacked TA further exploits the multi-level feature attention. In addition, the CFR module boosts the accuracy of localization without excessive computation cost. Experimental results on the validation set of KITTI dataset demonstrate that, in the challenging noisy cases, i.e., adding additional random noisy points around each object, the presented approach goes far beyond state-of-the-art approaches. Furthermore, for the 3D object detection task of the KITTI benchmark, our approach ranks the first place on Pedestrian class, by using the point clouds as the only input. The running speed is around 29 frames per second. Zhe Liu 0033, Xin Zhao 0012, Tengteng Huang, Ruolan Hu, Yu Zhou 0016, Xiang Bai |
AAAI | 1 |
| 2020 | EPNet: Enhancing Point Features with Image Semantics for 3D Object Detection
Tengteng Huang, Zhe Liu 0033, Xiwu Chen, Xiang Bai |
ECCV (15) | 2 |
| 2019 | 3D Object Detection Using Scale Invariant and Feature Reweighting Networksabstract3D object detection plays an important role in a large number of real-world applications. It requires us to estimate the localizations and the orientations of 3D objects in real scenes. In this paper, we present a new network architecture which focuses on utilizing the front view images and frustum point clouds to generate 3D detection results. On the one hand, a PointSIFT module is utilized to improve the performance of 3D segmentation. It can capture the information from different orientations in space and the robustness to different scale shapes. On the other hand, our network obtains the useful features and suppresses the features with less information by a SENet module. This module reweights channel features and estimates the 3D bounding boxes more effectively. Our method is evaluated on both KITTI dataset for outdoor scenes and SUN-RGBD dataset for indoor scenes. The experimental results illustrate that our method achieves better performance than the state-of-the-art methods especially when point clouds are highly sparse. Xin Zhao 0012, Zhe Liu 0033, Ruolan Hu, Kaiqi Huang |
AAAI | 2 |