Haizhuang Liu

dblp:304/4786 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-5492-2831ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Implicit alignment and query refinement for RGB-T semantic segmentation
Chang Liu 0136, Haizhuang Liu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Qianchuan Zhao, Huimin Ma 0001
Pattern Recognit.2
2025 Occlusion-guided multi-modal fusion for vehicle-infrastructure cooperative 3D object detection
Huazhen Chu, Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.2
2025 SparseComm: An Efficient Sparse Communication Framework for Vehicle-Infrastructure Cooperative 3D Detection
Haizhuang Liu, Huazhen Chu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.1
2024 PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes
abstract
3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target dataset, particularly towards object sizes. State-of-the-art unsupervised domain adaptation approaches explicitly address the variation in object sizes by appropriately scaling the source data. However, such methods require additional target domain statistics information, which contradicts the original unsupervised assumption. In this work, we present PLS, a novel unsupervised domain adaptation method for 3D object detection to overcome the object sizes bias via Pseudo-Label Sizes, which utilizes only source domain annotations. PLS alternates between generating high-quality pseudo-label sizes through the detector and model training with the pseudo-label sizes to scale and augment the source data. This iterative process enables the detector to be trained with augmented data that resembles the target domain sizes, thereby improving the performance of detector in cross-domain scenarios. Our experimental results show the outstanding performance of our PLS in various scenarios. In addition, PLS is a plug-and-play module that can be used to directly replace existing weakly-supervised scaling methods. Experimental results show that existing excellent architectures with PLS are able to achieve better performance, and making them completely unsupervised.
Rongquan Wang, Xin Li 0034, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
ICASSP5
2024 EMo Transformer: Transformer-Based Depression Detection via Eye Movements
abstract
Depressive disorder has become a prevalent psychological illness that significantly impacts individuals’ daily lives. Traditional questionnaire assessment and clinical interviews suffer from issues such as subjectivity and a high consumption of medical resources. With the advancement of artificial intelligence, there is a growing number of depression detection methods based on statistical features. However, these methods have problems of insufficient stimulus extraction and neglecting temporal information. In order to solve these problems, we propose a transformer-based model named EMo Transformer, designed for detecting depression by effectively extracting features from stimuli and combining them with eye movements. Additionally, due to challenge in collecting data from depression patients, we design a simple and effective data augmentation method to solve this challenge. Subsequently, we design an ensemble model using the models with and without data augmentation. The experimental results of accuracy 91.95% demonstrate that our method is effective.
Xin Li 0034, Haizhuang Liu, Rongquan Wang, Bochao Zou, Huimin Ma 0001
ICME2
2024 CMT: Co-training Mean-Teacher for Unsupervised Domain Adaptation on 3D Object Detection
Junbao Zhuo, Xin Li 0034, Haizhuang Liu, Rongquan Wang, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia4
2024 Affinity3D: Propagating Instance-Level Semantic Affinity for Zero-Shot Point Cloud Semantic Segmentation
abstract
Zero-shot point cloud semantic segmentation aims to recognize novel classes at the point level. Previous methods mainly transfer excellent zero-shot generalization capabilities from images to point clouds. However, directly transferring knowledge from images to point clouds faces two ambiguous problems. On the one hand, 2D models will generate wrong predictions when the image changes. On the other hand, directly mapping 3D points to 2D pixels by perspective projection fails to consider the visibility of 3D points in camera view. The wrong geometric alignment of 3D points and 2D pixels causes semantic ambiguity. To tackle these two problems, we propose a framework named Affinity3D that intends to empower 3D semantic segmentation models to perceive novel samples. Our framework aggregates instances in 3D and recognizes them in 2D, leveraging the excellent geometric separation in 3D and the zero-shot capabilities of 2D models. Affinity3D involves an affinity module that rectifies the wrong predictions by comparing them with similar instances and a visibility module preventing knowledge transfer from visible 2D pixels to invisible 3D points. Extensive experiments have been conducted on the SemanticKITTI and nuScenes datasets. Our framework achieves state-of-the-art performance on both two datasets. Code is available at https://github.com/opjang5/Affinity3D.
Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia1
2024 Enhancing pseudo label quality for pedestrian and cyclist in weakly supervised 3D object detection
Haizhuang Liu, Huazhen Chu, Bochao Zou, Huimin Ma 0001
Neurocomputing1
2023 Temporal Information Fusion Network for Driving Behavior Prediction
abstract
Since enormous hazards are caused by traffic crashes every year, ensuring safe driving is a hot topic in transportation. Technologies related to the Advanced Driver Assistance System (ADAS) are evolving rapidly. But without an adequate understanding of driving intention, ADAS usually can’t help the driver prepare for the danger in advance. This paper focuses on the fusion strategy of driver and environment information and proposes a lightweight end-to-end model, temporal information fusion network (TIFN). Driving behavior is the interactive result of the driver and the external world. To better understand the driver’s intention, the state update cell (STU) is proposed to introduce the influence of environment information into the driver’s state modeling, inspired by the selective attention of the human cognition process. Meanwhile, semantic segmentation features are extracted to offer clear clues affecting driver attention in place of motion optical flow images and binary value vectors. Finally, the driver’s intention and environment state are combined to make a joint prediction. The experiments evaluated on Brain4cars and IESDD show that the proposed approach has superior performance than other approaches that only use camera data.
Chenghao Guo, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial Labels
abstract
Previous weakly-supervised methods of 3D object detection in driving scenes mainly rely on spatial labels, which provide the location, dimension, or orientation information. The annotation of 3D spatial labels is time-consuming. There also exist methods that do not require spatial labels, but their detections may fall on object parts rather than entire objects or backgrounds. In this paper, a novel cross-modal weakly-supervised 3D progressive refinement framework (WS3DPR) for 3D object detection that only needs image-level class annotations is introduced. The proposed framework consists of two stages: 1) classification refinement for potential objects localization and 2) regression refinement for spatial pseudo labels reasoning. In the first stage, a region proposal network is trained by cross-modal class knowledge transferred from 2D image to 3D point cloud and class information propagation. In the second stage, the locations, dimensions, and orientations of 3D bounding boxes are further refined with geometric reasoning based on 2D frustum and 3D region. When only image-level class labels are available, proposals with different 3D locations become overlapped in 2D, leading to the misclassification of foreground objects. Therefore, a 2D-3D semantic consistency block is proposed to disentangle different 3D proposals after projection. The overall framework progressively learns features in a coarse to fine manner. Comprehensive experiments on the KITTI3D dataset demonstrate that our method achieves competitive performance compared with previous methods with a lightweight labeling process.
Haizhuang Liu, Huimin Ma 0001, Bochao Zou, Rongquan Wang, Jiansheng Chen 0001
ACM Multimedia1
2021 PLNL-3DSSD: Part-Aware 3D Single Stage Detector Using Local And Non-Local Attention
abstract
3D object detection in the real crowded scene is still a challenging task due to occlusion and density change. We propose a part-aware 3D single-stage detector with local and non-local attention (PLNL-3DSSD) to fully use part information and inter-object relation. A primary part feature fusion is proposed for encoding the entire box feature vector by introducing semantic parts dividing. We develop a parallel part branch for robust and accurate object detection. We also develop 10-cal and non-local attention in set abstraction for enhancing data flow transfer between objects. Our method ranks second in single-stage 3D object detector on the KITTI 3D car detection benchmark while ensuring satisfactory efficiency.
Haizhuang Liu, Huimin Ma 0001, Yanxian Chen, Xi Li 0010
ICIP1
2021 LiDAR-Based Symmetrical Guidance for 3D Object Detection
Huazhen Chu, Huimin Ma 0001, Haizhuang Liu, Rongquan Wang
PRCV (4)3