Huaici Zhao

dblp:80/5232 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-7772-8652ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
abstract
Generative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized evaluation. We present LiDARCrafter, a unified framework for controllable 4D LiDAR generation and editing. Free-form language instructions are converted into ego-centric scene graphs that guide a tri-branch diffusion model to generate object geometry, motion, and structural priors. An autoregressive module further produces temporally coherent and stable LiDAR sequences with improved global consistency. To enable fair comparison, we introduce a comprehensive benchmark covering scene-, object-, and sequence-level metrics for rigorous and reproducible evaluation. Experiments on nuScenes show that LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, paving the way for scalable data augmentation and realistic simulation in diverse scenarios. Code have been publicly available at https://lidarcrafter.github.io.
Alan Liang, Youquan Liu, Dongyue Lu, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi
AAAI7
2026 TA-OEM : Multimodal sentiment analysis using two-step hierarchical attention and self-supervised label generation
Zhijia Zhang, Jian Fang 0003, Huaici Zhao
Knowl. Based Syst.5
2026 MIIGAN: Mambas make strong GAN for infrared image generation
Fuchao Wang, Huaici Zhao, Yuhuai Peng
Neural Networks2
2025 Perspective-Invariant 3D Object Detection
abstract
With the rise of robotics, LiDAR-based 3D object detection has garnered significant attention in both academia and industry. However, existing datasets and methods predominantly focus on vehicle-mounted platforms, leaving other autonomous platforms underexplored. To bridge this gap, we introduce Pi3DET, the first benchmark featuring LiDAR data and 3D bounding box annotations collected from multiple platforms: vehicle, quadruped, and drone, thereby facilitating research in 3D object detection for non-vehicle platforms as well as cross-platform 3D detection. Based on Pi3DET, we propose a novel cross-platform adaptation framework that transfers knowledge from the well-studied vehicle platform to other platforms. This framework achieves perspective-invariant 3D detection through robust alignment at both geometric and feature levels. Additionally, we establish a benchmark to evaluate the resilience and robustness of current 3D detectors in cross-platform scenarios, providing valuable insights for developing adaptive 3D perception systems. Extensive experiments validate the effectiveness of our approach on challenging cross-platform tasks, demonstrating substantial gains over existing adaptation methods. We hope this work paves the way for generalizable and unified 3D perception systems across diverse and complex environments. Our Pi3DET dataset, cross-platform benchmark suite, and annotation toolkit have been made publicly available.
Ao Liang, Lingdong Kong, Dongyue Lu, Youquan Liu, Huaici Zhao, Wei Tsang Ooi
ICCV6
2025 Multi-Modal BEV Enhancement Fusion for 3D Object Detection in Autonomous Driving
abstract
Recent success in 3D object detection have underscored its importance in autonomous driving, particularly through the integration of diverse sensor modalities like RGB images and LiDAR point clouds. With the benefits of Lift-Splat Shift (LSS) paradigm, different data modalities can be effectively fused in Bird’s-Eye-View (BEV), significantly improving the detection performance. Although BEV-based fusion has significantly advanced 3D detection technology, the limited enhancement of image features and the inconsistency between different modalities still hinder the overall performance of the detector. In this work, we focus on effective enhancement strategies and design 3D object detection pipeline named ECL3D to further push the detection performance boundary. The first strategy, Depth-Semantic Feature Enhancement (DSE), aims to improve input features for the view transformer during training without adding computational burden during inference. This approach leverages low-resolution depth distribution supervision to maintain the accuracy of frustum generation and high-resolution depth supervision to provide richer clues for front-end features. The second strategy, Instance BEV Feature Enhancement (IBFE), introduces a mechanism to enhance instance-relevant features for multi-modal fusion. This strategy effectively suppresses background noise and enhances object-region features. Comprehensive experiments on nuScenes dataset demonstrate the effectiveness of our approach. Without any test-time-augmentation strategy, our detector achieves state-of-the-art performance in camera-LiDAR fusion 3D object detection task with mAP and NDS of 72.8% and 75.2%, respectively. The code is coming soon at: https://github.com/muchen2019/ECL3D
Yuan Zhang 0023, Xinchi Li, Huaici Zhao
IEEE Trans. Intell. Transp. Syst.6
2024 SuPrNet: Super Proxy for 4D occupancy forecasting
Ao Liang, Huaici Zhao
Knowl. Based Syst.4
2023 LiDAR-camera fusion: Dual transformer enhancement for 3D object detection
Mu Chen 0002, Huaici Zhao
Eng. Appl. Artif. Intell.3
2023 SPSNet: Boosting 3D point-based object detectors with stable point sampling
Ao Liang, Haiyang Hua, Whenyu Chen, Huaici Zhao
Eng. Appl. Artif. Intell.5
2023 BCAF-3D: Bilateral Content Awareness Fusion for cross-modal 3D object detection
Mu Chen 0002, Huaici Zhao
Knowl. Based Syst.3
2022 MSL3D: 3D object detection from monocular, stereo and point cloud for autonomous driving
Huaici Zhao
Neurocomputing3
2021 RTS3D: Real-time Stereo 3D Detection from 4D Feature-Consistency Embedding Space for Autonomous Driving
abstract
Although the recent image-based 3D object detection methods using Pseudo-LiDAR representation have shown great capabilities, a notable gap in efficiency and accuracy still exist compared with LiDAR-based methods. Besides, over-reliance on the stand-alone depth estimator, requiring a large number of pixel-wise annotations in the training stage and more computation in the inferencing stage, limits the scaling application in the real world. In this paper, we propose an efficient and accurate 3D object detection method from stereo images, named RTS3D. Different from the 3D occupancy space in the Pseudo-LiDAR similar methods, we design a novel 4D feature-consistent embedding (FCE) space as the intermediate representation of the 3D scene without depth supervision. The FCE space encodes the object's structural and semantic information by exploring the multi-scale feature consistency warped from stereo pair. Furthermore, a semantic-guided RBF (Radial Basis Function) and a structure-aware attention module are devised to reduce the influence of FCE space noise without instance mask supervision. Experiments on KITTI benchmark show that RTS3D is the first true real-time system (FPS>24) for stereo image 3D detection meanwhile achieves 10% improvement in average precision comparing with the previous state-of-the-art method.
Shun Su, Huaici Zhao
AAAI3
2021 Monocular 3D object detection using dual quadric for autonomous driving
Huaici Zhao
Neurocomputing2
2020 RTM3D: Real-Time Monocular 3D Detection from Object Keypoints for Autonomous Driving
Huaici Zhao, Feidao Cao
ECCV (3)2
2017 Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy Diagnosis
abstract
Successful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data.
Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao
ACM Trans. Multim. Comput. Commun. Appl.8
2016 Scalable gastroscopic video summarization via similar-inhibition dictionary selection
Shuai Wang 0003, Yang Cong, Jun Cao 0002, Yunsheng Yang, Yandong Tang, Huaici Zhao
Artif. Intell. Medicine6
2016 Multimodal image matching based on Multimodality Robust Line Segment Descriptor
Huaici Zhao, Jinfeng Lv
Neurocomputing2
2016 Accurate and robust feature-based homography estimation using HALF-SIFT and feature localization error weighting
Huaici Zhao
J. Vis. Commun. Image Represent.2
2015 Computer aided endoscope diagnosis via weakly labeled data mining
abstract
In comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model.
Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao
ICIP6
2010 Reexamination of CBR Hypothesis
Xifeng Zhou, Zelin Shi, Huaici Zhao
ICCBR3