VLDB 2026 Research / reviewers in the wild / expert
Huaici Zhao
dblp:80/5232
· DBLP profile ↗
19ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-7772-8652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesabstractGenerative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized evaluation. We present LiDARCrafter, a unified framework for controllable 4D LiDAR generation and editing. Free-form language instructions are converted into ego-centric scene graphs that guide a tri-branch diffusion model to generate object geometry, motion, and structural priors. An autoregressive module further produces temporally coherent and stable LiDAR sequences with improved global consistency. To enable fair comparison, we introduce a comprehensive benchmark covering scene-, object-, and sequence-level metrics for rigorous and reproducible evaluation. Experiments on nuScenes show that LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, paving the way for scalable data augmentation and realistic simulation in diverse scenarios. Code have been publicly available at https://lidarcrafter.github.io. Alan Liang, Youquan Liu, Dongyue Lu, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi |
AAAI | 7 |
| 2026 | TA-OEM : Multimodal sentiment analysis using two-step hierarchical attention and self-supervised label generation
Zhijia Zhang, Jian Fang 0003, Huaici Zhao |
Knowl. Based Syst. | 5 |
| 2026 | MIIGAN: Mambas make strong GAN for infrared image generation
Fuchao Wang, Huaici Zhao, Yuhuai Peng |
Neural Networks | 2 |
| 2025 | Perspective-Invariant 3D Object DetectionabstractWith the rise of robotics, LiDAR-based 3D object detection has garnered significant attention in both academia and industry. However, existing datasets and methods predominantly focus on vehicle-mounted platforms, leaving other autonomous platforms underexplored. To bridge this gap, we introduce Pi3DET, the first benchmark featuring LiDAR data and 3D bounding box annotations collected from multiple platforms: vehicle, quadruped, and drone, thereby facilitating research in 3D object detection for non-vehicle platforms as well as cross-platform 3D detection. Based on Pi3DET, we propose a novel cross-platform adaptation framework that transfers knowledge from the well-studied vehicle platform to other platforms. This framework achieves perspective-invariant 3D detection through robust alignment at both geometric and feature levels. Additionally, we establish a benchmark to evaluate the resilience and robustness of current 3D detectors in cross-platform scenarios, providing valuable insights for developing adaptive 3D perception systems. Extensive experiments validate the effectiveness of our approach on challenging cross-platform tasks, demonstrating substantial gains over existing adaptation methods. We hope this work paves the way for generalizable and unified 3D perception systems across diverse and complex environments. Our Pi3DET dataset, cross-platform benchmark suite, and annotation toolkit have been made publicly available. Ao Liang, Lingdong Kong, Dongyue Lu, Youquan Liu, Huaici Zhao, Wei Tsang Ooi |
ICCV | 6 |
| 2025 | Multi-Modal BEV Enhancement Fusion for 3D Object Detection in Autonomous DrivingabstractRecent success in 3D object detection have underscored its importance in autonomous driving, particularly through the integration of diverse sensor modalities like RGB images and LiDAR point clouds. With the benefits of Lift-Splat Shift (LSS) paradigm, different data modalities can be effectively fused in Bird’s-Eye-View (BEV), significantly improving the detection performance. Although BEV-based fusion has significantly advanced 3D detection technology, the limited enhancement of image features and the inconsistency between different modalities still hinder the overall performance of the detector. In this work, we focus on effective enhancement strategies and design 3D object detection pipeline named ECL3D to further push the detection performance boundary. The first strategy, Depth-Semantic Feature Enhancement (DSE), aims to improve input features for the view transformer during training without adding computational burden during inference. This approach leverages low-resolution depth distribution supervision to maintain the accuracy of frustum generation and high-resolution depth supervision to provide richer clues for front-end features. The second strategy, Instance BEV Feature Enhancement (IBFE), introduces a mechanism to enhance instance-relevant features for multi-modal fusion. This strategy effectively suppresses background noise and enhances object-region features. Comprehensive experiments on nuScenes dataset demonstrate the effectiveness of our approach. Without any test-time-augmentation strategy, our detector achieves state-of-the-art performance in camera-LiDAR fusion 3D object detection task with mAP and NDS of 72.8% and 75.2%, respectively. The code is coming soon at: https://github.com/muchen2019/ECL3D Yuan Zhang 0023, Xinchi Li, Huaici Zhao |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | SuPrNet: Super Proxy for 4D occupancy forecasting
Ao Liang, Huaici Zhao |
Knowl. Based Syst. | 4 |
| 2023 | LiDAR-camera fusion: Dual transformer enhancement for 3D object detection
Mu Chen 0002, Huaici Zhao |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | SPSNet: Boosting 3D point-based object detectors with stable point sampling
Ao Liang, Haiyang Hua, Whenyu Chen, Huaici Zhao |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | BCAF-3D: Bilateral Content Awareness Fusion for cross-modal 3D object detection
Mu Chen 0002, Huaici Zhao |
Knowl. Based Syst. | 3 |
| 2022 | MSL3D: 3D object detection from monocular, stereo and point cloud for autonomous driving
Huaici Zhao |
Neurocomputing | 3 |
| 2021 | RTS3D: Real-time Stereo 3D Detection from 4D Feature-Consistency Embedding Space for Autonomous DrivingabstractAlthough the recent image-based 3D object detection methods using Pseudo-LiDAR representation have shown great capabilities, a notable gap in efficiency and accuracy still exist compared with LiDAR-based methods. Besides, over-reliance on the stand-alone depth estimator, requiring a large number of pixel-wise annotations in the training stage and more computation in the inferencing stage, limits the scaling application in the real world. In this paper, we propose an efficient and accurate 3D object detection method from stereo images, named RTS3D. Different from the 3D occupancy space in the Pseudo-LiDAR similar methods, we design a novel 4D feature-consistent embedding (FCE) space as the intermediate representation of the 3D scene without depth supervision. The FCE space encodes the object's structural and semantic information by exploring the multi-scale feature consistency warped from stereo pair. Furthermore, a semantic-guided RBF (Radial Basis Function) and a structure-aware attention module are devised to reduce the influence of FCE space noise without instance mask supervision. Experiments on KITTI benchmark show that RTS3D is the first true real-time system (FPS>24) for stereo image 3D detection meanwhile achieves 10% improvement in average precision comparing with the previous state-of-the-art method. Shun Su, Huaici Zhao |
AAAI | 3 |
| 2021 | Monocular 3D object detection using dual quadric for autonomous driving
Huaici Zhao |
Neurocomputing | 2 |
| 2020 | RTM3D: Real-Time Monocular 3D Detection from Object Keypoints for Autonomous Driving
Huaici Zhao, Feidao Cao |
ECCV (3) | 2 |
| 2017 | Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy DiagnosisabstractSuccessful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data. Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2016 | Scalable gastroscopic video summarization via similar-inhibition dictionary selection
Shuai Wang 0003, Yang Cong, Jun Cao 0002, Yunsheng Yang, Yandong Tang, Huaici Zhao |
Artif. Intell. Medicine | 6 |
| 2016 | Multimodal image matching based on Multimodality Robust Line Segment Descriptor
Huaici Zhao, Jinfeng Lv |
Neurocomputing | 2 |
| 2016 | Accurate and robust feature-based homography estimation using HALF-SIFT and feature localization error weighting
Huaici Zhao |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Computer aided endoscope diagnosis via weakly labeled data miningabstractIn comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model. Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao |
ICIP | 6 |
| 2010 | Reexamination of CBR Hypothesis
Xifeng Zhou, Zelin Shi, Huaici Zhao |
ICCBR | 3 |