EDBT 2026 Demo / reviewers in the wild / expert
Haimei Zhao
dblp:231/1005
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-1139-4183ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A robust multi-view clustering framework with instance and feature-level self-paced learning
Haikun Xu, Haimei Zhao |
Knowl. Based Syst. | 4 |
| 2026 | WA-SAND: Wavelet attention diffusion with spatially adaptive noising for multi-view face generation
Weibo Zhong, Shichao Hu, Haimei Zhao |
Neural Networks | 4 |
| 2025 | MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving PerceptionabstractMulti-sensor fusion models play a crucial role in autonomous driving perception, particularly in tasks like 3D object detection and HD map construction. These models provide essential and comprehensive static environmental information for autonomous driving systems. While camera-LiDAR fusion methods have shown promising results by integrating data from both modalities, they often depend on complete sensor inputs. This reliance can lead to low robustness and potential failures when sensors are corrupted or missing, raising significant safety concerns. To tackle this challenge, we introduce the Multi-Sensor Corruption Benchmark (MSC-Bench), the first comprehensive benchmark aimed at evaluating the robustness of multi-sensor autonomous driving perception models against various sensor corruptions. Our benchmark includes 16 combinations of corruption types that disrupt both camera and LiDAR inputs, either individually or concurrently. Extensive evaluations of six 3D object detection models and four HD map construction models reveal substantial performance degradation under adverse weather conditions and sensor failures, underscoring critical safety issues. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible. Project website: MSC-Bench. Xiaoshuai Hao, Guanqun Liu 0008, Yuheng Ji, Mengchuan Wei, Haimei Zhao, Lingdong Kong, Rong Yin 0001, Yu Liu 0023 |
ICME | 6 |
| 2025 | STViT+: improving self-supervised multi-camera depth estimation with spatial-temporal context and adversarial geometry regularizationabstractAbstract Multi-camera depth estimation has gained significant attention in autonomous driving due to its importance in perceiving complex environments. However, extending monocular self-supervised methods to multi-camera setups introduces unique challenges that existing techniques often fail to address. In this paper, we propose STViT+ , a novel Transformer-based framework for self-supervised multi-camera depth estimation. Our key contributions include: 1) the Spatial-Temporal Transformer (STTrans) , which integrates local spatial connectivity and global context to capture enriched spatial-temporal cross-view correlations, resulting in more accurate 3D geometry reconstruction; 2) the Spatial-Temporal Photometric Consistency Correction (STPCC) strategy that mitigates the impact of varying illumination, ensuring brightness consistency across frames during photometric loss calculation; 3) the Adversarial Geometry Regularization (AGR) module, which employs Generative Adversarial Networks to impose spatial constraints by using unpaired depth maps, enhancing performance under adverse conditions such as rain and nighttime driving. Extensive evaluations on large-scale autonomous driving datasets, including Nuscenes and DDAD, confirm that STViT+ sets a new benchmark for multi-camera depth estimation. Zhuo Chen 0040, Haimei Zhao, Xiaoshuai Hao, Bo Yuan 0003, Xiu Li 0001 |
Appl. Intell. | 2 |
| 2024 | Dual Mapping of 2D StyleGAN for 3D-Aware Image Generation and Manipulation (Student Abstract)abstract3D-aware GANs successfully solve the problem of 3D-consistency generation and furthermore provide a 3D shape of the generated object. However, the application of the volume renderer disturbs the disentanglement of the latent space, which makes it difficult to manipulate 3D-aware GANs and lowers the image quality of style-based generators. In this work, we devise a dual-mapping framework to make the generated images of pretrained 2D StyleGAN consistent in 3D space. We utilize a tri-plane representation to estimate the 3D shape of the generated object and two mapping networks to bridge the latent space of StyleGAN and the 3D tri-plane space. Our method does not alter the parameters of the pretrained generator, which means the interpretability of latent space is preserved for various image manipulations. Experiments show that our method lifts the 3D awareness of pretrained 2D StyleGAN to 3D-aware GANs and outperforms the 3D-aware GANs in controllability and image quality. Zhuo Chen 0040, Haimei Zhao, Bo Yuan 0003, Xiu Li 0001 |
AAAI | 2 |
| 2024 | STViT: Improving Self-Supervised Multi-Camera Depth Estimation with Spatial-Temporal Context and Adversarial Geometry Regularization (Student Abstract)abstractMulti-camera depth estimation has recently garnered significant attention due to its substantial practical implications in the realm of autonomous driving. In this paper, we delve into the task of self-supervised multi-camera depth estimation and propose an innovative framework, STViT, featuring several noteworthy enhancements: 1) we propose a Spatial-Temporal Transformer to comprehensively exploit both local connectivity and the global context of image features, meanwhile learning enriched spatial-temporal cross-view correlations to recover 3D geometry. 2) to alleviate the severe effect of adverse conditions, e.g., rainy weather and nighttime driving, we introduce a GAN-based Adversarial Geometry Regularization Module (AGR) to further constrain the depth estimation with unpaired normal-condition depth maps and prevent the model from being incorrectly trained. Experiments on challenging autonomous driving datasets Nuscenes and DDAD show that our method achieves state-of-the-art performance. Zhuo Chen 0040, Haimei Zhao, Bo Yuan 0003, Xiu Li 0001 |
AAAI | 2 |
| 2024 | SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionabstractMulti-view camera-based 3D object detection has become popular due to its low cost, but accurately inferring 3D geometry solely from camera data remains challenging and may lead to inferior performance. Although distilling precise 3D geometry knowledge from LiDAR data could help tackle this challenge, the benefits of LiDAR information could be greatly hindered by the significant modality gap between different sensory modalities. To address this issue, we propose a Simulated multi-modal Distillation (SimDistill) method by carefully crafting the model architecture and distillation strategy. Specifically, we devise multi-modal architectures for both teacher and student models, including a LiDAR-camera fusion-based teacher and a simulated fusion-based student. Owing to the ``identical'' architecture design, the student can mimic the teacher to generate multi-modal features with merely multi-view images as input, where a geometry compensation module is introduced to bridge the modality gap. Furthermore, we propose a comprehensive multi-modal distillation scheme that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space. Incorporating them together, our SimDistill can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment. Extensive experiments validate the effectiveness and superiority of SimDistill over state-of-the-art methods, achieving an improvement of 4.8% mAP and 4.1% NDS over the baseline detector. The source code will be released at https://github.com/ViTAE-Transformer/SimDistill. Haimei Zhao, Qiming Zhang 0001, Shanshan Zhao 0001, Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
AAAI | 1 |
| 2024 | UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse WeatherabstractLiDAR semantic segmentation (LSS) is a critical task in autonomous driving and has achieved promising progress. However, prior LSS methods are conventionally investigated and evaluated on datasets within the same domain in clear weather. The robustness of LSS models in unseen scenes and all weather conditions is crucial for ensuring safety and reliability in real applications. To this end, we propose UniMix, a universal method that enhances the adaptability and generalizability of LSS models. UniMix first leverages physically valid adverse weather simulation to construct a Bridge Domain, which serves to bridge the domain gap between the clear weather scenes and the adverse weather scenes. Then, a Universal Mixing operator is defined regarding spatial, intensity, and semantic distributions to create the intermediate domain with mixed samples from given domains. Integrating the proposed two techniques into a teacher-student framework, UniMix efficiently mitigates the domain gap and enables LSS models to learn weather-robust and domain-invariant representations. We devote UniMix to two main setups: 1) unsupervised domain adaption, adapting the model from the clear weather source domain to the adverse weather target domain; 2) domain generalization, learning a model that generalizes well to unseen scenes in adverse weather. Extensive experiments validate the effectiveness of UniMix across different tasks and datasets, all achieving superior performance over state-of-the-art methods. The code will be released. Haimei Zhao, Jing Zhang 0037, Zhuo Chen 0040, Shanshan Zhao 0001, Dacheng Tao |
CVPR | 1 |
| 2024 | MapDistill: Boosting Efficient Camera-Based HD Map Construction via Camera-LiDAR Fusion Model Distillation
Xiaoshuai Hao, Ruikai Li, Hui Zhang 0093, Dingzhe Li, Rong Yin 0001, Sangil Jung, Seung In Park, ByungIn Yoo, Haimei Zhao, Jing Zhang 0037 |
ECCV (3) | 9 |
| 2024 | Is Your HD Map Constructor Reliable under Sensor Corruptions?abstractDriving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is not well understood, raising safety concerns. This work introduces MapBench, the first comprehensive benchmark designed to evaluate the robustness of HD map construction methods against various sensor corruptions. Our benchmark encompasses a total of 29 types of corruptions that occur from cameras and LiDAR sensors. Extensive evaluations across 31 HD map constructors reveal significant performance degradation of existing methods under adverse weather conditions and sensor failures, underscoring critical safety concerns. We identify effective strategies for enhancing robustness, including innovative approaches that leverage multi-modal fusion, advanced data augmentation, and architectural techniques. These insights provide a pathway for developing more reliable HD map construction methods, which are essential for the advancement of autonomous driving technology. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible. Xiaoshuai Hao, Mengchuan Wei, Yifan Yang 0007, Haimei Zhao, Hui Zhang 0093, Yi Zhou 0020, Lingdong Kong, Jing Zhang 0037 |
NeurIPS | 4 |
| 2024 | SPH-Net: Hyperspectral Image Super-Resolution via Smoothed Particle Hydrodynamics ModelingabstractReconstructing a high-resolution hyperspectral image (HSI) from a low-resolution HSI is significant for many applications, such as remote sensing and aerospace. Most deep learning-based HSI super-resolution methods pay more attention to developing novel network structures but rarely study the HSI super-resolution problem from the perspective of image dynamic evolution. In this article, we propose that the HSI pixel motion during the super-resolution reconstruction process can be analogized to the particle movement in the smoothed particle hydrodynamics (SPH) field. To this end, we design an SPH network (SPH-Net) for HSI super-resolution in light of the SPH theory. Specifically, we construct a smooth function based on SPH and design a smooth convolution in multiscales to exploit spectral correlation and preserve the spectral information in the super-resolved image. In addition, we apply the SPH approximation method to discretize the Navier-Stokes motion equation into SPH equation form, which can guide the HSI pixel motion in the desired direction during super-resolution reconstruction, thereby producing clear edges in the spatial domain. Experiments on three public hyperspectral datasets demonstrate that the proposed SPH-Net outperforms the state-of-the-art methods in terms of objective metrics and visual quality. Mingjin Zhang, Jiamin Xu, Jing Zhang 0037, Haimei Zhao, Wenteng Shang, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | Deep CornerabstractAbstract Recent studies have shown promising results on joint learning of local feature detectors and descriptors. To address the lack of ground-truth keypoint supervision, previous methods mainly inject appropriate knowledge about keypoint attributes into the network to facilitate model learning. In this paper, inspired by traditional corner detectors, we develop an end-to-end deep network, named Deep Corner, which adds a local similarity-based keypoint measure into a plain convolutional network. Deep Corner enables finding reliable keypoints and thus benefits the learning of the distinctive descriptors. Moreover, to improve keypoint localization, we first study previous multi-level keypoint detection strategies and then develop a multi-level U-Net architecture, where the similarity of features at multiple levels can be exploited effectively. Finally, to improve the invariance of descriptors, we propose a feature self-transformation operation, which transforms the learned features adaptively according to the specific local information. The experimental results on several tasks and comprehensive ablation studies demonstrate the effectiveness of our method and the involved components. Shanshan Zhao 0001, Mingming Gong, Haimei Zhao, Jing Zhang 0037, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2022 | JPerceiver: Joint Perception Network for Depth, Pose and Layout Estimation in Driving Scenes
Haimei Zhao, Jing Zhang 0037, Sen Zhang 0006, Dacheng Tao |
ECCV (38) | 1 |
| 2022 | D2Animator: Dual Distillation of StyleGAN For High-Resolution Face AnimationabstractThe style-based generator architectures (e.g. StyleGAN v1, v2) largely promote the controllability and explainability of Generative Adversarial Networks (GANs). Many researchers have applied the pretrained style-based generators to image manipulation and video editing by exploring the correlation between linear interpolation in the latent space and semantic transformation in the synthesized image manifold. However, most previous studies focused on manipulating separate discrete attributes, which is insufficient to animate a still image to generate videos with complex and diverse poses and expressions. In this work, we devise a dual distillation strategy (D2Animator) for generating animated high-resolution face videos conditioned on identities and poses from different images. Specifically, we first introduce a Clustering-based Distiller (CluDistiller) to distill diverse interpolation directions in the latent space, and synthesize identity-consistent faces with various poses and expressions, such as blinking, frowning, looking up/down, etc. Then we propose an Augmentation-based Distiller (AugDistiller) that learns to encode arbitrary face deformation into a combination of interpolation directions via training on augmentation samples synthesized by CluDistiller. Through assembling the two distillation methods, D2Animator can generate high-resolution face animation videos without training on video sequences. Extensive experiments on self-driving, cross-identity and sequence-driving tasks demonstrate the superiority of the proposed D2Animator over existing StyleGAN manipulation and face animation methods in both generation quality and animation fidelity. Zhuo Chen 0040, Haimei Zhao, Bo Yuan 0003, Xiu Li 0001 |
ACM Multimedia | 3 |
| 2020 | Collaborative Learning of Depth Estimation, Visual Odometry and Camera Relocalization from Monocular VideosabstractScene perceiving and understanding tasks including depth estimation, visual odometry (VO) and camera relocalization are fundamental for applications such as autonomous driving, robots and drones. Driven by the power of deep learning, significant progress has been achieved on individual tasks but the rich correlations among the three tasks are largely neglected. In previous studies, VO is generally accurate in local scope yet suffers from drift in long distances. By contrast, camera relocalization performs well in the global sense but lacks local precision. We argue that these two tasks should be strategically combined to leverage the complementary advantages, and be further improved by exploiting the 3D geometric information from depth data, which is also beneficial for depth estimation in turn. Therefore, we present a collaborative learning framework, consisting of DepthNet, LocalPoseNet and GlobalPoseNet with a joint optimization loss to estimate depth, VO and camera localization unitedly. Moreover, the Geometric Attention Guidance Model is introduced to exploit the geometric relevance among three branches during learning. Extensive experiments demonstrate that the joint learning scheme is useful for all tasks and our method outperforms current state-of-the-art techniques in depth estimation and camera relocalization with highly competitive performance in VO. Haimei Zhao, Wei Bian 0003, Bo Yuan 0003, Dacheng Tao |
IJCAI | 1 |
| 2018 | Towards a Compact and Effective Representation for Datasets with Inhomogeneous Clusters
Haimei Zhao, Zhuo Chen 0040, Qiuhui Tong, Bo Yuan 0003 |
ICONIP (4) | 1 |