VLDB 2026 Research / reviewers in the wild / expert
Zhenghao Zhao
dblp:319/2723
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
0009-0000-1934-4661ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distilling Long-tailed DatasetsabstractDataset distillation aims to synthesize a small, information-rich dataset from a large one for efficient model training. However, existing dataset distillation methods struggle with long-tailed datasets, which are prevalent in real-world scenarios. By investigating the reasons behind this unexpected result, we identified two main causes: 1) The distillation process on imbalanced datasets develops biased gradients, leading to the synthesis of similarly imbalanced distilled datasets. 2) The experts trained on such datasets perform suboptimally on tail classes, resulting in misguided distillation supervision and poor-quality soft-label initialization. To address these issues, we first propose Distribution-agnostic Matching to avoid directly matching the biased expert trajectories. It reduces the distance between the student and the biased expert trajectories and prevents the tail class bias from being distilled to the synthetic dataset. Moreover, we improve the distillation guidance with Expert Decoupling, which jointly matches the decoupled backbone and classifier to improve the tail class performance and initialize reliable soft labels. This work pioneers the field of long-tailed dataset distillation, marking the first effective effort to distill long-tailed datasets. Our code will be made public at https://github.com/ichbill/LTDD. Zhenghao Zhao, Haoxuan Wang 0002, Yuzhang Shang, Kai Wang 0036, Yan Yan 0002 |
CVPR | 1 |
| 2025 | CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
Haoxuan Wang 0002, Zhenghao Zhao, Junyi Wu 0002, Yuzhang Shang, Gaowen Liu, Yan Yan 0002 |
ICCV | 2 |
| 2025 | SSDL: Sensor-to-Skeleton Diffusion Model with Lipschitz Regularization for Human Activity Recognition
Changchang Sun, Zhenghao Zhao, Anne H. H. Ngu, Hugo Latapie, Yan Yan 0002 |
MMM (4) | 3 |
| 2025 | Efficient Multimodal Dataset Distillation via Generative ModelsabstractDataset distillation aims to synthesize a small dataset from a large dataset, enabling the model trained on it to perform well on the original dataset. With the blooming of large language models and multimodal large language models, the importance of multimodal datasets, particularly image-text datasets, has grown significantly. However, existing multimodal dataset distillation methods are constrained by the Matching Training Trajectories algorithm, which significantly increases the computing resource requirement, and takes days to process the distillation. In this work, we introduce EDGE, a generative distillation method for efficient multimodal dataset distillation. Specifically, we identify two key challenges of distilling multimodal datasets with generative models: 1) The lack of correlation between generated images and captions. 2) The lack of diversity among generated samples.
To address the aforementioned issues, we propose a novel generative model training workflow with a bi-directional contrastive loss and a diversity loss. Furthermore, we propose a caption synthesis strategy to further improve text-to-image retrieval performance by introducing more text information. Our method is evaluated on Flickr30K, COCO, and CC3M datasets, demonstrating superior performance and efficiency compared to existing approaches. Notably, our method achieves results 18$\times$ faster than the state-of-the-art method. Our code will be made public at https://github.com/ichbill/EDGE. Zhenghao Zhao, Haoxuan Wang 0002, Junyi Wu 0002, Yuzhang Shang, Gaowen Liu, Yan Yan 0002 |
NeurIPS | 1 |
| 2024 | Dataset Quantization with Active Learning Based Adaptive Sampling
Zhenghao Zhao, Yuzhang Shang, Junyi Wu 0002, Yan Yan 0002 |
ECCV (60) | 1 |
| 2024 | Supplementing Missing Visions Via Dialog for Scene Graph GenerationsabstractMost AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reasons (e.g., restricted view range and occlusions). To this end, we investigate a task setting with incomplete visual input data. Specifically, we exploit the Scene Graph Generation (SGG) task with various levels of visual data missingness as input. While insufficient visual input naturally leads to performance drop, we propose to supplement the missing visions via natural language dialog interactions to better accomplish the task objective. We design a model-agnostic Supplementary Interactive Dialog (SI-Dial) framework that can be jointly learned with most existing models, endowing the current AI systems with the ability of question-answer interactions in natural language. We demonstrate the feasibility of such task setting with missing visual input and the effectiveness of our proposed dialog module as the supplementary information source through extensive experiments, by achieving promising performance improvement over multiple baselines. Zhenghao Zhao, Xiaoguang Zhu, Yuzhang Shang, Yan Yan 0002 |
ICASSP | 1 |
| 2024 | Audio-Visual Navigation with Anti-Backtracking
Zhenghao Zhao, Hao Tang 0005, Yan Yan 0002 |
ICPR (18) | 1 |
| 2024 | Monocular Expressive 3D Human Reconstruction of Multiple PeopleabstractWhole-body pose estimation aims to regress human pose models that include the body, hand, and facial details from RGB images. While the task of whole-body mesh recovery has been extensively studied in recent literature, the focus has predominantly been on human mesh recovery for a single person, despite the frequent occurrence of multiple people in practical scenarios. Similar to body-only cases, such single-person whole-body pose estimation methods often fail in the multiple-people problem for two reasons: (i) Given the ambiguous bounding box, which could contain more than one instance, it is difficult for single-person-oriented methods to regress the body mesh model of the target person. (ii) Single-person pose estimation approaches neglect the person-person occlusions and the depth order among instances, thus generating interpenetrated models. In this paper, we propose the Multi-person Expressive POse (MEPO) model, which exploits expressive 3D human model reconstruction for multiple people. To our best knowledge, our model is the first multi-person whole-body mesh reconstruction model, which is intensified by heatmap, depthmap, and depth order loss. We propose the Heatmap Enhancement Net (HENet) to leverage the heatmap information to assist the model in concentrating on the target person in crowded multi-person cases, while the depthmap delivers depth information of the image. Furthermore, we impose a depth order loss to recover human mesh precisely for overlapped people. In our experiments, we evaluate our model on multiple challenging datasets, including AGORA, which consists of complex occlusions similar to real-world scenarios. Our method has a significant performance improvement compared with the state-of-the-art pose estimation methods. Zhenghao Zhao, Hao Tang 0005, Joy Wan, Yan Yan 0002 |
ICMR | 1 |