VLDB 2026 Research / reviewers in the wild / expert
Hongguang Zhu
dblp:290/1484
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-1356-5153ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 24% Image recognition and object detection · 22% Representation and self-supervised learning · 18% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 64% Image and video processing · 36% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object detection
prohibited item detection |
1.0 | 1 | 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026 |
Visual content generation and editing
image generation |
1.0 | 1 | 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026 |
Visual content generation and editing › image generation
text-to-image generation |
1.0 | 1 | 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
0.8 | 1 | 2024 | Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation · ECCV (33) 2024 |
Computer vision › Vision and language › multimodal representation
video-language representation learning |
0.8 | 1 | 2024 | Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation · ECCV (33) 2024 |
Machine learning › Learning paradigms
continual learning |
0.7 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Natural language and speech › Language models and text generation › large language model training
continual pre-training |
0.7 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Machine learning › Representation and self-supervised learning
topology preservation |
0.7 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Computer vision › Vision and language
vision-language pretraining |
0.7 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Image and video processing › super-resolution › image super-resolution
depth super-resolution |
0.5 | 1 | 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021 |
Image and video processing › super-resolution › image super-resolution › depth super-resolution
RGB-guided depth super-resolution |
0.5 | 1 | 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021 |
Computer vision › Image recognition and object detection
object detection |
0.3 | 1 | 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
momentum contrastive learning |
0.2 | 1 | 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023 |
Image and video processing › super-resolution
image super-resolution |
0.1 | 1 | 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.0cross-attention refinement · 2.0background occlusion modeling · 2.0contrastive learning · 0.8topology preservation · 0.7momentum contrast · 0.7high-frequency component decomposition · 0.5convolutional neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Taming Generative Synthetic Data for X-Ray Prohibited Item DetectionabstractTraining prohibited item detection models requires a large amount of X-ray security images, but collecting and annotating these images is time-consuming and laborious. To address data insufficiency, X-ray security image synthesis methods composite images to scale up datasets. However, previous methods primarily follow a two-stage pipeline, where they implement labor-intensive foreground extraction in the first stage and then composite images in the second stage. Such a pipeline introduces inevitable extra labor cost and is not efficient. In this paper, we propose a one-stage X-ray security image synthesis pipeline (Xsyn) based on text-to-image generation, which incorporates two effective strategies to improve the usability of synthetic images. The Cross-Attention Refinement (CAR) strategy leverages the cross-attention map from the diffusion model to refine the bounding box annotation. The Background Occlusion Modeling (BOM) strategy explicitly models background occlusion in the latent space to enhance imaging complexity. To the best of our knowledge, compared with previous methods, Xsyn is the first to achieve high-quality X-ray security image synthesis without extra labor cost. Experiments demonstrate that our method outperforms all previous methods with 1.2% mAP improvement, and the synthetic images generated by our method are beneficial to improve prohibited item detection performance across various X-ray security datasets and detectors. Code is available at https://github.com/pILLOW-1/Xsyn/. Jialong Sun, Hongguang Zhu, Weizhe Liu, Yunda Sun, Renshuai Tao, Yunchao Wei |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
Siyu Jiao, Hongguang Zhu, Jiannan Huang 0002, Yao Zhao 0001, Yunchao Wei, Humphrey Shi |
ECCV (33) | 2 |
| 2023 | CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology PreservationabstractVision-Language Pretraining (VLP) has shown impressive results on diverse downstream tasks by offline training on large-scale datasets. Regarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learning ability to accumulate knowledge constantly. However, most continual learning studies are limited to uni-modal classification and existing multimodal datasets cannot simulate continual non-stationary data stream scenarios. To support the study of Vision-Language Continual Pretraining (VLCP), we first contribute a comprehensive and unified benchmark dataset P9D which contains over one million product image-text pairs from 9 industries. The data from each industry as an independent task supports continual learning and conforms to the real-world long-tail nature to simulate pretraining on web data. We comprehensively study the characteristics and challenges of VLCP, and propose a new algorithm: Compatible momentum contrast with Topology Preservation, dubbed CTP. The compatible momentum model absorbs the knowledge of the current and previous-task models to flexibly update the modal feature. Moreover, Topology Preservation transfers the knowledge of embedding across tasks while preserving the flexibility of feature adjustment. The experimental results demonstrate our method not only achieves superior performance compared with other baselines but also does not bring an expensive training burden. Dataset and codes are available at https://github.com/KevinLight831/CTP. Hongguang Zhu, Yunchao Wei, Xiaodan Liang, Chunjie Zhang 0001, Yao Zhao 0001 |
ICCV | 1 |
| 2023 | ESA: External Space Attention Aggregation for Image-Text RetrievalabstractDue to the large gap between vision and language modalities, effective and efficient image-text retrieval is still an unsolved problem. Recent progress devotes to unilaterally pursuing retrieval accuracy by either entangled image-text interaction or large-scale vision-language pre-training in a brute force way. However, the former often leads to unacceptable retrieval time explosion when deploying on large-scale databases. The latter heavily relies on the extra corpus to learn better alignment in the feature space while obscuring the contribution of the network architecture. In this work, we aim to investigate a trade-off to balance effectiveness and efficiency. To this end, on the premise of efficient retrieval, we propose the plug-and-play External Space attention Aggregation (ESA) module to enable element-wise fusion of modal features under spatial dimensional attention. Based on flexible spatial awareness, we further propose the Self-Expanding triplet Loss (SEL) to expand the representation space of samples and optimize the alignment of embedding space. The extensive experiments demonstrate the effectiveness of our method on two benchmark datasets. With identical visual and textual backbones, our single model has outperformed the ensemble modal of similar methods, and our ensemble model can further expand the advantage. Meanwhile, compared with the vision-language pre-training embedding-base method that used$83\times $image-text pairs than ours, our approach not only surpasses in performance but also accelerates$3\times $on retrieval time. Codes and pre-trained models are available athttps://github.com/KevinLight831/ESA. Hongguang Zhu, Chunjie Zhang 0001, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | AMC: Adaptive Multi-expert Collaborative Network for Text-guided Image RetrievalabstractText-guided image retrieval integrates reference image and text feedback as a multimodal query to search the image corresponding to user intention. Recent approaches employ multi-level matching, multiple accesses, or multiple subnetworks for better performance regardless of the heavy burden of storage and computation in the deployment. Additionally, these models not only rely on expert knowledge to handcraft image-text composing modules but also do inference by the static computational graph. It limits the representation capability and generalization ability of networks in the face of challenges from complex and varied combinations of reference image and text feedback. To break the shackles of the static network concept, we introduce the dynamic router mechanism to achieve data-dependent expert activation and flexible collaboration of multiple experts to explore more implicit multimodal fusion patterns. Specifically, we construct AMC, our A daptive M ulti-expert C ollaborative network, by using the proposed router to activate the different experts with different levels of image-text interaction. Since routers can dynamically adjust the activation of experts for the current samples, AMC can achieve the adaptive fusion mode for the different reference image and text combinations and generate dynamic computational graphs according to varied multimodal queries. Extensive experiments on two benchmark datasets demonstrate that due to benefits from the image-text composing representation produced by an adaptive multi-expert collaboration mechanism, AMC has better retrieval performance and zero-shot generalization ability than the state-of-the-art method while keeping the lightweight model and fast retrieval speed. Moreover, we analyze the visualization of path activation, attention map, and retrieval results to further understand the routing decisions and semantic localization ability of AMC. The codes and pretrained models are available at https://github.com/KevinLight831/AMC . Hongguang Zhu, Yunchao Wei, Yao Zhao 0001, Chunjie Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and BaselineabstractDepth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors. Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
CVPR | 2 |