Hongguang Zhu

dblp:290/1484 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-1356-5153ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 24% Image recognition and object detection · 22% Representation and self-supervised learning · 18%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 64% Image and video processing · 36%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
prohibited item detection
1.012026
Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026
Visual content generation and editing
image generation
1.012026
Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026
Visual content generation and editing › image generation
text-to-image generation
1.012026
Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.812024
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation · ECCV (33) 2024
Computer vision › Vision and language › multimodal representation
video-language representation learning
0.812024
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation · ECCV (33) 2024
Machine learning › Learning paradigms
continual learning
0.712023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Natural language and speech › Language models and text generation › large language model training
continual pre-training
0.712023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Machine learning › Representation and self-supervised learning
topology preservation
0.712023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Image and video processing › super-resolution › image super-resolution
depth super-resolution
0.512021
Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021
Image and video processing › super-resolution › image super-resolution › depth super-resolution
RGB-guided depth super-resolution
0.512021
Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021
Computer vision › Image recognition and object detection
object detection
0.312026
Taming Generative Synthetic Data for X-Ray Prohibited Item Detection · IEEE Trans. Inf. Forensics Secur. 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
momentum contrastive learning
0.212023
CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation · ICCV 2023
Image and video processing › super-resolution
image super-resolution
0.112021
Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline · CVPR 2021

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.0cross-attention refinement · 2.0background occlusion modeling · 2.0contrastive learning · 0.8topology preservation · 0.7momentum contrast · 0.7high-frequency component decomposition · 0.5convolutional neural network · 0.5
YearPublicationVenuePosition
2026 Taming Generative Synthetic Data for X-Ray Prohibited Item Detection
abstract
Training prohibited item detection models requires a large amount of X-ray security images, but collecting and annotating these images is time-consuming and laborious. To address data insufficiency, X-ray security image synthesis methods composite images to scale up datasets. However, previous methods primarily follow a two-stage pipeline, where they implement labor-intensive foreground extraction in the first stage and then composite images in the second stage. Such a pipeline introduces inevitable extra labor cost and is not efficient. In this paper, we propose a one-stage X-ray security image synthesis pipeline (Xsyn) based on text-to-image generation, which incorporates two effective strategies to improve the usability of synthetic images. The Cross-Attention Refinement (CAR) strategy leverages the cross-attention map from the diffusion model to refine the bounding box annotation. The Background Occlusion Modeling (BOM) strategy explicitly models background occlusion in the latent space to enhance imaging complexity. To the best of our knowledge, compared with previous methods, Xsyn is the first to achieve high-quality X-ray security image synthesis without extra labor cost. Experiments demonstrate that our method outperforms all previous methods with 1.2% mAP improvement, and the synthetic images generated by our method are beneficial to improve prohibited item detection performance across various X-ray security datasets and detectors. Code is available at https://github.com/pILLOW-1/Xsyn/.
Jialong Sun, Hongguang Zhu, Weizhe Liu, Yunda Sun, Renshuai Tao, Yunchao Wei
IEEE Trans. Inf. Forensics Secur.2
2024 Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
Siyu Jiao, Hongguang Zhu, Jiannan Huang 0002, Yao Zhao 0001, Yunchao Wei, Humphrey Shi
ECCV (33)2
2023 CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation
abstract
Vision-Language Pretraining (VLP) has shown impressive results on diverse downstream tasks by offline training on large-scale datasets. Regarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learning ability to accumulate knowledge constantly. However, most continual learning studies are limited to uni-modal classification and existing multimodal datasets cannot simulate continual non-stationary data stream scenarios. To support the study of Vision-Language Continual Pretraining (VLCP), we first contribute a comprehensive and unified benchmark dataset P9D which contains over one million product image-text pairs from 9 industries. The data from each industry as an independent task supports continual learning and conforms to the real-world long-tail nature to simulate pretraining on web data. We comprehensively study the characteristics and challenges of VLCP, and propose a new algorithm: Compatible momentum contrast with Topology Preservation, dubbed CTP. The compatible momentum model absorbs the knowledge of the current and previous-task models to flexibly update the modal feature. Moreover, Topology Preservation transfers the knowledge of embedding across tasks while preserving the flexibility of feature adjustment. The experimental results demonstrate our method not only achieves superior performance compared with other baselines but also does not bring an expensive training burden. Dataset and codes are available at https://github.com/KevinLight831/CTP.
Hongguang Zhu, Yunchao Wei, Xiaodan Liang, Chunjie Zhang 0001, Yao Zhao 0001
ICCV1
2023 ESA: External Space Attention Aggregation for Image-Text Retrieval
abstract
Due to the large gap between vision and language modalities, effective and efficient image-text retrieval is still an unsolved problem. Recent progress devotes to unilaterally pursuing retrieval accuracy by either entangled image-text interaction or large-scale vision-language pre-training in a brute force way. However, the former often leads to unacceptable retrieval time explosion when deploying on large-scale databases. The latter heavily relies on the extra corpus to learn better alignment in the feature space while obscuring the contribution of the network architecture. In this work, we aim to investigate a trade-off to balance effectiveness and efficiency. To this end, on the premise of efficient retrieval, we propose the plug-and-play External Space attention Aggregation (ESA) module to enable element-wise fusion of modal features under spatial dimensional attention. Based on flexible spatial awareness, we further propose the Self-Expanding triplet Loss (SEL) to expand the representation space of samples and optimize the alignment of embedding space. The extensive experiments demonstrate the effectiveness of our method on two benchmark datasets. With identical visual and textual backbones, our single model has outperformed the ensemble modal of similar methods, and our ensemble model can further expand the advantage. Meanwhile, compared with the vision-language pre-training embedding-base method that used$83\times $image-text pairs than ours, our approach not only surpasses in performance but also accelerates$3\times $on retrieval time. Codes and pre-trained models are available athttps://github.com/KevinLight831/ESA.
Hongguang Zhu, Chunjie Zhang 0001, Yunchao Wei, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 AMC: Adaptive Multi-expert Collaborative Network for Text-guided Image Retrieval
abstract
Text-guided image retrieval integrates reference image and text feedback as a multimodal query to search the image corresponding to user intention. Recent approaches employ multi-level matching, multiple accesses, or multiple subnetworks for better performance regardless of the heavy burden of storage and computation in the deployment. Additionally, these models not only rely on expert knowledge to handcraft image-text composing modules but also do inference by the static computational graph. It limits the representation capability and generalization ability of networks in the face of challenges from complex and varied combinations of reference image and text feedback. To break the shackles of the static network concept, we introduce the dynamic router mechanism to achieve data-dependent expert activation and flexible collaboration of multiple experts to explore more implicit multimodal fusion patterns. Specifically, we construct AMC, our A daptive M ulti-expert C ollaborative network, by using the proposed router to activate the different experts with different levels of image-text interaction. Since routers can dynamically adjust the activation of experts for the current samples, AMC can achieve the adaptive fusion mode for the different reference image and text combinations and generate dynamic computational graphs according to varied multimodal queries. Extensive experiments on two benchmark datasets demonstrate that due to benefits from the image-text composing representation produced by an adaptive multi-expert collaboration mechanism, AMC has better retrieval performance and zero-shot generalization ability than the state-of-the-art method while keeping the lightweight model and fast retrieval speed. Moreover, we analyze the visualization of path activation, attention map, and retrieval results to further understand the routing decisions and semantic localization ability of AMC. The codes and pretrained models are available at https://github.com/KevinLight831/AMC .
Hongguang Zhu, Yunchao Wei, Yao Zhao 0001, Chunjie Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline
abstract
Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors.
Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001
CVPR2