Zonghui Guo

dblp:300/5288 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-7830-6606ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Face, body and person analysis · 48% Video understanding and tracking · 41% Deep learning architectures and training · 11%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 56% Computational photography and imaging · 23% Image and video processing · 21%
Network and information security
1 paper
Digital forensics and information hiding · 87% Biometric security · 13%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image editing › image compositing
image harmonization
1.732023
Transformer for Image Harmonization and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Image Harmonization with Transformer · ICCV 2021
Intrinsic Image Harmonization · CVPR 2021
Computer vision › Face, body and person analysis
face recognition
1.012026
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Face, body and person analysis › face recognition › face representation
face representation learning
1.012026
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Digital forensics and information hiding
digital forensics
1.012026
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Digital forensics and information hiding › forgery detection
face forgery detection
1.012026
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Face, body and person analysis
face forgery detection
0.912025
Face Forgery Video Detection via Temporal Forgery Cue Unraveling · CVPR 2025
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal feature fusion
0.912025
Face Forgery Video Detection via Temporal Forgery Cue Unraveling · CVPR 2025
Computer vision › Video understanding and tracking
temporal modeling
0.912025
Face Forgery Video Detection via Temporal Forgery Cue Unraveling · CVPR 2025
Visual content generation and editing › video editing
video harmonization
0.812024
Video Harmonization with Triplet Spatio-Temporal Variation Patterns · CVPR 2024
Image and video processing
image enhancement
0.722023
Image Harmonization with Transformer · ICCV 2021
Transformer for Image Harmonization and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.712023
Transformer for Image Harmonization and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computational photography and imaging
intrinsic image decomposition
0.512021
Intrinsic Image Harmonization · CVPR 2021
Computational photography and imaging › intrinsic image decomposition
reflectance and shading
0.512021
Intrinsic Image Harmonization · CVPR 2021
Biometric security
anti-spoofing
0.312026
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Image and video processing › image restoration
image inpainting
0.212023
Transformer for Image Harmonization and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

threshold optimization · 2.0self-supervised pretraining · 2.0competitive learning · 2.0transformer · 1.8triplet transformer · 1.5temporal consistency metric · 1.5encoder-decoder architecture · 1.3temporal correlation · 0.9momentum accumulation · 0.9patch relation modeling · 0.5material-consistency penalty · 0.5encoder-decoder · 0.5disentanglement · 0.5autoencoder · 0.5
YearPublicationVenuePosition
2026 Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection
abstract
Face Forgery Detection (FFD), or Deepfake detection, aims to determine whether a digital face is real or fake. Due to different face synthesis algorithms with diverse forgery patterns, FFD models often overfit specific patterns in training datasets, resulting in poor generalization to other unseen forgeries. Existing FFD methods primarily leverage pre-trained backbones with general image representation capabilities and fine-tune them to identify facial forgery cues. However, these backbones lack domain-specific facial knowledge and insufficiently capture complex facial features, thus hindering effective implicit forgery cue identification and limiting generalization. Therefore, it is essential to revisit FFD workflow across the pre-training and fine-tuning stages, achieving an elaborate integration from facial representation to forgery detection to improve generalization. Specifically, we develop an FFD-specific pre-trained backbone with superior facial representation capabilities through self-supervised pre-training on real faces. We then propose a competitive fine-tuning framework that stimulates the backbone to identify implicit forgery cues through a competitive learning mechanism. Moreover, we devise a threshold optimization mechanism that utilizes prediction confidence to improve the inference reliability. Comprehensive experiments demonstrate that our method achieves excellent performance in FFD and extra face-related tasks, i.e., presentation attack detection.
Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 G2LFormer: Global-to-local token mixing transformer for blind image inpainting and beyond
Haoru Zhao, Zonghui Guo, Shishi Qiao, Zhaorui Gu, Junyu Dong, Haiyong Zheng
Pattern Recognit.2
2025 Face Forgery Video Detection via Temporal Forgery Cue Unraveling
abstract
Face Forgery Video Detection (FFVD) is a critical yet challenging task in determining whether a digital facial video is authentic or forged. Existing FFVD methods typically focus on isolated spatial or coarsely fused spatiotemporal information, failing to leverage temporal forgery cues thus resulting in unsatisfactory performance. We strive to unravel these cues across three progressive levels: momentary anomaly, gradual inconsistency, and cumulative distortion. Accordingly, we design a consecutive correlate module to capture momentary anomaly cues by correlating interactions among consecutive frames. Then, we devise a future guide module to unravel inconsistency cues by iteratively aggregating historical anomaly cues and gradually propagating them into future frames. Finally, we introduce a historical review module that unravels distortion cues via momentum accumulation from future to historical frames. These three modules form our Temporal Forgery Cue Unraveling (TFCU) framework, sequentially highlighting spatial discriminative features by unraveling temporal forgery cues bidirectionally between historical and future frames. Extensive experiments and ablation studies demonstrate the effectiveness of our TFCU method, achieving state-of-the-art performance across diverse unseen datasets and manipulation methods. Code is available at https://github.com/zhenglab/TFCU.
Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan
CVPR1
2024 Video Harmonization with Triplet Spatio-Temporal Variation Patterns
abstract
Video harmonization is an important and challenging task that aims to obtain visually realistic composite videos by automatically adjusting the foreground's appearance to harmonize with the background. Inspired by the short-term and long-term gradual adjustment process of manual har-monization, we present a Video Triplet Transformer frame-work to model three spatio-temporal variation patterns within videos, i.e., short-term spatial as well as long-term global and dynamic, for video-to-video tasks like video har-monization. Specifically, for short-term harmonization, we adjust foreground appearance to consist with background in spatial dimension based on the neighbor frames; for long-term harmonization, we not only explore global ap-pearance variations to enhance temporal consistency but also alleviate motion offset constraints to align similar con-textual appearances dynamically. Extensive experiments and ablation studies demonstrate the effectiveness of our method, achieving state-of-the-art performance in video harmonization, video enhancement, and video demoireing tasks. We also propose a temporal consistency metric to better evaluate the harmonized videos. Code is available at https://github.com/zhenglablVideoTripletTransformer.
Zonghui Guo, Jie Zhang 0071, Shiguang Shan, Haiyong Zheng
CVPR1
2023 Transformer for Image Harmonization and Beyond
abstract
Image harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it should look like in the foreground referring to surrounding background. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our two vision Transformer frameworks and corresponding methods, as well as comprehensive experiments and empirical study, demonstrating the power of Transformer and investigating the Transformer for vision. Our methods achieve state-of-the-art performance on the image harmonization as well as four additional vision and graphics tasks, i.e., image enhancement, image inpainting, white-balance editing, and portrait relighting, indicating the superiority of our work. Code, models, more results and details can be found at the project website http://ouc.ai/project/HarmonyTransformer.
Zonghui Guo, Zhaorui Gu, Junyu Dong, Haiyong Zheng
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Intrinsic Image Harmonization
abstract
Compositing an image usually inevitably suffers from inharmony problem that is mainly caused by incompatibility of foreground and background from two different images with distinct surfaces and lights, corresponding to material-dependent and light-dependent characteristics, namely, reflectance and illumination intrinsic images, respectively. Therefore, we seek to solve image harmonization via separable harmonization of reflectance and illumination, i.e., intrinsic image harmonization. Our method is based on an autoencoder that disentangles composite image into reflectance and illumination for further separate harmonization. Specifically, we harmonize reflectance through material-consistency penalty, while harmonize illumination by learning and transferring light from background to foreground, moreover, we model patch relations between foreground and background of composite images in an inharmony-free learning way, to adaptively guide our intrinsic image harmonization. Both extensive experiments and ablation studies demonstrate the power of our method as well as the efficacy of each component. We also contribute a new challenging dataset for benchmarking illumination harmonization. Code and dataset are at https://github.com/zhenglab/IntrinsicHarmony.
Zonghui Guo, Haiyong Zheng, Zhaorui Gu
CVPR1
2021 Image Harmonization with Transformer
abstract
Image harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it looks like in the surrounding background near the foreground. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our harmonization Transformer frameworks without and with disentanglement, as well as comprehensive experiments and ablation study, demonstrating the power of Transformer and investigating the Transformer for vision. Our method achieves state-of-the-art performance on both image harmonization and image inpainting/enhancement, indicating its superiority. Our code and models are available at https://github.com/zhenglab/HarmonyTransformer.
Zonghui Guo, Haiyong Zheng, Zhaorui Gu, Junyu Dong
ICCV1