EDBT 2026 Demo / reviewers in the wild / expert
Bo Peng 0002
dblp:03/5954-2
· DBLP profile ↗
38ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0002-9014-7369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 17 since 2021Security and privacy · 7 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dark Miner: Towards combating residuals in concept erasure for text-to-image diffusion models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 2 |
| 2026 | DREAM: A Benchmark Study for Deepfake PhotoRealism AssessMentabstractDeep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its computational modeling, imitation, is also a significant problem but lacks adequate study. In this paper, we focus on the photorealism assessment of deepfakes, which is defined as the automatic assessment of deepfake photorealism that approximates human perception of deepfakes. It is important for evaluating the quality, deceptiveness of deepfakes which can be used for predicting the influence of deepfakes on Internet, it also has potentials in improving the deepfake generation process by serving as a critic. This paper promotes this new direction by presenting a comprehensive benchmark called DREAM, which stands for Deepfake photoREalism AssessMent. It is comprised of a deepfake video dataset of diverse quality, a large scale annotation that includes 140, 000 photorealism scores, textual descriptions obtained from 3, 500 human annotators, a comprehensive evaluation, analysis of 18 representative photorealism assessment methods, including recent large vision language model based methods, a newly proposed description-aligned CLIP method. The benchmark, insights included in this study can lay the foundation for future research in this direction, other related areas. Bo Peng 0002, Zichuan Wang, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Unveiling Deepfakes with Latent Diffusion Counterfactual ExplanationsabstractDeepfake technology, driven by deep learning, produces highly convincing synthetic media, raising concerns about misuse. While DeepFake detection models have achieved impressive accuracy, but due to the difficulty of distinguishing fake from real, interpretability remains challenging that humans cannot understand or trust the detection results. We propose a novel approach to enhance interpretability by generating counterfactual explanations. By integrating ensemble classifier loss and text instructions into the fine-tuning of a Latent Diffusion Model, our method effectively improves the quality and efficiency of generated counterfactual explanations. Experiments on DeepFake datasets validate the effectiveness of our approach, contributing the interpretability of Deepfake detection. Bo Peng 0002, Jing Dong 0003, Xiaoyu Zhang 0002 |
ICASSP | 2 |
| 2025 | Partial Reconstruction Error for Deepfake DetectionabstractThe rapid development of deepfake technology poses a formidable challenge to personal privacy and security, underscoring the urgent need for deepfake detection. Recently, the methods based on the reconstruction error, such as DIRE and RECCE, achieve impressive performance in forgery detection. However, their performance on facial forgery datasets is relatively poor. The reconstruction process is performed on the whole images, neglecting contextual information for reconstruction. In this paper, we propose Partial Reconstruction Error to perform deepfake detection based on the reconstruction of masked regions in an image. In this way, contextual information helps to reveal the inconsistencies between the original and reconstructed regions thereby improving the detection performance. This method outperforms the best global reconstruction-based approaches on the FF++, Celeb-DF, and DiFF datasets by 4.00%, 2.83%, and 2.67%, respectively. Zheling Meng, Bo Peng 0002, Jing Dong 0003, Beilin Chu, Wei Wang 0025 |
ICASSP | 3 |
| 2025 | Image-level Memorization Detection via Inversion-based Inference PerturbationabstractRecent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may need to verify whether their proprietary or personal images have been memorized by the model, even in the absence of paired prompts or related metadata. We refer to this challenge as image-level memorization detection, where current methods relying on original prompts fall short. In this work, we uncover two characteristics of memorized images after perturbing the inference procedure: lower similarity of the original images and larger magnitudes of TCNP.
Building on these insights, we propose Inversion-based Inference Perturbation (IIP), a new framework for image-level memorization detection. Our approach uses unconditional DDIM inversion to derive latent codes that contain core semantic information of original images and optimizes random prompt embeddings to introduce effective perturbation. Memorized images exhibit distinct characteristics within the proposed pipeline, providing a robust basis for detection. To support this task, we construct a comprehensive setup for the image-level memorization detection, carefully curating datasets to simulate realistic memorization scenarios. Using this setup, we evaluate our IIP framework across three different memorization settings, demonstrating its state-of-the-art performance in identifying memorized images in various settings, even in the presence of data augmentation attacks. Haokun Lin, Bo Peng 0002, Zhili Liu, Yueming Lyu, Xing Zheng, Jing Dong 0003 |
ICLR | 4 |
| 2025 | Concept Corrector: Erase Concepts on the Fly for Text-to-Image Diffusion Models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Yueming Lyu, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 2 |
| 2025 | Noise-Informed Diffusion-Generated Image Detection With Anomaly AttentionabstractWith the rapid development of image generation technologies, especially the advancement of Diffusion Models, the quality of synthesized images has significantly improved, raising concerns among researchers about information security. To mitigate the malicious abuse of diffusion models, diffusion-generated image detection has proven to be an effective countermeasure. However, a key challenge for forgery detection is generalising to diffusion models not seen during training. In this paper, we address this problem by focusing on image noise. We observe that images from different diffusion models share similar noise patterns, distinct from genuine images. Building upon this insight, we introduce a novel Noise-Aware Self-Attention (NASA) module that focuses on noise regions to capture anomalous patterns. To implement a SOTA detection model, we incorporate NASA into Swin Transformer, forming an novel detection architecture NASA-Swin. Additionally, we employ a cross-modality fusion embedding to combine RGB and noise images, along with a channel mask strategy to enhance feature learning from both modalities. Extensive experiments demonstrate the effectiveness of our approach in enhancing detection capabilities for diffusion-generated images. When encountering unseen generation methods, our approach achieves the state-of-the-art performance. Weinan Guan, Wei Wang 0025, Bo Peng 0002, Ziwen He, Jing Dong 0003, Haonan Cheng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Latent Watermark: Inject and Detect Watermarks in Latent Diffusion SpaceabstractWatermarking is a tool for actively identifying and attributing the images generated by latent diffusion models. Existing methods face the dilemma of image quality and watermark robustness. Watermarks with superior image quality usually have inferior robustness against attacks such as blurring and JPEG compression, while watermarks with superior robustness usually significantly damage image quality. This dilemma stems from the traditional paradigm where watermarks are injected and detected in pixel space, relying on pixel perturbation for watermark detection and resilience against attacks. In this paper, we highlight that an effective solution to the problem is to both inject and detect watermarks in the latent diffusion space, and propose Latent Watermark with a progressive training strategy. It weakens the direct connection between quality and robustness and thus alleviates their contradiction. We conduct evaluations on two datasets and against 10 watermark attacks. Six metrics measure the image quality and watermark robustness. Results show that compared to the recently proposed methods such as StableSignature, StegaStamp, RoSteALS, LaWa, TreeRing, and DiffuseTrace, LW not only surpasses them in terms of robustness but also offers superior image quality. Zheling Meng, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Multim. | 2 |
| 2024 | AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisabstractAudio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations. Wei Wang 0025, Bo Peng 0002, Yingya Zhang, Jing Dong 0003, Tieniu Tan |
AAAI | 4 |
| 2024 | Learning Dense Correspondence for NeRF-Based Face ReenactmentabstractFace reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment in photo-realism and 3D consistency. However, establishing dense correspondence between different face NeRFs is non-trivial, because implicit representations lack ground-truth correspondence annotations like mesh-based 3D parametric models (e.g., 3DMM) with index-aligned vertexes. Although aligning 3DMM space with NeRF-based face representations can realize motion control, it is sub-optimal for their limited face-only modeling and low identity fidelity. Therefore, we are inspired to ask: Can we learn the dense correspondence between different NeRF-based face representations without a 3D parametric model prior? To address this challenge, we propose a novel framework, which adopts tri-planes as fundamental NeRF representation and decomposes face tri-planes into three components: canonical tri-planes, identity deformations, and motion. In terms of motion control, our key contribution is proposing a Plane Dictionary (PlaneDict) module, which efficiently maps the motion conditions to a linear weighted addition of learnable orthogonal plane bases. To the best of our knowledge, our framework is the first method that achieves one-shot multi-view face reenactment without a 3D parametric model prior. Extensive experiments demonstrate that we produce better results in fine-grained motion control and identity preservation than previous methods. Wei Wang 0025, Yushi Lan, Bo Peng 0002, Jing Dong 0003 |
AAAI | 5 |
| 2024 | Mumpy: Multilateral Temporal-view Pyramid Transformer for Video Inpainting Detection
Yuezun Li, Bo Peng 0002, Jiaran Zhou, Huiyu Zhou 0001, Junyu Dong |
BMVC | 3 |
| 2024 | S3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis
Wei Wang 0025, Yifeng Ma 0001, Bo Peng 0002, Yingya Zhang, Jing Dong 0003 |
ECCV (10) | 5 |
| 2024 | Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of ArtifactsabstractHighly realistic AI generated face forgeries known as deepfakes have raised serious social concerns. Although DNN-based face forgery detection models have achieved good performance, they are vulnerable to latest generative methods that have less forgery traces and adversarial attacks. This limitation of generalization and robustness hinders the credibility of detection results and requires more explanations. In this work, we provide counterfactual explanations for face forgery detection from an artifact removal perspective. Specifically, we first invert the forgery images into the StyleGAN latent space, and then adversarially optimize their latent representations with the discrimination supervision from the target detection model. We verify the effectiveness of the proposed explanations from two aspects: (1) Counterfactual Trace Visualization: the enhanced forgery images are useful to reveal artifacts by visually contrasting the original images and two different visualization methods; (2) Transferable Adversarial Attacks: the adversarial forgery images generated by attacking the detection model are able to mislead other detection models, implying the removed artifacts are general. Extensive experiments demonstrate that our method achieves over 90% attack success rate and superior attack transferability. Compared with naive adversarial noise methods, our method adopts both generative and discriminative model priors, and optimize the latent representations in a synthesis-by-analysis way, which forces the search of counterfactual explanations on the natural face manifold. Thus, more general counterfactual traces can be found and better adversarial attack transferability can be achieved. Our code is available at https://github.com/yangli-lab/Artifact-Eraser/. Yang Li 0255, Wei Wang 0025, Ziwen He, Bo Peng 0002, Jing Dong 0003 |
ICME | 5 |
| 2024 | SPI2I: Structure-Preserved Image-to-Image Translation with Diffusion Models
Beibei Dong, Bo Peng 0002, Jing Dong 0003 |
ICPR (21) | 2 |
| 2024 | Mitigating Social Biases in Text-to-Image Diffusion Models via Linguistic-Aligned Attention GuidanceabstractRecent advancements in text-to-image generative models have showcased remarkable capabilities across various tasks. However, these powerful models have revealed the inherent risks of social biases. Such biases can propagate distorted real-world perspectives and spread unforeseen prejudice and discrimination. Current debiasing methods are primarily designed for scenarios with a single individual in the image and exhibit homogenous race or gender when multiple individuals are involved, harming the diversity of social groups within the image. To address this problem, we consider the semantic consistency between text prompts and generated images in text-to-image diffusion models to identify how biases are generated. We propose a novel method to locate where the biases are based on different tokens and then mitigate them for each individual. Specifically, we introduce a Linguistic-aligned Attention Guidance module consisting of Block Voting and Linguistic Alignment, to effectively locate the semantic regions related to biases. Additionally, we employ Fair Inference in these regions to generate fair attributes across arbitrary distributions while preserving the original structural and semantic information. Extensive experiments and analyses demonstrate our method outperforms existing methods for debiasing with multiple individuals across various scenarios. Yueming Lyu, Ziwen He, Bo Peng 0002, Jing Dong 0003 |
ACM Multimedia | 4 |
| 2024 | ST-SBV: Spatial-Temporal Self-Blended Videos for Deepfake Detection
Weinan Guan, Wei Wang 0025, Bo Peng 0002, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 3 |
| 2024 | Artifact feature purification for cross-domain detection of AI-generated images
Zheling Meng, Bo Peng 0002, Jing Dong 0003, Tieniu Tan, Haonan Cheng |
Comput. Vis. Image Underst. | 2 |
| 2024 | InfoStyler: Disentanglement Information Bottleneck for Artistic Style TransferabstractArtistic style transfer aims to transfer the style of an artwork to a photograph while maintaining its original overall content. Many prior works focus on designing various transfer modules to transfer the style statistics to the content image. Although effective, ignoring the clear disentanglement of the content features and the style features from the first beginning, they have difficulty in balancing between content preservation and style transferring. To tackle this problem, we propose a novel information disentanglement method, named InfoStyler, to capture the minimal sufficient information for both content and style representations from the pre-trained encoding network. InfoStyler formulates the disentanglement representation learning as an information compression problem by eliminating style statistics from the content image and removing the content structure from the style image. Besides, to further facilitate disentanglement learning, a cross-domain Information Bottleneck (IB) learning strategy is proposed by reconstructing the content and style domains. Extensive experiments demonstrate that our InfoStyler can synthesize high-quality stylized images while balancing content structure preservation and style pattern richness. Yueming Lyu, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Improving Generalization of Deepfake Detectors by Imposing Gradient RegularizationabstractThe rapid development of face forgery technology has posed a significant threat to information security. While deepfake detection has proven to be an effective countermeasure, it often struggles to detect fake images generated by unknown forgery methods. Thus, the generalization ability of deepfake detectors to unseen forgery data is a critical concern. Despite many efforts aimed at discovering new forgery artifacts, they often fail to generalize to new manipulation technologies. In this paper, we tackle this challenge by focusing on the difference in texture patterns between training forgeries and unseen forgeries, which can lead to a degradation of generalization. Based on this principle, we propose a new conjecture that encourages deepfake detectors to reduce their sensitivity to forgery texture patterns, thereby improving the detection performance. To this end, we introduce an additional gradient regularization term to the original empirical loss during training. However, computing the Hessian matrix in the gradient calculation process of the regularization term poses a computational complexity. In order to overcome this issue, we optimize the formulation of the gradient regularization term using a first-order approximation method based on Taylor expansion and design a Perturbation Injection Module (PIM) to simplify the implementation process. Additionally, we provide a theoretical analysis from an optimization perspective and explore an interesting aspect of our method. Extensive experiments demonstrate the effectiveness of our approach in improving the generalization ability of deepfake detectors. Importantly, our method is orthogonal to recent advancements in powerful backbones and training data augmentation techniques. When combined with other effective techniques, our method achieves state-of-the-art experimental results. Weinan Guan, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | DRAN: Detailed Region-Adaptive Normalization for Conditional Image SynthesisabstractIn recent years, conditional image synthesis has attracted growing attention due to its controllability in the image generation process. Although recent works have achieved realistic results, most of them have difficulty handling fine-grained styles with subtle details. To address this problem, a novel normalization module, named Detailed Region-Adaptive Normalization (DRAN), is proposed. It adaptively learns both fine-grained and coarse-grained style representations. Specifically, we first introduce a multi-level structure, Spatiality-aware Pyramid Pooling, to guide the model to learn coarse-to-fine features. Then, to adaptively fuse different levels of styles, we propose Dynamic Gating, making it possible to adaptively fuse different levels of styles according to different spatial regions. Finally, we collect a new makeup dataset (Makeup-Complex dataset) that contains a wide range of complex makeup styles with diverse poses and expressions. To evaluate the effectiveness and show the general use of our method, we conduct a set of experiments on makeup transfer and semantic image synthesis. Quantitative and qualitative experiments show that equipped with DRAN, simple baseline models are able to achieve promising improvements in complex style transfer and detailed texture synthesis. Yueming Lyu, Peibin Chen, Jingna Sun, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | Designing A 3d-Aware Stylenerf Encoder for Face EditingabstractGAN inversion has been exploited in many face manipulation tasks, but 2D GANs often fail to generate multi-view 3D consistent images. The encoders designed for 2D GANs are not able to provide sufficient 3D information for the inversion and editing. Therefore, 3D-aware GAN inversion is proposed to increase the 3D editing capability of GANs. However, the 3D-aware GAN inversion remains under-explored. To tackle this problem, we propose a 3D-aware (3Da) encoder for GAN inversion and face editing based on the powerful StyleNeRF model. Our proposed 3Da encoder combines a parametric 3D face model with a learnable detail representation model to generate geometry, texture and view direction codes. For more flexible face manipulation, we then design a dual-branch StyleFlow module to transfer the StyleNeRF codes with disentangled geometry and texture flows. Extensive experiments demonstrate that we realize 3D consistent face manipulation in both facial attribute editing and texture transfer. Furthermore, for video editing, we make the sequence of frame codes share a common canonical manifold, which improves the temporal consistency of the edited attributes. Wei Wang 0025, Bo Peng 0002, Jing Dong 0003 |
ICASSP | 3 |
| 2023 | DFGC-VRA: DeepFake Game Competition on Visual Realism AssessmentabstractThis paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark). Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao |
IJCB | 1 |
| 2023 | Visual Realism Assessment for Face-Swap Videos
Xianyun Sun, Beibei Dong, Caiyong Wang, Bo Peng 0002, Jing Dong 0003 |
ICIG (1) | 4 |
| 2023 | Context-Aware Talking-Head Video EditingabstractTalking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head video clips which simultaneously have accurate lip synchronization and motion smoothness. Previous approaches, including 3DMM-based (3D Morphable Model) methods and NeRF-based (Neural Radiance Field) methods, are sub-optimal in that they either require minutes of source videos and days of training time or lack the disentangled control of verbal (e.g., lip motion) and non-verbal (e.g., head pose and expression) representations for video clip insertion. In this work, we fully utilize the video context to design a novel framework for talking-head video editing, which achieves efficiency, disentangled motion control, and sequential smoothness. Specifically, we decompose this framework to motion prediction and motion-conditioned rendering: (1) We first design an animation prediction module that efficiently obtains smooth and lip-sync motion sequences conditioned on the driven speech. This module adopts a non-autoregressive network to obtain context prior and improve the prediction efficiency, and it learns a speech-animation mapping prior with better generalization to novel speech from a multi-identity video dataset. (2) We then introduce a neural rendering module to synthesize the photo-realistic and full-head video frames given the predicted motion sequence. This module adopts a pre-trained head topology and uses only few frames for efficient fine-tuning to obtain a person-specific rendering model. Extensive experiments demonstrate that our method efficiently achieves smoother editing results with higher image quality and lip accuracy using less data than previous methods. Wei Wang 0025, Jun Ling, Bo Peng 0002, Xu Tan 0003, Jing Dong 0003 |
ACM Multimedia | 4 |
| 2023 | 3D-Aware Adversarial Makeup Generation for Facial Privacy ProtectionabstractThe privacy and security of face data on social media are facing unprecedented challenges as it is vulnerable to unauthorized access and identification. A common practice for solving this problem is to modify the original data so that it could be protected from being recognized by malicious face recognition (FR) systems. However, such "adversarial examples" obtained by existing methods usually suffer from low transferability and poor image quality, which severely limits the application of these methods in real-world scenarios. In this paper, we propose a 3D-Aware Adversarial Makeup Generation GAN (3DAM-GAN). which aims to improve the quality and transferability of synthetic makeup for identity information concealing. Specifically, a UV-based generator consisting of a novel Makeup Adjustment Module (MAM) and Makeup Transfer Module (MTM) is designed to render realistic and robust makeup with the aid of symmetric characteristics of human faces. Moreover, a makeup attack mechanism with an ensemble training strategy is proposed to boost the transferability of black-box models. Extensive experiment results on several benchmark datasets demonstrate that 3DAM-GAN could effectively protect faces against various FR models, including both publicly available state-of-the-art models and commercial face verification APIs, such as Face++, Baidu, and Aliyun. Yueming Lyu, Ziwen He, Bo Peng 0002, Yunfan Liu 0001, Jing Dong 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | DFGC 2022: The Second DeepFake Game CompetitionabstractThis paper presents the summary report on our DFGC 2022 competition. The DeepFake is rapidly evolving, and realistic face-swaps are becoming more deceptive and difficult to detect. On the other hand, methods for detecting DeepFakes are also improving. There is a two-party game between DeepFake creators and defenders. This competition provides a common platform for benchmarking the game between the current state-of-the-arts in Deep-Fake creation and detection methods. The main research question to be answered by this competition is the current state of the two adversaries when competed with each other. This is the second edition after the last year's DFGC 2021, with a new, more diverse video dataset, a more realistic game setting, and more reasonable evaluation metrics. With this competition, we aim to stimulate research ideas for building better defenses against the DeepFake threats. We also release our DFGC 2022 dataset contributed by both our participants and ourselves to enrich the DeepFake data resources for the research community (https://github.com/NiCE-X/DFGC-2022). Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Zhen Lei 0001, Siwei Lyu |
IJCB | 1 |
| 2022 | Defending Against Deepfakes with Ensemble Adversarial PerturbationabstractMaliciously manipulated images and videos, represented by prevalent deepfakes, can easily deceive human and mislead the public opinions. A great deal of effort was spent on detecting these fake images or videos. However, these detection methods always encounter various problems in practical applications. Do we have other ways to block the spread of fake image or videos? This motivates us to focus on an emerging interesting topic, disruption of deepfake generation. We propose the ensemble attacks of various types of deepfake models including facial attribute editing, face swapping and face reenactment models. With the help of hard model mining, we boost the attack success rate significantly comparing with the straightforward average ensemble. Extensive experiments demonstrate the proposed approach can successfully disrupt multiple deepfake models simultaneously under white-box or gray-box attack protocols. Weinan Guan, Ziwen He, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
ICPR | 5 |
| 2022 | DesignerGAN: Sketch Your Own PhotoabstractPerson image generation is a challenging problem due to the complexity of human body structure and the richness of clothing texture. Recent works have made great progress on pose transfer by using keypoints, but cannot characterize the personalized shape attributes. Hence, they have limited person image editing ability, especially in respect of shape editing. In this paper, we propose to use sketches as the expression of the target image, which can not only represent the pose and shape simultaneously but is also flexible to manipulate at the semantic level. We propose DesignerGAN, a novel two-stage model for pose transfer and shape-related attributes editing. The first stage predicts the target semantic parsing using the target sketch and obtains parsing feature maps. In the second stage, with the parsing feature maps and the scaled target sketch, we devise a domain-matching spatially-adaptive normalization method to guide target image generation in multi-level. Qualitative and quantitative comparison results demonstrate our method’s superiority over state-of-the-arts on pose transfer. Besides, we achieve flexible person image editing through simple hand-drawings on sketches. Binghao Zhao, Tianxiang Ma, Bo Peng 0002, Jing Dong 0003 |
ICPR | 3 |
| 2022 | Counterfactual Image Enhancement for Explanation of Face Swap Deepfakes
Bo Peng 0002, Siwei Lyu, Wei Wang 0025, Jing Dong 0003 |
PRCV (2) | 1 |
| 2022 | A Unified Framework for High Fidelity Face Swap and Expression ReenactmentabstractFace manipulation techniques improve fast with the development of powerful image generation models. Two particular face manipulation methods, namely face swap and expression reenactment attract much attention for their flexibility and ease to generate high quality synthesis results. Recently, these two subjects are actively studied. However, most existing methods treat the two tasks separately, ignoring their underlying similarity. In this paper, we propose to tackle the two problems within a unified framework that achieves high quality synthesis results. The enabling component for our unified framework is the clean disentanglement of 3D pose, shape, and expression factors and then recombining them for different tasks accordingly. We then use the same set of 2D representations for face swap and expression reenactment tasks that are input to a common image translation model to directly generate the final synthetic images. Once trained, the proposed model can accomplish both face swap and expression reenactment tasks for previously unseen subjects. Comprehensive experiments and comparisons show that the proposed method achieves high fidelity results in multiple aspects, and it is especially good at faithfully preserving source facial shape in the face swap task, and accurately transferring facial movements in the expression reenactment task. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationabstractPose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem, we propose a novel multi-level statistics transfer model, which disentangles and transfers multi-level appearance features from person images and merges them with pose features to reconstruct the source person images themselves. So that the source images can be used as supervision for self-driven person image generation. Specifically, our model extracts multi-level features from the appearance encoder and learns the optimal appearance representation through attention mechanism and attributes statistics. Then we transfer them to a pose-guided generator for re-fusion of appearance and pose. Our approach allows for flexible manipulation of person appearance and pose properties to perform pose transfer and clothes style transfer tasks. Experimental results on the DeepFashion dataset demonstrate our method’s superiority compared with state-of-the-art supervised and unsupervised methods. In addition, our approach also performs well in the wild. Tianxiang Ma, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
CVPR | 2 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 1 |
| 2021 | SOGAN: 3D-Aware Shadow and Occlusion Robust GAN for Makeup TransferabstractIn recent years, virtual makeup applications have become more and more popular. However, it is still challenging to propose a robust makeup transfer method in the real-world environment. Current makeup transfer methods mostly work well on good-conditioned clean makeup images, but transferring makeup that exhibits shadow and occlusion is not satisfying. To alleviate it, we propose a novel makeup transfer method, called 3D-Aware Shadow and Occlusion Robust GAN (SOGAN). Given the source and the reference faces, we first fit a 3D face model and then disentangle the faces into shape and texture. In the texture branch, we map the texture to the UV space and design a UV texture generator to transfer the makeup. Since human faces are symmetrical in the UV space, we can conveniently remove the undesired shadow and occlusion from the reference image by carefully designing a Flip Attention Module (FAM). After obtaining cleaner makeup features from the reference image, a Makeup Transfer Module (MTM) is introduced to perform accurate makeup transfer. The qualitative and quantitative experiments demonstrate that our SOGAN not only achieves superior results in shadow and occlusion situations but also performs well in large pose and expression variations. Yueming Lyu, Jing Dong 0003, Bo Peng 0002, Wei Wang 0025, Tieniu Tan |
ACM Multimedia | 3 |
| 2021 | Learning pose-invariant 3D object reconstruction from single-view images
Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 1 |
| 2019 | An Accurate LSTM Based Video Heart Rate Estimation Method
Mingyun Bian, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
PRCV (3) | 2 |
| 2018 | Image Forensics Based on Planar Contact Constraints of 3D ObjectsabstractStanding objects on planar surfaces are common to see in images, e.g., people on the ground. For most objects to stay stable on the plane, planar contact is a necessary requirement. However, 2D image splicing usually disregards this physical constraint of 3D world, leading to a potential artifact of object not attached to the plane. This paper is the first attempt to use the contact constraint of standing objects as a new clue for image forensics. Accordingly, we propose a novel approach to first reconstruct the 3D poses of standing objects and their supporting plane and then measure the contact conditions for splicing detection. To tackle the problem of unknown object shape for pose estimation, we effectively employ the prior knowledge of 3D morphable model to simultaneously estimate both shape and pose parameters by fitting to image observations. The 3D normal orientation of the supporting plane is estimated given its vanishing line. Dealing with uncertainty factors in estimations, we approximate a distribution of estimates using sampling strategies and then make the final decision. Particularly, we focused our method on the important scenario of human figure splicing detection, and comprehensive experiments on multiple data sets and typical images proved the encouraging effectiveness of the new forensic clue and the proposed approach. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Optimized 3D Lighting Environment Estimation for Image Forgery DetectionabstractImage forgery is becoming a growing threat to information credibility. Among all kinds of image forgeries, photographic composites of human faces have very serious impacts. To combat this kind of forgery, some forensic methods propose to estimate the 3D lighting environments from different faces and investigate the consistency between them. Although they are very effective, existing 3D lighting-based forensic methods are limited by many simplifying assumptions about the surface reflection model, among which convexity and constant reflectance are two critical ones. In this paper, we propose an optimized 3D lighting estimation method by incorporating a more general surface reflection model. In this model, we relax the convexity and constant reflectance assumptions by taking the occlusion geometry and surface texture information into consideration. The proposed reflection model is more general and accurate; hence, it can achieve better lighting estimation accuracy and more reliable discrimination performance. Comprehensive experiments on both synthetic and real data sets validate the correctness and efficacy of the proposed method. Comparisons with two existing 3D lighting-based forensic methods also demonstrate the superiority of the proposed method for detecting face splicing. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Automatic detection of 3D lighting inconsistencies via a facial landmark based morphable modelabstractExisting 3D lighting consistency based forensic methods have some practical problems. They usually require additional images and human labor to reconstruct the 3D face model for lighting estimation, and furthermore, they cannot deal with expressional faces effectively. These drawbacks make them unusable in many practical cases. In this paper, we propose a more practical 3D lighting based forensic method by incorporating a facial landmark based 3D morphable model to efficiently fit the face shape. We also introduce a residual error based algorithm to automatically exclude outliers in lighting estimation. Our proposed method is fully automatic and very efficient compared to previous ones. Also, it does not depend on additional images and has better performance for expressional faces. Experiments on a realistic face dataset with variational lighting conditions indicate the efficacy and superiority of our method. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 1 |