EDBT 2026 Demo / reviewers in the wild / expert
Yesheng Chai
dblp:146/0464
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2024
0000-0003-1307-1601ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generative Transferable Universal Adversarial Perturbation for Combating DeepfakesabstractRecently, Deepfake has posed a significant threat to our digital society. This technology allows for the modification of facial identity, expression, and attributes in facial images and videos. The misuse of Deepfake can invade personal privacy, damage individuals’ reputations, and have serious consequences. To counter this threat, researchers have proposed active defense methods using adversarial perturbation to distort Deepfake products which can hinder the dissemination of false information. However, the existing methods are primarily based on image-specific approaches, which are inefficient for large-scale data. To address these issues, we propose an end-to-end approach to generate universal perturbations for combating Deepfake. To further cope with diverse Deepfakes, we introduce an adaptive balancing strategy to combat multiple models simultaneously. Specifically, for different scenarios, we propose two types of universal perturbations. Disrupting Universal Perturbation (DUP) leads Deepfake models to generate distorted outputs. In contrast, Lapsing Universal Perturbation (LUP) tries to make the output consistent with the original image, allowing the correct information to continue propagating. Experiments demonstrate the effectiveness and better generalization of our proposed perturbation compared with state-of-the-art methods. Consequently, our proposed method offers a powerful and efficient solution for combating Deepfake, which can help preserve personal privacy and prevent reputational damage. Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Yesheng Chai, Jizhong Han |
CSCWD | 6 |
| 2024 | Explainable Deepfake Detection with Human PromptsabstractFacial manipulation techniques pose a significant threat to society due to the prevalence of deepfake content on the internet. While previous efforts have focused on developing accurate deepfake detection models, these models may be limited in real-world scenarios due to the lack of confidence that human analysts have in their results. Therefore, this study presents a novel approach to improve the practicality of deepfake detection models by incorporating human understanding. We propose a human prompt based deepfake detection framework that overlays Human-enhanced artifacts attention onto image artifact attention, which utilizes vision prompts to improve the model’s responsiveness and feedback ability while preserving its precision and generalizability. The deepfake detection model achieves an AUC score of 0.99 on the FaceForensics++ dataset and exhibits graceful generalization when evaluated on the Celeb-DF dataset. Furthermore, the model generates "possible area of manipulation" that provides an intuitive signal to facilitate interpretation of the detection process, bridging the gap between machine and human perception of "fake". Our proposed approach can potentially mitigate the harm caused by deepfakes and provide a more reliable solution for real-world applications. Xiaorong Ma, Zhaoxing Li, Yesheng Chai, Liangjun Zang, Jizhong Han |
CSCWD | 4 |
| 2024 | Real Appearance Modeling for More General Deepfake Detection
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jiao Dai, Jizhong Han, Yesheng Chai |
ECCV (52) | 8 |
| 2024 | ConfR: Conflict Resolving for Generalizable Deepfake DetectionabstractDeepfake detectors often encounter performance degradation when tested on unseen forgery methods. Existing literature tries to capture common features among multiple source forgery domains. However, we show that conflict arises in the shared feature space when each domain expresses domain-specific bias. If left unresolved, this conflict might mislead the model to learn domain-specific features and lead to inferior generalization. In this paper, we propose a new learning approach, Conflict Resolving (ConfR), designed to minimize conflict and learn features that generalize across forgeries. ConfR incorporates two key elements: the Intra-Domain Consistency Preserving (ICP) loss ensures updating consistency within forgery types, and the Inter-Domain Conflict Resolving (ICR) Module resolves updating conflicts between different forgery types. Extensive experiments demonstrate that ConfR significantly improves upon the state-of-the-art method, highlighting its potential for more generalizable deepfake detection. Cai Yu, Xi Wang 0014, Zhaoxing Li, Yesheng Chai, Jiao Dai, Jizhong Han |
ICME | 6 |
| 2024 | HIDD: Human-perception-centric Incremental Deepfake DetectionabstractFacial manipulation techniques pose a significant societal threat due to the widespread dissemination of deepfake content on the internet. Existing efforts for deepfake detection exhibit inadequate generalization performance when encountering unseen or degraded samples. We attribute this limitation to the overfitting of minor forgery patterns and variations in data distribution among disparate datasets. To tackle this issue, we introduce an innovative human-perception-centric incremental deepfake detection framework to enhance the generalization capabilities of deepfake detection models through continuous learning from a limited set of new samples. Firstly, the model leverages human perceptual salience to discern and comprehend significant artifacts, thereby mitigating overfitting to minor features. Subsequently, in the incremental learning process, we utilize multi-perspective knowledge distillation and a replay strategy to maintain the performance of the old model and minimize the feature distance between old and new samples. This comprehensive approach mitigates feature-level overfitting and addresses distribution differences among various datasets in the incremental phase. We conducted thorough experiments on four benchmark datasets (FF++, DFDC-P, CDF2, and DFD), and the experimental results demonstrate the superior performance of our method. Xiaorong Ma, Yesheng Chai, Zhaoxing Li, Jiao Dai, Liangjun Zang, Jizhong Han |
ICME | 4 |
| 2024 | HDDA: Human-perception-centric Deepfake Detection AdapterabstractFacial manipulation techniques pose a significant societal threat due to the prevalent presence of deepfake content online. Current deepfake detection methods demonstrate subpar generalization performance when applied to unseen samples. The cause of this limitation lies in the overfitting of minor forgery patterns and variations in data distribution across different datasets. To tackle this issue, we introduce an innovative Human-perception-centric Deepfake Detection Adapter, namely HDDA, to enhance the generalization ability of deepfake detection models. This adaptation primarily involves two stages. During the pre-training stage, the model utilizes human perception salience to spot significant artifacts, thus reducing overfitting to minor features. In the subsequent fine-tuning stage, we introduce an efficient parameter tuning module named Deepfake Detection Adapter. The Adapter introduces two types of lightweight yet specialized adapter modules to the pre-trained model while keeping the backbone network frozen. It fine-tunes the pre-trained model through the adapter to adapt new and unseen datasets, thereby enhancing generalization. We conducted comprehensive experiments on various standard deepfake detection benchmarks to validate the effectiveness of our approach, particularly in showcasing a compelling advantage under cross-dataset and cross-manipulation settings. Xiaorong Ma, Yesheng Chai, Jiao Dai, Zhaoxing Li, Liangjun Zang, Jizhong Han |
IJCNN | 3 |
| 2024 | Dynamic Mixed-Prototype Model for Incremental Deepfake DetectionabstractThe rapid advancement of deepfake technology poses significant threats to social trust. Although recent deepfake detectors have exhibited promising results on deepfakes of the same type as those present in training, their effectiveness degrades significantly on novel deepfakes crafted by unseen algorithms due to the gap in forgery patterns. Some studies have enhanced detectors by adapting to the continuously emerging deepfakes through incremental learning. Despite the progress, they overlooked the scarcity of novel samples that can easily lead to insufficient learning of forgery patterns. To mitigate this issue, we introduce the Dynamic Mixed-Prototype (DMP) model, which dynamically increases prototypes to adapt to novel deepfakes efficiently. Specifically, the DMP model adopts multiple prototypes to represent both real and fake classes, enabling learning novel patterns by expanding prototypes and jointly retaining knowledge learned in previous prototypes. Furthermore, we propose the Prototype-Guided Replay strategy and Prototype Representation Distillation loss, both of which effectively prevent forgetting learned knowledge based on the prototypical representation of samples. Our method surpasses existing incremental deepfake detectors across four datasets and can generalize to novel deepfakes by learning limited deepfake samples. Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jizhong Han, Yesheng Chai |
ACM Multimedia | 7 |
| 2024 | OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head MotionsabstractOne-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | OPT: One-shot Pose-Controllable Talking Head GenerationabstractOne-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICASSP | 4 |
| 2023 | FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsabstractOne-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICME | 4 |
| 2023 | MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelabstractFace-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ACM Multimedia | 4 |
| 2015 | Formal consistency checking over specifications in natural languages
Rongjie Yan, Chih-Hong Cheng, Yesheng Chai |
DATE | 3 |