EDBT 2026 Demo / reviewers in the wild / expert
Cai Yu
dblp:60/6046
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2024
0000-0002-0659-0967ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Real Appearance Modeling for More General Deepfake Detection
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jiao Dai, Jizhong Han, Yesheng Chai |
ECCV (52) | 2 |
| 2024 | ConfR: Conflict Resolving for Generalizable Deepfake DetectionabstractDeepfake detectors often encounter performance degradation when tested on unseen forgery methods. Existing literature tries to capture common features among multiple source forgery domains. However, we show that conflict arises in the shared feature space when each domain expresses domain-specific bias. If left unresolved, this conflict might mislead the model to learn domain-specific features and lead to inferior generalization. In this paper, we propose a new learning approach, Conflict Resolving (ConfR), designed to minimize conflict and learn features that generalize across forgeries. ConfR incorporates two key elements: the Intra-Domain Consistency Preserving (ICP) loss ensures updating consistency within forgery types, and the Inter-Domain Conflict Resolving (ICR) Module resolves updating conflicts between different forgery types. Extensive experiments demonstrate that ConfR significantly improves upon the state-of-the-art method, highlighting its potential for more generalizable deepfake detection. Cai Yu, Xi Wang 0014, Zhaoxing Li, Yesheng Chai, Jiao Dai, Jizhong Han |
ICME | 3 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 1 |
| 2024 | Dynamic Mixed-Prototype Model for Incremental Deepfake DetectionabstractThe rapid advancement of deepfake technology poses significant threats to social trust. Although recent deepfake detectors have exhibited promising results on deepfakes of the same type as those present in training, their effectiveness degrades significantly on novel deepfakes crafted by unseen algorithms due to the gap in forgery patterns. Some studies have enhanced detectors by adapting to the continuously emerging deepfakes through incremental learning. Despite the progress, they overlooked the scarcity of novel samples that can easily lead to insufficient learning of forgery patterns. To mitigate this issue, we introduce the Dynamic Mixed-Prototype (DMP) model, which dynamically increases prototypes to adapt to novel deepfakes efficiently. Specifically, the DMP model adopts multiple prototypes to represent both real and fake classes, enabling learning novel patterns by expanding prototypes and jointly retaining knowledge learned in previous prototypes. Furthermore, we propose the Prototype-Guided Replay strategy and Prototype Representation Distillation loss, both of which effectively prevent forgetting learned knowledge based on the prototypical representation of samples. Our method surpasses existing incremental deepfake detectors across four datasets and can generalize to novel deepfakes by learning limited deepfake samples. Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jizhong Han, Yesheng Chai |
ACM Multimedia | 2 |
| 2024 | OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head MotionsabstractOne-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning to Discover Forgery Cues for Face Forgery DetectionabstractLocating manipulation maps,i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objects have also been widely adopted as auxiliary tasks to improve the classification performance of detectors whereas they require comparisons between paired real and forged faces to obtain manipulation maps as supervision. This requirement restricts their applicability to unpaired faces and contradicts real-world scenarios. Moreover, the used comparison methods annotate all changed pixels, including noise introduced by compression and upsampling. Using such maps as supervision hinders the learning of exploitable cues and makes models prone to overfitting. To address these issues, we introduce a weakly supervised model in this paper, named Forgery Cue Discovery (FoCus), to locate forgery cues in unpaired faces. Unlike some detectors that claim to locate forged regions in attention maps, FoCus is designed to sidestep their shortcomings of capturing partial and inaccurate forgery cues. Specifically, we propose a classification attentive regions proposal module to locate forgery cues during classification and a complementary learning module to facilitate the learning of richer cues. The produced manipulation maps can serve as better supervision to enhance face forgery detectors. Visualization of the manipulation maps of the proposed FoCus exhibits superior interpretability and robustness compared to existing methods. Experiments on five datasets and four multi-task models demonstrate the effectiveness of FoCus in both in-dataset and cross-dataset evaluations. Cai Yu, Xiaomeng Fu, Xi Wang 0014, Jiao Dai, Jizhong Han |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | OPT: One-shot Pose-Controllable Talking Head GenerationabstractOne-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICASSP | 5 |
| 2023 | FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsabstractOne-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICME | 5 |
| 2023 | MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelabstractFace-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ACM Multimedia | 5 |
| 2022 | Focus by Prior: Deepfake Detection Based on Prior-AttentionabstractNowadays advanced facial manipulation techniques produce deepfake videos more realistically, which makes deepfake detection more difficult. To capture subtle and intricate artifacts, recent works attempt to enhance low-level textural information by attention-based framework. However, these methods require complex simulated data or extra supervision. Highly dependent on training settings, these methods not only have high training costs but also are prone to overfitting. To address this issue, we propose a novel perspective of deepfake detection via so-called prior-attention. Specifically, we introduce prior textural information, such as edge and noise, to model the attention maps explicitly. Benefiting from these natural “attention maps”, our model significantly enhances discriminative information without additional supervision. Furthermore, we design a Feature Abstraction Block (FAB) to facilitate cross-layer features interaction and insert it into distinct layers of CNN to detect the inconsistencies at multiple spatial levels. Extensive experiments demonstrate that our method achieves performance comparable to state-of-the-art methods. Cai Yu, Jiao Dai, Xi Wang 0014, Weibo Zhang, Jin Liu 0020, Jizhong Han |
ICME | 1 |
| 2021 | Li-Net: Large-Pose Identity-Preserving Face Reenactment NetworkabstractFace reenactment is a challenging task, as it is difficult to maintain accurate expression, pose and identity simultaneously. Most existing methods directly apply driving facial landmarks to reenact source faces and ignore the intrinsic gap between two identities, resulting in the identity mismatch issue. Besides, they neglect the entanglement of expression and pose features when encoding driving faces, leading to inaccurate expressions and visual artifacts on large-pose reenacted faces. To address these problems, we propose a Large-pose Identity-preserving face reenactment network, LI-Net. Specifically, the Landmark Transformer is adopted to adjust driving landmark images, which aims to narrow the identity gap between driving and source landmark images. Then the Face Rotation Module and the Expression Enhancing Generator decouple the transformed landmark image into pose and expression features, and reenact those attributes separately to generate identity-preserving faces with accurate expressions and poses. Both qualitative and quantitative experimental results demonstrate the superiority of our method. Jin Liu 0020, Zhaoxing Li, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 5 |
| 2021 | DLFMNet: End-to-End Detection and Localization of Face Manipulation Using Multi-Domain FeaturesabstractRecently, more and more realistic facial manipulation images and videos, known as DeepFakes, have been created and rapidly circulated in social media. Therefore, it is crucial to develop effective and efficient methods to detect the malicious DeepFakes. Previous approaches all adopt a two-step pipeline with multiple separate models, i.e., first face detection and then face forensics, and lacks robustness against compressed data. In this paper, we propose an end-to-end framework for detection and localization of face manipulation, named DLFMNet, which effectively integrates face detection and face forensics into one model, avoiding intermediate processes like image cropping and feature re-extraction. In addition, to capture richer and more robust manipulated clues, we exploit multi-domain features that takes advantages of two different but complementary domains (i.e., RGB and noise). The evaluations on FaceForensics++ dataset demonstrate the effectiveness of our proposed DLFMNet. https://github.com/LightningChan/DLFMNet. Jin Liu 0020, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 4 |