Tianyi Wang 0006

dblp:88/8398-6 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0003-2920-6099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation Localization
abstract
The rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in sub-optimal performance. To address this limitation, we propose the Focus Region Discrepancy Network (FRD-Net), a novel and efficient framework that significantly enhances manipulation localization by amplifying discrepancies at both macro- and micro-levels. Specifically, our proposed Iterative Clustering Module (ICM) groups features into two discriminative clusters and refines representations via backward propagation from cluster centers, improving the distinction between tampered and authentic regions at the macro level. Thereafter, our Differential Progressive Module (DPM) is constructed to capture fine-grained structural inconsistencies within local neighborhoods and integrate them into a Central Difference Convolution, increasing sensitivity to subtle manipulation details at the micro level. Finally, these complementary modules are seamlessly integrated into a compact architecture that achieves a favorable balance between accuracy and efficiency. Extensive experiments on multiple benchmarks demonstrate that FRD-Net consistently surpasses state-of-the-art methods in terms of manipulation localization performance while maintaining a lower computational cost.
Shenghao Chen, Yibo Zhao 0001, Tianyi Wang 0006, Chunjie Ma, Weili Guan, Ming Li 0083, Zan Gao 0001
AAAI3
2026 Scaffolding thought: Imposing logical structure on LLMs with knowledge graphs for counterfactual generation
Jiasheng Si, Yingjie Zhu, Yeqing Teng, Rui Wang 0043, Tianyi Wang 0006, Weiyu Zhang 0001, Chaoqun Zheng, Wenpeng Lu
Knowl. Based Syst.5
2026 Prior-knowledge guidance and dual-domain representation refinement for deepfake detection
Zhiyong Cheng 0001, Tianyi Wang 0006, Chunjie Ma, Yibo Zhao 0001, Zan Gao 0001
Pattern Recognit.3
2026 Contrastive Representation Learning for Cross-Domain Blood Cell Image Classification With Denoising Mechanism
abstract
Accurate identification and classification of white blood cells are essential for diagnosing hematological malignancies and analyzing blood disorders. Existing approaches predominantly leverage masked autoencoders (MAEs) to extract intrinsic blood cell features through image reconstruction as a pretext task. However, these methods encounter two critical challenges: (1) their generalization performance deteriorates under domain shifts caused by variations in staining techniques, illumination conditions, and microscope settings, and (2) the learned data distribution often deviates from the true distribution of blood cell features. To overcome these limitations, we propose CD-CBC, a novel framework for cross-domain blood cell image classification that integrates contrastive representation learning with a denoising mechanism. CD-CBC consists of two key components: a LoRA-based segmentation anything model (LoRA-SAM) and a contrastive masked autoencoder (CMAE). LoRA-SAM mitigates shortcut learning in contrastive learning by eliminating background noise and platelet interference, while CMAE captures fine-grained semantic features and models spatial relationships, enhancing cross-domain robustness. Additionally, we introduce a denoising mechanism in the latent space, which guides the model to focus on unmasked patches during reconstruction, allowing it to better capture the true distribution of blood cell features. Extensive experiments on two benchmark blood cell datasets demonstrate that CD-CBC achieves superior cross-domain performance, reaching an average accuracy of 62.47%, which is 3.17% higher than the current state-of-the-art, thereby confirming its strong generalization capability.
Renyu Fu, Chengfu Ji, Sen Xiang, Guanghui Yue 0001, Tianyi Wang 0006, Chang Tang
IEEE J. Biomed. Health Informatics7
2026 Towards Generalizable Deepfake Detection by Primary Region Regularization
abstract
The existing deepfake detection methods have reached a bottleneck in generalizing to unseen forgeries and manipulation approaches. Based on the observation that the deepfake detectors exhibit a preference for overfitting specific primary regions in input, this article enhances the generalization capability from a novel regularization perspective. This can be simply achieved by augmenting the images through primary region removal, thereby preventing the detector from over-relying on data bias. Our method consists of two stages, namely the static localization for primary region maps, as well as the dynamic exploitation of primary region masks. The proposed method can be seamlessly integrated into different backbones without affecting their inference efficiency. We conduct extensive experiments over five widely used deepfake datasets—DFDC, DF-1.0, Celeb-DF, WildDF, and FFIW with seven backbones. Our method demonstrates an average performance improvement of 6% across different backbones and performs competitively with several state-of-the-art baselines.
Harry Cheng 0002, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.3
2025 NullSwap: Proactive Identity Cloaking Against Deepfake Face Swapping
Tianyi Wang 0006, Shuaicheng Niu, Harry Cheng 0002, Yinglong Wang 0001
ICCV1
2025 Self-Bootstrapping for Versatile Test-Time Adaptation
abstract
In this paper, we seek to develop a versatile test-time adaptation (TTA) objective for a variety of tasks — classification and regression across image-, object-, and pixel-level predictions. We achieve this through a self-bootstrapping scheme that optimizes prediction consistency between the test image (as target) and its deteriorated view. The key challenge lies in devising effective augmentations/deteriorations that: i) preserve the image’s geometric information, e.g., object sizes and locations, which is crucial for TTA on object/pixel-level tasks, and ii) provide sufficient learning signals for TTA. To this end, we analyze how common distribution shifts affect the image’s information power across spatial frequencies in the Fourier domain, and reveal that low-frequency components carry high power and masking these components supplies more learning signals, while masking high-frequency components can not. In light of this, we randomly mask the low-frequency amplitude of an image in its Fourier domain for augmentation. Meanwhile, we also augment the image with noise injection to compensate for missing learning signals at high frequencies, by enhancing the information power there. Experiments show that, either independently or as a plug-and-play module, our method achieves superior results across classification, segmentation, and 3D monocular detection tasks with both transformer and CNN models.
Shuaicheng Niu, Peilin Zhao, Tianyi Wang 0006, Zhiqi Shen 0001
ICML4
2025 Learning Real Facial Concepts for Independent Deepfake Detection
abstract
Deepfake detection models often struggle with generalization to unseen datasets, manifesting as misclassifying real instances as fake in target domains. This is primarily due to an overreliance on forgery artifacts and a limited understanding of real faces. To address this challenge, we propose a novel approach RealID to enhance generalization by learning a comprehensive concept of real faces while assessing the probabilities of belonging to the real and fake classes independently. RealID comprises two key modules: the Real Concept Capture Module (RealC^2) and the Independent Dual-Decision Classifier (IDC). With the assistance of a Multi-Real Memory, RealC^2 maintains various prototypes for real faces, allowing the model to capture a comprehensive concept of real class. Meanwhile, IDC redefines the classification strategy by making independent decisions based on the concept of the real class and the presence of forgery artifacts. Through the combined effect of the above modules, the influence of forgery-irrelevant patterns is alleviated, and extensive experiments on five widely used datasets demonstrate that RealID significantly outperforms existing state-of-the-art methods, achieving a 1.74% improvement in average accuracy.
Minghui Liu 0001, Harry Cheng 0002, Tianyi Wang 0006, Xin Luo 0006, Xin-Shun Xu
IJCAI3
2025 ALSA: Context-Sensitive Prompt Privacy Preservation in Large Language Models
abstract
The remarkable prompting capability of large language models (LLMs) offers substantial convenience to users across diverse backgrounds. Nevertheless, as the sensitive information within prompts is inevitably exposed to LLMs, caution must be exercised to preserve privacy. Among various studies, text anonymization is considered an effective approach to preventing privacy leakage in prompts through text substitution. However, existing works overemphasize privacy while overlooks preserving contextual integrity, degrading semantic consistency. To address these concerns, this paper introduces a context-sensitive prompt privacy-preserving framework, namely Adaptive Linguistic Sanitization and Anonymization (ALSA). In specific, ALSA incorporates a three-dimensional scoring mechanism to dynamically quantify the substitutability of each word within a prompt by integrating the Privacy Leakage Risk Score (PLRS), the Contextual Information Importance Score (CIIS), and the Task Relevance Score (TRS). Subsequently, a clustering technique is adopted to dynamically determine the threshold for assigning an anonymization action (i.e., Retain, Replace, Encrypt, or Delete) by balancing privacy, semantics, and task relevance. Extensive experiments on five benchmark datasets validate the superiority of ALSA over state-of-the-art baselines in terms of accuracy, privacy preservation, and semantic integrity.
Hongru Ma, Wenpeng Lu, Tianyi Wang 0006, Qi Zhang 0020, Yingjie Zhu, Jiasheng Si
KDD (2)4
2025 FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
abstract
Proactive Deepfake detection via robust watermarks has seen interest ever since passive Deepfake detectors encountered challenges in identifying high-quality synthetic images. However, while demonstrating reasonable detection performance, they lack localization functionality and explainability in detection results. Additionally, the unstable robustness of watermarks can significantly affect the detection performance. In this study, we propose novel fractal watermarks for proactive Deepfake detection and localization, namely FractalForensics. Benefiting from the characteristics of fractals, we devise a parameter-driven watermark generation pipeline that derives fractal-based watermarks and performs one-way encryption of the selected parameters. Subsequently, we propose a semi-fragile watermarking framework for watermark embedding and recovery, trained to be robust against benign image processing operations and fragile when facing Deepfake manipulations in a black-box setting. Moreover, we introduce an entry-to-patch strategy that implicitly embeds the watermark matrix entries into image patches at corresponding positions, achieving localization of Deepfake manipulations. Extensive experiments demonstrate satisfactory robustness and fragility of our approach against common image processing operations and Deepfake manipulations, outperforming state-of-the-art semi-fragile watermarking algorithms and passive detectors for Deepfake detection. Furthermore, by highlighting the areas manipulated, our method provides explainability for the proactive Deepfake detection results.
Tianyi Wang 0006, Harry Cheng 0002, Minghui Liu 0001, Mohan Kankanhalli
ACM Multimedia1
2025 Fair Deepfake Detectors Can Generalize
abstract
Deepfake detection models face two critical challenges: generalization to unseen manipulations and demographic fairness among population groups. However, existing approaches often demonstrate that these two objectives are inherently conflicting, revealing a trade-off between them. In this paper, we, for the first time, uncover and formally define a causal relationship between fairness and generalization. Building on the back-door adjustment, we show that controlling for confounders (data distribution and model capacity) enables improved generalization via fairness interventions. Motivated by this insight, we propose Demographic Attribute-insensitive Intervention Detection (DAID), a plug-and-play framework composed of: i) Demographic-aware data rebalancing, which employs inverse-propensity weighting and subgroup-wise feature normalization to neutralize distributional biases; and ii) Demographic-agnostic feature aggregation, which uses a novel alignment loss to suppress sensitive-attribute signals. Across three cross-domain benchmarks, DAID consistently achieves superior performance in both fairness and generalization compared to several state-of-the-art detectors, validating both its theoretical foundation and practical effectiveness.
Harry Cheng 0002, Minghui Liu 0001, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
NeurIPS4
2025 A Spatial-Frequency Aware Multi-scale Fusion Network for Real-Time Deepfake Detection
Libo Lv, Tianyi Wang 0006, Mengxiao Huang, Ruixia Liu, Yinglong Wang 0001
PRCV (7)2
2025 Traffic-Associated Link Delay Learning for Industrial Internet of Things
abstract
Link delay is a key factor to evaluate and ensure the stringent network service quality required by the Industrial Internet of Things (IIoT). Because link delay is seriously affected by traffic, obtaining link delay features associated with network traffic is important. This article presents a traffic-associated link delay learning solution for the IIoT. In our solution, the network of the IIoT is divided into many local networks and a software-defined network (SDN). Our solution uses low-loaded methods to collect traffic-delay samples, and uses a traffic-interval-based mechanism to solve the traffic-associated delay statistics problem. We present a link traffic-delay model learning method for local networks of the IIoT. This method uses path traffic-delay samples, independent from specific network paradigms. Our solution uses a particular deep neural network structure to explore the information implied in path traffic-delay samples. We also propose a link traffic-delay model learning method for the SDN, which selects source links by a feature-similarity-based method and generates link traffic-delay models based on transfer learning. Our solution evaluates the accuracy of link traffic-delay models, and further improves the models with low accuracy.
Xinchang Zhang 0001, Maoli Wang, Tianyi Wang 0006, Qingliang Liu 0003
IEEE Internet Things J.4
2025 Scene generalization for biomedical fact verification via hierarchical mixture of experts
Jiasheng Si, Yibo Zhao 0007, Weiyu Zhang 0001, Tianyi Wang 0006, Wenpeng Lu
Inf. Sci.5
2025 MFCLIP: Multi-Modal Fine-Grained CLIP for Generalizable Diffusion Face Forgery Detection
abstract
The rapid development of photo-realistic face generation methods has raised significant concerns in society and academia, highlighting the urgent need for robust and generalizable face forgery detection (FFD) techniques. Although existing approaches mainly capture face forgery patterns using image modality, other modalities like fine-grained noises and texts are not fully explored, which limits the generalization capability of the model. In addition, most FFD methods tend to identify facial images generated by GAN, but struggle to detect unseen diffusion-synthesized ones. To address the limitations, we aim to leverage the cutting-edge foundation model, contrastive language-image pre-training (CLIP), to achieve generalizable diffusion face forgery detection (DFFD). In this paper, we propose a novel multi-modal fine-grained CLIP (MFCLIP) model, which mines comprehensive and fine-grained forgery traces across image-noise modalities via language-guided face forgery representation learning, to facilitate the advancement of DFFD. Specifically, we devise a fine-grained language encoder (FLE) that extracts fine global language features from hierarchical text prompts. We design a multi-modal vision encoder (MVE) to capture global image forgery embeddings as well as fine-grained noise forgery patterns extracted from the richest patch, and integrate them to mine general visual forgery traces. Moreover, we build an innovative plug-and-play sample pair attention (SPA) method to emphasize relevant negative pairs and suppress irrelevant ones, allowing cross-modality sample pairs to conduct more flexible alignment. Extensive experiments and visualizations show that our model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Our code will be available at https://github.com/Jenine-321/MFCLIP.
Tianyi Wang 0006, Zitong Yu, Zan Gao 0001, LinLin Shen, Shengyong Chen
IEEE Trans. Inf. Forensics Secur.2
2025 Multi-View Clustering via High-Order Bipartite Graph Learning and Tensor Low-Rank Representation
abstract
Graph-based multi-view clustering methods have demonstrated satisfying performance by effectively capturing relationships among data samples. However, most existing methods primarily emphasize direct pairwise relationships, neglecting the exploration of high-order correlations present within each view. To this end, a novel approach, called multiview clustering via high-order bipartite graph learning and tensor low-rank representation (HBGTLRR), is proposed. Specifically, we first construct high-order bipartite graphs to capture latent relationships and concatenate them into a tensor. By applying tensor nuclear norm (TNN) minimization, we obtain a low-rank representation that reduces noise and preserves high-order consistency. Subsequently, a consensus graph is constructed by adaptively fusing the high-order bipartite graphs with corresponding weights, and then a Laplacian low-rank constraint is imposed on it to effectively capture the intrinsic data structure. Finally, extensive experimental results show that HBGTLRR significantly outperforms existing methods, thereby validating the effectiveness of our proposed method.
Chuan Tang, Miaomiao Li 0001, Jun Wang 0118, Chang Tang, Jiahe Jiang, Tianyi Wang 0006, En Zhu, Xinwang Liu 0002
IEEE Trans. Knowl. Data Eng.6
2025 IoT-Dedup: Device Relationship-Based IoT Data Deduplication Scheme
abstract
The cyclical and continuous working characteristics ofInternet of Things(IoT) devices make a large amount of the same or similar data, which can significantly consume storage space. To solve this problem, various secure data deduplication schemes have been proposed. However, existing deduplication schemes only perform deduplication based on data similarity, ignoring the internal connection among devices, making the existing schemes not directly applicable to parallel and distributed scenarios like IoT. Furthermore, since secure data deduplication leads to multiple users sharing same encryption key, which may lead to security issues. To this end, we propose a device relationship-based IoT data deduplication scheme that fully considers the IoT data characteristics and devices internal connections. Specifically, we propose a device relationship prediction approach, which can obtain device collaborative relationships by clustering the topology of their communication graph, and classifies the data types based on device relationships to achieve data deduplication with different security levels. Then, we design a similarity-preserving encryption algorithm, so that the security level of encryption key is determined by the data type, ensuring the security of the deduplicated data. In addition, two different data deduplication methods, identical deduplication and similar deduplication, have been designed to meet the privacy requirement of different data types, improving the efficiency of deduplication while ensuring data privacy as much as possible. We evaluate the performance of our scheme using five real datasets, and the results show that our scheme has favorable results in terms of both deduplication performance and computational cost.
Yuan Gao 0034, Liquan Chen, Jianchang Lai, Tianyi Wang 0006, Shui Yu 0001
IEEE Trans. Parallel Distributed Syst.4
2024 Diffusion Facial Forgery Detection
abstract
Detecting diffusion-generated images has recently developed as an emerging research area. Existing diffusion-based datasets predominantly focus on general image generation. However, facial forgeries, which pose severe social risks, have remained less explored thus far. To address this gap, this paper introduces DiFF, a comprehensive dataset dedicated to face-focused diffusion-generated images. DiFF comprises over 500,000 images that are synthesized using thirteen distinct generation methods under four conditions. In particular, this dataset utilizes 30,000 carefully collected textual and visual prompts, ensuring the synthesis of images with both high fidelity and semantic consistency. We conduct extensive experiments on the DiFF dataset via human subject tests and several representative forgery detection methods. The results demonstrate that the binary detection accuracies of both human observers and automated detectors often fall below 30%, revealing insights on the challenges in detecting diffusion-generated facial forgeries. Moreover, our experiments demonstrate that DiFF, compared to previous facial forgery datasets, contains a more diverse and realistic range of forgeries, showcasing its potential to aid in the development of more generalized detectors. Finally, we propose an edge graph regularization approach to effectively enhance the generalization capability of existing detectors.
Harry Cheng 0002, Tianyi Wang 0006, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia3
2024 LampMark: Proactive Deepfake Detection via Training-Free Landmark Perceptual Watermarks
abstract
Deepfake facial manipulation has garnered significant public attention due to its impacts on enhancing human experiences and posing privacy threats. Despite numerous passive algorithms that have been attempted to thwart malicious Deepfake attacks, they mostly struggle with the generalizability challenge when confronted with hyper-realistic synthetic facial images. To tackle the problem, this paper proposes a proactive Deepfake detection approach by introducing a novel training-free landmark perceptual watermark, LampMark for short. We first analyze the structure-sensitive characteristics of Deepfake manipulations and devise a secure and confidential transformation pipeline from the structural representations, i.e. facial landmarks, to binary landmark perceptual watermarks. Subsequently, we present an end-to-end watermarking framework that imperceptibly and robustly embeds and extracts watermarks concerning the images to be protected. Relying on promising watermark recovery accuracies, Deepfake detection is accomplished by assessing the consistency between the content-matched landmark perceptual watermark and the robustly recovered watermark of the suspect image. Experimental results demonstrate the superior performance of our approach in watermark recovery and Deepfake detection compared to state-of-the-art methods across in-dataset, cross-dataset, and cross-manipulation scenarios.
Tianyi Wang 0006, Mengxiao Huang, Harry Cheng 0002, Zhiqi Shen 0001
ACM Multimedia1
2024 GenFace: A Large-Scale Fine-Grained Face Forgery Benchmark and Cross Appearance-Edge Learning
abstract
The rapid advancement of photorealistic generators has reached a critical juncture where the discrepancy between authentic and manipulated images is increasingly indistinguishable. Thus, benchmarking and advancing techniques detecting digital manipulation become an urgent issue. Although there have been a number of publicly available face forgery datasets, the forgery faces are mostly generated using GAN-based synthesis technology, which does not involve the most recent technologies like diffusion. The diversity and quality of images generated by diffusion models have been significantly improved and thus a much more challenging face forgery dataset shall be used to evaluate SOTA forgery detection literature. In this paper, we propose a large-scale, diverse, and fine-grained high-fidelity dataset, namely GenFace, to facilitate the advancement of deepfake detection, which contains a large number of forgery faces generated by advanced generators such as the diffusion-based model and more detailed labels about the manipulation approaches and adopted generators. In addition to evaluating SOTA approaches on our benchmark, we design an innovative Cross Appearance-Edge Learning (CAEL) detector to capture multi-grained appearance and edge global representations, and detect discriminative and general forgery traces. Moreover, we devise an Appearance-Edge Cross-Attention (AECA) module to explore the various integrations across two domains. Extensive experiment results and visualizations show that our detection model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Code and datasets will be available athttps://github.com/Jenine-321/GenFace.
Zitong Yu, Tianyi Wang 0006, Xiaobin Huang, LinLin Shen, Zan Gao 0001, Jianfeng Ren
IEEE Trans. Inf. Forensics Secur.3
2024 An Efficient Attribute-Preserving Framework for Face Swapping
abstract
By leveraging deep neural networks, recent face swapping techniques have performed admirably in generating faces that maintain consistent identities. Nevertheless, while these methods accurately transfer source identities, they often struggle to preserve important attributes (such as head poses, expressions, and gaze directions) in the target faces. As a consequence, the current research in this domain has not resulted in satisfactory performance. In this paper, we propose an efficient attribute-preserving framework, called AP-Swap, for short, for face swapping. Our approach incorporates two innovative modules designed specifically to preserve critical facial attributes. First, we propose a global residual attribute-preserving encoder (GRAPE), which adaptively extracts globally complete attribute features from target faces. Second, in addition to the regular network streams for the source and target facial images, we introduce a network stream that takes into account the facial landmarks of the target faces. This additional stream enables our landmark-guided feature entanglement module (LFEM), which efficiently preserves fine-grained facial attributes by conducting a landmark-based attribute-preserving (LBAP) operation. Through extensive quantitative and qualitative experiments, we demonstrate the superiority of AP-Swap over other state-ofthe-art methods in terms of facial attribute preservation and model efficiency, along with satisfactory identity consistency performance
Tianyi Wang 0006, Zian Li, Ruixia Liu, Yinglong Wang 0001, Liqiang Nie
IEEE Trans. Multim.1
2024 Voice-Face Homogeneity Tells Deepfake
abstract
Detecting forgery videos is highly desirable due to the abuse of deepfake. Existing detection approaches contribute to exploring the specific artifacts in deepfake videos and fit well on certain data. However, the growing technique on these artifacts keeps challenging the robustness of traditional deepfake detectors. As a result, the development of these approaches has reached a blockage. In this article, we propose to perform deepfake detection from an unexplored voice-face matching view. Our approach is founded on two supporting points: first, there is a high degree of homogeneity between the voice and face of an individual (i.e., they are highly correlated), and second, deepfake videos often involve mismatched identities between the voice and face due to face-swapping techniques. To this end, we develop a voice-face matching method that measures the matching degree between these two modalities to identify deepfake videos. Nevertheless, training on specific deepfake datasets makes the model overfit certain traits of deepfake algorithms. We instead advocate a method that quickly adapts to untapped forgery, with a pre-training then fine-tuning paradigm. Specifically, we first pre-train the model on a generic audio-visual dataset, followed by the fine-tuning on downstream deepfake data. We conduct extensive experiments over three widely exploited deepfake datasets: DFDC, FakeAVCeleb, and DeepfakeTIMIT. Our method obtains significant performance gains as compared to other state-of-the-art competitors. For instance, our method outperforms the baselines by nearly 2%, achieving an AUC of 86.11% on FakeAVCeleb. It is also worth noting that our method already achieves competitive results when fine-tuned on limited deepfake data.
Harry Cheng 0002, Tianyi Wang 0006, Xiaojun Chang, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Noise Based Deepfake Detection via Multi-Head Relative-Interaction
abstract
Deepfake brings huge and potential negative impacts to our daily lives. As the real-life Deepfake videos circulated on the Internet become more authentic, most existing detection algorithms have failed since few visual differences can be observed between an authentic video and a Deepfake one. However, the forensic traces are always retained within the synthesized videos. In this study, we present a noise-based Deepfake detection model, NoiseDF for short, which focuses on the underlying forensic noise traces left behind the Deepfake videos. In particular, we enhance the RIDNet denoiser to extract noise traces and features from the cropped face and background squares of the video image frames. Meanwhile, we devise a novel Multi-Head Relative-Interaction method to evaluate the degree of interaction between the faces and backgrounds that plays a pivotal role in the Deepfake detection task. Besides outperforming the state-of-the-art models, the visualization of the extracted Deepfake forensic noise traces has further displayed the evidence and proved the robustness of our approach.
Tianyi Wang 0006, Kam-Pui Chow
AAAI1
2023 TransFS: Face Swapping Using Transformer
abstract
This paper proposes a Transformer based face swapping model, namely, TransFS. The proposed model mainly solves two current problems of face swapping: 1) the face swapping result does not fully preserve pose and expression of the target face as expected; 2) most of the existing models fail to accomplish high-quality face swapping on high-resolution images. To address these two challenges, we first propose a Cross- Window Face Encoder based on Swin Transformer that learns rich facial features including poses and expressions. Then, we devise an Identity Generator to reconstruct high-resolution images of specific identity with high quality while utilizing the Transformer attention mechanism to increase identity information retention. Finally, a Face Conversion Module is proposed to transform the source identity reconstructed image into the target face image to synthesize the final face swapping result while maintaining the details of pose and expression of the target face. Through extensive experiments, our method not only accomplishes face swapping for low-resolution images with arbitrary identities, but also accomplishes face swapping for high-resolution images. Furthermore, our method achieves the state-of-the-art performance in pose and expression controls compared to other methods.
Tianyi Wang 0006, Anming Dong, Minglei Shu
FG2
2023 Dynamic Facial Expression Recognition in Unconstrained Real-World Scenarios Leveraging Dempster-Shafer Evidence Theory
Tianyi Wang 0006, Shuwang Zhou, Minglei Shu
ICANN (2)2
2023 FAMM: Facial Muscle Motions for Detecting Compressed Deepfake Videos Over Social Networks
abstract
As a face manipulation technique, the misuse of Deepfakes poses potential threats to the state, society, and individuals. Several countermeasures have been proposed to reduce the negative effects produced by Deepfakes. Current detection methods achieve satisfactory performance in dealing with uncompressed videos. However, videos are generally compressed when spread over social networks because of limited bandwidth and storage space, which generates compression artifacts and the detection performance inevitably decreases. Hence, how to effectively identify compressed Deepfake videos over social networks becomes a significant problem in video forensics. In this paper, we propose a facial-muscle-motions-based (FAMM) framework to solve the problem of compressed Deepfake video detection. Specifically, we first locate faces from consecutive frames and extract landmarks from the face images. Then, continuous facial landmarks are utilized to construct facial muscle motion features by modeling the five sensory and face regions. Finally, we fuse the diverse forensic knowledge using Dempster-Shafer theory and provide the final detection results. Furthermore, we demonstrate the effectiveness of FAMM through analyzing mutual information, compression procedure, and facial landmarks for compressed Deepfake videos. Theoretical analyses illustrate that compression does not affect facial muscle motion feature construction and the differences in designed features exist between the real and Deepfake videos. Extensive experimental results conclude that the proposed method outperforms the state-of-the-art methods in detecting compressed Deepfake videos. More importantly, FAMM achieves comparable detection performance on compressed videos that are over real-world social networks.
Xin Liao 0001, Yumei Wang, Tianyi Wang 0006, Xiaoshuai Wu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Deep Convolutional Pooling Transformer for Deepfake Detection
abstract
Recently, Deepfake has drawn considerable public attention due to security and privacy concerns in social media digital forensics. As the wildly spreading Deepfake videos on the Internet become more realistic, traditional detection techniques have failed in distinguishing between real and fake. Most existing deep learning methods mainly focus on local features and relations within the face image using convolutional neural networks as a backbone. However, local features and relations are insufficient for model training to learn enough general information for Deepfake detection. Therefore, the existing Deepfake detection methods have reached a bottleneck to further improve the detection performance. To address this issue, we propose a deep convolutional Transformer to incorporate the decisive image features both locally and globally. Specifically, we apply convolutional pooling and re-attention to enrich the extracted features and enhance efficacy. Moreover, we employ the barely discussed image keyframes in model training for performance improvement and visualize the feature quantity gap between the key and normal image frames caused by video compression. We finally illustrate the transferability with extensive experiments on several Deepfake benchmark datasets. The proposed solution consistently outperforms several state-of-the-art baselines on both within- and cross-dataset experiments.
Tianyi Wang 0006, Harry Cheng 0002, Kam-Pui Chow, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.1
2022 A Robust Lightweight Deepfake Detection Network Using Transformers
Tianyi Wang 0006, Minglei Shu, Yinglong Wang 0001
PRICAI (1)2
2022 Elastic and Reliable Bandwidth Reservation Based on Distributed Traffic Monitoring and Control
abstract
Bandwidth reservation can effectively improve the service quality for data transfers because of dedicated network resources. However, it is difficult to achieve a desired tradeoff between resource utilization and reliable bandwidth guarantees for data transfers with time-varying traffic. In this article, we study a novel bandwidth reservation solution based on distributed traffic monitoring and control for applications that require reliable bandwidth guarantees. In the proposed solution, designated bandwidth is allocated for an application in advance according to its maximum traffic peak, and idle reserved bandwidth resources are dynamically shared according to regular traffic. Dynamic resource sharing evidently improves resource utilization and effectively eliminates the potential congestion caused by sudden traffic bursts. To ensure that the congestion that occurs occasionally can dissipate rapidly, our solution monitors and manages traffic by a distributed monitoring and control strategy. Hence, we study a delay-constrained and proxy-assisted traffic monitoring structure construction problem and propose an algorithm to solve it. The proposed algorithm can also be used to build a delay-constrained traffic control structure. In addition to the above algorithm, we propose a dynamic traffic control algorithm that can achieve a desirable tradeoff between resource utilization and congestion avoidance capability.
Xinchang Zhang 0001, Tianyi Wang 0006
IEEE Trans. Parallel Distributed Syst.2
2019 Automatic Tagging of Cyber Threat Intelligence Unstructured Data using Semantics Extraction
abstract
Threat intelligence, information about potential or current attacks to an organization, is an important component in cyber security territory. As new threats consecutively occurring, cyber security professionals always keep an eye on the latest threat intelligence in order to continuously lower the security risks for their organizations. Cyber threat intelligence is usually conveyed by structured data like CVE entities and unstructured data like articles and reports. Structured data are always under certain patterns that can be easily analyzed, while unstructured data have more difficulties to find fixed patterns to analyze. There exists plenty of methods and algorithms on information extraction from structured data, but no current work is complete or suitable for semantics extraction upon unstructured cyber threat intelligence data. In this paper, we introduce an idea of automatic tagging applying JAPE feature within GATE framework to perform semantics extraction upon cyber threat intelligence unstructured data such as articles and reports. We extract token entities from each cyber threat intelligence article or report and evaluate the usefulness of them. A threat intelligence ontology then can be constructed with the useful entities extracted from related resources and provide convenience for professionals to find latest useful threat intelligence they need.
Tianyi Wang 0006, Kam-Pui Chow
ISI1