Ke Sun 0016

dblp:69/476-16 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0003-1868-225XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Towards General Visual-Linguistic Face Forgery Detection
abstract
Face manipulation techniques have achieved significant advances, presenting serious challenges to security and social trust. Recent works demonstrate that leveraging multimodal models can enhance the generalization and interpretability of face forgery detection. However, existing annotation approaches, whether through human labeling or direct Multimodal Large Language Model (MLLM) generation, often suffer from hallucination issues, leading to inaccurate text descriptions, especially for high-quality forgeries. To address this, we propose Face Forgery Text Generator (FFTG), a novel annotation pipeline that generates accurate text descriptions by leveraging forgery masks for initial region and type identification, followed by a comprehensive prompting strategy to guide MLLMs in reducing hallucination. We validate our approach through fine-tuning both CLIP with a three-branch training framework combining unimodal and multimodal objectives, and MLLMs with our structured annotations. Experimental results demonstrate that our method not only achieves more accurate annotations with higher region identification accuracy, but also leads to improvements in model performance across various forgery detection benchmarks. Our Codes are available in https://github.com/skJack/VLFFD.git.
Ke Sun 0016, Shen Chen 0004, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, Rongrong Ji
CVPR1
2025 Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
abstract
The rapid development of AI-generated content (AIGC) technology has led to the misuse of highly realistic AI-generated images (AIGI) in spreading misinformation, posing a threat to public information security. Although existing AIGI detection techniques are generally effective, they face two issues: 1) a lack of human-verifiable explanations, and 2) a lack of generalization in the latest generation technology. To address these issues, we introduce a large-scale and comprehensive dataset, Holmes-Set, which includes the Holmes-SFTSet, an instruction-tuning dataset with explanations on whether images are AI-generated, and the Holmes-DPOSet, a human-aligned preference dataset. Our work introduces an efficient data annotation method called the Multi-Expert Jury, enhancing data generation through structured MLLM explanations and quality control via cross-model evaluation, expert defect filtering, and human preference modification. In addition, we propose Holmes Pipeline, a meticulously designed three-stage training framework comprising visual expert pre-training, supervised fine-tuning, and direct preference optimization. Holmes Pipeline adapts multimodal large language models (MLLMs) for AIGI detection while generating human-verifiable and human-aligned explanations, ultimately yielding our model AIGI-Holmes. During the inference stage, we introduce a collaborative decoding strategy that integrates the model perception of the visual expert with the semantic reasoning of MLLMs, further enhancing the generalization capabilities. Extensive experiments on three benchmarks validate the effectiveness of our AIGI-Holmes.
Ziyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun 0016, Jiayi Ji, Shouhong Ding, Xiaoshuai Sun, Yunsheng Wu, Rongrong Ji
ICCV4
2025 Continual Face Forgery Detection via Historical Distribution Preserving
Ke Sun 0016, Shen Chen 0004, Taiping Yao, Xiaoshuai Sun, Shouhong Ding, Rongrong Ji
Int. J. Comput. Vis.1
2025 Correction: Continual Face Forgery Detection via Historical Distribution Preserving
Ke Sun 0016, Shen Chen 0004, Taiping Yao, Xiaoshuai Sun, Shouhong Ding, Rongrong Ji
Int. J. Comput. Vis.1
2025 Conditional Diffusion Models for Camouflaged and Salient Object Detection
abstract
Camouflaged Object Detection (COD) poses a significant challenge in computer vision, playing a critical role in applications. Existing COD methods often exhibit challenges in accurately predicting nuanced boundaries with high-confidence predictions. In this work, we introduce CamoDiffusion, a new learning method that employs a conditional diffusion model to generate masks that progressively refine the boundaries of camouflaged objects. In particular, we first design an adaptive transformer conditional network, specifically designed for integration into a Denoising Network, which facilitates iterative refinement of the saliency masks. Second, based on the classical diffusion model training, we investigate a variance noise schedule and a structure corruption strategy, which aim to enhance the accuracy of our denoising model by effectively handling uncertain input. Third, we introduce a Consensus Time Ensemble technique, which integrates intermediate predictions using a sampling mechanism, thus reducing overconfidence and incorrect predictions. Finally, we conduct extensive experiments on three benchmark datasets that show that: 1) the efficacy and universality of our method is demonstrated in both camouflaged and salient object detection tasks. 2) compared to existing state-of-the-art methods, CamoDiffusion demonstrates superior performance 3) CamoDiffusion offers flexible enhancements, such as an accelerated version based on the VQ-VAE model and a skip approach.
Ke Sun 0016, Zhongxi Chen, Xianming Lin, Xiaoshuai Sun, Hong Liu 0009, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Camouflaged Object Detection via Dual-branch Fusion and Dual Self-similarity constraints
Haozhe Yang, Ke Sun 0016, Haoyang Ding, Xianming Lin
Pattern Recognit.3
2024 CamoDiffusion: Camouflaged Object Detection via Conditional Diffusion Models
abstract
Camouflaged Object Detection (COD) is a challenging task in computer vision due to the high similarity between camouflaged objects and their surroundings. Existing COD methods struggle with nuanced object boundaries and overconfident incorrect predictions. In response, we propose a new paradigm that treats COD as a conditional mask-generation task leveraging diffusion models. Our method, dubbed CamoDiffusion, employs the denoising process to progressively refine predictions while incorporating image conditions. Due to the stochastic sampling process of diffusion, our model is capable of sampling multiple possible predictions, avoiding the problem of overconfident point estimation. Moreover, we develop specialized network architecture, training, and sampling strategies, to enhance the model’s expressive power, refinement capabilities and suppress overconfident mis-segmentations, thus aptly tailoring the diffusion model to the demands of COD. Extensive experiments on three COD datasets attest to the superior performance of our model compared to existing state-of-the-art methods, particularly on the most challenging COD10K dataset, where our approach achieves 0.019 in terms of MAE. Codes and models are available at https://github.com/Rapisurazurite/CamoDiffusion.
Zhongxi Chen, Ke Sun 0016, Xianming Lin
AAAI2
2024 Enhancing Tampered Text Detection Through Frequency Feature Fusion and Decomposition
Zhongxi Chen, Shen Chen 0004, Taiping Yao, Ke Sun 0016, Shouhong Ding, Xianming Lin, Liujuan Cao, Rongrong Ji
ECCV (33)4
2024 Towards Video-Text Retrieval Adversarial Attack
abstract
Video-text retrieval has widespread applications in economic and security domains, making it crucial to evaluate its robustness through adversarial attack. However, the existing research in this field is inadequate. In this paper, we first introduce adversarial attack to this task. By leveraging the concept of metric learning, we propose novel attack methods Cross-modal Dual Level Contrastive Attack (CDCA) and Cross-modal Rank Pairing Attack (CRPA). In the white-box scenario, CDCA utilizes the distribution of head and tail examples in the retrieval list to form positive and negative example sets, employing both coarse and fine-grained features. In the black-box scenario, CRPA employs the rank difference in retrieval list as example pairs and utilizes the Rank Difference Loss (RDL) as the attack objective function. Experiments validate the superiority of our methods. Furthermore, we contribute a benchmark, which lays a foundation for understanding the vulnerability of multi-modal models.
Haozhe Yang, Yuhan Xiang, Ke Sun 0016, Jianlong Hu, Xianming Lin
ICASSP3
2024 StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
abstract
The rapid progress in generative models has given rise to the critical task of AI-Generated Content Stealth (AIGC-S), which aims to create AI-generated images that can evade both forensic detectors and human inspection. This task is crucial for understanding the vulnerabilities of existing detection methods and developing more robust techniques. However, current adversarial attacks often introduce visible noise, have poor transferability, and fail to address spectral differences between AI-generated and genuine images. To address this, we propose StealthDiffusion, a framework based on stable diffusion that modifies AI-generated images into high-quality, imperceptible adversarial examples capable of evading state-of-the-art forensic detectors. StealthDiffusion comprises two main components: Latent Adversarial Optimization, which generates adversarial perturbations in the latent space of stable diffusion, and Control-VAE, a module that reduces spectral differences between the generated adversarial images and genuine images without affecting the original diffusion model's generation process. Extensive experiments show that StealthDiffusion is effective in both white-box and black-box settings, transforming AI-generated images into high-quality adversarial forgeries with frequency spectra similar to genuine images. These forgeries are classified as genuine by advanced forensic classifiers and are difficult for humans to distinguish.
Ziyin Zhou, Ke Sun 0016, Zhongxi Chen, Huafeng Kuang, Xiaoshuai Sun, Rongrong Ji
ACM Multimedia2
2024 DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion
abstract
The rapid progress of Deepfake technology has made face swapping highly realistic, raising concerns about the malicious use of fabricated facial content. Existing methods often struggle to generalize to unseen domains due to the diverse nature of facial manipulations. In this paper, we revisit the generation process and identify a universal principle: Deepfake images inherently contain information from both source and target identities, while genuine faces maintain a consistent identity. Building upon this insight, we introduce DiffusionFake, a novel plug-and-play framework that reverses the generative process of face forgeries to enhance the generalization of detection models. DiffusionFake achieves this by injecting the features extracted by the detection model into a frozen pre-trained Stable Diffusion model, compelling it to reconstruct the corresponding target and source images. This guided reconstruction process constrains the detection network to capture the source and target related features to facilitate the reconstruction, thereby learning rich and disentangled representations that are more resilient to unseen forgeries. Extensive experiments demonstrate that DiffusionFake significantly improves cross-domain generalization of various detector architectures without introducing additional parameters during inference. The code are available in https://github.com/skJack/DiffusionFake.git.
Ke Sun 0016, Shen Chen 0004, Taiping Yao, Hong Liu 0009, Xiaoshuai Sun, Shouhong Ding, Rongrong Ji
NeurIPS1
2023 InterFormer Real-time Interactive Image Segmentation
abstract
Interactive image segmentation enables annotators to efficiently perform pixel-level annotation for segmentation tasks. However, the existing interactive segmentation pipeline suffers from inefficient computations of interactive models because of the following two issues. First, annotators’ later click is based on models’ feedback of annotators’ former click. This serial interaction is unable to utilize model’s parallelism capabilities. Second, in each interaction step, the model handles the invariant image along with the sparse variable clicks, resulting in a process that’s highly repetitive and redundant. For efficient computations, we propose a method named InterFormer that follows a new pipeline to address these issues. In-terFormer extracts and preprocesses the computationally time-consuming part i.e. image processing from the existing process. Specifically, InterFormer employs a large vision transformer (ViT) on high-performance devices to prepro-cess images in parallel, and then uses a lightweight module called interactive multi-head self attention (I-MSA) for interactive segmentation. Furthermore, the I-MSA module’s deployment on low-power devices extends the practical application of interactive segmentation. The I-MSA module utilizes the preprocessed features to efficiently response to the annotator inputs in real-time. The experiments on several datasets demonstrate the effectiveness of Inter-Former, which outperforms previous interactive segmentation models in terms of computational efficiency and segmentation quality, achieve real-time high-quality interactive segmentation on CPU-only devices. The code is available at https://github.com/YouHuang67/InterFormer.
You Huang, Ke Sun 0016, Shengchuan Zhang, Liujuan Cao, Guannan Jiang, Rongrong Ji
ICCV3
2022 Dual Contrastive Learning for General Face Forgery Detection
abstract
With various facial manipulation techniques arising, face forgery detection has drawn growing attention due to security concerns. Previous works always formulate face forgery detection as a classification problem based on cross-entropy loss, which emphasizes category-level differences rather than the essential discrepancies between real and fake faces, limiting model generalization in unseen domains. To address this issue, we propose a novel face forgery detection framework, named Dual Contrastive Learning (DCL), which specially constructs positive and negative paired data and performs designed contrastive learning at different granularities to learn generalized feature representation. Concretely, combined with the hard sample selection strategy, Inter-Instance Contrastive Learning (Inter-ICL) is first proposed to promote task-related discriminative features learning by especially constructing instance pairs. Moreover, to further explore the essential discrepancies, Intra-Instance Contrastive Learning (Intra-ICL) is introduced to focus on the local content inconsistencies prevalent in the forged faces by constructing local region pairs inside instances. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art competitors. Our Code is available at https://github.com/Tencent/TFace.git.
Ke Sun 0016, Taiping Yao, Shen Chen 0004, Shouhong Ding, Rongrong Ji
AAAI1
2022 An Information Theoretic Approach for Attention-Driven Face Forgery Detection
Ke Sun 0016, Hong Liu 0009, Taiping Yao, Xiaoshuai Sun, Shen Chen 0004, Shouhong Ding, Rongrong Ji
ECCV (14)1
2021 Domain General Face Forgery Detection by Learning to Weight
abstract
In this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making it challenging to detect fake faces in unseen domains. We argue that different faces contribute differently to a detection model trained on multiple domains, making the model likely to fit domain-specific biases. As such, we propose the LTW approach based on the meta-weight learning algorithm, which configures different weights for face images from different domains. The LTW network can balance the model's generalizability across multiple domains. Then, the meta-optimization calibrates the source domain's gradient enabling more discriminative features to be learned. The detection ability of the network is further improved by introducing an intra-class compact loss. Extensive experiments on several commonly used deepfake datasets to demonstrate the effectiveness of our method in detecting synthetic faces. Code and supplemental material are available at https://github.com/skJack/LTW.
Ke Sun 0016, Hong Liu 0009, Qixiang Ye, Yue Gao 0002, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji
AAAI1