VLDB 2026 Research / reviewers in the wild / expert
Yuezun Li
dblp:136/5524
· DBLP profile ↗
47ranked-venue papers
13as first author
36since 2021 · last 2026
0000-0001-9299-1945ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 7 first-author · 17 since 2021Security and privacy · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Active Defense Persistence: A Two-Stage Defense Framework Combining Interruption and Poisoning Against DeepfakeabstractActive defense strategies have been developed to counter the threat of deepfake technology. However, a primary challenge is their lack of persistence, as their effectiveness is often short-lived. Attackers can bypass these defenses by simply collecting protected samples and retraining their models. This means that static defenses inevitably fail when attackers retrain their models, which severely limits practical use. We argue that an effective defense not only distorts forged content but also blocks the model’s ability to adapt, which occurs when attackers retrain their models on protected images. To achieve this, we propose an innovative Two-Stage Defense Framework (TSDF). Benefiting from the intensity separation mechanism designed in this paper, the framework uses dual-function adversarial perturbations to perform two roles. First, it can directly distort the forged results. Second, it acts as a poisoning vehicle that disrupts the data preparation process essential for an attacker’s retraining pipeline. By poisoning the data source, TSDF aims to prevent the attacker’s model from adapting to the defensive perturbations, thus ensuring the defense remains effective long-term. Comprehensive experiments show that the performance of traditional interruption methods degrades sharply when these methods are subjected to adversarial retraining. However, our framework shows a strong dual defense capability, which can improve the persistence of active defense. Our code will be available at https://github.com/vpsg-research/TSDF. Hongrui Zheng, Yuezun Li, Yunfeng Diao, Zhiqing Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionabstractWe describe the Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is nontrivial as forgery-related knowledge is entangled with a wide range of unrelated knowledge. Existing methods treat CLIP merely as a feature extractor, lacking task-specific adaptation, which limits their effectiveness. To address this, we introduce an adapter to learn face forgery traces – the blending boundaries unique to forged faces, guided by task-specific objectives. Then we enhance the CLIP visual tokens with a dedicated interaction strategy that communicates knowledge across CLIP and the adapter. Since the adapter is alongside CLIP, its versatility is highly retained, naturally ensuring strong generalizability in face forgery detection. With only 5.7M trainable parameters, our method achieves a significant performance boost, improving by approximately 7% on average across five standard datasets. We believe the proposed method can serve as a baseline for future CLIP-based face forgery detection methods. The code is available at https://github.com/OUCVAS/ForensicsAdapter. Xinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou, Junyu Dong |
CVPR | 2 |
| 2025 | Where the Devil Hides: Deepfake Detectors Can No Longer Be TrustedabstractWith the advancement of AI generative techniques, Deepfake faces have become incredibly realistic and nearly indistinguishable to the human eye. To counter this, Deep-fake detectors have been developed as reliable tools for assessing face authenticity. These detectors are typically developed on Deep Neural Networks (DNNs) and trained using third-party datasets. However, this protocol raises a new security risk that can seriously undermine the trustfulness of Deepfake detectors: Once the third-party data providers insert poisoned (corrupted) data maliciously, Deepfake detectors trained on these datasets will be injected "backdoors" that cause abnormal behavior when presented with samples containing specific triggers. This is a practical concern, as third-party providers may distribute or sell these triggers to malicious users, allowing them to manipulate detector performance and escape accountability.This paper investigates this risk in depth and describes a solution to stealthily infect Deepfake detectors. Specifically, we develop a trigger generator, that can synthesize passcode-controlled, semantic-suppression, adaptive, and invisible trigger patterns, ensuring both the stealthiness and effectiveness of these triggers. Then we discuss two poisoning scenarios, dirty-label poisoning and clean-label poisoning, to accomplish the injection of backdoors. Extensive experiments demonstrate the effectiveness, stealthiness, and practicality of our method compared to several baselines. Shuaiwei Yuan, Junyu Dong, Yuezun Li |
CVPR | 3 |
| 2025 | HRGR: Enhancing Image Manipulation Detection via Hierarchical Region-aware Graph ReasoningabstractImage manipulation detection is to identify the authenticity of each pixel in images. One typical approach to uncover manipulation traces is to model image correlations. The previous methods commonly adopt the grids, which are fixed-size squares, as graph nodes to model correlations. However, these grids, being independent of image content, struggle to retain local content coherence, resulting in imprecise detection. To address this issue, we describe a new method named Hierarchical Region-aware Graph Reasoning (HRGR) to enhance image manipulation detection. Unlike existing grid-based methods, we model image correlations based on content-coherence feature regions with irregular shapes, generated by a novel Differentiable Feature Partition strategy. Then we construct a Hierarchical Region-aware Graph based on these regions within and across different feature layers. Subsequently, we describe a structural-agnostic graph reasoning strategy tailored for our graph to enhance the representation of nodes. Our method is fully differentiable and can seamlessly integrate into mainstream networks in an end-to-end manner, without requiring additional supervision. Extensive experiments demonstrate the effectiveness of our method in image manipulation detection, exhibiting its great potential as a plug-and-play component for existing architectures. Codes and models are available at https://github.com/OUC-VAS/HRGR-IMD. Jiaran Zhou, Huiyu Zhou 0001, Junyu Dong, Yuezun Li |
ICME | 5 |
| 2025 | Texture, Shape and Order Matter: A New Transformer Design for Sequential DeepFake DetectionabstractSequential DeepFake detection is an emerging task that predicts the manipulation sequence in order. Existing methods typically formulate it as an image-to-sequence problem, employing conventional Transformer architectures. However, these methods lack dedicated design and consequently result in limited performance. As such, this paper describes a new Transformer design, called TSOM, by exploring three perspectives: Texture, Shape, and Order of Manipulations. Our method features four major improvements: we describe a new texture-aware branch that effectively captures subtle manipulation traces with a Diversiform Pixel Difference Attention module. Then we introduce a Multi-source Cross-attention module to seek deep correlations among spatial and sequential features, enabling effective modeling of complex manipulation traces. To further enhance the cross-attention, we describe a Shape-guided Gaussian mapping strategy, providing initial priors of the manipulation shape. Finally, observing that the subsequent manipulation in a sequence may influence traces left in the preceding one, we intriguingly invert the prediction order from forward to backward, leading to notable gains as expected. Extensive experimental results demonstrate that our method outperforms others by a large margin, highlighting the superiority of our method. Yuezun Li, Xin Wang 0068, Baoyuan Wu, Jiaran Zhou, Junyu Dong |
WACV | 2 |
| 2025 | UWStereo: A Large Synthetic Dataset for Underwater Stereo MatchingabstractDespite recent advances in stereo matching, the extension to intricate underwater settings remains unexplored, primarily owing to: 1) the reduced visibility, low contrast, and other adverse effects of underwater images; 2) the difficulty in obtaining ground truth data for training deep learning models, i.e. simultaneously capturing an image and estimating its corresponding pixel-wise depth information in underwater environments. To enable further advance in underwater stereo matching, we introduce a large synthetic dataset called UWStereo. Our dataset includes 29,568 synthetic stereo image pairs with dense and accurate disparity annotations for left view. We design four distinct underwater scenes filled with diverse objects such as corals, ships and robots. We also induce additional variations in camera model, lighting, and environmental effects. In comparison with existing underwater datasets, UWStereo is superior in terms of scale, variation, annotation, and photo-realistic image quality. To substantiate the efficacy of the UWStereo dataset, we undertake a comprehensive evaluation compared with eleven state-of-the-art algorithms as benchmarks. The results indicate that current models still struggle to generalize to new domains. Hence, we design a new strategy that learns to reconstruct cross domain masked images before stereo matching training and integrate a cross view attention enhancement module that aggregates long-range content information to enhance the generalization ability. Qingxuan Lv, Junyu Dong, Yuezun Li, Sheng Chen 0001, Hui Yu 0001, Shu Zhang 0002, Wenhan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PhyTracker: An Online Tracker for PhytoplanktonabstractPhytoplankton, a crucial component of aquatic ecosystems, requires efficient monitoring to understand marine ecological processes and environmental conditions. Traditional phytoplankton monitoring methods, relying on non-in situ observations, are time-consuming and resource-intensive, limiting timely analysis. To address these limitations, we introduce PhyTracker, an intelligent in situ tracking framework designed for automatic tracking of phytoplankton. PhyTracker overcomes significant challenges unique to phytoplankton monitoring, such as constrained mobility within water flow, inconspicuous appearance, and the presence of impurities. Our method incorporates three innovative modules: a Texture-enhanced Feature Extraction (TFE) module, an Attention-enhanced Temporal Association (ATA) module, and a Flow-agnostic Movement Refinement (FMR) module. These modules enhance feature capture, differentiate between phytoplankton and impurities, and refine movement characteristics, respectively. Extensive experiments on the PMOT dataset validate the superiority of PhyTracker in phytoplankton tracking, and additional tests on the MOT dataset demonstrate its general applicability, outperforming conventional tracking methods. This work highlights key differences between phytoplankton and traditional objects, offering an effective solution for phytoplankton monitoring. Qingxuan Lv, Yuezun Li, Zhiqiang Wei 0002, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face DetectionabstractFace-swapping DeepFakes have become an escalating societal concern, attracting increasing attention in recent years. To counter this, we investigate a new proactive defense framework to prevent individuals from being victimized in DeepFake videos. The core idea of this framework is to contaminate the inputs of DeepFake models by disrupting face detectors, based on the observation that face detectors are commonly used to automatically extract victim faces in most DeepFake techniques. Once the face detectors malfunction, the faces will not be correctly extracted, thereby impairing the training or synthesis stages of DeepFake models. To achieve this, we describe a strategy named FacePoison, which fools face detectors by adding dedicated adversarial perturbations to video frames. Building upon this, we introduce VideoFacePoison, an extended strategy that can efficiently propagate FacePoison across video frames instead of applying it individually to each frame, thus significantly reducing the computational overhead while retaining favorable attack performance. This framework is validated on five face detectors, and extensive experiments against eleven different DeepFake models demonstrate the effectiveness of disrupting face detectors to hinder DeepFake generation. The source code is publicly available at: https://github.com/OUC-VAS/FacePoison. Delong Zhu 0002, Yuezun Li, Baoyuan Wu, Jiaran Zhou, Zhibo Wang 0001, Siwei Lyu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | HG-SFDA: HyperGraph Learning Meets Source-Free Unsupervised Domain AdaptationabstractSource-Free unsupervised Domain Adaptation (SFDA) aims to classify target samples by only accessing a pre-trained source model and unlabelled target samples. Since no source data is available, transferring the knowledge from the source domain to the target domain is challenging. Existing methods normally exploit the pair-wise relation among target samples and attempt to discover their correlations by clustering these samples based on semantic features. The drawbacks of these methods include: 1) the pair-wise relation is limited to exposing the underlying correlations of two more samples, hindering the exploration of the structural information embedded in the target domain; and 2) the clustering process only relies on the semantic feature, while overlooking the critical effect of domain shift, i.e., the distribution differences between the source and target domains. To address these issues, we propose a new SFDA method that exploits the high-order neighborhood relation and explicitly takes the domain shift effect into account. Specifically, we formulate the SFDA as a hypergraph learning problem and construct hyperedges to explore the deep structural and context information among multiple samples. Moreover, we integrate a self-loop strategy into the constructed hypergraph to elegantly introduce the domain uncertainty of each sample. By clustering these samples based on hyperedges, both the semantic feature and domain shift effects are considered. We then describe an adaptive relation-based objective to tune the model with soft attention levels for all samples. Extensive experiments are conducted on Office-31, Office-Home, VisDA, DomainNet-126 and PointDA-10 datasets. The results demonstrate the superiority of our method over state-of-the-art counterparts. Our code is avaliable at https://github.com/OUC-POVA/HG-SFDA. Jinkun Jiang, Qingxuan Lv, Yuezun Li, Yong Du 0003, Junyu Dong, Sheng Chen 0001, Hui Yu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | DPL: Cross-Quality DeepFake Detection via Dual Progressive Learning
Jiaran Zhou, Yuezun Li |
ACCV (6) | 4 |
| 2024 | FastForensics: Efficient Two-Stream Design for Real-Time Image Manipulation Detection
Yangxiang Zhang, Yuezun Li, Ao Luo, Jiaran Zhou, Junyu Dong |
BMVC | 2 |
| 2024 | Mumpy: Multilateral Temporal-view Pyramid Transformer for Video Inpainting Detection
Yuezun Li, Bo Peng 0002, Jiaran Zhou, Huiyu Zhou 0001, Junyu Dong |
BMVC | 2 |
| 2024 | Fake It till You Make It: Curricular Dynamic Forgery Augmentations Towards General Deepfake Detection
Yuzhen Lin, Wentang Song, Bin Li 0011, Yuezun Li, Jiangqun Ni, Qiushi Li 0001 |
ECCV (86) | 4 |
| 2024 | Enhancing Adversarial Robustness of DNNS Via Weight Decorrelation in TrainingabstractDeep Neural Networks (DNNs) are vulnerable to adversarial perturbations, raising significant concerns about their security. Numerous methods have been proposed to enhance DNN robustness. However, many methods, including adversarial training and noise injection, improve robustness by incorporating external data into the network. Exploring the network’s inherent potential is crucial to improve adversarial robustness. Inspired by principles in physical chemistry, where increased disorder leads to greater energetic stability, we introduce the Weight Decorrelation Loss. This method is simple but effective, enhancing robustness by disrupting the feature space’s ordered structure. The proposed loss achieves substantial performance improvements and state-of-the-art performance after being combined with Gaussian noise. We conduct comprehensive experiments on five datasets, comparing our approach to state-of-the-art defense methods. The results demonstrate our method’s effectiveness against several powerful white-box and black-box attacks. Yuezun Li, Honggang Qi, Siwei Lyu |
ICASSP | 2 |
| 2024 | Dynamic Soft Labeling for Visual Semantic EmbeddingabstractVisual Semantic Embedding (VSE) is a prominent approach in image-text retrieval, aiming to learn a deep embedding space that aligns visual data with semantic text labels. However, current VSE methods oversimplify the retrieval task, treating it as a binary classification problem with triplet loss constraints. This ignores the semantic correlation between pairs of mismatched samples and fails to capture the similarity gradient between samples. In addition, hard constraints on negative samples with high semantic relevance can be detrimental to the model's representational capabilities. To address these limitations, we propose a novel training strategy that introduces dynamic soft labels without additional annotations. This captures the correlation between positive and negative sample pairs and guides feature representation learning using the Soft Negative Alignment Loss (SNAL). SNAL fully takes into account the influence by similar negative samples, enhancing the representation of cross-modal data. In addition, we propose the Stepwise Negative Decoupling Loss (SNDL) to increase the distance between positive and negative samples. Stepwise decoupling of negative samples can be adaptively distanced based on their semantic relevance to the anchor, resulting in a wider distribution of sample features in the common space. Experiments on Flickr30K and MS-COCO datasets validate the effectiveness of our dynamic soft labeling (DSL) methods, demonstrating the importance of considering complex relationships between sample pairs and the limitations of rigid negative sample categorization based on subjective annotations. Jiaao Yu 0001, Yunlai Ding, Junyu Dong, Yuezun Li |
ICMR | 4 |
| 2024 | FreqBlender: Enhancing DeepFake Detection by Blending Frequency KnowledgeabstractGenerating synthetic fake faces, known as pseudo-fake faces, is an effective way to improve the generalization of DeepFake detection. Existing methods typically generate these faces by blending real or fake faces in spatial domain. While these methods have shown promise, they overlook the simulation of frequency distribution in pseudo-fake faces, limiting the learning of generic forgery traces in-depth. To address this, this paper introduces {\em FreqBlender}, a new method that can generate pseudo-fake faces by blending frequency knowledge. Concretely, we investigate the major frequency components and propose a Frequency Parsing Network to adaptively partition frequency components related to forgery traces. Then we blend this frequency knowledge from fake faces into real faces to generate pseudo-fake faces. Since there is no ground truth for frequency components, we describe a dedicated training strategy by leveraging the inner correlations among different frequency knowledge to instruct the learning process. Experimental results demonstrate the effectiveness of our method in enhancing DeepFake detection, making it a potential plug-and-play strategy for other methods. Jiaran Zhou, Yuezun Li, Baoyuan Wu, Bin Li 0011, Junyu Dong |
NeurIPS | 3 |
| 2024 | LandmarkBreaker: A proactive method to obstruct DeepFakes via disrupting facial landmark extraction
Yuezun Li, Pu Sun 0001, Honggang Qi, Siwei Lyu |
Comput. Vis. Image Underst. | 1 |
| 2024 | AdaNI: Adaptive Noise Injection to improve adversarial robustness
Yuezun Li, Honggang Qi, Siwei Lyu |
Comput. Vis. Image Underst. | 1 |
| 2024 | Multiview adaptive attention pooling for image-text retrieval
Yunlai Ding, Jiaao Yu 0001, Qingxuan Lv, Junyu Dong, Yuezun Li |
Knowl. Based Syst. | 6 |
| 2024 | COMICS: End-to-End Bi-Grained Contrastive Learning for Multi-Face Forgery DetectionabstractDeepFakes have raised serious societal concerns, leading to a great surge in detection-based forensics methods in recent years. Face forgery recognition is a standard detection method that usually follows a two-phase pipeline,i.e., it extracts the face first and then determines its authenticity by classification. While those methods perform well in ideal experimental environment, they face challenges when dealing with DeepFakes in the wild involving complex background and multiple faces of varying sizes. Moreover, most face forgery recognition methods can only process one face at a time. One straightforward way to address this issue is to simultaneous process multi-face by integrating face extraction and forgery detection in an end-to-end fashion by adapting advanced object detection architectures. However, as these object detection architectures are designed to capture the discriminative features of different object categories rather than the subtle forgery traces among the faces, the direct adaptation suffers from limited representation ability. In this paper, we propose Contrastive Multi-FaceForensics (COMICS), an end-to-end framework for multi-face forgery detection. COMICS integrates face extraction and forgery detection in a seamless manner and adapts to the advanced object detection architectures. The core of the proposed framework is a bi-grained contrastive learning approach that explores face forgery traces at both the coarse- and fine-grained levels. Specifically, coarse-grained level contrastive learning captures the discriminative features among positive and negative proposal pairs at multiple layers produced by the proposal generator, and the fine-grained level contrastive learning captures the pixel-wise discrepancy between the forged and original areas of the same face and the pixel-wise content inconsistency among different faces. Extensive experiments on the OpenForensics and FFIW datasets demonstrate that our method outperforms other counterparts and shows great potential for being integrated into various architectures. Codes are available at https://github.com/zhangconghhh/COMICS. Honggang Qi, Shuhui Wang, Yuezun Li, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | ForensicsForest Family: A Series of Multi-Scale Hierarchical Cascade Forests for Detecting GAN-Generated FacesabstractThe prominent progress in generative models has significantly improved the authenticity of generated faces, raising serious concerns in society. To combat GAN-generated faces, many countermeasures based on Convolutional Neural Networks (CNNs) have been spawned due to their strong learning capabilities. In this paper, we rethink this problem and explore a new approach based on forest models instead of CNNs. Concretely, we describe a simple and effective forest-based method set, termed ForensicsForest Family, to detect GAN-generate faces. The ForensicsForest family is composed of three variants: ForensicsForest, Hybrid ForensicsForest and Divide-and-Conquer ForensicsForest. ForenscisForest is a novel Multi-scale Hierarchical Cascade Forest that takes appearance, frequency, and biological features as input, hierarchically cascades different levels of features for authenticity prediction, and employs a multi-scale ensemble scheme to consider different levels of information comprehensively for further performance improvement. Building upon ForensicsForest, we create Hybrid ForensicsForest, an extended version that integrates the CNN layers into models, to further enhance the efficacy of augmented features. Furthermore, to reduce memory usage during training, we introduce Divide-and-Conquer ForensicsForest, which can construct a forest model using only a portion of training samplings. In the training stage, we train several candidate forest models using the subsets of training samples. Then, a ForensicsForest is assembled by selecting suitable components from these candidate forest models. Our method is validated on state-of-the-art GAN-generated face datasets and compared with several CNN models, demonstrating the surprising effectiveness of our method in detecting GAN-generated faces. Jiucui Lu, Jiaran Zhou, Junyu Dong, Bin Li 0011, Siwei Lyu, Yuezun Li |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | DomainForensics: Exposing Face Forgery Across Domains via Bi-Directional AdaptationabstractRecent DeepFake detection methods have shown excellent performance on public datasets but are significantly degraded on new forgeries. Solving this problem is important, as new forgeries emerge daily with the continuously evolving generative techniques. Many efforts have been made for this issue by seeking the commonly existing traces empirically on data level. In this paper, we rethink this problem and propose a new solution from the unsupervised domain adaptation perspective. Our solution, called DomainForensics, aims to transfer the forgery knowledge from known forgeries (fully labeled source domain) to new forgeries (label-free target domain). Unlike recent efforts, our solution does not focus on data view but on learning strategies of DeepFake detectors to capture the knowledge of new forgeries through the alignment of domain discrepancies. In particular, unlike the general domain adaptation methods which consider the knowledge transfer in the semantic class category, thus having limited application, our approach captures the subtle forgery traces. We describe a new bi-directional adaptation strategy dedicated to capturing the forgery knowledge across domains. Specifically, our strategy considers both forward and backward adaptation, to transfer the forgery knowledge from the source domain to the target domain in forward adaptation and then reverse the adaptation from the target domain to the source domain in backward adaptation. In forward adaptation, we perform supervised training for the DeepFake detector in the source domain and jointly employ adversarial feature adaptation to transfer the ability to detect manipulated faces from known forgeries to new forgeries. In backward adaptation, we further improve the knowledge transfer by coupling adversarial adaptation with self-distillation on new forgeries. This enables the detector to expose new forgery features from unlabeled data and avoid forgetting the known knowledge of known forgery. Extensive experiments demonstrate that our method is surprisingly effective in exposing new forgeries, and can be plug-and-play on other DeepFake detection architectures. Qingxuan Lv, Yuezun Li, Junyu Dong, Sheng Chen 0001, Hui Yu 0001, Huiyu Zhou 0001, Shu Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Face Poison: Obstructing DeepFakes by Disrupting Face DetectionabstractRecent years have seen fast development in synthesizing realistic human faces using AI-based forgery technique called DeepFake, which can be weaponized to cause negative personal and social impacts. In this work, we develop a defense method, namely FacePosion, to prevent individuals from becoming victims of DeepFake videos by sabotaging would-be training data. This is achieved by disrupting face detection, a prerequisite step to prepare victim faces for training DeepFake model. Once the training faces are wrongly extracted, the DeepFake model can not be well trained. Specifically, we propose a multi-scale feature-level adversarial attack to disrupt the intermediate features of face detectors using different scales. Extensive experiments are conducted on seven various DeepFake models using six face detection methods, empirically showing that disrupting face detectors using our method can effectively obstruct DeepFakes. Yuezun Li, Jiaran Zhou, Siwei Lyu |
ICME | 1 |
| 2023 | Forensics Forest: Multi-scale Hierarchical Cascade Forest for Detecting GAN-generated FacesabstractWe describe a simple and effective method called ForensicsForest to detect GAN-generate faces. Instead of using the commonly used CNN models, we describe a novel multi-scale hierarchical cascade forest, which takes semantic and frequency features as input, and hierarchically cascades different levels of features for authenticity prediction. We then propose a multi-scale ensemble, which comprehensively considers different levels of information to improve the performance further. Our method is validated on state-of-the-art GAN-generated face datasets in comparison with several CNN models, which demonstrates the surprising effectiveness of our method in detecting GAN-generated faces. Jiucui Lu, Yuezun Li, Jiaran Zhou, Bin Li 0011, Siwei Lyu |
ICME | 2 |
| 2023 | Improving transferable adversarial attack via feature-momentum
Xianglong He, Yuezun Li, Haipeng Qu, Junyu Dong |
Comput. Secur. | 2 |
| 2023 | Watching the BiG artifacts: Exposing DeepFake videos via Bi-granularity artifacts
Yuezun Li, Dongdong Lin, Bin Li 0011, Junqiang Wu |
Pattern Recognit. | 2 |
| 2023 | LaFea: Learning Latent Representation Beyond Feature for Universal Domain AdaptationabstractUniversal Domain Adaptation (UniDA) is a recent advent problem that aims to transfer the knowledge from the source domain to the target domain without any prior knowledge on label sets. The main challenge is to separate common samples from private samples in the target domain. In general, existing methods achieve this goal by performing domain adaptation only on the features extracted by the backbone networks. However, solely relying on the learning of the backbone network may not fully exploit the effectiveness of features, due to that 1) the discrepancy between two domains can naturally distract the learning of backbone network and 2) the irrelevant content of samples (e. g., backgrounds) likely goes through the backbone network, and accordingly may hinder the learning of domain-informative features. To this end, we describe a new method to provide extra guidance to the learning of the backbone network based on the latent representation beyond features (LaFea). We are motivated by the fact that the latent representation can be learned to contain the domain-relevant information scattered in features, and the learning of this latent representation can naturally promote the effectiveness of corresponding features in return. To achieve this goal, we develop a simple GAN-style architecture to transform features into the latent representation and propose new objectives to adversarially learn this representation. It should be noted that the latent representation only serves as an auxiliary in training, but it is not needed in inference. Extensive experiments on four datasets corroborate the superiority of our method compared to the state-of-the-arts. Qingxuan Lv, Yuezun Li, Junyu Dong, Ziqian Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Robust Scene Parsing by Mining Supportive Knowledge From DatasetabstractScene parsing, or semantic segmentation, aims at labeling all pixels in an image with the predefined categories of things and stuff. Learning a robust representation for each pixel is crucial for this task. Existing state-of-the-art (SOTA) algorithms employ deep neural networks to learn (discover) the representations needed for parsing from raw data. Nevertheless, these networks discover desired features or representations only from the given image (content), ignoring more generic knowledge contained in the dataset. To overcome this deficiency, we make the first attempt to explore the meaningful supportive knowledge, including general visual concepts (i.e., the generic representations for objects and stuff) and their relations from the whole dataset to enhance the underlying representations of a specific scene for better scene parsing. Specifically, we propose a novel supportive knowledge mining module (SKMM) and a knowledge augmentation operator (KAO), which can be easily plugged into modern scene parsing networks. By taking image-specific content and dataset-level supportive knowledge into full consideration, the resulting model, called knowledge augmented neural network (KANN), can better understand the given scene and provide greater representational power. Experiments are conducted on three challenging scene parsing and semantic segmentation datasets: Cityscapes, Pascal-Context, and ADE20K. The results show that our KANN is effective and achieves better results than all existing SOTA methods. Ao Luo, Fan Yang 0054, Xin Li 0079, Yuezun Li, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Faketracer: Exposing Deepfakes with Training Data ContaminationabstractWe describe a proactive defense method to expose Deep-Fakes with training data contamination. Note that the existing methods usually focus on defending from general DeepFakes, which are synthesized by GAN using random noise. In contrast, our method is dedicated to defending from native Deep-Fakes, which is synthesized by auto-encoder that involves face swapping and encoding-decoding process that general DeepFakes do not have. Specifically, we design two types of traces namely sustainable traces and erasable traces, which are added on the faces to manipulate the training of DeepFake models. Consequently, the trained DeepFake model can synthesize faces with sustainable traces but no erasable traces. With the help of these two traces, we can expose DeepFakes proactively. Our method is compared with recent passive and proactive defense methods, which corroborates the efficacy of our method. Pu Sun 0001, Yuezun Li, Honggang Qi, Siwei Lyu |
ICIP | 2 |
| 2022 | Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
Pattern Recognit. | 4 |
| 2022 | LandmarkGAN: Synthesizing faces from landmarks
Pu Sun 0001, Yuezun Li, Honggang Qi, Siwei Lyu |
Pattern Recognit. Lett. | 2 |
| 2021 | Exposing GAN-Generated Faces Using Inconsistent Corneal Specular HighlightsabstractSophisticated generative adversary network (GAN) models are now able to synthesize highly realistic human faces that are difficult to discern from real ones visually. In this work, we show that GAN synthesized faces can be exposed with the inconsistent corneal specular highlights between two eyes. The inconsistency is caused by the lack of physical/physiological constraints in the GAN models. We show that such artifacts exist widely in high-quality GAN synthesized faces and further describe an automatic method to extract and compare corneal specular highlights from two eyes. Qualitative and quantitative evaluations of our method suggest its simplicity and effectiveness in distinguishing GAN synthesized faces. Shu Hu 0001, Yuezun Li, Siwei Lyu |
ICASSP | 2 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 5 |
| 2021 | Invisible Backdoor Attack with Sample-Specific TriggersabstractRecently, backdoor attacks pose a new security threat to the training process of deep neural networks (DNNs). Attackers intend to inject hidden backdoors into DNNs, such that the attacked model performs well on benign samples, whereas its prediction will be maliciously changed if hidden backdoors are activated by the attacker-defined trigger. Existing backdoor attacks usually adopt the setting that triggers are sample-agnostic, i.e., different poisoned samples contain the same trigger, resulting in that the attacks could be easily mitigated by current backdoor defenses. In this work, we explore a novel attack paradigm, where backdoor triggers are sample-specific. In our attack, we only need to modify certain training samples with invisible perturbation, while not need to manipulate other training components (e.g., training loss, and model structure) as required in many existing attacks. Specifically, inspired by the recent advance in DNN-based image steganography, we generate sample-specific invisible additive noises as backdoor triggers by encoding an attacker-specified string into benign images through an encoder-decoder network. The mapping from the string to the target label will be generated when DNNs are trained on the poisoned dataset. Extensive experiments on benchmark datasets verify the effectiveness of our method in attacking models with or without defenses. The code will be available at https://github.com/yuezunli/ISSBA. Yuezun Li, Yiming Li 0004, Baoyuan Wu, Longkang Li, Ran He 0001, Siwei Lyu |
ICCV | 1 |
| 2021 | Imperceptible Adversarial Examples For Fake Image DetectionabstractFooling people with highly realistic fake images generated with Deepfake or GANs brings a great social disturbance to our society. Many methods have been proposed to detect fake images, but they are vulnerable to adversarial perturbations – intentionally designed noises that can lead to the wrong prediction. Existing methods of attacking fake image detectors usually generate adversarial perturbations to perturb almost the entire image. This is redundant and increases the perceptibility of perturbations. In this paper, we propose a novel method to disrupt the fake image detection by determining key pixels to a fake image detector and attacking only the key pixels, which results in the L0and the L2norms of adversarial perturbations much less than those of existing works. Experiments on two public datasets with three fake image detectors indicate that our proposed method achieves state-of the-art performance in both white-box and black-box attacks. Quanyu Liao, Yuezun Li, Xin Wang 0045, Bin Kong 0001, Bin B. Zhu, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICIP | 2 |
| 2021 | TransRPN: Towards the Transferable Adversarial Perturbations using Region Proposal Networks and Beyond
Yuezun Li, Ming-Ching Chang, Pu Sun 0001, Honggang Qi, Junyu Dong, Siwei Lyu |
Comput. Vis. Image Underst. | 1 |
| 2020 | Celeb-DF: A Large-Scale Challenging Dataset for DeepFake ForensicsabstractAI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for datasets of DeepFake videos. However, current DeepFake datasets suffer from low visual quality and do not resemble DeepFake videos circulated on the Internet. We present a new large-scale challenging DeepFake video dataset, Celeb-DF, which contains 5,639 high-quality DeepFake videos of celebrities generated using improved synthesis process. We conduct a comprehensive evaluation of DeepFake detection methods and datasets to demonstrate the escalated level of challenges posed by Celeb-DF. Yuezun Li, Xin Yang 0008, Pu Sun 0001, Honggang Qi, Siwei Lyu |
CVPR | 1 |
| 2020 | Fast Portrait Segmentation With Highly Light-Weight NetworkabstractIn this paper, we describe a fast and light-weight portrait segmentation method based on a new highly light-weight backbone (HLB) architecture. The core element of HLB is a bottleneck-based factorized block (BFB) that has much fewer parameters than existing alternatives while keeping good learning capacity. Consequently, the HLB-based portrait segmentation method can run faster than the existing methods yet retaining the competitive accuracy performance with state-of-the-arts. Experiments conducted on two benchmark datasets demonstrate the effectiveness and efficiency of our method. Yuezun Li, Ao Luo, Siwei Lyu |
ICIP | 1 |
| 2019 | Graph-to-Graph Energy Minimization for Video Object SegmentationabstractWe describe a new unsupervised video object segmentation (VOS) method based on the graph-to-graph energy minimization, which focuses on exploiting the mutual bootstrapping information between bottom-up (i.e., using pixel/superpixel attributes) and top-down (i.e., using learned appearance and motion cues) processes in a unified framework. Specifically, we construct a graph-to-graph energy function to encode the spatial similarities among superpixels (superpixel-graph) and temporal consistency among regions (region-graph). An efficient heuristic iterative algorithm is used to minimize the energy function to get the optimal assignment of superpixel and region labels to complete the VOS task. Experiments on two challenging benchmarks (i.e., SegTrack v2 and DAVIS) show that the proposed method achieves favorable performance against the state-of-the-art unsupervised VOS methods and comparable performance with the state-of-the-art semi-supervised methods. Yuezun Li, Longyin Wen, Ming-Ching Chang, Siwei Lyu |
AVSS | 1 |
| 2019 | Exploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
Yuezun Li, Xiao Bian, Ming-Ching Chang, Siwei Lyu |
BMVC | 1 |
| 2019 | Exposing Deep Fakes Using Inconsistent Head PosesabstractIn this paper, we propose a new method to expose AI-generated fake face images or videos (commonly known as the Deep Fakes). Our method is based on the observations that Deep Fakes are created by splicing synthesized face region into the original image, and in doing so, introducing errors that can be revealed when 3D head poses are estimated from the face images. We perform experiments to demonstrate this phenomenon and further develop a classification method based on this cue. Using features based on this cue, an SVM classifier is evaluated using a set of real face images and Deep Fakes. Xin Yang 0008, Yuezun Li, Siwei Lyu |
ICASSP | 2 |
| 2019 | De-identification Without Losing FacesabstractTraining of deep learning models for computer vision requires large image or video datasets from real world. Often, in collecting such datasets, we also need to protect the privacy of the people captured in the images or videos, while still preserve useful attributes such as facial expressions. In this work, we describe a new face de-identification method to achieve this, which is based on a face attribute transfer model (FATM). FATM is a deep neural network model trained to map non-identity related facial attributes to the face of donors, who are a small number of consented subjects. Using the donors' faces ensures the natural appearance of the synthesized faces, and FATM blends the donors' facial attributes to those of the original faces to diversify the appearance of the synthesized faces. Experimental results on several sets of images and videos demonstrate the effectiveness of our face de-ID algorithm. Yuezun Li, Siwei Lyu |
IH&MMSec | 1 |
| 2019 | Exposing GAN-synthesized Faces Using Landmark LocationsabstractGenerative adversary networks (GANs) have recently led to highly realistic image synthesis results. In this work, we describe a new method to expose GAN-synthesized images using the locations of the facial landmark points. Our method is based on the observations that the facial parts configuration generated by GAN models are different from those of the real faces, due to the lack of global constraints. We perform experiments demonstrating this phenomenon, and show that an SVM classifier trained using the locations of facial landmark points is sufficient to achieve good classification performance for GAN-synthesized faces. Xin Yang 0008, Yuezun Li, Honggang Qi, Siwei Lyu |
IH&MMSec | 2 |
| 2018 | Pixel Offset Regression (POR) for Single-shot Instance SegmentationabstractState-of-the-art instance segmentation methods including Mask-RCNN and MNC are multi-shot, as multiple region of interest (ROI) forward passes are required to distinguish candidate regions. Multi-shot architectures usually achieve good performance on public benchmarks. However, hundreds of ROI forward passes in sequel limits their running efficiency, which is a critical point in several utilities such as vehicle surveillance. As such, we arrange our focus on seeking a well trade-off between performance and efficiency. In this paper, we introduce a novel Pixel Offset Regression (POR) scheme which can simply extend single-shot object detector to single-shot instance segmentation system, i.e., segmenting all instances in a single pass. Our framework is based on VGG161with following four parts: (1) a single-shot detection branch to generate object detections, (2) a segmentation branch to estimate foreground masks, (3) a pixel offset regression branch to effectively estimate the distance and orientation from each pixel to the respective object center and (4) a merging process combining output of each branch to obtain instances. Our framework is evaluated on Berkeley-BDD, KITTI and PASCAL VOC2012 validation set, with comparison against several VGG16 based multi-shot methods. Without whistles and bells, our framework exhibits decent performance, which shows good potential for fast speed required applications. Yuezun Li, Xiao Bian, Ming-Ching Chang, Longyin Wen, Siwei Lyu |
AVSS | 1 |
| 2018 | Robust Adversarial Perturbation on Deep Proposal-based Models
Yuezun Li, Daniel Tian, Ming-Ching Chang, Xiao Bian, Siwei Lyu |
BMVC | 1 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 6 |
| 2013 | A Background Correction Method Based on Lazy SnappingabstractInteractive segmentation is greatly practical importance in image processing and very useful for selecting objects of interest in images. This is still a topic of much study. In this paper we propose a simple background correction method, it can eliminate some regions which is confused as objects of interest. Our method is suitable for the case that background is rich and foreground has small color differences. This method is based on Lazy Snapping and combining Gaussian Mixture Model (GMM) with K-means. We show that the proposed method increases segmentation accuracy in the same user-provided scribbles and reduce effort on the part of the user. Yuezun Li |
ICIG | 1 |