EDBT 2026 Demo / reviewers in the wild / expert
Wei Wang 0025
dblp:35/7092-25
· DBLP profile ↗
83ranked-venue papers
5as first author
61since 2021 · last 2026
0000-0002-8598-0831ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 34 since 2021Artificial intelligence and machine learning · 27 · 24 since 2021Security and privacy · 19 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-branch hierarchical feature fusion network for video source camera identification
Bo Wang 0024, Jiaqi Chi, Zhuocheng Wu, Wei Wang 0025 |
Expert Syst. Appl. | 5 |
| 2026 | Dark Miner: Towards combating residuals in concept erasure for text-to-image diffusion models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 5 |
| 2026 | Scalable and Robust Watermarking for Diffusion Models via Task-Decoupled Mixture-of-WatermarksabstractWith the growing deployment of text-to-image diffusion models in real-world applications, concerns about model copyright have become increasingly prominent. Model watermarking has emerged as a promising solution. However, existing methods often suffer from degradation of model fidelity and poor scalability in model distribution scenarios. Moreover, recent studies have shown that some watermarking approaches lack robustness against ambiguity attacks. To address these challenges, we propose TD-MoW, a lightweight and plug-and-play watermarking framework that enables robust and scalable watermark embedding. TD-MoW consists of a two-stage process: In the first stage, we independently optimize the word embeddings of the watermark trigger along with a Low-Rank Adaptation (LoRA) module, effectively decoupling watermark learning from the original generation objective without modifying the base model. In the second stage, inspired by the Mixture-of-Experts (MoE), we introduce the MoW mechanism, which integrates multiple specialized watermark experts into the model. Through encoding optimization for watermarking, we enablenwatermark experts to support up to2nusers, reducing the complexityO(2n) toO(n). Moreover, by leveraging gated activation, the watermark experts remain inactive during standard inference, thereby preserving generation fidelity. Extensive experiments demonstrate that our method not only achieves superior efficiency and generation fidelity compared to prior approaches, but also exhibits strong robustness. Xiaorui Dai, Wei Wang 0025, Bo Wang 0024 |
IEEE Internet Things J. | 2 |
| 2026 | Stealthy Backdoor Carriers: The Threat of Visual Prompts to CLIPabstractVisual Prompt (VP) learning has rapidly emerged as a popular paradigm for parameter-efficient task adaptation in CLIP-based models. However, while VP optimizes pixel-space vectors without altering CLIP’s internal weights, this decoupled design inadvertently introduces critical security vulnerabilities. Attackers can exploit VP to implant covert backdoors using imperceptible trigger patterns that bypass traditional anomaly detection mechanisms, posing significant risks to real-world applications. Existing backdoor techniques, however, rely on visually noticeable patterns and exhibit inconsistencies in representation alignment, which make them prone to detection and ineffective against robust defenses. In response to these limitations, we propose Stealthy Backdoor Carriers (SBC), a novel attack framework that leverages CLIP’s inherent vulnerabilities to covertly and persistently inject backdoors.SBCadopts a dual-constrained optimization strategy that balances imperceptibility—minimizing trigger perturbations for visual stealth—and cross-modal embedding alignment—ensuring poisoned and target samples share consistent representations within CLIP’s multimodal space. Experimental results across five benchmark datasets demonstrateSBC’s exceptional effectiveness, achieving a +49.63% improvement in attack success rate relative to existing methods while maintaining robustness against advanced defenses like Neural Cleanse. Our work highlights the need for reevaluating the security implications of VP learning frameworks and provides valuable insights for mitigating prompt-based vulnerabilities in AI systems. Our code is available at https://github.com/Maozhen-Zhang/sbc.git. Maozhen Zhang, Mengnan Zhao 0001, Wei Wang 0025, Bo Wang 0024 |
IEEE Internet Things J. | 3 |
| 2026 | AMF-CFL: Anomaly model filtering based on clustering in federated learning
Bo Wang 0024, Xiaorui Dai, Wei Wang 0025, Zhaoning Wang, Maozhen Zhang |
J. Inf. Secur. Appl. | 3 |
| 2026 | DualVeil: Persistent and invisible backdoor attacks in federated learning via dual optimization
Maozhen Zhang, Mengnan Zhao 0001, Wei Wang 0025, Bo Wang 0024 |
Knowl. Based Syst. | 3 |
| 2026 | DREAM: A Benchmark Study for Deepfake PhotoRealism AssessMentabstractDeep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its computational modeling, imitation, is also a significant problem but lacks adequate study. In this paper, we focus on the photorealism assessment of deepfakes, which is defined as the automatic assessment of deepfake photorealism that approximates human perception of deepfakes. It is important for evaluating the quality, deceptiveness of deepfakes which can be used for predicting the influence of deepfakes on Internet, it also has potentials in improving the deepfake generation process by serving as a critic. This paper promotes this new direction by presenting a comprehensive benchmark called DREAM, which stands for Deepfake photoREalism AssessMent. It is comprised of a deepfake video dataset of diverse quality, a large scale annotation that includes 140, 000 photorealism scores, textual descriptions obtained from 3, 500 human annotators, a comprehensive evaluation, analysis of 18 representative photorealism assessment methods, including recent large vision language model based methods, a newly proposed description-aligned CLIP method. The benchmark, insights included in this study can lay the foundation for future research in this direction, other related areas. Bo Peng 0002, Zichuan Wang, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Probing unlearned diffusion models: A transferable adversarial attack perspective
Xiaoxuan Han, Wei Wang 0025, Yang Li 0255, Jing Dong 0003 |
Pattern Recognit. | 3 |
| 2026 | Beyond data dependency: FedPET enables robust federated learning via data-free dual-teacher knowledge distillation
Bo Wang 0024, Zhaoning Wang, Wei Wang 0025 |
Pattern Recognit. Lett. | 5 |
| 2026 | Model Backdoor Attack on Federated Learning Based on Parameter AnalysisabstractWith the increasingly widespread application of federated learning (FL) in various fields, the issue of backdoor attacks against FL has garnered significant attention from both academia and industry. While there has been some progress in researching backdoor attacks against FL, data backdoor attacks are easily mitigated in federated environments, and model backdoor attacks are susceptible to detection by defense mechanisms. Therefore, we propose a refined backdoor attack method tailored for FL under the image classification task. Our method involves the collaborative operation of three key modules. Firstly, the parameter importance analysis module identifies parameters with minimal impact on model performance, and creates parameter importance masks to provide precise targets for subsequent operations. Subsequently, the activation difference computation module calculates the activation differences between backdoor samples and benign samples to locate trigger-sensitive parameters. Our method implants the backdoor by flipping and zeroing the precisely located layer parameters, while maintaining the model's classification performance on benign samples. Experimental results show that our method is feasible to achieve an average attack success rate more than 99% across the three victim models. This demonstrates the effectiveness of our method in FL environments and its robustness against various FL defense mechanisms. Bo Wang 0024, Maozhen Zhang, Wei Wang 0025, Hongwei Yao |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | MAP-Mamba: Multi-Artifacts Perception Mamba for Generalizable Face Forgery DetectionabstractFace forgery detection suffers from cross-dataset generalization challenges, where performance degradation occurs due to distribution shifts between training and testing data. Recently, pseudo-fake face generation strategy has mitigated models overfitting to specific forgery traces. However, detectors based on this strategy exhibit an overreliance on blending boundary artifacts for their classification decisions. This overreliance significantly limits their ability to generalize to more advanced face manipulation algorithms, such as FaceDancer and InSwap, which are designed to produce smooth and natural transitions in the blending boundary region. To address this, we propose MAP-Mamba, a novel Multi-Artifacts Perception Mamba framework for modeling generalizable artifact representations from “Generation” to “Enrichment” to “Strengthening”. First, we design an attribute-level face blending method that generate pseudo-fake faces containing fine-grained artifacts via three attribute generators. These pseudo-fakes mimic subtle local inconsistencies in advanced forgery algorithms, guiding the MAP-Mamba to learn diverse forgery features beyond the blending boundary artifacts. Second, considering the variability of face artifacts distribution caused by different forgery algorithms, an artifact style mixing strategy is designed to enrich the artifact style distribution in the training phase by mixing and reorganizing the artifact style features, and to enhance the model’s ability to handle unknown forgery methods. Finally, an adaptive artifact guidance mechanism is proposed to dynamically amplify the artifact-related feature to further strengthen the model’s sensitivity to key artifacts. Extensive experiments on several benchmarks show that MAP-Mamba achieves superior robustness and generalization performance. Ziwen He, Xinjue Hu, Weinan Guan, Wei Wang 0025, Zhangjie Fu 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Fast Adversarial Training With Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on VideosabstractAdversarial Training (AT) has been shown to significantly enhance adversarial robustness via a min-max optimization approach. However, its effectiveness in video recognition tasks is hampered by two main challenges. First, fast adversarial training for video models remains largely unexplored, which severely impedes its practical applications. Specifically, most video adversarial training methods are computationally costly, with long training times and high expenses. Second, existing methods struggle with the trade-off between clean accuracy and adversarial robustness. To address these challenges, we introduce Video Fast Adversarial Training with Weak-to-Strong consistency (VFAT-WS), the first fast adversarial training method for video data. Specifically, VFAT-WS incorporates the following key designs: First, it integrates a straightforward yet effective temporal frequency augmentation (TF-AUG), and its spatial-temporal enhanced form STF-AUG, along with Fast Gradient Sign Method (FGSM) to boost training efficiency and robustness. Second, it devises a weak-to-strong spatial-temporal consistency regularization, which seamlessly integrates the simple TF-AUG and the more complex STF-AUG. Leveraging the consistency regularization, it steers the learning process from simple to complex augmentations. Both of them work together to achieve a better trade-off between clean accuracy and robustness. Extensive experiments on UCF-101 and HMDB-51 with both CNN and Transformer-based models demonstrate that VFAT-WS achieves great improvements in adversarial robustness and corruption robustness, while accelerating training by nearly 490%. Songping Wang, Yueming Lyu, Xiantao Hu, Ziwen He, Wei Wang 0025, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Partial Reconstruction Error for Deepfake DetectionabstractThe rapid development of deepfake technology poses a formidable challenge to personal privacy and security, underscoring the urgent need for deepfake detection. Recently, the methods based on the reconstruction error, such as DIRE and RECCE, achieve impressive performance in forgery detection. However, their performance on facial forgery datasets is relatively poor. The reconstruction process is performed on the whole images, neglecting contextual information for reconstruction. In this paper, we propose Partial Reconstruction Error to perform deepfake detection based on the reconstruction of masked regions in an image. In this way, contextual information helps to reveal the inconsistencies between the original and reconstructed regions thereby improving the detection performance. This method outperforms the best global reconstruction-based approaches on the FF++, Celeb-DF, and DiFF datasets by 4.00%, 2.83%, and 2.67%, respectively. Zheling Meng, Bo Peng 0002, Jing Dong 0003, Beilin Chu, Wei Wang 0025 |
ICASSP | 6 |
| 2025 | Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference TimeabstractText-to-image (T2I) diffusion models have raised concerns about generating inappropriate content, such as "nudity". Despite efforts to erase undesirable concepts through unlearning techniques, these unlearned models remain vulnerable to adversarial inputs that can potentially regenerate such content. To safeguard unlearned models, we propose a novel inference-time defense strategy that mitigates the impact of adversarial inputs. Specifically, we first reformulate the challenge of ensuring robustness in unlearned diffusion models as a robust regression problem. Building upon the naive median smoothing for regression robustness, which employs isotropic Gaussian noise, we develop a generalized median smoothing framework that incorporates anisotropic noise. Based on this framework, we introduce a token-wise Adaptive Median Smoothing method that dynamically adjusts noise intensity according to each token’s relevance to target concepts. Furthermore, to improve inference efficiency, we explore implementations of this adaptive method at the text-encoding stage. Extensive experiments demonstrate that our approach enhances adversarial robustness while preserving model utility and inference efficiency, outperforming baseline defense techniques. Xiaoxuan Han, Wei Wang 0025, Yang Li 0255, Jing Dong 0003 |
ICML | 3 |
| 2025 | Concept Corrector: Erase Concepts on the Fly for Text-to-Image Diffusion Models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Yueming Lyu, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 5 |
| 2025 | Robust Adaptive Tracking Control for Aerial Transporting a Cable-Suspended Payload Using Backstepping Sliding Mode TechniquesabstractAerial transportation technology is the lifeline of air disaster rescue. In this article, a robust adaptive tracking control scheme using backstepping sliding mode techniques is proposed for a quadrotor-based aerial transportation system with a cable-suspended payload in disaster rescue, where the payload is ensured to be driven to predefined trajectories in the presence of strong coupling, uncertainties, and external disturbances. The quadrotor and the payload are modeled as a rigid body and a point mass, respectively, and the two coupling terms between the virtual input of the payload position loop and the payload attitude error as well as between the input force and the quadrotor attitude error are analyzed owing to the underactuated of the quadrotor-based transportation system. Then, adaptive backstepping sliding mode control strategies are designed for the position and swing dynamics of the payload to guarantee payload trajectory tracking, and an observer-based geometric attitude control method is presented for the quadrotor attitude dynamics to ensure the global attitude stability of the system, where prior information about disturbances is not required. The closed-loop stability of the whole system is strictly proven. Finally, real-world experiments are conducted to verify the feasibility and robustness of the proposed control scheme.Note to Practitioners—The motivation of this article is to investigate a robust and adaptive control tracking scheme for aerial transportation systems with a cable-suspended payload in disaster rescue. In most of the existing aerial transportation control schemes with a cable-suspended payload, the payload is driven to follow a desired trajectory while only considering the coupling effect between the aerial platform and the payload. However, in practical disaster rescue applications, the aerial transportation system is inevitably affected by strong coupling, uncertainties, and external disturbances. Meanwhile, to the authors’ best knowledge, there exist few studies that investigate payload following issues while considering strong coupling, uncertainties, and external disturbances simultaneously. Therefore, this article proposes a robust and adaptive tracking control scheme using backstepping sliding mode techniques for a quadrotor-based aerial transportation system with a cable-suspended payload to ensure the stable and accurate payload following control under strong coupling, uncertainties, and external disturbances, where prior information about disturbances is not required under the proposed scheme. The closed-loop stability of the whole system is strictly and mathematically analyzed as well as real-world experiments provide promising results. Moreover, the proposed scheme provides a more realistic setup for autonomous aerial transportation with cable-suspended supplies in disaster rescue. Jiacheng Liang, Yaonan Wang 0001, Hang Zhong, Hongwen Li, Hean Hua, Wei Wang 0025 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Beyond Inserting: Learning Subject Embedding for Semantic-Fidelity Personalized Diffusion GenerationabstractText-to-Image (T2I) personalization based on advanced diffusion models (e.g., Stable Diffusion), which aims to generate images of target subjects given various prompts, has drawn huge attention. However, when users require personalized image generation for specific subjects such as themselves or their pet cat, the T2I models fail to accurately generate their subject-preserved images. The main problem is that pre-trained T2I models do not learn the T2I mapping between the target subjects and their corresponding visual contents. Even if multiple target subject images are provided, previous personalization methods either failed to accurately fit the subject region or lost the interactive generative ability with other existing concepts in T2I model space. For example, they are unable to generate T2I-aligned and semantic-fidelity images for the given prompts with other concepts such as scenes (“Eiffel Tower”), actions (“holding a basketball”), and facial attributes (“eyes closed”). In this paper, we focus on inserting accurate and interactive subject embedding into the Stable Diffusion Model for semantic-fidelity personalized generation using one image. We address this challenge from two perspectives: subject-wise attention loss and semantic-fidelity token optimization. Specifically, we propose a subject-wise attention loss to guide the subject embedding onto a manifold with high subject identity similarity and diverse interactive generative ability. Then, we optimize one subject representation as multiple per-stage tokens, and each token contains two disentangled features. This expansion of the textual conditioning space enhances the semantic control, thereby improving semantic-fidelity. We conduct extensive experiments on the most challenging subjects, face identities, to validate that our results exhibit superior subject accuracy and fine-grained manipulation ability. We further validate the generalization of our methods on various non-face subjects. Yang Li 0255, Wei Wang 0025, Jing Dong 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Noise-Informed Diffusion-Generated Image Detection With Anomaly AttentionabstractWith the rapid development of image generation technologies, especially the advancement of Diffusion Models, the quality of synthesized images has significantly improved, raising concerns among researchers about information security. To mitigate the malicious abuse of diffusion models, diffusion-generated image detection has proven to be an effective countermeasure. However, a key challenge for forgery detection is generalising to diffusion models not seen during training. In this paper, we address this problem by focusing on image noise. We observe that images from different diffusion models share similar noise patterns, distinct from genuine images. Building upon this insight, we introduce a novel Noise-Aware Self-Attention (NASA) module that focuses on noise regions to capture anomalous patterns. To implement a SOTA detection model, we incorporate NASA into Swin Transformer, forming an novel detection architecture NASA-Swin. Additionally, we employ a cross-modality fusion embedding to combine RGB and noise images, along with a channel mask strategy to enhance feature learning from both modalities. Extensive experiments demonstrate the effectiveness of our approach in enhancing detection capabilities for diffusion-generated images. When encountering unseen generation methods, our approach achieves the state-of-the-art performance. Weinan Guan, Wei Wang 0025, Bo Peng 0002, Ziwen He, Jing Dong 0003, Haonan Cheng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Exploiting Backdoors of Face Synthesis Detection with Natural TriggersabstractDeep neural networks have enhanced face synthesis detection in discriminating Artificial Intelligence Generated Content (AIGC). However, their security is threatened by the injection of carefully crafted triggers during model training (i.e., backdoor attacks). Although existing backdoor defenses and manual data selection are able to mitigate those using human-eye-sensitive triggers, such as patches or adversarial noises, the more challenging natural backdoor triggers remain insufficiently researched. To further investigate natural triggers, we propose a novel analysis-by-synthesis backdoor attack against face synthesis detection models, which embeds natural triggers in the latent space. We study such backdoor vulnerability from two perspectives: (1) Model Discrimination (Optimization-Based Trigger) : we adopt a substitute detection model and find the trigger by minimizing the cross-entropy loss; (2) Data Distribution (Custom Trigger): we manipulate the uncommon facial attributes in the long-tailed distribution to generate poisoned samples without the supervision from detection models. Furthermore, to evaluate the detection models toward the latest AIGC, we utilize both the state-of-the-art StyleGAN and Stable Diffusion for trigger generation. Finally, these backdoor triggers introduce specific semantic features to the generated poisoned samples (e.g., skin textures and smile), which are more natural and robust. Extensive experiments show that our method is superior over existing pixel space backdoor attacks on three levels: (1) Attack Success Rate : achieving an attack success rate exceeding 99 \(\%\) , comparable to baseline methods, with less than 0.1 \(\%\) model accuracy drop and under 3 \(\%\) poisoning rate; (2) Backdoor Defense : showing superior robustness when faced with existing backdoor defenses (e.g., surpassing baseline methods by over 30 \(\%\) after a 15 \({}^{\circ}\) rotation); (3) Human Inspection : being less human-eye-sensitive from a user study with 46 participants and a collection of 2,300 data points. Xiaoxuan Han, Wei Wang 0025, Ziwen He, Jing Dong 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisabstractAudio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations. Wei Wang 0025, Bo Peng 0002, Yingya Zhang, Jing Dong 0003, Tieniu Tan |
AAAI | 3 |
| 2024 | Learning Dense Correspondence for NeRF-Based Face ReenactmentabstractFace reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment in photo-realism and 3D consistency. However, establishing dense correspondence between different face NeRFs is non-trivial, because implicit representations lack ground-truth correspondence annotations like mesh-based 3D parametric models (e.g., 3DMM) with index-aligned vertexes. Although aligning 3DMM space with NeRF-based face representations can realize motion control, it is sub-optimal for their limited face-only modeling and low identity fidelity. Therefore, we are inspired to ask: Can we learn the dense correspondence between different NeRF-based face representations without a 3D parametric model prior? To address this challenge, we propose a novel framework, which adopts tri-planes as fundamental NeRF representation and decomposes face tri-planes into three components: canonical tri-planes, identity deformations, and motion. In terms of motion control, our key contribution is proposing a Plane Dictionary (PlaneDict) module, which efficiently maps the motion conditions to a linear weighted addition of learnable orthogonal plane bases. To the best of our knowledge, our framework is the first method that achieves one-shot multi-view face reenactment without a 3D parametric model prior. Extensive experiments demonstrate that we produce better results in fine-grained motion control and identity preservation than previous methods. Wei Wang 0025, Yushi Lan, Bo Peng 0002, Jing Dong 0003 |
AAAI | 2 |
| 2024 | S3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis
Wei Wang 0025, Yifeng Ma 0001, Bo Peng 0002, Yingya Zhang, Jing Dong 0003 |
ECCV (10) | 3 |
| 2024 | Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of ArtifactsabstractHighly realistic AI generated face forgeries known as deepfakes have raised serious social concerns. Although DNN-based face forgery detection models have achieved good performance, they are vulnerable to latest generative methods that have less forgery traces and adversarial attacks. This limitation of generalization and robustness hinders the credibility of detection results and requires more explanations. In this work, we provide counterfactual explanations for face forgery detection from an artifact removal perspective. Specifically, we first invert the forgery images into the StyleGAN latent space, and then adversarially optimize their latent representations with the discrimination supervision from the target detection model. We verify the effectiveness of the proposed explanations from two aspects: (1) Counterfactual Trace Visualization: the enhanced forgery images are useful to reveal artifacts by visually contrasting the original images and two different visualization methods; (2) Transferable Adversarial Attacks: the adversarial forgery images generated by attacking the detection model are able to mislead other detection models, implying the removed artifacts are general. Extensive experiments demonstrate that our method achieves over 90% attack success rate and superior attack transferability. Compared with naive adversarial noise methods, our method adopts both generative and discriminative model priors, and optimize the latent representations in a synthesis-by-analysis way, which forces the search of counterfactual explanations on the natural face manifold. Thus, more general counterfactual traces can be found and better adversarial attack transferability can be achieved. Our code is available at https://github.com/yangli-lab/Artifact-Eraser/. Yang Li 0255, Wei Wang 0025, Ziwen He, Bo Peng 0002, Jing Dong 0003 |
ICME | 3 |
| 2024 | ST-SBV: Spatial-Temporal Self-Blended Videos for Deepfake Detection
Weinan Guan, Wei Wang 0025, Bo Peng 0002, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 2 |
| 2024 | PPNet : pooling position attention network for semantic segmentation
Hai-Xia Xu 0001, Wei Wang 0025, Shuailong Wang, Wei Zhou 0027 |
Multim. Tools Appl. | 2 |
| 2024 | TCNet: tensor and covariance attention network for semantic segmentation
Hai-Xia Xu 0001, Yanbang Liu, Wei Wang 0025, Wei Zhou 0027, Fanxun Ding |
Soft Comput. | 3 |
| 2024 | Robust Adversarial Watermark Defending Against GAN Synthesization AttackabstractThe proliferation of facial manipulation has been propelled by generative adversarial networks (GAN), severely threatening to the personal privacy and reputation. Accordingly, one such countermeasure is adversarial watermark, which is embedded into the protected image prior to GAN synthesization attack, resulting into the distorted fake content obtained by malicious attackers. However, in practice, JPEG compression usually causes a remarkable degradation on the performance of adversarial watermark. To address this challengeable issue, this letter presents a novel robust adversarial watermark, which can effectively defend against GAN synthesization attack, even though suffering from JPEG compression. Extensive experiments verify the superiority of our proposed method in the benchmark dataset; more importantly, the robustness of the proposed adversarial watermark is comprehensively evaluated on the both simulated transmission channel and the realism social network platform. Shengwang Xu, Ming Xu 0001, Wei Wang 0025, Ning Zheng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Deep Stereo Network With MRF-Based Cost AggregationabstractDespite the remarkable progress made in learning-based stereo-matching algorithms, it is an open challenge for stereo-matching in disparity discontinuities and textureless regions. In this paper, we propose the deep Markov Random Field based cost aggregation network (DMCA-Net) for stereo matching, which is an end-to-end model-driven network architecture. This architecture introduces an efficient feature extraction network to extract richer textual and contextual feature information for stereo feature similarity representation at multi-stages and levels. Furthermore, with the aim of alleviating the edge-fattening phenomenon at disparity discontinuities and generating accurate disparities in textureless regions, we proposed the differentiable Markov Random Field model for cost aggregation, where the model’s data term utilizes image detail information, such as boundary and contour features, to guide matching cost aggregation, and the model’s smoothness term penalizes the adjacency similarity of the cost between the four-nearest neighboring pixel pairs to predict the disparity in textureless regions. The detailed experiment demonstrates that DMCA network achieves competitive performance on the SceneFlow, KITTI 2012, KITTI 2015, and Middlebury 2014 datasets. Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Improving Generalization of Deepfake Detectors by Imposing Gradient RegularizationabstractThe rapid development of face forgery technology has posed a significant threat to information security. While deepfake detection has proven to be an effective countermeasure, it often struggles to detect fake images generated by unknown forgery methods. Thus, the generalization ability of deepfake detectors to unseen forgery data is a critical concern. Despite many efforts aimed at discovering new forgery artifacts, they often fail to generalize to new manipulation technologies. In this paper, we tackle this challenge by focusing on the difference in texture patterns between training forgeries and unseen forgeries, which can lead to a degradation of generalization. Based on this principle, we propose a new conjecture that encourages deepfake detectors to reduce their sensitivity to forgery texture patterns, thereby improving the detection performance. To this end, we introduce an additional gradient regularization term to the original empirical loss during training. However, computing the Hessian matrix in the gradient calculation process of the regularization term poses a computational complexity. In order to overcome this issue, we optimize the formulation of the gradient regularization term using a first-order approximation method based on Taylor expansion and design a Perturbation Injection Module (PIM) to simplify the implementation process. Additionally, we provide a theoretical analysis from an optimization perspective and explore an interesting aspect of our method. Extensive experiments demonstrate the effectiveness of our approach in improving the generalization ability of deepfake detectors. Importantly, our method is orthogonal to recent advancements in powerful backbones and training data augmentation techniques. When combined with other effective techniques, our method achieves state-of-the-art experimental results. Weinan Guan, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Adversarial Attacks on Scene Graph GenerationabstractScene graph generation (SGG) effectively improves semantic understanding of the visual world. However, the recent interest of researchers focuses on enhancing SGG in non-adversarial settings, which raises our curiosity about the adversarial robustness of SGG models. To bridge this gap, we perform adversarial attacks on two typical SGG tasks, Scene Graph Detection (SGDet) and Scene Graph Classification (SGCls). Specifically, we initially propose a bounding box relabeling method to reconstruct reasonable attack targets for SGCls. It solves the inconsistency between the specified bounding boxes and the scene graphs selected as attack targets. Subsequently, we introduce a two-step weighted attack by removing the predicted objects and relational triples that affect attack performance, which significantly increases the success rate of adversarial attacks on two SGG tasks. Extensive experiments demonstrate the effectiveness of our methods on five popular SGG models and four adversarial attacks. The Pytorch® implementation can be downloaded from an open-source Github project https://github.com/Dlut-lab-zmn/SGG_Attack. Mengnan Zhao 0001, Lihe Zhang, Wei Wang 0025, Yuqiu Kong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Robust Variable Impedance Control for Aerial Compliant Interaction With Stability GuaranteeabstractThis article investigates a robust variable impedance control methodology for aerial manipulators to realize compliant and safe interaction tasks. Considering that the stability characteristics are generally overlooked in existing variable impedance controllers of the aerial manipulator, state-independent stability conditions are applied for time-varying impedance profiles to ensure the exponential stability of the desired variable impedance dynamics (DVID) as well as the boundedness of the state variables in the DVID. A command trajectory variable is introduced for converting the impedance control issue to a particular tracking issue, and then, a robust variable impedance controller based on the wrench estimator is designed to guarantee the exponential convergence of the translational states and impedance error of the aerial manipulator. The designed impedance controller is structurally simple and results in low implementation costs. Next, an improved attitude control approach with the command filter is developed for global flight attitude stability without any singularities or ambiguities, where the filter is introduced to avoid computing the derivative signals of the generalized force input. Finally, the effectiveness of the proposed control method is illustrated via numerical simulations and interaction experiments with different targets in real scenarios. Jiacheng Liang, Yaonan Wang 0001, Hang Zhong, Hongwen Li, Jianxu Mao, Wei Wang 0025 |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | Invisible Intruders: Label-Consistent Backdoor Attack Using Re-Parameterized Noise TriggerabstractAremarkable number of backdoor attack methods have been proposed in the literature on deep neural networks (DNNs). However, it hasn't been sufficiently addressed in the existing methods of achieving true senseless backdoor attacks that are visually invisible and label-consistent. In this paper, we propose a new backdoor attack method where the labels of the backdoor images are perfectly aligned with their content, ensuring label consistency. Additionally, the backdoor trigger is meticulously designed, allowing the attack to evade DNN model checks and human inspection. Our approach employs an auto-encoder (AE) to conduct representation learning of benign images and interferes with salient classification features to increase the dependence of backdoor image classification on backdoor triggers. To ensure visual invisibility, we implement a method inspired by image steganography that embeds trigger patterns into the image using the DNN and enable sample-specific backdoor triggers. We conduct comprehensive experiments on multiple benchmark datasets and network architectures to verify the effectiveness of our proposed method under the metric of attack success rate and invisibility. The results also demonstrate satisfactory performance against a variety of defense methods. Bo Wang 0024, Fei Wei, Yi Li 0018, Wei Wang 0025 |
IEEE Trans. Multim. | 5 |
| 2023 | Designing A 3d-Aware Stylenerf Encoder for Face EditingabstractGAN inversion has been exploited in many face manipulation tasks, but 2D GANs often fail to generate multi-view 3D consistent images. The encoders designed for 2D GANs are not able to provide sufficient 3D information for the inversion and editing. Therefore, 3D-aware GAN inversion is proposed to increase the 3D editing capability of GANs. However, the 3D-aware GAN inversion remains under-explored. To tackle this problem, we propose a 3D-aware (3Da) encoder for GAN inversion and face editing based on the powerful StyleNeRF model. Our proposed 3Da encoder combines a parametric 3D face model with a learnable detail representation model to generate geometry, texture and view direction codes. For more flexible face manipulation, we then design a dual-branch StyleFlow module to transfer the StyleNeRF codes with disentangled geometry and texture flows. Extensive experiments demonstrate that we realize 3D consistent face manipulation in both facial attribute editing and texture transfer. Furthermore, for video editing, we make the sequence of frame codes share a common canonical manifold, which improves the temporal consistency of the edited attributes. Wei Wang 0025, Bo Peng 0002, Jing Dong 0003 |
ICASSP | 2 |
| 2023 | DFGC-VRA: DeepFake Game Competition on Visual Realism AssessmentabstractThis paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark). Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao |
IJCB | 4 |
| 2023 | Context-Aware Talking-Head Video EditingabstractTalking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head video clips which simultaneously have accurate lip synchronization and motion smoothness. Previous approaches, including 3DMM-based (3D Morphable Model) methods and NeRF-based (Neural Radiance Field) methods, are sub-optimal in that they either require minutes of source videos and days of training time or lack the disentangled control of verbal (e.g., lip motion) and non-verbal (e.g., head pose and expression) representations for video clip insertion. In this work, we fully utilize the video context to design a novel framework for talking-head video editing, which achieves efficiency, disentangled motion control, and sequential smoothness. Specifically, we decompose this framework to motion prediction and motion-conditioned rendering: (1) We first design an animation prediction module that efficiently obtains smooth and lip-sync motion sequences conditioned on the driven speech. This module adopts a non-autoregressive network to obtain context prior and improve the prediction efficiency, and it learns a speech-animation mapping prior with better generalization to novel speech from a multi-identity video dataset. (2) We then introduce a neural rendering module to synthesize the photo-realistic and full-head video frames given the predicted motion sequence. This module adopts a pre-trained head topology and uses only few frames for efficient fine-tuning to obtain a person-specific rendering model. Extensive experiments demonstrate that our method efficiently achieves smoother editing results with higher image quality and lip accuracy using less data than previous methods. Wei Wang 0025, Jun Ling, Bo Peng 0002, Xu Tan 0003, Jing Dong 0003 |
ACM Multimedia | 2 |
| 2023 | Deep Stereo Matching with Superpixel Based Feature and Cost
Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
PRCV (2) | 3 |
| 2023 | Protecting by attacking: A personal information protecting method with cross-modal adversarial examples
Mengnan Zhao 0001, Bo Wang 0024, Weikuo Guo, Wei Wang 0025 |
Neurocomputing | 4 |
| 2023 | Identification of image global processing operator chain based on feature decoupling
Xin Liao 0001, Wei Wang 0025, Zheng Qin 0001 |
Inf. Sci. | 3 |
| 2023 | Few-shot learning with unsupervised part discovery and part-aligned similarity
Zhang Zhang 0001, Wei Wang 0025, Liang Wang 0001, Zilei Wang, Tieniu Tan |
Pattern Recognit. | 3 |
| 2023 | Temporal sparse adversarial attack on sequence-based gait recognition
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Pattern Recognit. | 2 |
| 2023 | SNIS: A Signal Noise Separation-Based Network for Post-Processed Image Forgery DetectionabstractImage forgery detection has aroused widespread research interest in both academia and industry because of its potential security threats. Existing forgery detection methods achieve excellent tampered regions localization performance when forged images have not undergone post-processing, which can be detected by observing changes in the statistical features of images. However, forged images may be carefully post-processed to conceal forgery boundaries in a particular scenario. It becomes tough challenging to these methods. In this paper, we perform an analogous analysis between image forgery detection and blind signal separation, and formulate the post-processed image forgery detection problem into a signal noise separation problem. We also propose a signal noise separation-based (SNIS) network to solve the problem of detecting post-processed image forgery. Specifically, we first adopt the signal noise separation module to separate tampered region from the complex background region with post-processing noise, which weakens or even eliminates the negative impact of post-processing on forgery detection. Then, the multi-scale feature learning module uses a parallel atrous convolution architecture to learn high-level global features from multiple perspectives. Besides, a feature fusion module is utilized to enhance the discriminability of tampered regions and real regions by strengthening the boundary information. Finally, the prediction module is designed to predict the tampered region and classify the type of tampering operation. Extensive experiments show that the proposed SNIS is not only effective for forgery detection on forged images without post-processing, but also promising in robustness against multiple post-processing attacks. Furthermore, SNIS is robust in detecting forged images from unknown sources. Xin Liao 0001, Wei Wang 0025, Zhenxing Qian, Zheng Qin 0001, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Deep Confidence Propagation Stereo NetworkabstractStereo matching depth estimation based on rectified image pairs is of great importance to many computer vision tasks such as vehicle navigation and autonomous driving. Confidence measures are typically used to refine stereo matching results, which provides robustness and efficiency for disparity estimation. However, previous learning-based confidence methods for stereo matching usually use the middle results or composition as a post-processing step to refine the stereo matching results. This cannot be optimized end-to-end and the performance is limited by the quality of the tri-modal output. To handle this issue, in this paper, we pursue an end-to-end hierarchical architecture and propose a differentiable confidence propagation (DCP) model of a cost aggregation network for stereo matching. The DCP model is integrated into an end-to-end neural network hierarchical architecture to guide matching cost volume aggregation. More specifically, to better represent the similarity of left and right feature maps, we extract unary context feature maps with an effective attention mechanism for matching cost construction. Moreover, we aggregate the cost volume with the multiple stacked DCP cost aggregation (DCPCA) networks to generate a more reliable and finer cost volume. This network suppresses multi-level disparity maps. Each output disparity is supervised with different training weights to learn in a coarse-to-fine way. Our method outperforms previous methods on the Sceneflow dataset by achieving the$0.6735px$EPE error, achieving 1.53% D1-all metric of Non-occluded pixels regions and 0.72% Non-occluded pixels of$5px$metric on KITTI 2015 and 2012 dataset. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our DCPCA-Net can significantly minimize the trade-off between accuracy and efficiency for stereo matching. Kai Zeng 0010, Yaonan Wang 0001, Wei Wang 0025, Hui Zhang 0023, Jianxu Mao, Qing Zhu 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | DFGC 2022: The Second DeepFake Game CompetitionabstractThis paper presents the summary report on our DFGC 2022 competition. The DeepFake is rapidly evolving, and realistic face-swaps are becoming more deceptive and difficult to detect. On the other hand, methods for detecting DeepFakes are also improving. There is a two-party game between DeepFake creators and defenders. This competition provides a common platform for benchmarking the game between the current state-of-the-arts in Deep-Fake creation and detection methods. The main research question to be answered by this competition is the current state of the two adversaries when competed with each other. This is the second edition after the last year's DFGC 2021, with a new, more diverse video dataset, a more realistic game setting, and more reasonable evaluation metrics. With this competition, we aim to stimulate research ideas for building better defenses against the DeepFake threats. We also release our DFGC 2022 dataset contributed by both our participants and ourselves to enrich the DeepFake data resources for the research community (https://github.com/NiCE-X/DFGC-2022). Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Zhen Lei 0001, Siwei Lyu |
IJCB | 4 |
| 2022 | Defending Against Deepfakes with Ensemble Adversarial PerturbationabstractMaliciously manipulated images and videos, represented by prevalent deepfakes, can easily deceive human and mislead the public opinions. A great deal of effort was spent on detecting these fake images or videos. However, these detection methods always encounter various problems in practical applications. Do we have other ways to block the spread of fake image or videos? This motivates us to focus on an emerging interesting topic, disruption of deepfake generation. We propose the ensemble attacks of various types of deepfake models including facial attribute editing, face swapping and face reenactment models. With the help of hard model mining, we boost the attack success rate significantly comparing with the straightforward average ensemble. Extensive experiments demonstrate the proposed approach can successfully disrupt multiple deepfake models simultaneously under white-box or gray-box attack protocols. Weinan Guan, Ziwen He, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
ICPR | 3 |
| 2022 | Contrastive Knowledge Transfer for Deepfake Detection with Limited DataabstractNowadays forensics methods have shown remarkable progress in detecting maliciously crafted fake images. However, without exception, the training process of deepfake detection models requires a large number of facial images. These models are usually unsuitable for real world applications because of their overlarge size and inferiority in speed. Thus, performing data-efficient deepfake detection is of great importance. In this paper, we propose a contrastive distillation method that maximizes the lower bound of mutual information between the teacher and the student to further improve student’s accuracy in a data-limited setting. We observe that models performing deepfake detection, different from other image classification tasks, have shown high robustness when there is a drop in data amount. The proposed knowledge transfer approach is of superior performance compared with vanilla few samples training baseline and other SOTA knowledge transfer methods. We believe we are the first to perform few-sample knowledge distillation on deepfake detection. Wenqi Zhuo, Wei Wang 0025, Jing Dong 0003 |
ICPR | 3 |
| 2022 | AdaDeId: Adjust Your Identity Attribute FreelyabstractFace de-identification has drawn increasing attention in recent years. It is important to protect people’s identity information meanwhile keeping the utility of the face data in many computer vision tasks. We propose a Adaptive De-identification (AdaDeId) method, a novel approach that can freely manipulate the identity attributes of given faces. We introduce an identity decoupling representation learning method, which is based on the autoencoder decoupling model as well as our proposed Identity Decoupling Representation (IDR) loss and Content Retention (CR) loss. Our method encodes the identity information of a face into a unit spherical space, where we can continuously manipulate the identity representation vector. Various de-identified faces derived from an original face can be generated through our method and maintain high similarity to the original image contents. Quantitative and qualitative experiments demonstrate our method achieves state-of-the-art on visual quality and de-identification validity. Tianxiang Ma, Wei Wang 0025, Jing Dong 0003 |
ICPR | 3 |
| 2022 | Defeating DeepFakes via Adversarial Visual ReconstructionabstractExisting DeepFake detection methods focus on passive detection, i.e., they detect fake face images by exploiting the artifacts produced during DeepFake manipulation. These detection-based methods have their limitation that they only work for ex-post forensics but cannot erase the negative influences of DeepFakes. In this work, we propose a proactive framework for combating DeepFake before the data manipulations. The key idea is to find a well defined substitute latent representation to reconstruct target facial data, leading the reconstructed face to disable the DeepFake generation. To this end, we invert face images into latent codes with a well trained auto-encoder, and search the adversarial face embeddings in their neighbor with the gradient descent method. Extensive experiments on three typical DeepFake manipulation methods, facial attribute editing, face expression manipulation, and face swapping, have demonstrated the effectiveness of our method in different settings. Ziwen He, Wei Wang 0025, Weinan Guan, Jing Dong 0003, Tieniu Tan |
ACM Multimedia | 2 |
| 2022 | Counterfactual Image Enhancement for Explanation of Face Swap Deepfakes
Bo Peng 0002, Siwei Lyu, Wei Wang 0025, Jing Dong 0003 |
PRCV (2) | 3 |
| 2022 | Revisiting ensemble adversarial attack
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Signal Process. Image Commun. | 2 |
| 2022 | Detecting Compressed Deepfake Videos in Social Networks Using Frame-Temporality Two-Stream Convolutional NetworkabstractThe development of technologies that can generate Deepfake videos is expanding rapidly. These videos are easily synthesized without leaving obvious traces of manipulation. Though forensically detection in high-definition video datasets has achieved remarkable results, the forensics of compressed videos is worth further exploring. In fact, compressed videos are common in social networks, such as videos from Instagram, Wechat, and Tiktok. Therefore, how to identify compressed Deepfake videos becomes a fundamental issue. In this paper, we propose a two-stream method by analyzing the frame-level and temporality-level of compressed Deepfake videos. Since the video compression brings lots of redundant information to frames, the proposed frame-level stream gradually prunes the network to prevent the model from fitting the compression noise. Aiming at the problem that the temporal consistency in Deepfake videos might be ignored, we apply a temporality-level stream to extract temporal correlation features. When combined with scores from the two streams, our proposed method performs better than the state-of-the-art methods in compressed Deepfake videos detection. Xin Liao 0001, Wei Wang 0025, Zheng Qin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Unified Framework for High Fidelity Face Swap and Expression ReenactmentabstractFace manipulation techniques improve fast with the development of powerful image generation models. Two particular face manipulation methods, namely face swap and expression reenactment attract much attention for their flexibility and ease to generate high quality synthesis results. Recently, these two subjects are actively studied. However, most existing methods treat the two tasks separately, ignoring their underlying similarity. In this paper, we propose to tackle the two problems within a unified framework that achieves high quality synthesis results. The enabling component for our unified framework is the clean disentanglement of 3D pose, shape, and expression factors and then recombining them for different tasks accordingly. We then use the same set of 2D representations for face swap and expression reenactment tasks that are input to a common image translation model to directly generate the final synthetic images. Once trained, the proposed model can accomplish both face swap and expression reenactment tasks for previously unseen subjects. Comprehensive experiments and comparisons show that the proposed method achieves high fidelity results in multiple aspects, and it is especially good at faithfully preserving source facial shape in the face swap task, and accurately transferring facial movements in the expression reenactment task. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Guided Erasable Adversarial Attack (GEAA) Toward Shared Data ProtectionabstractIn recent years, there has been increasing interest in studying the adversarial attack, which poses potential risks to deep learning applications and has stimulated numerous researches, e.g. improving the robustness of deep neural networks. In this work, we propose a novel double-stream architecture – Guided Erasable Adversarial Attack (GEAA) – for protecting high-quality labeled data with high commercial values under data-sharing scenarios. GEAA contains three phases, the double-stream adversarial attack, denoising reconstruction, and watermark extraction. Specifically, the double-stream adversarial attack injects erasable perturbations into the training data to avoid database abuse. The denoising reconstruction rebuilds the traceable denoising data from adversarial examples. The watermark extraction recovers identity information from the denoised data for copyright protection. Additionally, we introduce the annealing optimization strategy to balance these phases and a boundary constraint to degrade the availability of adversarial examples. Through extensive experiments, we demonstrate the effectiveness of the proposed framework in data protection. The Pytorch® implementations of GEAA can be downloaded from an open-source Github project https://github.com/Dlut-lab-zmn/ GEAA-for-data-protection. Mengnan Zhao 0001, Bo Wang 0024, Wei Wang 0025, Yuqiu Kong, Tianhang Zheng, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Exploring Adversarial Fake Images on Face ManifoldabstractImages synthesized by powerful generative adversarial network (GAN) based methods have drawn moral and privacy concerns. Although image forensic models have reached great performance in detecting fake images from real ones, these models can be easily fooled with a simple adversarial attack. But, the noise adding adversarial samples are also arousing suspicion. In this paper, instead of adding adversarial noise, we optimally search adversarial points on face manifold to generate anti-forensic fake face images. We iteratively do a gradient-descent with each small step in the latent space of a generative model, e.g. Style-GAN, to find an adversarial latent vector, which is similar to norm-based adversarial attack but in latent space. Then, the generated fake images driven by the adversarial latent vectors with the help of GANs can defeat main-stream forensic models. For examples, they make the accuracy of deepfake detection models based on Xception or EfficientNet drop from over 90% to nearly 0%, mean-while maintaining high visual quality. In addition, we find manipulating noise vectors n at different levels have different impacts on attack success rate, and the generated adversarial images mainly have changes on facial texture or face attributes. Wei Wang 0025, Hongxing Fan, Jing Dong 0003 |
CVPR | 2 |
| 2021 | MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationabstractPose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem, we propose a novel multi-level statistics transfer model, which disentangles and transfers multi-level appearance features from person images and merges them with pose features to reconstruct the source person images themselves. So that the source images can be used as supervision for self-driven person image generation. Specifically, our model extracts multi-level features from the appearance encoder and learns the optimal appearance representation through attention mechanism and attributes statistics. Then we transfer them to a pose-guided generator for re-fusion of appearance and pose. Our approach allows for flexible manipulation of person appearance and pose properties to perform pose transfer and clothes style transfer tasks. Experimental results on the DeepFashion dataset demonstrate our method’s superiority compared with state-of-the-art supervised and unsupervised methods. In addition, our approach also performs well in the wild. Tianxiang Ma, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
CVPR | 3 |
| 2021 | A Features Decoupling Method for Multiple Manipulations Identification in Image Operation ChainsabstractRecently, many forensic techniques have been developed to detect the use of a certain processing operation. When utilizing several manipulations to alter an image, artifacts left by manipulations that have been applied later can potentially disguise traces left by manipulations that were applied earlier. Therefore, the detection of manipulations become difficult. In this paper, we focus on identifying the manipulations in an image operation chain composed of multiple manipulations in a certain order. To address this issue, we analyze the relationship between manipulations identification and blind signal separation. Then, we propose a features decoupling method based on blind signal separation, which decouples the coupled features due to the superimposed processing artifacts and exploits the decoupled features to identify multiple operations. The experiments carried out on two image operation chains confirm the effectiveness of the proposed method. Xin Liao 0001, Wei Wang 0025, Zheng Qin 0001 |
ICASSP | 3 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 3 |
| 2021 | SOGAN: 3D-Aware Shadow and Occlusion Robust GAN for Makeup TransferabstractIn recent years, virtual makeup applications have become more and more popular. However, it is still challenging to propose a robust makeup transfer method in the real-world environment. Current makeup transfer methods mostly work well on good-conditioned clean makeup images, but transferring makeup that exhibits shadow and occlusion is not satisfying. To alleviate it, we propose a novel makeup transfer method, called 3D-Aware Shadow and Occlusion Robust GAN (SOGAN). Given the source and the reference faces, we first fit a 3D face model and then disentangle the faces into shape and texture. In the texture branch, we map the texture to the UV space and design a UV texture generator to transfer the makeup. Since human faces are symmetrical in the UV space, we can conveniently remove the undesired shadow and occlusion from the reference image by carefully designing a Flip Attention Module (FAM). After obtaining cleaner makeup features from the reference image, a Makeup Transfer Module (MTM) is introduced to perform accurate makeup transfer. The qualitative and quantitative experiments demonstrate that our SOGAN not only achieves superior results in shadow and occlusion situations but also performs well in large pose and expression variations. Yueming Lyu, Jing Dong 0003, Bo Peng 0002, Wei Wang 0025, Tieniu Tan |
ACM Multimedia | 4 |
| 2021 | Learning pose-invariant 3D object reconstruction from single-view images
Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 2 |
| 2021 | A new globally adaptive k-nearest neighbor classifier based on local mean optimization
Zhibin Pan, Yiwei Pan, Wei Wang 0025 |
Soft Comput. | 4 |
| 2021 | Adversarial Analysis for Source Camera IdentificationabstractRecent studies highlight the vulnerability of convolutional neural networks (CNNs) to adversarial attacks, which also calls into question the reliability of forensic methods. Existing adversarial attacks generate one-to-one noise, which means these methods have not learned the fingerprint information. Therefore, we introduce two powerful attacks, fingerprint copy-move attack, and joint feature-based auto-learning attack. To validate the performance of attack methods, we move a step ahead and introduce the higher possible defense mechanism relation mismatch. which expands the characterization differences of classifiers in the same classification network. Extensive experiments show that relation mismatch is superior in recognizing adversarial examples and prove that the proposed fingerprint-based attacks are more powerful. Both proposed attacks show excellent attack transferability to unknown samples. The Pytorch® implementations of these methods can download from an open-source GitHub projecthttps://github.com/Dlut-lab-zmn/Source-attack. Bo Wang 0024, Mengnan Zhao 0001, Wei Wang 0025, Xiaorui Dai, Yi Li 0018, Yanqing Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Are You Confident That You Have Successfully Generated Adversarial Examples?abstractDeep neural networks (DNNs) have seen extensive studies on image recognition and classification, image segmentation, and related topics. However, recent studies show that DNNs are vulnerable in defending adversarial examples. The classification network can be deceived by adding a small amount of perturbation to clean samples. There are challenges when researchers want to design a general approach to defend against a wide variety of adversarial examples. To solve this problem, we introduce a defensive method to prevent adversarial examples from generating. Instead of designing a stronger classifier, we built a more robust classification system that can be viewed as a structural black box. After adding a buffer to the classification system, attackers can be efficiently deceived. The real evaluation results of the generated adversarial examples are often contrary to what the attacker thinks. Additionally, we do not assume a specific attack method premise. This incognizance to underlying attacks demonstrates the generalizability of the buffer to potential adversarial attacks. Extensive experiments indicate that the defense method greatly improves the security performance of DNNs. Bo Wang 0024, Mengnan Zhao 0001, Wei Wang 0025, Fei Wei, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | A new fast search algorithm for exact k-nearest neighbors based on optimal triangle-inequality-based check strategy
Yiwei Pan, Zhibin Pan, Wei Wang 0025 |
Knowl. Based Syst. | 4 |
| 2019 | An Accurate LSTM Based Video Heart Rate Estimation Method
Mingyun Bian, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
PRCV (3) | 3 |
| 2018 | DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm RetrievalabstractThere are great demands for automatically regulating inappropriate appearance of shocking firearm images in social media or identifying firearm types in forensics. Image retrieval techniques have great potential to solve these problems. To facilitate research in this area, we introduce Firearm 14k, a large dataset consisting of over 14,000 images in 167 categories. It can be used for both fine-grained recognition and retrieval of firearm images. Recent advances in image retrieval are mainly driven by fine-tuning state-of-the-art convolutional neural networks for retrieval task. The conventional single margin contrastive loss, known for its simplicity and good performance, has been widely used. We find that it performs poorly on the Firearm 14k dataset due to: (1) Loss contributed by positive and negative image pairs is unbalanced during training process. (2) A huge domain gap exists between this dataset and ImageNet. We propose to deal with the unbalanced loss by employing a double margin contrastive loss. We tackle the domain gap issue with a two-stage training strategy, where we first fine-tune the network for classification, and then fine-tune it for retrieval. Experimental results show that our approach outperforms the conventional single margin approach by a large margin (up to 88.5% relative improvement) and even surpasses the strong triplet-loss-based approach. Jiedong Hao, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICPR | 3 |
| 2018 | Ensemble Reversible Data HidingabstractThe conventional reversible data hiding (RDH) algorithms often consider the host as a whole to embed a secret payload. In order to achieve satisfactory rate-distortion performance, the secret bits are embedded into the noise-like component of the host such as prediction errors. From the rate-distortion optimization view, it may be not optimal since the data embedding units use the identical parameters. This motivates us to present a segmented data embedding strategy for efficient RDH in this paper, in which the raw host could be partitioned into multiple subhosts such that each one can freely optimize and use the data embedding parameters. Moreover, it enables us to apply different RDH algorithms within different subhosts, which is defined as ensemble. Notice that, the ensemble defined here is different from that in machine learning. Accordingly, the conventional operation corresponds to a special case of the proposed work. Since it is a general strategy, we combine some state-of-the-art algorithms to construct a new system using the proposed embedding strategy to evaluate the rate-distortion performance. Experimental results have shown that, the ensemble RDH system could outperform the original versions in most cases, which has shown the superiority and applicability. Hanzhou Wu, Wei Wang 0025, Jing Dong 0003, Hongxia Wang 0001 |
ICPR | 2 |
| 2018 | Feature learning for steganalysis using convolutional neural networks
Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
Multim. Tools Appl. | 3 |
| 2018 | Image Forensics Based on Planar Contact Constraints of 3D ObjectsabstractStanding objects on planar surfaces are common to see in images, e.g., people on the ground. For most objects to stay stable on the plane, planar contact is a necessary requirement. However, 2D image splicing usually disregards this physical constraint of 3D world, leading to a potential artifact of object not attached to the plane. This paper is the first attempt to use the contact constraint of standing objects as a new clue for image forensics. Accordingly, we propose a novel approach to first reconstruct the 3D poses of standing objects and their supporting plane and then measure the contact conditions for splicing detection. To tackle the problem of unknown object shape for pose estimation, we effectively employ the prior knowledge of 3D morphable model to simultaneously estimate both shape and pose parameters by fitting to image observations. The 3D normal orientation of the supporting plane is estimated given its vanishing line. Dealing with uncertainty factors in estimations, we approximate a distribution of estimates using sampling strategies and then make the final decision. Particularly, we focused our method on the important scenario of human figure splicing detection, and comprehensive experiments on multiple data sets and typical images proved the encouraging effectiveness of the new forensic clue and the proposed approach. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Matrix Separation Based on LMaFit-SeedabstractMatrix separation has a wide range of potential applications and many approaches have been devised to solve it. Especially, for applications involving large-scale data, such as vision tasks, improving the scalability of algorithms has attracted much attention. Reviewing these methods, they mainly involve convex optimization and factorization optimization. Convex optimization models become increasingly costly as the matrix size and rank grow. Factorization optimization models, to a large extent, reduce the computational complexity of matrix separation. l1-filtering applied the generalized Nyström to matrix separation, and proposed the seed-based convex optimization. In this paper, to benefit from the seed strategy, we propose matrix separation based on Low-Rank Matrix Fitting (LMaFit)-Seed, which is an algorithm on low-rank factorization optimization, to enhance the scalability in solving the problems of large-scale matrix separation and to be less time-consuming. We evaluate the proposed method and demonstrate comparisons with several state-of-the-art methods on synthetic data simulations and real sequences experiments. Hai-Xia Xu 0001, Wei Zhou 0027, Yaonan Wang 0001, Wei Wang 0025, Yan Mo |
Comput. J. | 4 |
| 2017 | Optimized 3D Lighting Environment Estimation for Image Forgery DetectionabstractImage forgery is becoming a growing threat to information credibility. Among all kinds of image forgeries, photographic composites of human faces have very serious impacts. To combat this kind of forgery, some forensic methods propose to estimate the 3D lighting environments from different faces and investigate the consistency between them. Although they are very effective, existing 3D lighting-based forensic methods are limited by many simplifying assumptions about the surface reflection model, among which convexity and constant reflectance are two critical ones. In this paper, we propose an optimized 3D lighting estimation method by incorporating a more general surface reflection model. In this model, we relax the convexity and constant reflectance assumptions by taking the occlusion geometry and surface texture information into consideration. The proposed reflection model is more general and accurate; hence, it can achieve better lighting estimation accuracy and more reliable discrimination performance. Comprehensive experiments on both synthetic and real data sets validate the correctness and efficacy of the proposed method. Comparisons with two existing 3D lighting-based forensic methods also demonstrate the superiority of the proposed method for detecting face splicing. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Automatic detection of 3D lighting inconsistencies via a facial landmark based morphable modelabstractExisting 3D lighting consistency based forensic methods have some practical problems. They usually require additional images and human labor to reconstruct the 3D face model for lighting estimation, and furthermore, they cannot deal with expressional faces effectively. These drawbacks make them unusable in many practical cases. In this paper, we propose a more practical 3D lighting based forensic method by incorporating a facial landmark based 3D morphable model to efficiently fit the face shape. We also introduce a residual error based algorithm to automatically exclude outliers in lighting estimation. Our proposed method is fully automatic and very efficient compared to previous ones. Also, it does not depend on additional images and has better performance for expressional faces. Experiments on a realistic face dataset with variational lighting conditions indicate the efficacy and superiority of our method. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2016 | Learning and transferring representations for image steganalysis using convolutional neural networkabstractThe major challenge of machine learning based image steganalysis lies in obtaining powerful feature representations. Recently, Qian et al. have shown that Convolutional Neural Network (CNN) is effective for learning features automatically for steganalysis. In this paper, we follow up this new paradigm in steganalysis, and propose a framework based on transfer learning to help the training of CNN for steganalysis, hence to achieve a better performance. We show that feature representations learned with a pre-trained CNN for detecting a steganographic algorithm with a high payload can be efficiently transferred to improve the learning of features for detecting the same steganographic algorithm with a low pay-load. By detecting representative WOW and S-UNIWARD steganographic algorithms, we demonstrate that the proposed scheme is effective in improving the feature learning in CNN models for steganalysis. Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 3 |
| 2015 | Robust steganalysis based on training set construction and ensemble classifiers weightingabstractThe cover source mismatch problem in steganalysis is a serious problem which keeps current steganalysis from practical use. It is mainly because of the high intra-class variation of cover and stego samples in the feature space, since current steganalytic features are inevitably affected much by the image content, size, quality and many other factors. Small training set often reflects only part of the real data distribution, hence the classifier (steganalyzer) may be undertrained and lack of robustness. In this paper, we propose a scheme to efficiently construct large representative training set for steganalysis. We also scheme out weighted ensemble classifiers which can be adaptive to testing data. Experimental results show that our method can improve the performance and robustness of ste-ganalysis under high intra-class variation. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 3 |
| 2014 | An effective watermarking method against valumetric distortionsabstractMost of the quantization based watermarking algorithms are very sensitive to valumetric distortions, while these distortions are regarded as common processing in audio/video analysis. In recent years, watermarking methods which can resist this kind of distortions have attracted a lot of interests. But still many proposed methods can only deal with one certain kind of valumetric distortion as amplitude scaling, and fail in other kinds of valumetric distortions like constant change attack or gamma correction. In this paper, we propose a very simple method to tackle all the three kinds of valumetric distortions. A constant change invariant domain is first constructed by spread transform, in which the watermark is embedded using a certain amplitude scaling invariant based watermarking scheme. Several typical watermarking methods and attacks have been implemented in our experiments to demonstrate the effectiveness of the proposed method. Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 3 |
| 2014 | Effects of Fragile and Semi-fragile Watermarking on Iris Recognition System
Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 3 |
| 2014 | Exploring DCT Coefficient Quantization Effects for Local Tampering DetectionabstractIn this paper, we focus on local image tampering detection. For a JPEG image, the probability distributions of its DCT coefficients will be disturbed by tampering operation. The tampered region and the unchanged region have different distributions, which is an important clue for locating tampering. Based on the assumption of Laplacian distribution of unquantized ac DCT coefficients, these two distributions as well as the size of tampered region can be estimated so that the probability of each DCT block being tampered is obtained. More accurate localization results could be got when we consider the prior knowledge of common tampered regions. We also design three kinds of features that can distinguish truly tampered regions from the false ones to reduce false alarm. For a tampered image which is saved in lossless compressed format, we also propose the specialized approach, which employs the quantization noise of high-frequency DCT coefficient, to improve the tampering localization performance. Extensive experiments on large scale databases prove the effectiveness of our proposed method and demonstrate that our method is suitable for locating tampered regions with different scales. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Video steganalysis based on the constraints of motion vectorsabstractIn this paper, we focus on detecting data hiding in motion vectors of compressed video and propose a new steganalytic algorithm based on the mutual constraints of motion vectors. The constraints of motion vectors from multiple frames are analyzed and formulized by three functions, then statistical features are extracted based on these functions. Moreover, we also incorporate calibration method to improve the detection accuracy. Experimental results demonstrate that the proposed method can effectively attack typical motion-vector-based video steganography. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 3 |
| 2012 | Continuum regression for cross-modal multimedia retrievalabstractUnderstanding the relationship among different modalities is a challenging task. The frequently used canonical correlation analysis (CCA) and its variants have proved effective for building a common space in which the correlation between different modalities is maximized. In this paper, we show that CCA and its variants may cause information dissipation when switching the modals, and thus propose to use the continuum regression (CR) model to handle this problem. In particular, the CR model with a fixed variance coefficient of 1/2 is adopted here. We also apply the multinomial logistic regression model for further classification task. To evaluate the CR model, we perform a series of cross-modal retrieval experiments in terms of two kinds of modals, namely image and text. Compared with previous methods, experimental results show that the CR model has achieved the best retrieval precision, which demonstrates the potential of our method for real internet search applications. Yongming Chen, Liang Wang 0001, Wei Wang 0025, Zhang Zhang 0001 |
ICIP | 3 |
| 2010 | Image tampering detection based on stationary distribution of Markov chainabstractIn this paper, we propose a passive image tampering detection method based on modeling edge information. We model the edge image of image chroma component as a finite-state Markov chain and extract low dimensional feature vector from its stationary distribution for tampering detection. The support vector machine (SVM) is utilized as classifier to evaluate the effectiveness of the proposed algorithm. The experimental results in a large scale of evaluation database illustrates that our proposed method is promising. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 1 |
| 2010 | Tampered Region Localization of Digital Color Images Based on JPEG Compression Noise
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 1 |
| 2009 | Effective image splicing detection based on image chromaabstractA color image splicing detection method based on gray level co-occurrence matrix (GLCM) of thresholded edge image of image chroma is proposed in this paper. Edge images are generated by subtracting horizontal, vertical, main and minor diagonal pixel values from current pixel values respectively and then thresholded with a predefined threshold T. The GLCMs of edge images along the four directions serve as features for image splicing detection. Boosting feature selection is applied to select optimal features and Support Vector Machine (SVM) is utilized as classifier in our approach. The effectiveness of the proposed method has been demonstrated by our experimental results. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 1 |
| 2009 | Multi-class Blind Steganalysis Based on Image Run-Length Analysis
Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 2 |
| 2009 | A Survey of Passive Image Tampering Detection
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 1 |
| 2008 | Run-Length and Edge Statistics Based Approach for Image Splicing Detection
Jing Dong 0003, Wei Wang 0025, Tieniu Tan, Yun Q. Shi 0001 |
IWDW | 2 |